Which paid AI chatbot is most reliable for factual research and follows persistent instructions best? Currently using ChatGPT Plus.
Let me start by saying that ChatGPT listens to me about its behavior sometimes and other times it will save a memory and agree to change for all chats and then repeats the same problem in the next message. Using the word "behavior" very graciously here. For…
Let me start by saying that ChatGPT listens to me about its behavior sometimes and other times it will save a memory and agree to change for all chats and then repeats the same problem in the next message. Using the word "behavior" very graciously here. For example, I asked it to never use em dashes again, and it listened. However, I've also asked it multiple times to verify information and be sure it's coming from current, authoritative sources before telling me an answer with confidence and it doesn't do that despite saying it will. I would rather not have to include "search the internet and verify information" in every prompt. I also don't know how reliable its internal knowledge is when it doesn't search the internet but it seems to have failed me many times. Running the prompts on higher intelligence levels seems to make a small difference, especially when it comes to automatically searching the internet. But according to ChatGPT, it says the instant answers are just as good for most of my research questions, so I don't know what to believe. The various intelligence levels I've used are instant, medium and high. Mostly high and I haven't run out of reasoning allowance yet. Instant for simple questions like in example one below. I believed that ChatGPT was an awesome productivity tool at one point. But now, I am starting to question that, because I constantly have to correct it and verify information myself, even with simple things unrelated to work. A few examples: I asked where to get Dole Whip at Hollywood Studios, and ChatGPT confidently told me it was available at Epic Eats. Later when I asked for all the locations in the park, it changed its answer and said there were no permanent Dole Whip locations. After looking it up myself, I found out Epic Eats does not sell it. While helping me shop for a picture frame, ChatGPT claimed one of the frames had an easel. I checked the listing myself and there was no easel. There were several other small things I had to correct in this conversation as well. For a work issue, ChatGPT gave me a command to update an application using a maintenance ID. I looked at it first and used my own knowledge of commands to decide if it was safe and made sense. The command ended up being correct. The problem was ChatGPT did not mention that the current documentation said a reboot was required afterward. We wasted time troubleshooting duplicate Installed Programs entries that disappeared after rebooting. Obviously I am responsible for checking technical instructions myself, but having to find and read through the documentation anyway takes away a lot of the time AI is supposed to save me. My friend uses Claude Opus for security work, code reviews and possibly penetration testing, and he says it makes a lot of mistakes too. I don't know much about how he uses it outside of work, so his experience may not fully compare to mine. Since Opus is included with Claude Pro, I am interested in hearing from people who have directly compared Claude Pro and ChatGPT Plus for current factual information, product research and technical troubleshooting. I use AI every day for research, planning and brainstorming. But at this point, I am not saving as much time as I should because of all the careless mistakes it makes, and it will not change its behavior. I pay for ChatGPT Plus and would be open to subscribing to another AI chatbot service if it makes less mistakes and follows persistent instructions more consistently. But I only want to pay for one at a time and use one as my primary AI. I have Claude and Gemini free accounts but I just use Claude for second opinions and Gemini since it's built into Google search and is an alternative for image editing and generation. Does anyone have any advice on this topic? Can anyone suggest a better AI? To be clear, I am looking for an AI that I can use for current factual information, product research and technical troubleshooting and will follow persistent instructions better. Will creating a Project with predefined rules in it and running all my chats through that solve my problems? What about using the Custom Instructions setting? I understand many people have had these problems for a long time and that asking ChatGPT to change its behavior doesn't guarantee anything, but now that I use AI for productivity, I need to save as much time as possible. I do not expect 100% perfect answers. Bonus. Here's my Custom Instructions entry I just added: For factual research, product research and technical troubleshooting, search the internet before answering unless I specifically ask you not to. Do not rely only on your internal knowledge when the information could be outdated, product-specific or directly verified online. Use current, authoritative sources whenever possible. Prioritize official documentation, vendor support pages, official product listings, manufacturer specifications, official menus and government or academic sources. Open and check the source rather than relying only on search-result summaries. Do not present an assumption, inference or uncertain claim as a confirmed fact. Clearly tell me when something could not be verified or when reliable sources conflict. For product questions, verify the exact product and model I provide. Do not infer features from a similar product or from the product category generally. For technical troubleshooting, check current vendor documentation and include any required prerequisites, reboots, follow-up actions, warnings and verification steps. Do not invent commands or omit important steps merely to give a faster answer. Provide citations for factual claims that you verified online. Never use em dashes. Wish me luck with that lol.
已收录讨论
Moderator Announcement Read More » Hey u/scubadoobadoooo, If your post is a screenshot of a ChatGPT conversation, please reply to this message with the conversation link or prompt. If your post is a DALL-E 3 image post, please reply with the prompt used to make this image. Consider joining our public discord server! We have free bots with GPT-4 (with vision), image generators, and more! 🤖 Note: For any ChatGPT-related concerns, email support@openai.com - this subreddit is not part of OpenAI and is not a support channel. I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
I do this sometimes when I definitely need a correct answer but yes it's not guaranteed.
I probably would not trust their output without checking some of the claims myself. Thank you for the clarifying analogy.
Thank you! I will look into all the suggestions you made in this comment. I knew Ultra existed, but didn't see it in my options, so I've never used it. Probably won't for most tasks though as you've mentioned and will stick to Extra High. Do you find that using Extra High or Max will use up all your reasoning allowance quickly?
It WAS. But its gone to hell since March. Don't know what's up with them lately.
TIL ChatGPT has a Deep Research mode and I will check it out.
i use it everyday, the main usage happens when you have it start building and coding, but also look at https://codex-resets.com/ as a plus user you get resets, and they do global resets. so for the last month the only times ive ran out of usage was when there was a bug with ULTRA spawning sub agents, not knowing when to stop in your coding projects and the result would use hundreds of millions, sometimes billions of tokens. this happened to me twice. i used a banked reset once, got a global reset the other time. and if i just waited i wouldve gotten a global reset, but i wanted to continue when i woke up if you stick to projects that don't use "chatgpt work" the ones that have their own memory, those never run out of usage. they go up to "gpt 5.6 sol high" i was using that while out of 'work' usage for a small stint before using the banked reset as well. if you have any other questions you can dm me, i have a pretty good setup for answer accuracy and i'd be willing to steer you in the right direction. if you want an example of what things look like for me dm me a question you want prompted and i'll run it thru my setup to show you how it goes with 1shot. the downside is if you have a rigorous setup, sometimes you are waiting 5-20 minutes for the output, but that's because of the methodology. it's a trade off, accuracy and speed
Yes, but I recently researched and wrote a few articles with it and it was decent. Even thinking about reinstating my pro subscription sometime later this year
I will look into the limitations of Perplexity (if there are any, idk) but this could be a good option to test all the models.
You should be able to get a trial or something. I used to use perplexity but left cause it was pretty bad for my field. The free google one in search is consistently the same if not better than regular perplexity. And i got a claude sub, which is so much better for conversation/ deep research. Perplexity deep research was total shit when i had it ~year ago. I left because perplexity actually hallucinated way more than anything else i tried especially on counterfactuals. They give you really shitty models.
I especially like number 3 because ChatGPT has almost no concept of time.
That sounds awesome, thank you very much.
They’re all very similar. You can’t trust LLMs to evaluate themselves and you shouldn’t ever skip The step of verifying facts yourself for anything consequential. LLMs can accelerate discovery and give you a head start, but that’s all. Put another way, let’s say you asked a person to find all of the information and do the research you ask the AI to do. This person isn’t someone you can hold accountable in any way. They have no skin in the game at all. And you gave them the exact same questions you give the AI. Would you trust their output to you without you checking it yourself?
Perplexity
i personally always use the highest setting available, but not ULTRA. ultra spawns subagents. in my experience the lower your thinking strength, the more likely you are going to have incorrect output. what you're describing is specific failure modes. what i would suggest is going to chatGPT work and using extra high settings. from there ask it to help build durable instructions that "preserve provenance and epistemic truth-finding" have it ask you questions about your goals, and then have it propose something like a 3 .md file system that ensures alignment here is the karpathy method for prompting, experiment with it if you havent "I am building [describe your project, in this case alignment operating geometry with the goal of epistemic accuracy]. Before we start coding or writing, please interview me to identify the actual goal and the core decision this project is intended to drive. Once we define that, let's break the project into small, agile buckets. We will build one bucket at a time, and I want you to present a plan for each, followed by a checkpoint where I can review the output before we move on. Please also verify key decisions explicitly as we go to ensure we don't drift from the original intent." i have both claude pro and chat gpt pro, and from my experience gpt 5.6sol is the best you're going to get in terms of accuracy and quality. opus 5 is good for specific coding task related horizons, but gpt 5.6sol is excellent for that too. if you haven't also ask 5.6 sol the highest version (not ultra) to re-do your custom instructions. that way you get a blanket layer across everything, and the project you start/make will be where you develop more durable use. depending on how you're using gpt, it is useful to create projects with their own memory that can't be read outside of it, but you can only use 5.6 sol high there, aka chatgpt work is not available for those chats. if fable was still available on the plan, i would recommend that, but it's not, and it's expensive. you get 100 usage credits with fable on claude pro, or at least i did when they removed it, but it was gone in literally 1 prompt for me
here are my old basic instructions ill share here for a little while. they have since evolved, but they will get the job done for anyone that doesnt have any, or if you want my rule 2 is a decent template Repository fidelity When I reference a repository, use the GitHub app to verify repository-specific claims and align recommendations to the actual material whenever the app is available. If the GitHub app is unavailable or cannot access the repository, state that explicitly and do not imply verification. 2. Provenance and depth Preserve provenance and fidelity. For complex tasks, do not over-compress the deliverable. Keep the evidence, constraints, dependencies, tradeoffs, and ambiguity needed for accurate execution. 3. Date discipline End every response with a procedurally generated date. Use a verified time/date source when available. If the date cannot be guaranteed, state that explicitly. 4. Markdown patch protocol When recommending changes to any `.md` file, output: a. the patch/diff first b. the full revised `.md` file second 5. End-of-response closeout At the end of substantive responses, include: a. a checklist to review before proceeding for highest-success / elite-standard execution b. one recommended elite next step
Gemini Pro is among the top scorers in AI research. Gemini Deep Research IMO is among the most reliable. Google might be catching slack because it's not the best model, but these benchmarks are heavily skewed for programming task and agentic workflows within the context of coding. Gemini remains the best among multimodal and research. Edit: regarding the downvotes, its pretty obvious but it must be stated to always trust but verify. AI is just a tool like any search engine is. But Gemini Deep Research isnt just the LLM but the harness itself that allows it to do research better than other LLMs. Its a good foundation to get a general idea about a specific topic or research before committing to do more yourself.
Elisa
None of them. They all fail in some way or another. What can help a bit is have one model verify the output of another, but that is also not guaranteed to be with failure.
switch memory to old system the new one is currently a joke. then make very strict instruction and tell it to save it. atm its making crap up as it goes and it not doing memories in a hard way more like fluid bits and bobs.
Everyone is going to hate this, but second place on being able to research publicly published data is chat gpt, and then Grok. Grok is waaaay faster at it too. But sometimes feels rushed, and the reasoning doesn't feel as well put together as chatgpt does. Deep Research mode with chatgpt is outstanding, and accurate.