Heavy AI Users Face 3x More Hallucinations and Spend 10x Longer to Get Answers
Source: Rev (Centiment survey) · 2025-08
Survey insights show heavy AI users face 3x more hallucinations, take 10x longer for satisfaction, and struggle most with AI prompting despite their experience.
In the three years since ChatGPT's debut, AI prompting has gone from an experimental novelty to an everyday staple. But using it frequently and using it well are two very different things. For example, the same question can get wildly different answers depending on how it's phrased.
To understand how people are navigating the gap, we surveyed over 1,000 artificial intelligence users across different generations, industries, and skill levels about their habits, challenges, and AI results.
A third (34%) of all users say phrasing prompts clearly is their top challenge.
Heavy AI users are nearly 3x more likely to experience frequent hallucinations.
Daily AI users are 14x more likely than casual users to double-check AI's work.
Long prompts are 32x more likely to trigger frequent revisions.
People who feel they're getting better at AI prompting are 64% more likely to say they never experience hallucinations.
For most people, using AI is quick — 77% get an answer they like in under two minutes. But for the power users who live in these tools for over six hours a week, that number drops to just 50%.
Notably, power users are 10x more likely than casual users to find themselves tweaking, rewriting, and wrestling with AI for over 11 minutes until they get an answer they're satisfied with (21% vs. just 2% of light users).
This isn't necessarily tied to inefficiency — complexity is the likely culprit. Heavy users are trying to do much harder things, like getting the AI to perform an in-depth analysis or perfectly mimic a specific writing style. Meanwhile, light users usually stick to simple requests that lead to quick victories.
Of course, it doesn't help that power users have higher standards — once you know how impactful the output can be, it's tough to settle for just "good enough."
Generational differences: 81% of baby boomers use one to three sentences for their prompts vs.
Longer prompt usage: Gen Z users write multi-paragraph prompts more than twice as often as boomers (34% vs.
Baby boomers reach satisfaction within two minutes slightly more often than Gen Z (79% vs. 74%) and report fewer hallucinations (21% never experience them vs. 13% of Gen Z). This suggests that brevity may be a quiet superpower (or that some users simply aren't checking the AI's work too closely).
34% of respondents pointed to "phrasing requests in a way the AI understands" as their biggest AI hurdle, ahead of other common stumbling blocks like knowing the right level of detail to provide (32%), or tailoring instructions to get the specific output they're looking for (26%).
You'd think practice would make perfect, but the data show the opposite. Users who spend over six hours each week with AI tools are far more likely to get stuck on phrasing than casual users (42% vs. 33%). And this isn't a one-off issue.
Simplicity might be an undervalued AI prompting strategy. More than 1 in 4 respondents who are frustrated with AI (28%) say their top challenge is "having to rewrite the same prompt multiple times" — compared to just 17% of efficient users.
According to our survey, only 17% of people said they never have to rewrite their prompts to correct for false or inaccurate information. If so many people are constantly having to rewrite prompts to get an answer they like, it begs the question of whether AI is really as helpful as it’s advertised to be.
Here's a surprising insight from our data: the more you use AI, the more flaws you see.
This doesn't necessarily mean tools are failing advanced users. Instead, it tells us that experienced users are either asking tougher questions or have just gotten much better at noticing when the AI gets something wrong.
When a session drags past 20 minutes, almost 9 out of 10 (88 %) say they "very often" have to revise for hallucinations.
People using very long prompts (e.g., entire documents or web pages, multi-source context, continual updating) are 32× more likely to revise for hallucinations "very often." (32 % vs.
With 700 million weekly users, it's no surprise that ChatGPT is king, and over half (51%) of all survey respondents say it's the AI tool that best understands their prompts.
Its closest competitor, Gemini, lags far behind at 16%. That lead spans every generation, but the gap narrows with age. Nearly three-quarters of Gen Z (72%) prefer ChatGPT, compared to just over a third of baby boomers (35%).
Diving deeper into this age gap, it seems younger users are also more willing to shop around. In contrast, older users tend to pick one tool and stick with it. More than a quarter of boomers (28%) admit they haven't tried enough AI tools to compare, something only 8% of Gen Z would say.
So which tool is actually the best? It depends on what you need. ChatGPT is seen as the most reliable all-rounder, reporting the fewest hallucinations. However, Gemini actually edges it out slightly in quick satisfaction rates (49% vs. 46%).
Additionally, Perplexity is the fastest AI tool to deliver a satisfying response, averaging just 1.9 minutes — the only platform with a sub-two-minute average.
The survey was conducted by Centiment for Rev. The survey was fielded between July 25 and July 28, 2025. The results are based on 1,038 completed surveys. In order to qualify, respondents were screened to be residents of the United States, over 18 years of age, and have used a prompt to ask AI a question. Data is unweighted, and the margin of error is approximately +/-2% for the overall sample with a 95% confidence level.
Lines of inquiry this paper opens 14
Research framings built by reading the notes related to this paper — the questions it feeds into.
Why do language models hallucinate and how can we prevent it?- Do certain news topics trigger more hallucinations than others?
- Does retrieval augmented generation actually eliminate hallucinations in any domain?
- Can architectural changes reduce hallucination without external retrieval or verification?
- Why do hallucination rates differ between vendor AI products and student-used models?
- Do legal AI tools marketed as hallucination-free actually hallucinate?
- How do different legal AI tools compare in accuracy across case eras?
- What happens when lawyers rely on AI citations that turn out false?
- What percentage of AI hallucination cases result in actual court sanctions?
- Do solo lawyers face different citation hallucination risks than large firms?
- How much do existing legal AI tools actually hallucinate in practice?