Did worries about calculators and medical aids really fade over time, and will AI worries fade the same way?
Have similar doubts about calculators or diagnostic aids actually disappeared over time?
This explores whether today's worries about relying on AI will fade the way worries about calculators and medical decision aids supposedly did, and whether the corpus can tell us if that comparison holds.
This reads the question as a test of a familiar argument: people once doubted calculators and diagnostic aids, those doubts faded, so doubts about AI will fade too. To be direct, the corpus has no historical material on calculators or clinical decision-support tools, so it can't say whether those doubts really disappeared. What it can show is why the comparison is shakier than it looks. A calculator earned trust because it is deterministic, narrow and easy to check. The AI research here describes tools that fail in ways that are much harder to notice.
Start with checkability. A calculator gives the same answer every time, and a wrong one tends to look wrong. Reasoning models don't behave that way. Spending more effort on a problem can make them drop a correct answer they had already reached Why does more reasoning sometimes make models worse?. They will also reason at length about questions that can't be answered, where a simpler model would just say the question is missing information Why do reasoning models overthink ill-posed questions?. Even the evidence that a model has improved can mislead: one widely used model could rebuild more than half of a standard math benchmark from partial prompts but scored 0% on problems published after its training Does RLVR success on math benchmarks reflect genuine reasoning improvement?. Nobody ever had to ask whether a calculator had memorized the test.
The diagnostic-aid comparison runs into a different problem. In medicine, AI performance depends mostly on accurate domain knowledge, not on reasoning skill. Models trained to reason well at math don't beat base models on medical tasks Why doesn't mathematical reasoning transfer to medicine?. So 'AI got good at reasoning' doesn't carry over to 'AI can be trusted at diagnosis' the way 'the calculator does arithmetic' carries over to every arithmetic task.
The most surprising evidence is about people rather than models. When users compared formats for showing a model's reasoning, they preferred the polished planning-and-decomposition style. But plain step-by-step reasoning helped them catch errors better, and the format they liked led to more false alarms and more misplaced trust Do people prefer the reasoning formats that help them verify?. Separately, AI suggestions can hurt a person's reasoning even when they are correct, because they break concentration Does AI assistance always help reasoning or does it carry hidden costs?. Both point the same way: doubts about AI may fade because people get comfortable with it, not because it becomes more reliable.
The takeaway is that the calculator story can't settle this question either way. The real test is whether a tool's mistakes can be seen. Doubts about calculators could fade because errors were rare and easy to spot. With AI, growing comfort may simply mean people stop noticing the errors. For the history itself, you would need sources outside this collection.
Sources 6 notes
Tracking flip events shows that extra reasoning tokens don't just hit diminishing returns—they actively cause models to second-guess and overwrite previously-correct answers, making accuracy non-monotonic in trace length.
Reasoning models generate redundant, lengthy responses to questions with missing premises while non-reasoning models correctly identify them as unanswerable. Training optimizes for producing reasoning steps but never teaches models when to disengage.
Qwen2.5-Math-7B reconstructs 54.6% of MATH-500 from partial prompts but scores 0.0% on post-release LiveMathBench, revealing dataset contamination. On clean benchmarks, only correct rewards improve performance; random and inverse rewards fail or degrade reasoning ability.
R1-distilled reasoning models fail to outperform base models on medical tasks because knowledge accuracy matters more than reasoning quality in medicine—the opposite of math. Fine-tuning cannot close this gap without domain-specific training data.
A controlled study found participants preferred planning and decomposition formats, yet simpler chain-of-thought traces better supported error detection, trust calibration, and interpretability. The favored formats increased false alarms and unwarranted trust.
Show all 6 sources
Well-intentioned AI suggestions can damage reasoning performance by severing cognitive immersion, forcing users to rebuild focus before continuing. Evaluation must measure flow preservation across entire tasks, not just local suggestion accuracy.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
- Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains
- Do Reasoning Representations Help Humans Evaluate LLM Outputs?