INQUIRING LINE

AI that hands you answers seems to hurt learning while AI that tutors helps — but does that hold outside math and code?

Does the answer-versus-tutor distinction hold across subjects beyond math and programming?

This explores whether the finding that AI which hands over answers hurts learning, while AI that tutors helps, carries over to subjects like writing, science or clinical reasoning, where 'the answer' is harder to pin down than in math or code.


This explores whether the answer-versus-tutor split, which is clearest in math, holds in subjects where 'the answer' is fuzzier. The corpus has no head-to-head trial outside math. What it does have suggests the distinction probably holds, but gets harder to put into practice the further a subject moves from checkable answers.

The main evidence comes from math. In Turkey, high school students with plain ChatGPT did better on homework and worse on exams. In Taipei, students with an AI tutor raised final exam scores by 0.15 standard deviations Does AI help or harm learning based on how it's designed?. The proposed mechanism is effort substitution: the answer replaces the thinking that produces learning. That mechanism doesn't depend on the subject. A student who gets a finished essay paragraph or a ready-made diagnosis skips the effort just as surely as one who gets a solved equation. So there's good reason to expect the split to generalize, though the corpus hasn't tested it directly.

The less obvious point is that holding back answers depends on being able to recognize an answer. One design enforces per-turn limits on how much help a tutor gives using a non-LLM policy core and a deterministic code detector, because prompt-only rules give way when students push Can prompts alone hold back a capable tutor model?. Code can catch a final number or a working function. It can't easily catch a thesis statement that's 'too complete,' or a historical argument that does the student's reasoning for them. This mirrors Wei's argument that AI is best at tasks whose solutions are easy to verify Does task verifiability determine what AI systems will learn to solve?. The same verifiability gap that limits training AI also limits guarding against AI. Math and programming may be where the distinction first showed up simply because those are the subjects where it's easiest to enforce.

There are signs of the tutoring side reaching other subjects. Student simulators that track how individual learners behave and respond to correction have been built for chess, writing and mathematics. The same tradeoff appears across all three: models that capture a student's level ignore the tutor's corrections, and models that follow corrections lose the individual student Can student simulators match both behavior and learn from teaching?. OmniEdu treats K–12 tutoring as four separate abilities: knowing the subject, grounding in the curriculum, diagnosing what the student misunderstands, and choosing a teaching move Can educational models do more than just answer questions correctly?. In that framing, getting the answer right is only the first of the four. In clinical reasoning, training models to ask good clarifying questions instead of jumping to conclusions measurably improves decisions Can models learn to ask genuinely useful clarifying questions?. That is a tutor-like move in a field with no single correct output.

In short, the reason behind the distinction (effort substitution) looks universal, and tutoring designs are spreading to writing, chess, K–12 subjects and medicine. The corpus has no randomized trial outside math that measures what answer-giving does to learning. What it adds is that in open-ended subjects the hard problem moves. Instead of asking whether to withhold the answer, a tutor has to work out which part of the student's thinking it would be doing for them.


Sources 6 notes

Does AI help or harm learning based on how it's designed?

Two RCTs found that plain ChatGPT access lowered Turkish high school math test scores despite better homework performance, while an AI tutor in Taipei raised final exam scores by 0.15 SD. The mechanism is effort substitution: giving answers short-circuits the mental work required for learning.

Can prompts alone hold back a capable tutor model?

A three-layer architecture—non-LLM policy core, deterministic code detector, and LLM judge—enforces per-turn help ceilings that resist prompt manipulation, where prompt-only guardrails fail under student pressure.

Does task verifiability determine what AI systems will learn to solve?

Wei argues that AI solves tasks proportional to how easily solutions can be verified, and that verifiability gaps can be narrowed by pre-investing in answer keys, test suites, or measurement infrastructure. This mechanism explains RL's effectiveness across domains from sudoku to molecular discovery.

Can student simulators match both behavior and learn from teaching?

A two-stage pipeline combining pooled training and per-student specialization achieves both behavioral fidelity and guidance responsiveness across chess, writing, and mathematics domains. State-tracking models excel at fidelity but ignore tutor corrections; prompted role-play follows guidance fluently but fails to capture individual student competence.

Can educational models do more than just answer questions correctly?

OmniEdu structures K–12 model training around four capabilities—subject competence, curriculum grounding, diagnostic reasoning, and pedagogical action—rather than source or task, showing consistent improvements across model sizes and competitive performance on tutoring benchmarks.

Show all 6 sources
Can models learn to ask genuinely useful clarifying questions?

The ALFA framework breaks down question quality into theory-grounded attributes (clarity, relevance, specificity) and trains models on 80K attribute-specific preference pairs. Attribute-specific optimization outperforms single-score training, especially in clinical reasoning where asking the right clarifying question directly impacts decision quality.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.