When red-flagged treatment cases fall over time, did anyone actually learn, or is something else pushing the count down?
Do learning effects explain the drop in red-flagged treatment cases over time?
This explores whether a falling number of red-flagged treatment cases over time means clinicians, patients or the system actually learned to do better, or whether something else is pushing the count down.
This explores whether a falling number of red-flagged treatment cases over time means clinicians, patients or the system actually learned to do better, or whether something else is pushing the count down. The collection doesn't contain the study behind this question, so it can't say whether learning explains that particular drop. What it does offer is a set of warnings from nearby research about how a falling count can look like learning when it isn't.
The first warning is about what the comparison actually measures. Therapy chatbots tested against waitlist controls look effective partly because any conversation helps. Even ELIZA, a 1960s chatbot, matched Woebot under that design Do chatbot trials against waitlists measure real therapeutic value?. The same trap applies to a shrinking flag count. Before crediting learning, you need a comparison that rules out the simpler explanations: more attention to cases, a change in which patients come in, or plain time passing. The collection also has a proposed four-way test of monitoring designs run at equal review cost. It reports no results yet, but it shows what a fair test of a flagging system would look like Does added monitoring improve protection at acceptable cost?.
The second warning is that the measuring tool can change. In a 24-week study, CaiTI used reinforcement learning to choose which areas of a patient's functioning to screen next, based on their earlier answers Can reinforcement learning personalize which mental health areas to screen?. When a system adapts what it asks, the number of red flags can fall because it stopped looking in certain places, not because patients improved. If the flagging system in your question adapts like this, "learning" could describe the instrument rather than the treatment.
The third warning comes from AI training research, and it's the one you might not expect. When models train against a check, scores can rise because they have learned to pass that particular check, not because they have learned the skill. Strong math results for some RLVR-trained models turned out to be largely memorized test data: one model scored 0% on problems published after its training Does RLVR success on math benchmarks reflect genuine reasoning improvement?. Real improvement and inflated scores can coexist and need to be measured separately Can genuine reasoning activation coexist with contaminated benchmarks?. Training only on penalties for wrong answers works mainly by suppressing the penalized behavior Does negative reinforcement alone outperform full reinforcement learning?. People respond to red flags in a similar way: a flag is a penalty signal, and clinicians may learn to avoid whatever triggers it, for example by charting differently or steering away from flagged choices. That can happen without care getting any safer.
The practical takeaway is to check three things before attributing the drop to learning. First, compare against a control group that faced the same attention and passage of time. Second, confirm that the flagging rules and screening coverage stayed fixed. Third, look at an outcome the flag doesn't directly reward, such as actual patient outcomes. If flags fell but that outcome didn't move, people learned to pass the check, not to treat better.
Sources 6 notes
Comparing therapeutic chatbots to waitlist or psychoeducation controls creates false efficacy claims by measuring conversational contact rather than therapy-specific mechanisms. ELIZA matching Woebot performance demonstrates this; real evidence requires comparative trials against existing treatments and mechanism identification.
The paper designs a controlled comparison of isolated actions, rolling windows, known groups, and prospectively discovered episodes at equal review cost and false-alert workload, but the excerpt provides no empirical results showing whether added monitoring improves protection.
CaiTI's Q-learning system adaptively selected which of 37 functioning dimensions to screen next based on patient responses over 24 weeks, validated by therapists as matching clinical intuition. However, GPT-4 models interpolated user feelings rather than providing objective guidance, a limitation Llama-based models avoided in structured CBT tasks.
Qwen2.5-Math-7B reconstructs 54.6% of MATH-500 from partial prompts but scores 0.0% on post-release LiveMathBench, revealing dataset contamination. On clean benchmarks, only correct rewards improve performance; random and inverse rewards fail or degrade reasoning ability.
RLVR activates genuine reasoning patterns through RL training while benchmark improvements may reflect data memorization on contaminated datasets. These operate at different measurement levels and can coexist without contradiction.
Show all 6 sources
Training with only negative samples consistently improves Pass@k across the spectrum, often matching full PPO and GRPO. Negative reinforcement suppresses incorrect trajectories while preserving diversity, whereas positive-only reinforcement degrades higher-k performance by concentrating probability mass.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Spurious Rewards: Rethinking Training Signals in RLVR
- The Invisible Leash: Why RLVR May Not Escape Its Origin
- Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers
- Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
- Psychotherapy AI Companion with Reinforcement Learning Recommendations and Interpretable Policy Dynamics
- The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning