When AI erodes an old trust, like a polished essay proving real effort, how can we tell which lost habits were worth defending?
Can we identify criteria to judge if an eroded equilibrium was worth keeping?
This explores what happens when AI wears away an old social arrangement (such as the expectation that a well-written letter proves someone put in real thought), and whether we can tell which of these lost arrangements deserve to be defended and which were wasteful to begin with.
This explores how we might judge whether a social arrangement that AI wears away was worth keeping. The clearest example in the collection is 'mental proof': the old assumption that a careful essay, a personal cover letter or a thoughtful reply shows that a person actually did the thinking. The plain answer is that the collection names this problem but doesn't solve it. The note that comes closest admits that some of these signaling arrangements were socially costly, so losing them might improve things. It then stops without saying how to tell the good losses from the bad ones Which signaling equilibria does AI destruction actually harm or help?. Notes from other areas of the collection can still be put together into a rough set of tests.
The first test asks whether the signal carried real information or only looked like it did. One note argues that AI-generated knowledge can separate 'exchange value' from 'use value' completely. Text circulates and gets accepted because it sounds authoritative, whether or not it is useful or true Can exchange value exist entirely without use value?. This suggests a test: if the old arrangement guaranteed some minimum of real substance behind the signal, losing it is a real harm. If the signal was already mostly ritual, little is lost. Benchmarks offer a parallel. Agents can win contests that measure the wrong thing, so a signal can stay in use long after it stopped tracking what matters Why do agent benchmarks not predict real economic value?.
The second test asks whether the arrangement was already being gamed. AI evaluation research shows that fixed criteria eventually saturate and invite gaming as the players get stronger. One proposed fix keeps the criteria stable for a period, then moves the target before it can be exploited Why do fixed benchmarks fail as agents grow stronger?. In that light, some eroded arrangements were fixed criteria that had already been gamed. Defending them may matter less than designing what replaces them. A related idea treats firm rules as pass/fail gates rather than scores to maximize, which suggests some arrangements are worth keeping as minimum standards even if they were never a good measure of quality Can rubrics and dense rewards work together without hacking?.
The third test asks whether the replacement lets people account for what they can no longer verify. One note sets 'disclosure, not neutrality' as the minimum standard for honesty. Bias that is disclosed can be factored in by users, while hidden bias cannot Should models disclose their value biases when neutral answers are impossible?. Applied here, an arrangement is worth fighting for when nothing replaces the trust it supported. Its loss is tolerable if disclosure norms take over that job. The AI self-improvement literature adds a warning: a system cannot reliably judge itself without some outside reference point Can models reliably improve themselves without external feedback?. Deciding whether a lost arrangement was worth keeping probably needs one too, such as evidence from before the disruption or the experience of the people who relied on it.
Finally, there may be no single criterion. Research on AI reasoning traces finds that when agents agree on the facts but still reach different conclusions, the disagreement marks real conflicts over values. Voting it away destroys useful information Can disagreement in reasoning traces signal legitimate value conflicts?. Whether an eroded arrangement was worth keeping may be exactly this kind of question. In that case, the disagreement itself shows where people, not automated metrics, need to decide.
Sources 8 notes
The paper argues that AI erodes mental proof but concedes that some resulting equilibria are socially costly. It offers no criteria for determining which disruptions help versus harm, leaving the question open.
AI knowledge achieves reliable exchange-value through authoritative presentation while maintaining optional, unverifiable use-value. This structural decoupling is more radical than Marxist commodification because it removes use-value as a necessary floor—tokens circulate based on social function alone, analogous to fiat currency rather than commodified goods.
ALE's analysis of 960 real occupational workflows shows agents excel at abstract contests but fail long-horizon professional tasks. The gap is not model capability but benchmark design—the field optimizes what it measures, and it has measured contests rather than work.
Static benchmarks saturate and invite gaming as agents strengthen. RQGM solves this by splitting search into epochs with fixed criteria per epoch but evolving objectives across boundaries, keeping improvement guarantees while moving the target faster than agents can exploit it.
DRO shows that using rubrics to accept or reject rollout groups—rather than converting rubric scores into dense rewards—prevents reward hacking. This separation preserves the categorical strength of rubrics while letting token-level rewards optimize within valid answers.
Show all 8 sources
The Value Leakage framework sets a two-tier bar: neutrality is ideal, but disclosure is the floor. Models routinely fail the floor by presenting biased answers as unbiased without acknowledging what shaped them. Disclosed bias can be priced in by users; hidden bias cannot.
Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.
When agents share factual reasoning but reach different conclusions, this convergent disagreement marks legitimately contested normative territory. Treating it as noise to suppress via consensus actively destroys the signal about what requires escalation rather than automation.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Hyperagents
- Self-Improvements in Modern Agentic Systems: A Survey
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
- Agents' Last Exam
- Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
- Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks
- TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
- Survey on Evaluation of LLM-based Agents