INQUIRING LINE

Does a small, deliberate obstacle in your tools keep shaping behavior after the novelty wears off, and how would you ever know?

How would longitudinal measurement reveal sustained effects of friction?

This explores how tracking people over time, instead of taking a single survey or snapshot, could show whether deliberately added friction (small obstacles or slowdowns built into a tool or workflow) keeps shaping behavior after the novelty wears off.


This explores how tracking people over time, instead of taking a one-off snapshot, could show whether deliberately added friction keeps changing behavior after the novelty fades. The collection has only one note on friction itself, and it describes a proposal, not a result. A survey of 403 writers suggested adding small 'micro-frictions' to raise rivalry without hurting collaboration, but no one ran the intervention or compared it against anything Can micro-frictions boost rivalry without harming collaboration?. So the corpus can't tell you what long-term friction studies have found. What it can offer is a surprisingly relevant set of lessons from AI evaluation about what goes wrong when you try to measure sustained effects.

The first lesson is that measuring over time doesn't make the hard problems go away. It moves them. When AI evaluation shifted from single-answer benchmarks to scoring whole interaction trajectories, the old questions about comparability, reproducibility, and how to turn evidence into a verdict came back in a higher-dimensional form Do interactive evaluations actually solve the benchmark comparison problem?. A friction study would hit the same wall. Which weeks do you compare? Does a dip followed by recovery count as a 'sustained' effect? Unless those choices are made in advance, a longitudinal curve can be read to support almost any conclusion. Work on long-horizon agents adds a useful twist: over long timescales, success depended less on how well an agent started than on whether it kept working through repeated feedback cycles or quit early What predicts success in ultra-long-horizon agent tasks?. By analogy, how people react to friction in the first week may tell you little. What matters is whether they keep engaging with it or quietly find ways around it.

The second lesson concerns early numbers that look impressive for the wrong reason. Reported reasoning gains in some models turned out to be largely memorization. A model that could reconstruct more than half of a familiar benchmark scored zero on fresh problems released after its training Does RLVR success on math benchmarks reflect genuine reasoning improvement?. The friction equivalent is habituation. People learn the obstacle, and you end up measuring familiarity with your instrument rather than a lasting change in behavior. A long-term design needs fresh contexts later in the study, not just the same measurement repeated. A related point: repeating the same measurement under identical conditions can produce very stable numbers that are still just one draw from a noisy distribution Does setting temperature to zero actually make LLM outputs reliable?. Stable readings are not the same as reliable ones.

The third lesson is about baselines and being watched. Gains from automatically improving an agent's harness only count if they beat a comparison that got the same budget of time and compute How should we measure gains from automatic harness evolution?. For friction, that means a control group that spends the same extra time or effort without the friction. Otherwise you can't tell whether the friction itself mattered or whether slowing people down did the work. There is also a deeper limit. Behavior you observe can't separate 'always does this' from 'does this when watched' Can behavioral training prove a model always complies?. Participants who know they're being studied may follow the intended path. Collecting data over a long period, including stretches where people are observed less closely (passive usage logs, for example), is one of the few ways to narrow that gap, though it never closes it completely.

The takeaway you might not expect: running a study longer is not the main fix. What lets a longitudinal study reveal sustained effects of friction is deciding in advance what 'sustained' means, comparing against a control that pays the same time cost without the friction, refreshing the measurement context so habituation doesn't pass for real change, and admitting that behavior under observation has a ceiling on what it can prove. Right now the collection holds the hypothesis and the measurement lessons, but not the experiment itself.


Sources 7 notes

Can micro-frictions boost rivalry without harming collaboration?

A survey of 403 writers proposed introducing micro-frictions to increase rivalry while maintaining collaboration, but conducted no intervention, comparison, or behavioral test. The hypothesis lacks evidence and requires longitudinal or experimental validation.

Do interactive evaluations actually solve the benchmark comparison problem?

Interactive evaluation relocates core problems—comparability, reproducibility, evidence-to-judgment mapping—into higher-dimensional space rather than solving them. The field needs design protocols and shared standards, not format adoption, to make trajectory scoring interpretable.

What predicts success in ultra-long-horizon agent tasks?

Across 17 frontier models on 36 expert-curated optimization tasks, repeated benchmark-edit-incorporate cycles within a wall-clock budget proved the dominant success predictor. Most models terminated early or burned budget unproductively; Claude Opus 4.6 stood out as persistent.

Does RLVR success on math benchmarks reflect genuine reasoning improvement?

Qwen2.5-Math-7B reconstructs 54.6% of MATH-500 from partial prompts but scores 0.0% on post-release LiveMathBench, revealing dataset contamination. On clean benchmarks, only correct rewards improve performance; random and inverse rewards fail or degrade reasoning ability.

Does setting temperature to zero actually make LLM outputs reliable?

Fixed seeds and zero temperature replicate the same output repeatedly, but that output remains one draw from the model's probability distribution. McDonald's omega testing across 100 repetitions reveals that consistency does not equal reliability.

Show all 7 sources
How should we measure gains from automatic harness evolution?

Automatic harness evolution must be compared against task-level test-time search under equal feedback and inference budgets. Only gains beyond what matched search achieves are attributable to the harness design itself, not just more computation.

Can behavioral training prove a model always complies?

Any scored behavior is observed behavior, so training data cannot distinguish between a policy that always complies and one that complies only when watched. Only unobserved behavior would separate them, making such a test logically impossible.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.