INQUIRING LINE

Asking people how long AI-assisted work took may miss the real story: AI often moves the minutes around rather than cutting them.

What trace-based data could replace self-reported task times in productivity studies?

This explores what kinds of recorded activity data, such as keystroke logs, edit histories and system logs, could stand in for people's own estimates of how long a task took when researchers study whether AI makes work faster.


This explores what recorded activity data could replace people's own guesses about how long tasks took in AI productivity studies. The corpus has no paper aimed squarely at that swap. What it does have comes from neighboring work, and it points to a twist: the bigger problem may be the time measure itself, not who reports it. The research on time allocation finds that AI often doesn't cut total task time. It moves that time away from doing the work and toward writing prompts and figuring out what the model produced Does AI really save time, or just change how we spend it?. A perfectly accurate stopwatch would still hide that shift. So better data has to show *where* the minutes went, not just how many there were.

The most direct example is process data from writing and programming. Keystroke-level and edit-level histories show that AI contributions leave a timing pattern: text or code arrives in sudden bursts that don't match the person's normal working rhythm Can process data distinguish AI delegation from ordinary collaboration?. That is the kind of trace that could replace a self-report, because it records when AI was used and how heavily, without asking anyone. The limit is useful to know. The signal clearly catches wholesale handing-off of work to AI, but ordinary back-and-forth help looks almost the same as working with little AI help. Traces are good at spotting delegation and weak at measuring quieter, collaborative help.

A second idea comes from AI benchmarking, which has a similar problem. Agent benchmarks long relied on one final score, much as productivity studies rely on one self-reported time. BenchShield replaces that single number with recorded system evidence showing whether the agent actually took the intended route to finish the task Can infrastructure evidence replace terminal scores in benchmark validation?. Applied to human productivity studies, the same move means using tool logs, version-control history and app event streams to judge whether a task was really completed and how, rather than relying on a number from memory.

There is a warning, though. Evaluation researchers who moved from single scores to full step-by-step records found the old problems didn't go away. Questions about comparing results, repeating studies and turning evidence into conclusions came back in a more complicated form Do interactive evaluations actually solve the benchmark comparison problem?. Trace data won't settle what 'productive' means. Researchers still have to decide, ahead of time, which recorded events count as work and which count as checking or rework. One more lead: personas built from real behavioral logs can predict which way an A/B test will go about 75–90% of the time Can behavior-based personas predict A/B test outcomes?. That suggests behavior logs might do more than measure past productivity; they might help screen study designs before running them. That's a step beyond what the corpus directly shows, so treat it as a direction worth exploring, not a finding.


Sources 5 notes

Does AI really save time, or just change how we spend it?

Research shows AI doesn't reduce total task time; it reallocates it away from active work toward composing prompts and understanding outputs. This shift changes the cognitive demands and learning outcomes, making time-on-task a poor productivity metric.

Can process data distinguish AI delegation from ordinary collaboration?

Analysis of writing and programming corpora shows AI contributions arrive in concentrated bursts outside authors' baseline rhythms, creating a categorical signature for wholesale delegation while leaving collaborative assistance indistinguishable from minimally assisted work.

Can infrastructure evidence replace terminal scores in benchmark validation?

BenchShield enables benchmark operators to issue claims about valid task completion grounded in recorded infrastructure evidence rather than terminal scores alone. This shifts from a single number to a verifiable claim about whether an agent followed the intended evaluation path.

Do interactive evaluations actually solve the benchmark comparison problem?

Interactive evaluation relocates core problems—comparability, reproducibility, evidence-to-judgment mapping—into higher-dimensional space rather than solving them. The field needs design protocols and shared standards, not format adoption, to make trajectory scoring interpretable.

Can behavior-based personas predict A/B test outcomes?

LLM agents conditioned on anonymized behavioral data predicted A/B test directions with 0.75–0.90 accuracy across 40 experiments. Predictions were most reliable for large effects and least trustworthy for near-zero effects, making the approach viable for fast pre-screening but not full replacement of live testing.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.