Can the rhythm of your writing flag when AI did the work, but not confirm you worked alongside it?
Does process data reliably distinguish delegation from genuine collaboration?
This explores whether traces of how work gets made (keystroke timing, edit bursts, session rhythms) can tell apart someone who handed a task to AI from someone who actually worked through it alongside AI.
This explores whether the *process* behind a piece of writing or code, rather than the finished product, can show whether a person handed the work to AI or actually worked with it. The short answer is that it works in one direction only. Studies of writing and programming sessions find that wholesale delegation leaves a clear temporal fingerprint: AI-generated content arrives in concentrated bursts that don't match the author's normal working rhythm. Lighter, back-and-forth assistance leaves no such trace and looks almost the same as work done with very little AI help at all Can process data distinguish AI delegation from ordinary collaboration?. So process data can flag delegation, but it can't confirm genuine collaboration. Its silence doesn't mean the person engaged with the work. It only means nothing large got pasted in.
This matters more than it might seem, because delegated work can go wrong without anyone noticing. In long relay workflows, even frontier models quietly corrupt about a quarter of a document's content, and spot-checks of the final output miss it Do frontier LLMs silently corrupt documents in long workflows?. If the finished artifact can't tell you whether it was carefully made, the process record is one of the few places left to look. Benchmark design is reaching the same conclusion: BenchShield argues that a final score is too thin to trust and that you need recorded evidence of *how* an agent completed a task before you can claim it was done properly Can infrastructure evidence replace terminal scores in benchmark validation?.
If timing traces can't see collaboration, what can? The more promising work in the corpus looks at the *content* of the interaction rather than its rhythm. In therapy research, the quality of the relationship between therapist and patient can be scored turn by turn from transcripts, and it shows where the two converge or stay misaligned Can we measure therapist-patient alliance from dialogue turns in real time?. In education, researchers don't wait for evidence of collaboration to show up. They have an AI teammate steer a natural conversation so that the student's collaborative skills become visible, then score it with agreement close to human raters Can AI teammates assess collaboration without losing naturalness?. The shared lesson is that collaboration is easier to *elicit and observe in dialogue* than to *infer afterward from activity logs*.
One more angle: whether delegation is even a problem depends on the task. A framework for intelligent delegation lists eleven properties of a task that decide how it should be handed off, and names verifiability (whether you can check the outcome at all) as the foundational one What makes delegation work beyond just splitting tasks?. If a task's output is easy to verify, detecting delegation matters less. If it isn't, as with long documents or subjective writing, process evidence becomes more important, and that is exactly where it is weakest at certifying real engagement.
The takeaway: process data works as an alarm for wholesale handoff, not as proof of partnership. To know whether someone genuinely collaborated with AI, the corpus points toward designing interactions that make the collaboration visible, rather than mining timestamps after the fact. The collection doesn't yet have work that directly validates content-based collaboration detection for writing or coding, so that remains an open gap.
Sources 6 notes
Analysis of writing and programming corpora shows AI contributions arrive in concentrated bursts outside authors' baseline rhythms, creating a categorical signature for wholesale delegation while leaving collaborative assistance indistinguishable from minimally assisted work.
Even the strongest models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) degrade documents by ~25% over long relay workflows across 52 domains. Degradation decelerates but never plateaus, and errors compound silently, remaining undetected in spot-checked outputs.
BenchShield enables benchmark operators to issue claims about valid task completion grounded in recorded infrastructure evidence rather than terminal scores alone. This shifts from a single number to a verifiable claim about whether an agent followed the intended evaluation path.
COMPASS maps dialogue turns onto WAI embeddings to produce 36-dimensional alliance scores per turn. Anxiety and depression show convergence in alliance metrics over time, while suicidality shows persistent misalignment between patient and therapist.
An LLM-based approach allows students to collaborate with AI teammates in human-like conversation while the system steers toward observable evidence of skill proficiency. The same LLM can also score the interaction against a rubric with inter-rater agreement matching human performance.
Show all 6 sources
Delegation requires matching tasks to agents across 11 dimensions: complexity, criticality, uncertainty, duration, cost, resource requirements, constraints, verifiability, reversibility, contextuality, and subjectivity. Verifiability is foundational—it determines whether outcomes can be evaluated at all.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- LLMs Corrupt Your Documents When You Delegate
- Does AI Assistance Leave a Temporal Fingerprint? Detecting Overreliance in AI-Assisted Writing and Programming
- ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
- Towards Scalable Measurement of Durable Skills
- COMPASS: Computational Mapping of Patient-Therapist Alliance Strategies with Language Modeling
- Working Alliance Transformer for Psychotherapy Dialogue Classification
- Psychotherapy AI Companion with Reinforcement Learning Recommendations and Interpretable Policy Dynamics
- A natural language processing approach reveals first-person pronoun usage and non-fluency as markers of therapeutic alliance in psychotherapy