INQUIRING LINE

When humans and AI build code together, is counting which lines each wrote the right way to say who did the work?

How should code authorship be measured in human-AI collaborative development?

This explores how to tell who wrote what (the human or the AI) when code is built together, and what 'authorship' should even mean once the work is mixed. The corpus has no study that measures code authorship directly, but nearby work on writing, coding conversations, and audit trails points to a useful reframing.


This explores how to tell who wrote what (the human or the AI) when code is built together, and what 'authorship' should even mean once the work is mixed. To be upfront: the collection has no paper that measures code authorship directly. It does have several neighboring findings, and together they suggest the usual approach of counting which lines came from whom may be the wrong unit.

The strongest clue comes from writing rather than code. When pairs of writers shared an editor, they consistently wanted to see more of each other's AI prompting: when the AI was used, how, and on which passages Do writers want to see each other's AI prompts in shared editors?. They didn't mainly want a percentage of AI-written text. They wanted the story of how the text came to be, so they could follow a collaborator's thinking and check the AI's contributions. Applied to code, that suggests authorship might be better recorded as a process (what was asked, what was accepted, what was rewritten) than as a ratio of characters. The same study found a tension that matters here: some writers felt full visibility was intrusive and made them self-conscious. Any authorship measure built on prompt logs will run into that privacy cost.

A second clue suggests that the conversation itself carries signal. Patterns pulled from developers' conversations with coding agents predicted coding outcomes better than the developers' prior skill did Can conversation patterns predict coding outcomes better than prior skill?. The catch is that these patterns weren't stable or transferable enough to count as real, learnable skills. In practice, how someone works with the AI tells you something about the result, but it isn't yet a reliable measure of that person's contribution. Tooling is starting to keep the record anyway: orchestration layers wrapped around coding agents keep persistent state and a traceable, recoverable trail of what happened Can orchestration layers make coding agents more auditable?. Systems like Magentic-UI build in checkpoints such as co-planning, action guards, and verification steps, each a moment where a human decision gets logged When should human-agent systems ask for human help?. These trails could serve as raw material for authorship records, though so far they're designed for auditing rather than for assigning credit.

The result that surprised me most comes from fiction detection. AI-written stories can be identified from structural choices, such as how characters act and how time is ordered, with 93% accuracy and no help from surface style Can AI stories be detected without analyzing writing style?. Those signals hold up because changing them requires rewriting, not light editing. If something similar holds for code, the real marks of authorship may sit in architecture and decomposition (who decided how the problem was split up) rather than in who typed each function. That would put meaningful authorship at the design level, which line-level tools like git blame can't see.

One caveat: AI context is constantly shifting (prompts, history, retrieved files, hidden state) How does AI context differ from conventional software context?, so the same request can produce different code on different days. 'The AI wrote this line' is therefore a weak, non-reproducible claim. Work on human–AI research teams frames the human contribution as direction and verification rather than output volume Can human-AI research teams improve faster than autonomous AI systems?. That may be the most defensible definition of authorship this collection offers: who set the direction, and who checked the result.


Sources 7 notes

Do writers want to see each other's AI prompts in shared editors?

Sixteen paired writers showed strong preference for higher levels of prompt visibility in shared editors, valuing awareness of when, how, and where AI was used. Benefits included understanding collaborators' thinking and verifying AI-generated text, though some found full sharing intrusive and self-conscious.

Can conversation patterns predict coding outcomes better than prior skill?

Machine learning identified interpretable traits from coding-agent conversations that explained outcomes beyond prior achievement. However, these traits lacked the stability and transferability required to qualify as learnable human-AI collaboration skills.

Can orchestration layers make coding agents more auditable?

Dr. Claw wraps existing coding agents in persistent state objects and skill libraries, reporting higher research completeness and a traceable, recoverable process trail while keeping the underlying executor unchanged.

When should human-agent systems ask for human help?

Magentic-UI identifies co-planning, co-tasking, action guards, verification, memory, and multitasking as mechanisms that work around the lack of ground truth for optimal deferral timing. Rather than solving the timing problem directly, these mechanisms distribute decision-making across multiple touchpoints.

Can AI stories be detected without analyzing writing style?

StoryScope achieved 93.2% accuracy separating AI from human fiction using only discourse-level features like character agency and chronological structure, retaining 97% of performance while eliminating stylistic cues. These structural choices resist humanization because they require rewrites, not surface edits.

Show all 7 sources
How does AI context differ from conventional software context?

AI interactions operate on a substrate of constantly shifting context—prompt, history, retrieved data, hidden state—that users cannot internalize like traditional UIs. This structural mutability demands a new design discipline centered on context engineering rather than interface design.

Can human-AI research teams improve faster than autonomous AI systems?

Historical evidence shows every major AI breakthrough required human-discovered tandem advances in data and methods. Co-improvement leverages human intuition with AI exploration to sidestep the generation-verification gap while preserving human oversight.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.