INQUIRING LINE

Can we actually measure what makes writing good enough to teach an AI to write well?

Can literary quality be measured precisely enough to train models?

This explores whether something as slippery as 'good writing' can be pinned down well enough to serve as a training signal for AI, and where those attempts break.


This explores whether literary quality can be turned into a reliable training signal for AI, and where those attempts break. The corpus suggests that precision isn't really the obstacle. Parts of literary writing are very easy to measure. The hard part is that the measurements capture the wrong things, and the judges doing the measuring are easy to sway.

Start with what measures well. Style is highly detectable: even GPT-2 can identify an author from stylistic patterns with about 95% accuracy Can language models truly understand literary style?. Narrative structure is measurable too. StoryScope separated AI fiction from human fiction with 93% accuracy using only choices like how much agency characters have and whether events are told in order, with no surface style involved Can AI stories be detected without analyzing writing style?. Some researchers have even turned reading efficiency into a number, counting unique pieces of knowledge per token, and found that LLM prose scores lower because models pad Can we measure reading efficiency as a quality metric?. These are precise signals. The catch is that detecting a pattern is not the same as knowing why it matters. A model that recognizes an author's style still can't explain what those choices accomplish, which amounts to cataloguing rather than criticism.

The weaker link is the judge. When AI models rate literary passages, simply labeling a passage as human-written raised their scores by 34 percentage points. That bias is 2.5 times stronger than in human readers, who show it too Do authorship labels bias how we judge literary quality?. LLM judges can also be fooled by fake citations and attractive formatting Can LLM judges be fooled by fake credentials and formatting?. If you train a model on these judgments, you risk teaching it to perform the signals of quality rather than produce quality.

The more promising results narrow the target instead of sharpening the ruler. Sun argues that LLMs stay weak as creative writers but become capable editors once they're given one writer's explicit taste rubric. Part of the reason is that general writing quality is so hard to measure, while a specific person's preferences can be written down Can LLMs become good editors by learning a writer's taste?. Argument quality points the same way. Fine-tuning on labeled examples teaches surface patterns that don't carry over to new kinds of argument, while spelling out an explicit theory of what makes an argument good does generalize Can models learn argument quality from labeled examples alone?.

One more idea from outside literature: 'scientific taste' turned out to be learnable when it was grounded in a community's long-run judgment. Models were trained on 700,000 paper pairs ranked by how often each paper was later cited Can models learn what makes research worth doing?. Literature has a rough equivalent in which books get reread, taught and quoted over decades, though the corpus doesn't yet include anyone trying it. Overall, quality seems trainable when someone commits to whose taste counts, whether that's one writer's rubric, an explicit theory, or a community's track record. Treating it as a single universal score is where it breaks.


Sources 8 notes

Can language models truly understand literary style?

GPT-2 achieves 95% accuracy identifying authorship through style patterns alone, but lacks the evaluative framework to explain why those stylistic choices carry meaning. Detection without interpretation remains cataloguing, not criticism.

Can AI stories be detected without analyzing writing style?

StoryScope achieved 93.2% accuracy separating AI from human fiction using only discourse-level features like character agency and chronological structure, retaining 97% of performance while eliminating stylistic cues. These structural choices resist humanization because they require rewrites, not surface edits.

Can we measure reading efficiency as a quality metric?

Knowledge Density (KD) operationalizes reading efficiency by dividing unique atomic knowledge units by text length. LLM-generated text scores lower on KD than human writing because retrieval redundancy and the model's tendency to elaborate inflate token count while holding knowledge content constant.

Do authorship labels bias how we judge literary quality?

Human judges rated identical passages 13.7 percentage points higher when labeled human-authored; AI models showed a 2.5-fold stronger bias at 34.3 points. The effect persists across AI architectures, suggesting evaluators respond to provenance cues rather than text quality alone.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Show all 8 sources
Can LLMs become good editors by learning a writer's taste?

Sun demonstrates that LLMs remain poor creative writers but can match human editors when trained on personalized rubrics. The gap traces to three factors: hard-to-measure writing quality, misaligned business incentives, and lack of lived grounding—none of which editing requires.

Can models learn argument quality from labeled examples alone?

Fine-tuning on labeled examples fails to transfer quality criteria to new argument types. Models learn surface patterns rather than principled criteria. Explicit instruction using frameworks like RATIO or QOAM significantly improves performance and generalization.

Can models learn what makes research worth doing?

Reinforcement learning trained on 700K citation-matched paper pairs successfully teaches models to predict research impact better than GPT-5.2 and generate higher-impact research ideas. Scientific taste emerges as a community-aligned capability distinct from execution skills.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.