Can counting repeated words and sentence shapes flag writing that feels templated, or is that feeling too subjective to measure?
Can lexical repetition and syntactic structure predict whether text feels templated?
This explores whether countable surface features of writing, like repeated words and recurring sentence shapes, can tell us when text reads as formulaic or 'AI-templated', and what the corpus says about where that templated feel comes from.
This explores whether countable surface features of writing, like repeated words and recurring sentence shapes, can tell us when text reads as formulaic. The corpus gives a partial yes. The clearest support comes from an effort to define 'AI slop.' Nineteen experts' definitions were coded and sorted into three dimensions, and one of them, style quality, is built around exactly the features in the question: repetition and templatedness What dimensions make text feel like AI slop?. The researchers map each dimension to proxies that can be measured automatically or annotated by people, so repetition is treated as something you can count. One caution: the corpus doesn't contain a study showing that lexical or syntactic counts alone reliably predict a reader's sense that text is templated. What's here is a framework that treats them as reasonable signals, not a validated detector.
The less obvious point is that the same note puts style next to two other dimensions: information utility (is the text dense and relevant?) and information quality (is it accurate and fair?). That suggests the templated feeling may not live only in the surface. Text can repeat itself in its structure and still be useful, or vary its wording and still say nothing. If you want to predict 'feels templated,' surface counts probably need to be combined with a measure of how much new information each sentence adds.
Where does the repetition come from in the first place? The corpus points to at least three sources, each operating at a different level. In training, when a reward signal barely varies between good and bad answers to the same prompt, the model drifts toward generic outputs that would fit almost any input. These are literal templates, and filtering for prompts that produce more varied rewards undoes much of the drift Why do language models collapse into generic templates?. In the architecture, transformer attention gives extra weight to content that's already repeated or prominent in the context, which creates a feedback loop where repetition breeds more repetition Does transformer attention architecture inherently favor repeated content?. Across society, many people relying on the same few models, each reflecting the same statistical regularities, pulls everyone's writing toward shared phrasings and framings, and co-writers adopt them without noticing Do large language models narrow human expression and thought?.
Syntax adds a twist. LLMs handle simple sentences well, but their accuracy drops steadily as sentences nest clauses inside clauses, which suggests they rely on surface patterns more than on rules of grammar Does LLM grammatical performance decline with structural complexity?. That finding is about how models understand sentences, not how they write them, so treat this as a hypothesis rather than a result: a narrow, predictable range of sentence shapes might be one of the syntactic fingerprints a templatedness measure could pick up. A more philosophical note proposes a reason the rhythm can feel off. AI text is produced token by token, with no pause to reconsider or revise, so it is ordered in sequence but lacks the reflective time that shapes how human writing unfolds Does AI text generation unfold through temporal reflection?.
The takeaway: repetition and sentence structure are credible, measurable signals of templated text, and experts already list them as such. But the corpus suggests they are symptoms of deeper causes, namely flat training rewards, attention that amplifies repetition, and population-wide convergence. A detector built only on surface counts would catch the symptom and could miss text that is varied on the surface but formulaic underneath.
Sources 6 notes
Coded definitions from 19 experts yield three axes: information utility (density and relevance), information quality (factuality and bias), and style quality (repetition and templatedness). Each axis maps to automatic or human-annotated proxies for assessment.
When within-prompt reward variance is low, task gradients weaken and regularization dominates, pushing policies toward generic outputs. SNR-Aware Filtering—selecting high-variance prompts before updates—recovers performance across tasks and scales.
Transformer soft attention systematically over-weights repeated and context-prominent tokens regardless of relevance, creating a positive feedback loop that amplifies opinions and framing before RLHF acts. System 2 Attention—regenerating context to remove irrelevant material—can interrupt this mechanism.
LLMs mirror skewed slices of human experience shaped by training data regularities, and widespread reliance on identical models amplifies convergence. Co-writing studies show users unconsciously adopt model stances and framings.
LLMs show systematic performance decline as syntactic depth and embedding increase. Simple sentences are handled well while complex structures with recursion and embedding fail consistently, suggesting LLMs learned surface heuristics rather than structural grammar rules.
Show all 6 sources
Token ordering in LLMs follows probabilistic selection without intervening reflection or revision. Human discourse gains meaning from temporal structure—time spent thinking changes what comes next—but AI text production lacks this duration-in-reflection despite appearing sequentially composed.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- From Human to Machine Psychology: A Conceptual Framework for Understanding Well-Being in Large Language Models
- Measuring AI "Slop" in Text
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- Six misconceptions about large language models: A minimal model and diagnostic taxonomy
- Argument Collapse: LLMs Flatten Long-Form Public Debate
- The Widespread Adoption of Large Language Model-Assisted Writing Across Society
- RAGEN-2: Reasoning Collapse in Agentic RL
- Semantic Structure in Large Language Model Embeddings