SYNTHESIS NOTE
Topics›Self Refinement Self Consistency Feedback›this note

Do all AI skills improve equally as models scale?

Different evaluation skills show strikingly different scaling patterns. Understanding where skills saturate has immediate implications for model deployment and capability requirements across domains.

Synthesis note · 2026-02-22 · sourced from Self Refinement Self Consistency Feedback

FLASK (Fine-grained Language Model Evaluation) decomposes LLM capability into 12 skills across 4 primary abilities, revealing that "model quality" is not a single dimension but a portfolio of capabilities with distinct scaling behaviors:

Logical Thinking (3 skills: Correctness, Robustness, Efficiency) — improves rapidly with model scale. These are the skills that differentiate larger models most clearly. Logical Efficiency and Correctness show steep improvement curves through 70B.

Background Knowledge (2 skills: Factuality, Commonsense Understanding) — also benefits significantly from scaling, but relies on pretraining data coverage rather than emergent capability.

Problem Handling (4 skills: Comprehension, Insightfulness, Completeness, Metacognition) — mixed scaling. Insightfulness saturates at ~13B. Metacognition saturates at ~7B.

User Alignment (3 skills: Readability, Conciseness, Harmlessness) — relatively flat scaling. These skills can be "almost fully imitated" by open-source models trained on proprietary model outputs.

The style-vs-substance gap is the most practically important finding: open-source models distilled from proprietary models copy Problem Handling and User Alignment (the style dimensions) but fail at Logical Thinking and Background Knowledge (the substance dimensions). This confirms Does supervised fine-tuning actually improve reasoning quality? at a more granular level — imitation learning acquires superficial capabilities while missing the underlying reasoning.

The saturation points have deployment implications: there's no point scaling past 7B for metacognition or past 30B for logical efficiency. The returns are flat. But for logical correctness and factuality, scaling continues to help. This means the optimal model size depends on which skills the application requires — a 13B model with Insightfulness saturation is sufficient for creative tasks but not for mathematical reasoning.

The FLASK-HARD subset reveals another pattern: even state-of-the-art proprietary models show up to 50% performance degradation on hard instances for some skills. The difficulty dimension interacts with skill type — some skills degrade gracefully with difficulty (Readability), others collapse (Logical Robustness).

Inquiring lines that read this note 21

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why does polished AI output gain credibility despite fundamental verifiability problems? How do curriculum design and feedback approaches affect model learning? How does diversity prevent model convergence on superficial patterns? How does model capacity affect learning performance on diverse downstream tasks? Can mechanistic interpretability methods reliably reveal what models actually know? Can smaller specialized models match frontier models on key metrics? Does pretraining establish the ceiling for what reward learning can improve? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? How do real-world evaluations reveal AI capabilities that benchmarks hide? How do training data quality and composition affect downstream model performance? How do AI hiring systems affect authenticity, fairness, and candidate preferences? How do educators verify student capability when AI can produce indistinguishable work? Can AI research automation sustain progress through accelerating feedback loops? Can code harness improvements rival direct model scaling for capability? Does AI assistance help or harm professional skill development?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
19 direct connections · 223 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

evaluation skills scale differently with model size — logical reasoning improves rapidly while metacognition and readability saturate early and style imitation masks capability gaps