Can models be smart without organized internal structure?
Explores whether linear feature decodability proves genuine compositional reasoning or merely indicates that the right features are present but poorly organized. Critical for understanding what performance metrics actually certify.
Two findings from mechanistic interpretability appear contradictory but operate at different levels of representational analysis:
Fractured Entangled Representations (FER): Since Can identical outputs hide broken internal representations?, SGD-trained models fail catastrophically under perturbation or distribution shift in ways that well-organized representations would not. The pathology is invisible to standard evaluation.
Compositional generalization at scale: Scaling data and model size produces representations where compositional features are linearly decodable — separable task constituents can be independently identified and manipulated. This has been taken as evidence for genuine compositional understanding.
The resolution: Linear decodability tests for the presence of features, not their organization. A fractured representation could contain every linearly decodable feature while being fractured in how those features relate to each other. The compositional parts are present but their composition is broken.
This connects directly to the "imposter intelligence" post angle: Can LLMs understand concepts they cannot apply?, Does supervised fine-tuning actually improve reasoning quality?, and Do foundation models learn world models or task-specific shortcuts?. All describe the same meta-pattern: surface metrics certify capability that internal structure analysis would disqualify.
The practical implication for model evaluation: passing compositional generalization tests does not guarantee robust compositional reasoning. Evaluation under distribution shift, perturbation, and novel recombination is required to distinguish genuine compositionality from fractured representations that happen to contain the right features.
Inquiring lines that read this note 200
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can smaller specialized models match frontier models on key metrics?- Why do only two of fourteen models improve when problem constraints are removed?
- What production constraints should determine paradigm selection?
- What performance trade-offs emerge when composing multiple independently trained model capabilities?
- Why do metric choices constrain which model capabilities get developed?
- Does model collapse occur across different architectures or only in specific conditions?
- Does model capability still matter once coordination infrastructure is optimized?
- What distinctive properties make open foundation models different from closed ones?
- What benefits do open foundation models create that closed systems cannot?
- Can end-to-end models maintain debuggability without modular components?
- How much does workflow architecture matter versus raw model capability?
- What structural constraints matter more than model depth for CF?
- How does fluent output mask the mythic function of a system?
- What distinguishes minimal-pair asymmetry from standard accuracy evaluation?
- How do unstated constraints become invisible to training data distributions?
- Can Kolmogorov complexity alone capture what makes intelligence general?
- When should model isolation be preferred over weight-averaging approaches?
- Why do power-law distributions make standard ML infrastructure assumptions fail?
- How do surface statistical regularities enable correct outputs while degrading robustness?
- Why do singular value experts compose better than low-rank adapter subspaces?
- What makes structured stochasticity more effective than unstructured randomness in reasoning?
- Do generic kernel-decay assumptions alone explain coarse-to-fine spectral ordering?
- How do spectral-norm constraints prevent divergence in world model rollouts?
- Can ensemble predictions be distilled back into a single deployable model?
- How does requential coding measure true simplicity without parameter count inflation?
- Why do parameter-based compressors fail to measure true model simplicity?
- Can parameter compression mechanically force value systems toward idealized centers?
- What makes the frame problem distinct from feature-level shortcuts?
- How do autonomous pipelines identify and fix silent bugs in data pipelines?
- How do unstated feasibility constraints affect model decision-making?
- Can a single SAE feature control reasoning behavior across model families?
- Why do models fail on logically equivalent tasks with different data distributions?
- What limits external scaling when a model lacks reasoning foundation?
- How does learnability at the observer's current state prevent novelty from breaking model reasoning?
- What design changes could make constraint inference more reliable without explicit cuing?
- What is the mechanistic signature when models chain facts never presented together?
- What architectural properties of deterministic models block multi-solution reasoning?
- Why do unresolved items cluster in structured patterns rather than randomly?
- Can mechanistic interpretability reveal how ideologies decompose into simpler features?
- Can a world model have rich representations without adequate data coverage?
- How do functional features differ from representational abstract features?
- What test distinguishes genuine compositionality from fractured feature presence?
- What happens when you remove core political features from a deep model?
- Why do models with less steerability have more abstract ideological features?
- How does LatentQA differ from predefined concept steering like representation engineering?
- What skills can large models identify and organize about their own abilities?
- What distinguishes conceptual understanding from statistical pattern matching in models?
- Can mechanistic interpretability explain explanation-execution disconnection?
- Why must world models be nested rather than flat and uniform?
- Can geometric structure in representations exist without supporting functional mechanisms?
- How does mechanistic interpretability complement learning mechanics in explaining deep learning?
- What distinguishes a representational feature from a causally inert correlation?
- Can interventions on model components prove mechanism without explaining encoding?
- Can representation analysis methods detect complex features models compute with?
- How do mechanistic features compare to natural language for interpretability?
- What makes representation engineering better than mechanistic interpretability for detecting hidden objectives?
- Why do feature visualizations alone fail to establish mechanistic claims?
- What makes some internal circuits more interpretable than others?
- How do mechanistic interpretability methods surface what models represent internally?
- What distinguishes new associations from existing ones at the computational level?
- Can mechanistic interpretability findings guide practical interventions in model design?
- How does representation engineering compare to mechanistic interpretability for auditing?
- How does external validation replace the need for model interpretability?
- How should disciplines evaluate theories built on opaque machine learning predictions?
- When does a model's lack of interpretability become a genuine epistemic problem?
- Can interpretability tools distinguish genuine reasoning from fabricated reasoning inside models?
- How should benchmarks test whether models fit algorithms or patterns?
- How does optimizing model performance decouple from optimizing user interpretability?
- How much do metric choices inflate claims about model capabilities?
- How do weight perturbations reveal what performance benchmarks cannot measure?
- Why do single function-calling benchmarks mask model weakness in specific areas?
- Can identical model performance mask fundamentally broken internal representations?
- Why do text-only benchmarks underestimate deployed model capability?
- How can benchmark accuracy scores mask the absence of interpretable reasoning structure?
- Can a single Elo ranking represent multidimensional model capability?
- How should single-axis benchmarks account for separable capability dimensions?
- How do embedding dimension limits constrain what concept models can represent?
- Does architectural discovery follow an empirical scaling law like neural networks?
- What makes multimodal conditioning effective when features are decomposed to the right granularity?
- What makes linear decodability a reliable signal of compositionality?
- Can steering vectors prove that representations are genuinely organized?
- Can fractured entangled representations hide undetected by standard analysis methods?
- Does the linear representation hypothesis reflect networks or reflect our analysis tools?
- Can representation engineering cleanly isolate single features in entangled semantic space?
- What are fractured entangled representations in neural networks?
- How do sparse circuits compare to the modular subnetworks that emerge naturally?
- Why does weight sparsity reduce superposition and force disentangled representations?
- Can sparse approximations reveal interpretable structure hidden in existing dense models?
- How do overparameterization and data size shift what attractors represent?
- Can we predict which tasks will decompose into modular subnetworks?
- Why does gradient descent discover compositional structure without explicit pressure?
- How can neural networks be interpretable by design rather than post-hoc?
- What physical structure does a Gaussian-regularized latent space actually encode?
- What makes regularization an implicit factor in embedding geometry?
- Do feature extraction methods systematically miss computationally important complex features?
- What makes a feature abstract versus concrete in neural network activations?
- How does scaling and training data enable compositional behavior without symbolic mechanisms?
- What prevents representation collapse in latent-prediction world models like JEPA?
- Can generative reconstruction preserve latent manifold structure better than geometric compression?
- How does representation-level reranking address residual gaps after decomposition?
- What makes a new representational primitive valuable enough to justify its representational cost?
- Can likelihood choice matter more than architectural depth for CF?
- What sparse high-rank patterns does the deep tower fail to capture?
- Why do cross-product features fail to generalize across unseen feature combinations?
- Why do structural signals across edges resist noise better than single-edge counts?
- Can granular function calling tasks learn composition from graph-sampled data?
- What compression explains why syntax fits in low-dimensional subspaces?
- Can steering vectors be combined with other compression techniques?
- How should product specifications measure alignment without naming the dimension?
- What happens when alignment targets measure only the preferred dimension of entangled properties?
- Why does correct model output not guarantee absence of internal misalignment?
- How does nesting optimization levels improve on traditional network depth?
- How does adjacent layer sharing differ from non-adjacent weight reuse?
- Why do standard transformers fail to encode recursive structure in their hidden states?
- Does Gemma's transformer explicitly exploit the inherited hierarchical geometry?
- How do pre-norm layers enable reliable fixed-point halting signals?
- What role do multi-dimensional quality frameworks play in assessing arguments versus single-metric approaches?
- How does vehicle causality differ from content causality in physical systems?
- What spectral signatures distinguish hierarchy-driven geometry from corpus-driven geometry?
- Is interpretive multiplicity a bug in language or a feature?
- Can structured decomposition fix evaluation gaps in other research tasks?
- What makes AI-discovered architectures reveal design principles invisible to humans?
- Can bilevel autoresearch succeed when the inner and outer loops use different models?
- How much does domain shift limit the mechanisms a bilevel system can autonomously discover?
- What interpretability challenges arise when algorithms are discovered rather than designed?
- What distinguishes a computational success like AlphaFold from a conceptual breakthrough?
- Do larger models develop more abstract features than smaller ones?
- What task structures benefit most from geometric parameter merging?
- Does scaling data automatically produce compositional reasoning or just better feature encoding?
- Why does the right structural prior matter more than raw model capacity?
- How can expensive models efficiently support cheap models in production?
- Does parameter composition work when adapter alignment is imperfect?
- Can a complexity-predictor be meaningful if models are redundant?
- Can dense models match specialized architectures by mixing data better?
- Can mathematical capability distributions be read as unified rather than separate?
- Why do text-to-image models fail at composing multiple concepts together?
- Can structural perturbations harm model accuracy more than semantic ones?
- What distinct structural signatures do model repetition and topic volatility create?
- Can we balance interpretability with the efficiency gains of compressed inter-model communication?
- When does internal model knowledge fail to appear in outputs?
- How does discretization make item representations more distinguishable?
- Do multi-vector or cross-encoder models escape these dimensional constraints?
- Why is a combinatorial framework better than family resemblance classification?
- Can spectral eigenvector ordering serve as a model-agnostic interpretability probe?
- What architectural alternatives can capture compositional structure beyond pooled cosine?
- Why do energy-based models generalize better on out-of-distribution data than standard transformers?
- Can seedless generation maintain explainability while scaling control?
- How should researchers evaluate whether correct model outputs reflect real structural learning?
- Why does the gap between theoretical expressiveness and learned capability matter?
- What features does a sample reinforce when it moves bands?
- What makes frozen model reasoning different from weight-based parameter updates?
- What makes some model capabilities reliable while others remain brittle?
- Can surface-level correctness hide failures in structural learning by LLMs?
- What makes well-formatted outputs misleading as evidence of model capability?
- How do local soundness signals work across different problem domains?
- Why do fluent predictions fail to capture reliable internal models?
- Why do accuracy scores alone miss important dimensions of model capability?
- Can you steer reasoning by directly manipulating SAE features?
- Why does structured stochasticity help reasoning more than naive randomness?
- Can we steer model reasoning by manipulating single features?
- Why do semi-formal templates improve verification accuracy over unstructured reasoning?
- How does MaxSim reranking differ from structural verification at the token level?
- What makes attractor-based probing better for third-party model auditing than alternatives?
- Can entropy signatures alone detect whether context was model-generated or externally prefilled?
- Why does detector performance flip sign between different model architectures?
- Which of the 167 numerical features carry the strongest detection signal?
- Does sparsity enforce compositional structure or merely amplify existing modularity?
- Can sparsity patterns reliably indicate how well a model knows its input?
- How do sparse weight patterns affect model interpretability?
- How do coverage and identifiability set separate performance ceilings?
- How does separating environment components make evaluation results more reproducible and analyzable?
- What makes the minimal-criterion check effective within a fixed representation?
Related concepts in this collection 1
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can we track and steer personality shifts during model finetuning?
This research explores whether personality traits in language models occupy specific linear directions in activation space, and whether we can detect and control unwanted personality changes during training using these geometric directions.
persona vectors demonstrate a case where linear decodability corresponds to genuine functional organization (steering works), providing a positive counterexample to FER's warning that decodability alone is insufficient
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
- Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
- Titans: Learning to Memorize at Test Time
- Break It Down: Evidence for Structural Compositionality in Neural Networks
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
- Computational structuralism: Toward a formal theory of meaning in the age of digital intelligence
- Farther the Shift, Sparser the Representation: Analyzing OOD Mechanisms in LLMs
- Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
Original note title
identical performance metrics can mask fundamentally different internal representations — feature linear decodability does not guarantee representational organization