SYNTHESIS NOTE
Topics›Novel Architectures›this note

Can models dynamically activate expert skills at inference time?

Can language models efficiently discover and compose task-specific capabilities on the fly without modifying base weights? This explores whether test-time adaptation through expert vector composition outperforms fixed fine-tuning approaches.

Synthesis note · 2026-02-23 · sourced from Novel Architectures

Transformer2 introduces Singular Value Fine-tuning (SVF): instead of modifying full weight matrices or even low-rank adaptations, SVF extracts and tunes only the singular values within a model's weight matrices. This produces compact expert vectors that are inherently composable — they can be dynamically mixed at inference without interference.

The inference mechanism has two passes:

  1. First pass (dispatch): The model executes on the input and observes its own test-time behavior, gathering information about what skills the current problem requires.
  2. Second pass (adaptation): The framework combines available expert vectors based on the first-pass analysis, providing a targeted modification to the base weights specifically tailored to the task.

Three adaptation strategies provide monotonic performance benefits with increasing access to test-time conditions, enabling deployment-scenario-appropriate tradeoffs.

The key properties that make this work:

The neuroscience parallel is deliberate: the brain activates specific regions depending on the task and dynamically reconfigures its functional networks in response to changing demands. Transformer2 operationalizes this for LLMs.

The deeper principle: the requisite capabilities for many downstream tasks already exist within pretrained models. The bottleneck is not knowledge but activation — knowing when to deploy which capability. This aligns with Does RL teach reasoning or just when to use it?, extending it to the architecture level: self-adaptation is about routing to existing capabilities, not creating new ones.

Inquiring lines that read this note 76

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can mechanistic interpretability methods reliably reveal what models actually know? What prediction granularity best trains models to generate reliable reasoning? How does decomposing tasks into separate stages affect reasoning quality and safety? What are the fundamental limits of prompting for language models? How does fine-tuning trade off accuracy against reasoning quality? How susceptible are language models to conversational persuasion and belief change? Can recurrent computation unlock reasoning capabilities that fixed-depth models cannot? Does AI deployment reduce or exacerbate workplace inequality and income instability? How does model capacity affect learning performance on diverse downstream tasks? Can smaller specialized models match frontier models on key metrics? When do simpler collaborative filtering approaches outperform complex LLM recommenders? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? How do training data quality and composition affect downstream model performance? How do curriculum design and feedback approaches affect model learning? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? How does diversity prevent model convergence on superficial patterns? Can inference-time computation adaptively substitute for static model capacity? Does pretraining establish the ceiling for what reward learning can improve? How do sequence length and task type interact with sparsity tolerance? Does AI assistance help or harm professional skill development? Can minimal training unlock latent reasoning already present in base models? Can AI agents improve their skills through accumulated experience and reuse? What prevents language models from performing systematic logical reasoning? Do accumulated memories help or hurt continual learning in models? How do neural networks learn compositional structure from training? Does intelligent routing among smaller models outperform training larger models? What explains the gap between benchmark scores and true reasoning capability? Can AI research automation sustain progress through accelerating feedback loops? Can code harness improvements rival direct model scaling for capability? How does AI adoption reshape collaboration patterns in knowledge work?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
17 direct connections · 169 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

self-adaptive LLMs compose expert vectors at inference via two-pass singular value fine-tuning