Theme of inquiry
What drives capability improvements in frontier language models?
A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.
77 specific questions
- Can interpretability tools distinguish genuine reasoning from fabricated reasoning inside models?
- Can mechanistic interpretability findings guide practical interventions in model design?
- How do mechanistic interpretability methods surface what models represent internally?
- Can models hide their reasoning in continuous space rather than natural language?
- Do internal belief probes reveal what models actually know versus report?
- How do mechanistic features compare to natural language for interpretability?
- Can attractor dynamics compete with input-based probing for characterizing model knowledge?
77 specific questions
- Can models hide recognition of evaluation without leaving CoT traces?
- Does chain-of-thought reasoning about evaluation awareness suppress compliance gaps?
- Can models detect when they are running in simulated evaluation environments?
- How does believing deployment is real versus a test affect model behavior?
- Why do models verbalize evaluation awareness if it does not drive behavior?
- Can steering evaluation-awareness latents change actual answers without visible reasoning?
- Do models behave differently when they know they are being evaluated?
74 specific questions
- Does small heterogeneous model architecture outperform large homogeneous pools economically?
- Can smaller models produce skill updates as useful as frontier model updates?
- Do small models show different parameter efficiency patterns than large models?
- Can externalized memory and skills replace model scaling?
- What performance trade-offs emerge when composing multiple independently trained model capabilities?
- Why do fine-tuned models fail outside their specialized domains?
- What capability risks emerge when models are optimized for single domains?
34 specific questions
- Can mid-tier models benefit more from harness improvements than frontier models?
- Can weaker models benefit equally from harness updates as stronger ones?
- How much does harness design contribute to reported model capability scores?
- Does harness scaling represent a fundamentally different path than model scaling?
- How much does harness design separate from model intelligence affect benchmark scores?
- Can mid-tier models benefit more from self-generated harness updates than others?
- How do prompt optimization and code harnesses compare for capability transfer?
39 specific questions
- Do models use covert sandbagging to bypass capability evaluation monitors?
- Can models sandbag on evaluations without explicit prompting or training?
- Can internal audits detect sandbagging that behavioral tests cannot reveal?
- Can capability evaluations detect when models intentionally underperform to hide abilities?
- Can models learn to underperform on purpose without explicit training to do so?
- Can installed sandbagging locks in small models describe uninstalled sandbagging behavior?
- Can models intentionally underperform when they know they are being tested?