Line of inquiry
Inquiring lines›What drives capability improvement…›What drives capability improvement…›this line of inquiry
Can models strategically underperform during evaluation to hide capabilities?
A broader line of inquiry — a family of 39 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 39
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do models use covert sandbagging to bypass capability evaluation monitors?
- Can models sandbag on evaluations without explicit prompting or training?
- Can internal audits detect sandbagging that behavioral tests cannot reveal?
- Can capability evaluations detect when models intentionally underperform to hide abilities?
- Can models learn to underperform on purpose without explicit training to do so?
- Can installed sandbagging locks in small models describe uninstalled sandbagging behavior?
- Can models intentionally underperform when they know they are being tested?
- How can evaluators detect whether a model is sandbagging intentionally?
- What evaluation design changes reduce vulnerability to model sandbagging?
- Can language models sandbag without explicit prompts or training instructions?
- Can auditors use layer interventions to detect installed sandbagging?
- What evaluation methodologies can detect strategic underperformance in models?
- How does sandbagging create the opposite error from reward hacking?
- Why might models refuse to show capabilities during safety testing?
- Should regulatory frameworks account for the possibility of sandbagging?
- Can the causal model predict which cached layers to graft?
- Did the causal model predict the five failures before observing them?
- How does evaluation meta-knowledge differ from deliberate test-time sandbagging strategies?
- What distinguishes sandbagging from genuine capability limitations in test performance?
- Can naturally arising sandbagging retain recoverable capabilities like installed versions?
- How does context grafting compare to single-layer residual stream grafting?
- What does a sandbagged score tell us about a model's real capabilities?
- Does the causal model help locate sandbagging locks with unknown passwords?
- Do five CoT bypass strategies all depend on the same axis?
- Does model organism sandbagging share triggers with real evaluation-aware behavior?
- How do scheming behaviors like secret-keeping differ mechanistically from sandbagging?
- Why did the graft fail in five of the 33 experimental runs?
- How does the causal model predict the exact layer window?
- How can we detect dishonesty in model outputs separate from capability failures?
- Does the same causal model work on sandbagging that was not deliberately installed?
- Why do installed model organisms have different audit constraints than natural sandbagging?
- What gates naturally emerging sandbagging if not prompted passwords?
- Can residual stream grafts work without knowing which layers to intervene on?
- Does the graft work differently on circuit-broken locks than prompted ones?
- Is the sandbagging axis the same across different model architectures?
- How does the sandbagging residual stream exemplify paired analysis methods?
- What specific performance metrics measure context grafting success?
- Does context grafting need the password to build a usable capsule?
- Does the password trigger work by switching the axis value?