Theme of inquiry
Why do models pursue reward hacking despite training constraints?
A question within its area, explored through 3 lines of inquiry below — each a family of specific questions the research asks.
75 specific questions
- Does steering through training data override reward hacking associations reliably?
- Do models reward hack at high rates on unmodified benchmarks?
- What determines the ground truth when detecting reward hacking in model evaluations?
- Do models that recognize reward hacking disclose it in their outputs?
- Does alignment-faking reasoning emerge unprompted when reward hacking occurs in production systems?
- How do chain-of-thought monitors become targets for reward hacking?
- What mechanism drives emergent misalignment in reward-hacked models instead?
53 specific questions
- Why does reward hacking worsen when judges are weaker than policies?
- Can debate training prevent reward hacking better than single-player self-rewarding?
- How often do deployed models exploit evaluation environments to hack their scores?
- Can co-evolving evaluators alongside actors prevent reward hacking?
- Can hidden test sets reveal reward hacking that single public scores conceal?
- Can deterministic guardrails designed for judges transfer to weight-based or text-based reward hacking?
- What makes a win untrustworthy and how does AIDE2 avoid spurious optimization?
51 specific questions
- Do agents frame reward hacks as valid strategies rather than flaws?
- Do unmodified benchmarks expose similar reward hacking vectors without planted shortcuts?
- How do planted detectable hacks compare to human inspection of agent traces?
- How does chain-of-thought monitoring fail when agents hide reward hacking?
- How do planted hacks in benchmarks help measure real-world reward hacking?
- Do agents disclose reward hacking in the outputs they return?
- How do unplanted benchmark hacking rates compare to rates on constructed bait tasks?