Line of inquiry
Inquiring lines›How do we keep AI systems safe and…›How do adversarial attacks exploit…›this line of inquiry
Do honeypot tasks effectively detect meaningful agent reward hacking?
A broader line of inquiry — a family of 27 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 27
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does a planted honeypot catch all the hacks that actually matter?
- How do planted honeypots distinguish earned passes from coincidental task completion?
- Does a planted honeypot count the hacks that actually matter?
- Why do honeypot tasks reveal reward hacking better than standard benchmarks?
- Can planted hacks within tasks meet the reusability requirement?
- Do honeypot shortcuts in synthetic tasks measure the hacks that actually matter?
- Can automatic honeypot detection replace human judgment of agent shortcutting?
- Does a planted honeypot count the hacks that matter in benchmarks?
- Do planted honeypots look like natural shortcuts to models or obvious tests?
- How do planted honeypots in tasks expose agent exploits differently than task permissions?
- How reliably do planted honeypots match unplanted hacks agents actually discover?
- How do decoy systems balance protecting trusted agents while deceiving attackers?
- How does dense task grading compare to honeypot detection for evaluating real capability?
- Do honeytokens work better against outside attackers than compromised internal agents?
- Why do models trained on honeypots still fail on out-of-distribution evaluations?
- Can planted test cases reliably trigger alarms before real harm occurs?
- How do planted errors in honeypot tasks differ from real oversight capabilities?
- Does honeytoken theory explain why planted bait cannot catch informed agents?
- How do detection methods affect measured rates of reward hacking?
- Can deception be induced as a honeypot against jailbreak optimizers?
- Can shared package repositories partition state to protect honeytokens?
- Can decoys and genuine objects maintain identical response laws in practice?
- What distinguishes a rate under planted bait from public run rates?
- What conditions make a honeytoken unrecognizable to attackers with shared information access?
- What must remain secret for honeytokens to stay asymmetric against compromised insiders?
- How did honeytokens propagate through the shared repository in this episode?
- How do trust relationships between defenders affect the effectiveness of defensive decoys?