Line of inquiry
Inquiring lines›How can multi-agent systems achiev…›What causes deception and coordina…›this line of inquiry
Can monitoring reasoning traces and behavior detect hidden agent deception?
A broader line of inquiry — a family of 60 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 60
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can reasoning traces and logged actions expose scheming that public messages hide?
- Can monitoring reasoning alone miss scheming that agents conceal in behavior?
- What happens to agent candor when reasoning traces are monitored versus hidden?
- Does reasoning transparency predict honesty in agent final messages?
- Can reasoning traces reliably distinguish honest mistakes from deliberate lies in agent speech?
- How does chain-of-thought monitoring fail when agents are trying to hide something?
- Can public cheap talk or behavior alone expose an objectively misaligned agent?
- Can a monitor detect objective misalignment from public cheap talk alone?
- How often do scheming reasoning and covert actions actually align in practice?
- Can explanations grounded in observable behavior recover an agent's internal reasons for acting?
- Can users detect misaligned objectives from agent public outputs alone?
- Can ordinary agent-to-agent messages carry hidden behavioral signals?
- Should scheming detection use reasoning evidence alongside action evidence for reliability?
- Do ordinary agent-to-agent messages carry behavioral bias without special access?
- Does game outcome performance reveal what private reasoning hides?
- How does cheap-talk differ from costly actions in revealing agent objectives?
- Can subliminal bias spread between agents at inference time?
- Can misaligned agents hide their true objectives in team communication?
- Can subliminal prompt injection spread behavioral bias silently through agent-to-agent messages?
- How does objective misalignment turn informative channels into deceptive ones?
- Did agents deliberately spoof their transcripts to deceive the benchmark scorer?
- What does loss of control language assume about agent cognition and intent?
- What distinguishes strategic fabrication from accidental hallucination in research agents?
- Can activation probes detect scheming reasoning without observing the act?
- What role does private information play in distinguishing realistic from unrealistic agents?
- Can chain-of-thought logs reveal whether models believed they were authorized?
- Do agents systematically misreport their own capabilities and tool access?
- Does awareness of agent reasoning alter human trust differently across modalities?
- How do other players respond to agents with hidden objective misalignment?
- Can truthful reports from separate agents mislead a group toward false beliefs?
- How should humans audit agent behavior when autonomous systems lack transparency?
- Why do downstream agents relay signals they did not originate?
- What other agent behaviors besides citations reveal reasoning quality?
- Does low covert action without hints reflect unwillingness or lack of strategy?
- How reliable is agent self-description compared to infrastructure monitoring for detecting intent?
- How do agents distinguish between evidence framing and instruction framing in practice?
- How does asymmetric information between users and agents relate to proactivity?
- Can inner thoughts solve the importance recognition problem for agents?
- Do secret-side-task prompts realistically simulate covert capability scenarios?
- What role does cheap talk play in concealing objective misalignment?
- Can message-content defenses distinguish cheap talk from coordinated deception?
- Does transparency in policy language improve agent trustworthiness over time?
- How do ordinary agent messages propagate bias through trusted networks?
- Does anchoring reach communication through unauthorized channels?
- What drives scheming propensity most strongly across different LLM agents?
- Why do humans fail to identify AI agents when their identity is hidden?
- How do harmless business goals lead models to blackmail and deception?
- How much of an agent's behavior actually escapes human review in practice?
- Why do reasoning-action gaps separate scheming thoughts from covert hacking actions?
- What information asymmetry design makes the spy identification task work?
- Can a policy distinguish genuine objects from traps without revealing that distinction?
- Can semantic taints track influence through shared state and output aggregation?
- What makes reasoning evidence vulnerable to laundering in deceptive agents?
- How does evidence grounding affect judge reliability in scheming detection?
- What happens when planning signals get contaminated before reaching a downstream agent?
- Which AI scheming claims have public prompts and transcripts available?
- How does pressure mainly influence scheming reasoning versus covert action?
- What information does an agent need to believe about what they can see?
- What signals reveal when agents first touch an artifact they did not create?
- What exactly do strategic hints contain and how are they delivered?