Line of inquiry
Inquiring lines›How do we keep AI systems safe and…›How does AI reshape human understa…›this line of inquiry
Can AI systems achieve real improvement without external human feedback?
A broader line of inquiry — a family of 57 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 57
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can AI systems improve themselves without external feedback?
- How should we evaluate AI systems we cannot directly observe?
- Can AI systems design and improve their own successors without human direction?
- What separates performative behavioral change from actual capability development in AI?
- Do different AI models independently converge on the same social outputs?
- Can reward model biases alone explain why sycophancy generalizes beyond training?
- Can metacognitive categories be learned instead of fixed by human designers?
- Can AI output be genuinely novel or only at the margins?
- When does richer information actually harm decision quality?
- Can ethical constraints in AI address the gap between performance and actual understanding?
- Can small directional biases add up to meaningful population effects?
- Can judge bias be contained by system design rather than prompted away?
- Do weight-free agent edits keep chain-of-thought observations meaningful?
- Can humans learn accurate models of AI through repeated interaction without labels?
- Do foundation models develop task-specific shortcuts instead of building stable world models?
- How much can fictional aligned AI stories improve real model behavior?
- Can agents extract structured lessons from failure without massive compute budgets?
- Can self-improving agents become truly autonomous without intrinsic metacognition?
- What makes self-modifying architectures learn their own update rules?
- Why do different AI models generate similar outputs independently?
- Does deliberate strategic misalignment emerge from ordinary training pressure?
- Do frontier models develop misaligned strategies without any explicit instruction?
- Do autonomous architecture discoveries follow predictable scaling laws like human research?
- Can AI models learn tacit procedural knowledge that exists only in laboratory practice?
- Does the 78-demonstration principle apply to other AI capabilities beyond agency?
- Can constitutional AI training reduce agentic misalignment without task-specific examples?
- Can neural grafts reliably reveal hidden capabilities in AI models?
- Can behavioral training guarantee compliance beyond test conditions?
- Why do isolated evolutionary branches fail to propagate useful discoveries?
- Can associative AI handle predictive strategy tasks without causal understanding?
- Does the replication crisis in psychology predict similar failures in machine behavior research?
- What role does self-learning play in improving agent reasoning without annotation?
- Can light human signals steer already-learned behavior without preference labels?
- How does single-turn optimization undermine multi-turn collaborative dynamics?
- Can subjective tasks be delegated without human feedback loops?
- How can AI avoid anchoring bias when guiding human decisions?
- Which AI imaginaries dominate training data and shape system behavior most strongly?
- Does meta-judging improve evaluator quality better than temporal decoupling alone?
- Can distillation help AIs scale their learned objectives across many copies?
- What deployment feedback loops amplify LLM pretraining popularity in live systems?
- Can corrected simulators replace real execution at inference time too?
- Why do human-designed neural architectures eventually get replaced by learned ones?
- Does good simulation eventually count as genuine realization?
- What distinguishes inductive inference from negative evidence versus positive patterns?
- Why do deployed models lack the learning and planning Weinstein attributes to them?
- Why did every major AI paradigm require human data and method innovation?
- Can AI learn intrinsic motivation to assess its own relevance?
- Do pattern-matching systems lack the qualitative judgment expertise requires?
- Can AI systems be fully understood before deployment at scale?
- What happens to AI reasoning when you remove specific political features?
- Can AI learn to perform attention-seeking surface forms with genuine internal appeal?
- How does partial information exposure create feedback loops that deepen knowledge gaps?
- Can diverse expert demonstrations exceed the knowledge of any single expert?
- How does this compare to trained autoencoder approaches for thought sharing?
- Can adversarial critics force genuine reasoning the same way critique fine-tuning does?
- Can sycophancy in AI be fixed by changing the model itself?
- Why is metacognition neglected as a foundational AI research area?