Atria Dawn: The Dawn of Agentic Superintelligence
As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn Preview, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world. This model is trained via a Verifiable Experience Pipeline that connects tool-mediated interactions to executable environments and externally verified outcomes. Across 16 benchmarks spanning real-world research, engineering, and digital work, Atria Dawn Preview is competitive with frontier agents and achieves the highest reported score on five of them. Beyond standalone performance, we examine the real research-anddevelopment process behind this model as a case study of human–AI collaboration, analyzing 769 task records from 56 participants together with agent logs. When asked to evaluate completed tasks under comparable conditions, participants rated about one-third of completed AI-assisted tasks as infeasible without AI. More strikingly, agents frequently propose methods and implement revisions, while humans retain most final decisions and guide exploration through judgment and feedback.
Introduction. As language-model agents take on tasks requiring sustained tool use (Dong et al., 2026; Schick et al., 2023; Shen et al., 2026; Xi et al., 2025, 2026), including software engineering (Jimenez et al., 2024; Yang et al., 2026a) and workplace document reasoning (Guo et al., 2026; Tang et al., 2026), they increasingly participate in the development of AI systems themselves. Studies on algorithm discovery and automated research show that agents can propose candidates, run experiments, and revise solutions in response to feedback (AlphaEvolve Team, 2025; Lu et al., 2026; Novikov et al., 2025). These capabilities raise a question that task performance alone cannot answer: in a real model-development project, who identifies worthwhile problems, chooses among proposed methods, interprets uncertain results, and decides what to pursue next? Understanding this division of responsibility is necessary to assess both the contribution of agents and the changing role of human researchers.
Discussion / Conclusion. While training our models, we observed a shift in the roles of AI agents and human researchers. AI agents took on greater responsibility in the research process, consistent with observations reported by OpenAI and Anthropic (Hitzig et al., 2026; OpenAI, 2026). Beyond executing tasks, agents also assumed responsibility for aspects of research planning, including designing workflows and deciding how to revise experimental plans across iterations. This shift points toward the possibility of recursive self-improvement (RSI): stronger models can contribute more effectively to research and development (R&D), resulting in stronger subsequent models, creating a self-reinforcing cycle of capability gains. Yet our experience with Atria Dawn suggests that human researchers remained important during this Every time AI crosses a major threshold, the human role would be redefined, and it is now moving from executing specific tasks to exercising judgment at critical decision points.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can brute-force automated research substitute for iterative depth and human research intuition?- What distinguishes artifact efficiency improvements from research process efficiency improvements?
- Do gains in optimization benchmark scores translate to gains in real research efficiency?
- Does delegating planning to agents change the speed of the research process?
- Can accumulated priors and outcome analysis speed up research automation?
- Can agents take on research planning tasks while humans focus on judgment?
- How should researchers operationalize and measure methodological guidance at different levels?
- How does this approach differ from AI research acceleration focused on insight distillation?
- How many acceptable rewrites can recursive self-improvement sustain before returns diminish?
- How does this scoped definition relate to the survey's open-ended recursive self-improvement?
- Can autonomous research agents outperform hand-tuned hyperparameter search?
- Does human-AI collaboration improve faster and safer than autonomous self-improvement?
- Can humans remain meaningfully in the loop as AI autonomy scales?
- Should human oversight capacity be designed as carefully as AI capability?
- Where should humans take over from AI during research tasks?
- How do different definitions of intelligence shape AI research priorities?
- Does greater inclusion of disciplines improve AI research goal alignment?