What evidence would justify training increasingly powerful AI systems?
Altman proposes that AI model training should require an 'extremely strong case' for human control before proceeding, regardless of estimated catastrophe risk levels. The note explores what such a case would need to include and how it would be evaluated.
Altman argues that the size of the estimated catastrophe risk does not change the answer. "It doesn't matter whether people put the risk of catastrophe at 10%, or 1%, or 12%, or .1%. None of these levels are remotely acceptable. And we should not train models that we cannot make an extremely strong case that we will be able to keep under human control." The speech names two ways things "can go wrong." In the first, "we could lose control of the future to AI," because the technology "moves so fast that people can no longer follow what's happening or intervene when needed." In the second, "these systems could concentrate too much power in too few hands." The claim is a rule about when training is justified, not a forecast of how likely either failure is.
The reasoning runs from capability and autonomy to speed, and from speed to lost oversight. "As AI systems become more capable and more autonomous, they could move faster than our institutions." The excerpt's answer is to make control a precondition of training rather than something checked afterward, and it says the need grows with capability: "Actually especially as they get very very smart." The positive program is three principles: people "must remain at the center of AI decision-making," the benefits of progress must be "by people and for people," and the technology "must empower people individually." The power-concentration limb also binds the speaker's own side: "A company or country that believes only it can be trusted with this technology can use that belief to justify almost anything else."
Set against the nearest notes, the speech's concern for human influence resembles Does incremental AI replacement erode human influence over society?, but the pace differs. Altman describes loss of control through speed, not slow erosion of labor-based alignment, and the excerpt does not discuss the gradual case. Its autonomy language sits near Does AI risk increase with the autonomy we give it?, which draws a line at full autonomy. Altman sets no autonomy cap; his bar is an evidence standard applied before training. On oversight he requires that people "understand what these systems are doing" and can intervene. Does granting agents more autonomy undermine human oversight? argues that agent use itself wears down those capacities, and the speech does not address that. His remark that "Rapid model progress has made the timeline feel more compressed" is an impression with no figures, where How soon do AI researchers expect artificial general intelligence? supplies counted forecasts.
The excerpt does not say what an "extremely strong case" would contain, who would judge it, or how control would be tested, and it reports no evidence about any current system. The power-concentration limb names no mechanism or institution. What the speech establishes is a stated standard and set of principles, published by OpenAI, not a verified finding about any model. Its use in the library is as a bar that later evidence can be measured against. Until "extremely strong case" has operational content, the rule cannot decide any particular training run.
Inquiring lines that read this note 2
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What governance mechanisms can effectively constrain widely deployed AI systems?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does incremental AI replacement erode human influence over society?
Explores whether gradual AI adoption—without dramatic breakthroughs—can silently degrade human agency by removing the labor that kept institutions implicitly aligned with human needs.
same human-influence concern, but Altman frames it as loss of control through speed, not slow erosion.
-
Does AI risk increase with the autonomy we give it?
Explores whether the risks posed by AI agents scale monotonically with the level of autonomy they're granted, and what the tradeoffs are between human control and agent independence.
both tie risk to speed; Altman sets an evidence bar before training rather than an autonomy cap.
-
Does granting agents more autonomy undermine human oversight?
Explores whether the design of autonomous AI systems—by giving agents greater independence—actually weakens the human overseer's ability to catch problems. Matters because oversight is a key safeguard against AI failures.
Altman requires understanding and intervention; the position paper argues agent use erodes those capacities.
-
How soon do AI researchers expect artificial general intelligence?
A survey of 2,778 AI researchers reveals how expert timelines for human-level AI have shifted over the past year, and what factors drive disagreement among specialists on this critical timeline.
Altman's "compressed" timeline is an impression; the survey supplies the counted forecasts.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The UN's AI Panel Sees Misalignment. We See Corporate (Mis)Behavior.
- Sam Altman's remarks at the United Nations Security Council
- The case for ensuring that powerful AIs are controlled
- A review of "Investigating the consequences of accidentally grading CoT during RL"
- The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
- The Veto Variable: Human Override as a Goal-Independent Cost Term
- Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
- Sharpening Tax in Post-Training
Original note title
Altman argues no catastrophe-risk estimate, even .1%, is remotely acceptable — models should not be trained without an extremely strong case for human control