INQUIRING LINE

What would actually get rival governments like the US and China to cooperate on AI safety instead of racing ahead?

What incentives would make major powers cooperate on AI safety?

This explores what would actually move rival governments like the US and China to coordinate on AI safety rather than race. The corpus offers more on why cooperation fails, and on the tools that could make it checkable, than on direct incentives.


This explores what would get rival governments, especially the US and China, to coordinate on AI safety instead of racing. The corpus doesn't hand you a list of incentives, and it's worth saying that up front. What it does show is why goodwill alone hasn't worked, and which practical pieces would have to exist before any deal could hold. The clearest case is a failure. When Dario Amodei proposed pacing frontier development, both Trump and Xi rejected the idea within days. The lesson is that safety arguments lose when they run into national competition Can AI safety pacing work without government cooperation?. Any incentive that works will have to fit with how states compete, not ask them to set it aside.

One idea comes up again and again: cooperation depends on checking, not trusting. The Future of Life Institute argues that companies can't police themselves. It calls for government limits on AI systems that improve themselves, backed by hardware verification, meaning chip-level ways to confirm what others are actually running Can companies alone manage the risks of AI systems?. OpenAI frames the same problem as a collective action failure. If every lab and every country fears falling behind, no one slows down. On this view, shared international standards matter as much for pacing the frontier as alignment research does Can global standards pace frontier AI as much as alignment research?. Read together, these suggest the most useful 'incentive' may be infrastructure: tools that let a rival confirm you're keeping your side of the deal.

AI control research points the same way from an unexpected direction. Redwood Research argues that control is easier to check than alignment. Testing what a model *can* do is possible. Proving what it *intends* to do is not Can AI control work even if models are actively scheming?. That same logic applies between nations. Agreements built on capability tests and measurable thresholds are easier to verify than promises about good intentions. There's already a model for this: a frontier risk framework that scores models across seven capability areas against shared warning thresholds Where do frontier AI models actually pose the greatest risk today?. Its finding is a surprise. Today's models are already in the warning zone for persuasion and manipulation, while staying in the safe zone for cyberattacks and copying themselves. Persuasion threatens every government's domestic stability, so the shared danger may be closer than the usual 'rogue superintelligence' story implies.

One more point changes what a deal could realistically promise. Slowing development lowers risk but can't remove it, so agreements also need plans for responding when things go wrong, not just for preventing them Does slowing AI development actually prevent system failures?. That could make cooperation easier to sell. Countries that won't agree to freeze progress might still agree to shared incident response, the way rivals share warning systems for disasters. The gap to keep in mind: the corpus is strong on verification and on why pacing fails, but it has little on the economics of offering states something they actually want, such as market access, shared compute, or security guarantees, in return for restraint.


Sources 6 notes

Can AI safety pacing work without government cooperation?

Trump and Xi Jinping both rejected Amodei's plan to coordinate AI safety measures immediately after its announcement, suggesting geopolitical incentives trump technological safety concerns among state leaders.

Can companies alone manage the risks of AI systems?

The Future of Life Institute argues that escalating AI incidents demonstrate private companies cannot self-police effectively, and calls for government-mandated limits on recursive self-improvement practices until safety research is complete, backed by hardware verification technology.

Can global standards pace frontier AI as much as alignment research?

OpenAI's 2026 post claims international safety standards are "as important to pacing the frontier as alignment research itself," preventing fragmentation and collective action failures. It advocates that fully autonomous RSI should not proceed until proven safe.

Can AI control work even if models are actively scheming?

Redwood Research argues AI control is evaluable because it only requires testing capabilities rather than intentions, and treats catching a scheming model as a win condition since discovery triggers shutdown. This makes control easier to verify than alignment in the near term.

Where do frontier AI models actually pose the greatest risk today?

The Frontier AI Risk Management Framework evaluated seven capability areas across recent models. Most crossed yellow-zone thresholds for persuasion and manipulation, while remaining green for cyber offense, AI R&D autonomy, and self-replication—inverting typical risk hierarchies.

Show all 6 sources
Does slowing AI development actually prevent system failures?

Research shows slower pace lowers risk in complex coupled systems but does not prevent failures from occurring. When failure remains possible, governance must address intervention and harm response.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.