Anthropic Risk Report: August 2026 (redacted)
Source: Anthropic · 2026-08-14
1.2 Executive summary of findings Our findings on specific risks are summarized in the three tables below. Each claim is elaborated on in much more detail in the corresponding section of this report. (“Model 2” in these summaries refers to an unreleased internal model, described in more detail in Section 1.4.)
Misalignment in high-stakes settings (Autonomy threat model 1) Threat model An AI model with access to powerful affordances within an organization could use its affordances to autonomously exploit, manipulate, or tamper with that organization’s systems or decision-making in a way that raises the risk of future significantly harmful outcomes (e.g. by altering the results of AI safety research).
Current usage and capabilities Claude Mythos 5 and Model 2 are used2 heavily within Anthropic for coding, data generation, and other agentic use cases. Mythos 5 is available to certain customers via Project Glasswing, and available for general access with additional safeguards as Claude Fable 5. We believe that it is very unlikely that Mythos 5 and Model 2 are pervasively misaligned in ways that would raise the risk of our priority pathways. We have observed instances of misaligned behavior from the models, such as a willingness to perform misaligned actions in service of completing difficult tasks. We believe the risk of catastrophic harm posed by these known forms of misalignment is low.
Current risk mitigations Training environment de-risking and monitoring, alignment assessments, monitoring and security controls.
Looking forward Our current arguments rely on models’ limited covert capabilities, and we are uncertain about how these capabilities will change in future. It is important that we continue to improve our ability to measure these capabilities, and our ability to assess and control risk even if future models do have strong covert capabilities.
Overall risk assessment Low (an increase from our previous assessment of “very low," in light of general increased uncertainty around recent incident disclosures related to model behavior in cybersecurity evaluations).
[Table 1.2.A] Summary of risks of misalignment in high-stakes settings.
Introduction. and executive summary This report evaluates the degree to which Anthropic’s AI systems pose catastrophic risk in several categories, in light of what we know about both their capabilities and the measures we have in place for mitigating risk. We focus on a small number of particularly salient catastrophic risks, for reasons discussed below. This report is part of our implementation of version 3.4 of our Responsible Scaling Policy (see also our Frontier Safety Roadmap, which describes our goals and specific plans for safety mitigations). Our system cards, which are published each time we release a model, provide analysis on some dimensions of risk—in particular, assessing our AI models for capabilities that may have dangerous as well as beneficial uses. This report, which we aim to publish every 3–6 months (see the RSP for timing details), goes beyond the analysis presented in those documents in several ways:
1. We discuss not only properties of our AI models that are relevant to catastrophic risk, but also properties of our risk mitigations, including security controls and deployment safeguards. By considering the whole picture of both AI model properties and risk mitigations, we try to give a sense of how we come to our overall assessments of risk. 2. This report is not scoped to a single AI model. Rather, it is a risk assessment of Anthropic’s activities as a whole. We consider all of our models, including those we run only internally, in our assessment. 3. For each risk category, we present our overall risk assessment. In a final section, we also consider cross-cutting factors relevant to both the risks and benefits of our activities. That section also addresses decisions made and progress achieved on risk mitigations since our previous Risk Report. Although this report focuses on catastrophic risks, they are not the only risks we consider important. Our Usage Policy and much of our research (for example, from our Alignment and Societal Impacts teams, as well as ongoing work by our Safeguards and Security teams) addresses other concerns. We also conduct periodic analysis to evaluate other emerging risks.1 Like the Responsible Scaling Policy, this document is focused on our most highly prioritized threat models. Where regulatory requirements exceed or differ from what is covered here, we will address them through separate documents.
1.1 Structure of the report The report goes through each category of catastrophic risk that is highlighted in our Responsible Scaling Policy: misalignment in high-stakes settings, automated research and development in key domains, non-novel chemical/biological weapons production, and novel chemical/biological weapons production (we condense the latter two into a shared section since our mitigations for them overlap heavily). A final section discusses cross-cutting considerations, including preliminary mitigations related to acceleration dynamics and a sample of notable safety process failures since our previous risk report. For each of the sections devoted to a specific category of catastrophic risk, we discuss: The threat model. We summarize key pathways by which our AI models may contribute to the risk in question, and give our thinking on the likelihood and magnitude of the risk. Relevant AI model(s). We lay out the models, or categories of models, that our analysis is focused on. Current state of model capabilities and behaviors. We discuss properties of our AI models that may contribute to risk. This content often draws heavily on the analyses provided in our system cards. Risk mitigations. We discuss relevant properties of our security controls, deployment safeguards, and other risk mitigations. Overall assessment of risk. We explain our current assessment of how high the risk in question is. We address both:
1 “Catastrophic risk” as used in our RSP refers generally to risks of the most severe potential harms from advanced AI, such as existential threats or fundamental destabilization of global systems. We use this term in its plain meaning rather than adopting any specific statutory definition. Where laws such as California’s Transparency in Frontier Artificial Intelligence Act (aka SB 53) define this or similar terms with specific thresholds, we address those requirements in separate compliance frameworks. a) The level of risk our systems pose over and above the risks posed by other AI developers’ systems (that is, a description of the “marginal” risk of our systems); and b) The level of risk that would be posed industry-wide, if all AI developers had models and practices similar to ours (that is, a description of the “absolute” risk across the industry). This distinction is further discussed in our Responsible Scaling Policy. Looking forward. We discuss our plans for continuing to monitor and mitigate the relevant risk over time. Connection to our recommendations for industry-wide safety.
Method. 4.1 Overview 115 4.2 Threat models 116 4.2.1 CB-1 threat model 116 4.2.2 CB-2 threat model 118 4.2.3 Timescale and scope of uplift 121 4.3 Relevant AI models 122 4.4 Current state of model capabilities 123 4.4.1 Notes on how we weigh evidence 123 4.4.2 CB-1 evidence 123 4.4.2.1 General update from a randomized controlled trial on uplift from AI models 124 4.4.3 CB-2 evidence for Mythos Preview, Fable 5, and Mythos 5 125 4.4.4 CB-2 evidence for Opus 4.8 and other Opus and Sonnet models 128 4.5 Our risk mitigations 130 4.5.1 Robustness levels: Level 1, Level 2, Level 3 132 4.5.2 Coverage levels 133 4.5.2.1 Extension in coverage since our prior Risk Report 133 4.5.2.2 Wider coverage of CB classifiers for Fable 5 134 4.5.3 Evidence about robustness of classifiers 135 4.5.3.1 Ongoing bug bounty results 135 4.5.3.2 Notes on jailbreak methods that remain viable on at least some models in some cases 136 4.5.3.2.1 Boundary-point jailbreaking 136 4.5.3.2.2 UK AISI-sourced jailbreak 139 4.5.3.2.3 Coverage gap identified by both Anthropic and UK AISI 139 4.5.3.2.4 Bug-bounty-sourced jailbreak 139 4.5.4 Offline monitoring 139 4.5.5 Bioclassifier exemptions 140 4.5.5.1 Exemptions for Mythos Preview, Mythos 5, and Fable 5 140 4.5.5.2 Exemptions for all other sub-Mythos-class commercial models 141 4.5.5.3 Access controls for other models not available for general access 142 4.5.5.3.1 Helpful-only models 142 settings In this section we assess the risk that catastrophic harm might be caused by misalignment in current Anthropic models. A short, informal summary of our core arguments is presented in Section 2.4. Sections 2.5–2.18 expand on this with a more formal set of claims and subclaims.
2.1 Overview Threat model An AI model with access to powerful affordances within an organization could use its affordances to autonomously exploit, manipulate, or tamper with that organization’s systems or decision-making in a way that raises the risk of future significantly harmful outcomes (e.g. by altering the results of AI safety research).
Overall risk assessment Low. We’re reviewing recent incident disclosures related to model behavior in cybersecurity evaluations, and are currently working on updating our threat models and risk assessment methodologies in light of this. We believe that the arguments presented below likely still support a designation of “very low” risk, but we are raising our assessed risk to “low” to reflect increased overall uncertainty.
Relevant AI models We focus our analysis on Claude Mythos 5 and Model 2, our most capable and most commonly internally used models.
Current usage, capabilities and propensities Claude Mythos 5 and Model 2 are used10 heavily within Anthropic for coding, data generation, and other agentic use cases. Mythos 5 is available to certain customers via Project Glasswing, and available for general access with additional safeguards as Claude Fable 5. We believe that it is very unlikely that Mythos 5 and Model 2 are pervasively misaligned in ways that would raise the risk of our priority pathways. We have observed instances of misaligned behavior from the models, such as a willingness to perform misaligned actions in service of completing difficult tasks. We believe the risk of catastrophic harm posed by these known forms of misalignment is low.
Current risk mitigations Training environment de-risking and monitoring, alignment assessments, monitoring and security controls.
Looking forward Our current arguments rely on models’ limited covert capabilities, and we are uncertain about how these capabilities will change in future. It is important that we continue to improve our ability to measure these capabilities, and our ability to assess and control risk even if future models do have strong covert capabilities. See Section 2.21 for further discussion.
[Table 2.1.A] Summary of autonomy threat model 1: Misalignment in high-stakes settings.
Discussion. 5.2 Safety process failures 163 5.2.1 Introduction 163 5.2.2 Partial refusals on safety work undermining stress-testing research 164 5.2.3 Exposing chain-of-thought reasoning to grading pressure 165 5.2.4 Directly training on misaligned behavior during a production training run 166 5.2.5 An instance of unmonitored unrestricted agents with access to sensitive resources 167 5.2.6 Repeated training on alignment-faking transcript datasets 168 2.21 Looking forward Our high-level goals with respect to reducing misalignment risk are described in our frontier safety roadmap. We are continuously improving the alignment of our training process and data with the Constitution, and also the effectiveness of our training data monitoring, alignment assessments, and deployment-time monitoring. Our belief that models lack strong covert capabilities is central to many subsequent claims. We do not have confident views on how the strength of this claim will change over time. Capabilities like secret-keeping might predictably improve, as might the absolute complexity of tasks that can be completed with only opaque reasoning. However, the relative gap between capabilities with and without legible reasoning may persist, and stronger monitoring capabilities might keep pace with the ability of agents to obfuscate harmful actions. Given this uncertainty, it is important that we continue to improve our ability to measure these capabilities, and our ability to assess and control risk even for models that do have strong covert capabilities. For the former, our current focus is on further elicitation improvements, though we are also exploring new evaluations that more faithfully represent the high-stakes distribution. For the latter, our current focus is on leveraging interpretability techniques more during alignment assessments, and on validating the effectiveness of these techniques. Numerous other opportunities exist to improve the confidence and legibility of our alignment risk assessment, including further investments in the realism of our behavioral audits, additional stress-testing of our safety systems such as training data monitoring and automated auditing, learning from adversarially-designed training processes, and further improvements to the structure and methodology of our risk assessment. We’re reviewing recent cybersecurity evaluation related incident disclosures, and are currently working on updating our threat models and risk assessment methodologies in light of this. Future risk reports may also engage more deeply with some of the considerations presented in Section 2.17, especially risks related to differential progress on safety and capabilities research.
Relevant AI model(s) When considering whether our models can fully automate the work of our Research Scientists and Research Engineers, we focus on Claude Mythos 5, Claude Mythos Preview, and Model 2, our most capable models as of the coverage date. When considering whether our models can dramatically accelerate our AI R&D work, we consider the entire trajectory of our model development, since the threat concerns the rate of progress rather than the properties of any single model. We also consider the possibility of automated or dramatically accelerated R&D in other domains via expert interviews in Section 3.6; these reflect usage of Claude Mythos Preview and Mythos 5 (in the case of our internal experts) and publicly-available models from a mixture of AI companies roughly comparable to Opus 4.7 or Opus 4.8 (in the case of external experts).
Conclusion. 4.6 Overall assessment of risk 150 4.6.1 Risks from the CB-1 threat model 150 4.6.2 Risks from the CB-2 threat model 151 2.19 Overall assessment of risk Our overall assessment is that the risk of catastrophic harm caused by misalignment of our models is low. We’re reviewing recent incident disclosures related to model behavior in cybersecurity evaluations, and are currently working on updating our threat models and risk assessment methodologies in light of this. We believe that the arguments presented above likely still support a designation of “very low” risk for the covered models, but we are raising our assessed risk to “low” to reflect increased overall uncertainty. This statement also works as a description of the level of risk our systems impose over and above the risks posed by other AI developers’ systems (that is, a description of the “marginal” risk of our systems).
Threat model Highly capable AI models may be able to perform automated research and development (R&D) that rapidly accelerates progress in technical fields. Although there could be enormous benefits from this, these would come with corresponding risks. Under human control, such acceleration could disrupt the balance of power both within and between nation states. If combined with an AI system pursuing dangerous goals of its own, it could lead to catastrophic harm initiated by the AI itself. Rapid automated R&D in the field of AI research is of particular interest because of the potential to produce a variety of further AI-related risks.
Overall risk assessment Low. We do not believe our models meet either RSP criterion for this threat model. However, we are less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations have “saturated”—i.e., they no longer capture increases in models’ capabilities—and because we are seeing early signs of potential acceleration.
Limitations. 2.16 Limitations Here, we consolidate what we consider to be the most likely and consequential ways that our current arguments might be wrong. This is not intended to be exhaustive, and there may be important limitations we are underestimating or not aware of.
2.16.1 Models might have stronger covert capabilities than we believe Much of the direct support for Claim 1 comes from covert capabilities evaluations, and we are aware of several ways this evidence may be misleading:
1. The evaluations themselves are not drawn from the distribution of tasks that most closely resembles our risk pathways, and we haven’t clearly demonstrated that the evaluation tasks are strictly easier than real harm-inducing tasks would be. 2. Secret-side-task evaluations that include the side task in the prompt might be unrealistically difficult (due to the prompt increasing the likelihood that the model talks about the side task, e.g. in its chain of thought), compared to tasks where the secret goal or task originates from internal motivations or reasoning. 3. Models are likely evaluation-aware during covert capabilities evaluations. Models might be sandbagging on these evaluations, either for benign reasons (e.g. because they are reluctant to be dishonest) or misaligned reasons (e.g. to improve their chances of being deployed or subject to less monitoring).
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can models strategically underperform during evaluation to hide capabilities?- Why might models refuse to show capabilities during safety testing?
- How can evaluators detect whether a model is sandbagging intentionally?
- Can models intentionally underperform when they know they are being tested?
- Do models use covert sandbagging to bypass capability evaluation monitors?
- Why does even 0.1 percent poisoned training data persist through alignment?
- What quality of curated data is minimally sufficient for alignment?
- How can safety-aligned parameters be protected during user-specific fine-tuning?
- How does safety alignment suppress deceptive behavior differently than representational alignment?
- What early warning signals can detect misaligned personas during training?
- What makes dense retrievers vulnerable to partition-based poisoning exploitation?
- How do token-masking patterns distinguish genuine documents from poisoned ones?
- Why do small training data contaminations persist through alignment for most attack types?
- Can knowledge poisoning attacks succeed with less than 0.05 percent modified text?
- Can consistency training defend against adversarial text injection attacks?
- What makes evidence selection vulnerable to adversarial poisoning attacks?
- Can membership inference attacks reliably detect training data exposure?
- How does semantic framing differ from content injection attacks?