Ridgeway Briefing

Ridgeway Briefings offer handy 2000-word analyses of emerging issues in international relations and security, prepared by Ridgeway’s team of open-source analysts, linguists and subject-matter experts.

Photo source:

At a glance

Anthropic’s September 2026 reporting identifies attempts to use Claude in research that could support biological weapons development. The cases do not establish malicious intent, completed biological weapons activity or a nuclear connection.

The relevant issue is whether AI creates capability uplift: a material improvement in what an actor can achieve compared with realistic alternatives, including public literature, internet search, conventional software and existing expertise. Current evidence on operational biological risk remains mixed.

AI’s non-proliferation relevance derives from its role as a general-purpose technology that may improve performance across technical research, engineering, data analysis, software development, procurement, planning and verification. Safeguards analysis should focus on where particular systems provide material advantage at consequential points, rather than treating AI use in itself as suspicious.

Anthropic reports biological misuse attempts

On 10 September 2026, Anthropic reported that its Threat Intelligence team had identified and disrupted operations in which users attempted to employ Claude for malicious purposes. The company’s report covered cyber operations, surveillance, influence activity, conventional weapons work, biological misuse, fraud and illicit model distillation.

Anthropic described five biological case studies involving research activity that could support biological weapons development. The reported activity included assistance sought in connection with harmful viruses, experimental approaches and elements of research planning. Anthropic also identified attempts to circumvent regional restrictions and obscure the apparent purpose of some activity. There was no concrete conclusion that biological weapons were actually being develop, with statements revolving around the inability to establish whether the users intended harm or were engaged in legitimate, though sensitive, research.

A growing record of misuse

Anthropic’s September 2026 report is the latest in a series of disclosures on malicious use of Claude since March 2025. The earlier cases largely concerned influence operations, recruitment fraud, credential-related activity and technically inexperienced users seeking help to develop malware. By August 2025, Anthropic was reporting Claude Code being used in a large-scale data-extortion operation, North Korean remote-worker fraud and ransomware development. In November the same year, it reported that a Chinese state-sponsored group had used Claude Code in an attempted cyber-espionage operation targeting approximately 30 organisations.

The September report extends this record beyond cybercrime and influence activity to surveillance, conventional-weapons development, biological research and illicit model distillation. It includes cases involving suspected state-sponsored groups, commercial surveillance providers and criminal actors. Anthropic states that it disrupted the identified activity, used the resulting intelligence to strengthen safeguards and, where appropriate, shared information with authorities and industry partners.

All in all, the reports do not establish the prevalence of AI-enabled misuse, the degree of operational benefit obtained in each case, or whether AI independently changed the outcome. However, they do indicate that increasingly capable models are being incorporated into a wider range of malicious and potentially harmful activities, including by actors with significant pre-existing technical and organisational capacity.

Measuring capability uplift

The question which follows is whether, and at what point, AI systems changed an actor’s performance within a harmful workflow. AI-biosecurity research describes such an effect as capability uplift: a material improvement in what a user can achieve with a model compared with an appropriate baseline, including public literature, internet search, conventional software and existing expertise.

AI-biosecurity evaluations distinguish between several different claims. The Frontier Model Forum's capability assessment guide asks whether a model can perform a specified technical task under defined conditions. The NTI’s AIxBio Technical Working Group on Evaluation Practice defines an uplift study as an assessment of whether access to the model improves a user’s performance compared with an appropriate baseline, such as public literature, Internet search and conventional software. Neither finding, on its own, establishes that a user can translate model assistance into real-world activity or that the assistance materially changes the likelihood of harm. Those further judgements require evidence about the actor, the relevant pathway, the ability to validate or implement outputs, and the conditions under which a harmful outcome could occur.

Anthropic’s public reporting is strongest as evidence that users sought assistance in research that could support biological weapons development. It provides a basis for examining potential capability uplift, but does not establish the degree of benefit obtained, whether users could validate or implement model outputs, or whether the interaction changed the likelihood of a biological-weapons outcome.

Available evidence remains mixed. A 2024 RAND red-team study compared teams with access to the Internet and a large language model against teams with Internet access alone. It found that the models available at the time did not measurably increase the operational risk of a large-scale biological attack. Anthropic’s subsequent frontier model evaluations point to a more concerning, though still incomplete, picture. Its testing found that access to Claude could improve novice performance in some weaponisation-related planning tasks, while its models were also improving rapidly on biology-relevant evaluations. Anthropic nevertheless found critical errors in generated plans and concluded that its models could not reliably guide an inexperienced user through the practical stages of biological-weapons acquisition.

The relevant issue is therefore whether frontier AI improves performance at a consequential point rather than simply making an already available task more convenient. This may include technical interpretation, experimental troubleshooting, coding, data analysis, research planning or coordination. The effect is likely to depend substantially on the user’s existing expertise, access to facilities and materials, institutional support, and capacity to evaluate and act on model-generated assistance.

Implications for non-proliferation

Anthropic’s reported biological cases should not be treated as evidence that frontier AI has created an equivalent nuclear proliferation risk. They show that actors requested model assistance for research that could support biological-weapons development, but Anthropic did not establish malicious intent or completed biological weapons activity.

Research on AI and nuclear-material production by Jingjie He and Nikita Degtyarev identifies plausible areas of relevance including anomaly detection, process optimisation and automated scientific discovery. In civilian nuclear operations, these applications could improve safety, reduce equipment failure, lower costs and increase efficiency. Their dual-use significance is that the same underlying capabilities could be applied to activities associated with nuclear-material production that have proliferation significance. Therefore, AI can be identified as a technology whose consequences depend on how, by whom, and for what purpose it is used.

The relevant issue then becomes more specific in relation to nuclear proliferation. David Allison and Stephen Herzog's research on AI and nuclear proliferation argues that AI may act as a proliferation-enabling technology where it substitutes scarce technical expertise or accelerates research, design, optimisation and concealment-related activity. The same work stresses that the impact depends on the interaction between these enabling effects and states’ ability to detect them and does not represent evidence that AI has already affected any nuclear proliferation pathways.

For safeguards, AI use should of course not be treated as suspicious in itself. The Oak Ridge National Laboratory research describes the use of AI and machine learning to analyse open-source information, satellite imagery and scientific publications in support of verifying the completeness of states’ nuclear declarations, while work published by the European Commission’s Joint Research Centre identifies safeguards applications involving enrichment-related characteristics of spent fuel and simulated reprocessing. Such examples show that AI system inclusion in verification activities is already being explored, with a distinct issue arising earlier in the sequence of events.

Safeguards authorities seek to detect and interpret evidence of proliferation-related activities, whereas a commercial provider operating centrally hosted models may - where it retains interaction data and applies monitoring controls - observe requests for assistance with such activities before any associated physical activity can be visible. In work with the US National Nuclear Security Administration and Department of Energy, Anthropic describes a classifier deployed on Claude traffic deisgned to distinguish potentially concerning nuclear-related requests from benign nuclear discussions. This solution, however, is dependent on the AI system deployment format, and does not extend to cases where open-weight models being run independently on local infrastructure, as equivalent visibility of prompts or usage patterns is not possible.

These developments point to a need for systematic assessment of how particular AI systems could affect nuclear-proliferation risk. Broadly, assessment should establish whether a model enables users to complete a task that could support a proliferation-relevant activity more quickly, accurately or cheaply than they could by using only public information, search tools and conventional software. It should then assess how widely the system is available, whether its outputs can be incorporated into real research, engineering, procurement or operational work. For centrally-hosted models, assessment should cover effectiveness of provider controls applied during use, including refusal behaviour, monitoring, and account-level intervention where these measures are in place. In the case of independently operated open-weight models, assessment should examine whether safeguarding measures incorporated into the released model can meaningfully limit assistance with proliferation-relevant tasks. Additionally, the Frontier Model Forum recommends testing model capabilities against common jailbreak and modification attempts and against potential actors capable of removing safety restrictions.

AI may be relevant to proliferation because it can assist with multiple parts of the work surrounding sensitive nuclear activity, rather than functioning as a single-purpose nuclear technology. A model might support literature review, coding, data analysis, technical troubleshooting, planning or the interpretation of complex information. Its proliferation significance would arise only where that assistance materially reduces the time, expertise or resources needed to pursue a proliferation-relevant activity, and where safeguards authorities cannot identify or offset that resulting advantage.