Abstract
Releasing open-weight models that can significantly help amateurs build bioweapons would cause roughly 100,000 expected fatalities per year, but the reduction such models provide in larger later risks — chiefly AI takeover — probably outweighs that cost; so those focused on existential risk should neither advocate for nor advocate against such releases, and should instead press companies to honour their existing commitments and apply mitigations.
Context and trigger
- Anthropic released Claude Opus 4 and said it could not rule out that the model triggered ASL-3 safeguards on CBRN grounds — i.e. “the ability to significantly help individuals or groups with basic technical backgrounds (e.g., undergraduate STEM degrees) create/obtain and deploy CBRN weapons” (Anthropic’s RSP), specifically bioweapons.
- Combined with results on the Virology Capabilities Test, the author judges it likely that several labs have or will soon have models that can significantly help amateurs make bioweapons.
- The post is written to inform readers focused on loss-of-control risk, to reduce poorly targeted advocacy, and to put the author’s views on public record.
Cost side: expected fatalities
- Estimate: open-weight models at the “significantly helping amateurs” threshold would kill around 100,000 people per year in expectation, relative to a counterfactual of no such open-weight release plus high-quality safeguards on closed models.
- Derivation: COVID killed around 30 million people; such models are judged to raise the annual probability of a comparably lethal pandemic by “somewhat more than 0.1%”.
- The estimate is highly uncertain and hinges on the number of at-least-slightly-competent bioterrorists.
- Salience effects could dominate: early publicised incidents might spread a “bioterrorism meme” the way mass shootings propagate in the US.
- Harm is heavy-tailed on two counts — pandemic fatalities are heavy-tailed, and bioterrorism may beget more bioterrorism.
- Monetised: annual expected damages of roughly $100 billion to $1 trillion (using $1–10 million per life, accounting for broader economic damage). Non-risk-reduction benefits would need to boost GDP by ~0.1–1% to match this.
- Risk grows with capability (aiding novel bioweapon R&D) and over time (as more dangerous viruses become designable or publicly derivable).
Benefit side: reduction in larger risks
- The author’s dominant concern is highly catastrophic or existential risk from systems able to fully automate human cognitive work — especially AI takeover, put at roughly a 30% chance, with ~20% expected fatalities from takeover and other AI sources.
- Open weights reduce these risks by:
- accelerating external alignment/safety research via arbitrary fine-tuning, helpful-only model access, and weights/activations access for model-internals work;
- increasing societal awareness of AI capabilities and risks, which also helps against other large risks such as AI-enabled coups.
- Countervailing effects: open weights may help non-lab researchers advance general capabilities (probably net bad); more speculatively, they might slow capabilities by reducing frontier-lab revenue and hence investment.
- The benefit shrinks if AI companies provide better model access to safety researchers, or if one believes future risks will be handled competently.
Implications for advocacy
- The author will not advocate for open-weight release at this capability level: it pays “a large tax in blood” for an uncertain future risk reduction, and other asks are more leveraged.
- Releasing a model with a reasonable chance of exceeding the threshold is characterised as a unilateral and aggressive action by a company, absent an impartial, unconflicted, legitimate third-party decision (e.g. a representative citizens’ assembly).
- It would also be a mistake for loss-of-control-focused people to argue such releases are net-harmful — the release is probably net-positive on the author’s view, and picking a fight with open-weight advocates is politically costly.
- Things the author thinks people should not say or advocate for:
- that safety policies should include a no-open-weights rule at the “substantially assist amateurs” CBRN threshold;
- that a given company’s open-weight release was very net-harmful on CBRN grounds;
- that a company’s CBRN open-weight threshold should trigger earlier.
- Things the author thinks it is good to say, where true:
- companies uphold their safety policies poorly, and released a model exceeding their own stated threshold;
- CBRN evals are shoddy and thresholds poorly justified;
- a company greatly weakened its safety framework without serious justification;
- companies should deploy API safeguards, filter datasets, and advocate for better DNA synthesis screening.
- Pressuring companies toward honesty and faithful adherence to their own policies is endorsed, as is documenting non-compliance to inform the case for regulation.
- Weakening a CBRN commitment is acceptable if accompanied by a clear public case that discusses the real downsides — though the author thinks that is unlikely to happen.
When the author’s view would change
- Open weights become probably bad once models can significantly accelerate AI R&D (roughly 1.5x faster algorithmic progress) or are capable of Autonomous Replication and Adaptation (ARA).
- Bioweapon risk becomes much higher once AIs can significantly accelerate bioweapons R&D (e.g. Anthropic’s CBRN-4); open weights are then more straightforwardly bad, though with less certainty than the AI R&D case.
- The author still doubts advocacy is the right lever here: large companies and governments are unlikely to want open releases at that capability level, so governance work mostly need not focus on preventing intentional releases above the bar.
- The precedent-setting argument — oppose early releases to establish a norm against later ones — is rejected as a costly fight over an ask one doesn’t believe in.
- Views could shift the other way if misalignment risk looked basically resolved.
- Open release is also bad when it leaks substantial state-of-the-art algorithmic advances not already known to relevant actors; currently most labs are leaky enough that this is not decisive.
Mitigations judged worthwhile
- For open-weight models: filtering problematic virology data out of training and/or robustly unlearning the relevant capabilities. This does not eliminate risk (fine-tuning can restore it) but might cut it by 2–3x. Downstream developers should not fine-tune those capabilities back in.
- Strong safeguards on non-open-weight models (e.g. constitutional classifiers), though the author is unsure current deployments suffice.
- Much better DNA synthesis screening — improved automatic screening software plus serious Know Your Customer policies — which could eliminate the majority of risk at this capability level, but not once AIs can substantially aid frontier bioweapons R&D.
- General biosecurity improvements.
- Companies releasing such models should at minimum filter virology data and publicly advocate for improved DNA synthesis screening; high-quality evaluations are a precondition.
- Failing to evaluate and seriously mitigate would be unacceptable regardless of the net-benefit judgement; mitigations and evals should be scrutinised by unconflicted third-party experts able to comment publicly.
Comment thread points
- Tao Lin: the post overestimates the impact of large open models on external safety research — the safety community has barely used DeepSeek R1/V3 weights, working instead with R1-distill-8B and QwQ-32B, so what matters is when small models cross the bioterrorism threshold; filtering biology data from small models is also cheaper because it does not cost customers.
- Greenblatt’s replies: benefits should grow as open-source tooling improves; people do use Llama-70B and will use bigger models; substantial chance further open models don’t matter, some chance they matter a lot.
- ChristianKl: Bruce Schneier’s movie-plot-threat contest surfaced many low-resource high-damage ideas that terrorists never actually pursue.
- Greenblatt’s reply: agreed the risk is reduced by the small number of scope-sensitive bioterrorists, but that suffices for the low probabilities in question, especially with meme-spread effects.
- anaguma: questions the 0.1% figure as too low.
- Against Moloch: pandemics may make society dumber and more reactive at a critical time; open weights bring near-SOTA capability to actors such as North Korea.
- Thane Ruthenis: the central extinction route is a dangerous model sandbagging internal evals, getting open-sourced, then going the rogue-replication route.
- Chris_Leong: value in some people saying in advance that it is a terrible idea, to have credibility later.