Abstract

Releasing open-weight models that can significantly help amateurs build bioweapons would cause roughly 100,000 expected fatalities per year, but the reduction such models provide in larger later risks — chiefly AI takeover — probably outweighs that cost; so those focused on existential risk should neither advocate for nor advocate against such releases, and should instead press companies to honour their existing commitments and apply mitigations.

Context and trigger

  • Anthropic released Claude Opus 4 and said it could not rule out that the model triggered ASL-3 safeguards on CBRN grounds — i.e. “the ability to significantly help individuals or groups with basic technical backgrounds (e.g., undergraduate STEM degrees) create/obtain and deploy CBRN weapons” (Anthropic’s RSP), specifically bioweapons.
  • Combined with results on the Virology Capabilities Test, the author judges it likely that several labs have or will soon have models that can significantly help amateurs make bioweapons.
  • The post is written to inform readers focused on loss-of-control risk, to reduce poorly targeted advocacy, and to put the author’s views on public record.

Cost side: expected fatalities

  • Estimate: open-weight models at the “significantly helping amateurs” threshold would kill around 100,000 people per year in expectation, relative to a counterfactual of no such open-weight release plus high-quality safeguards on closed models.
  • Derivation: COVID killed around 30 million people; such models are judged to raise the annual probability of a comparably lethal pandemic by “somewhat more than 0.1%”.
  • The estimate is highly uncertain and hinges on the number of at-least-slightly-competent bioterrorists.
  • Salience effects could dominate: early publicised incidents might spread a “bioterrorism meme” the way mass shootings propagate in the US.
  • Harm is heavy-tailed on two counts — pandemic fatalities are heavy-tailed, and bioterrorism may beget more bioterrorism.
  • Monetised: annual expected damages of roughly $100 billion to $1 trillion (using $1–10 million per life, accounting for broader economic damage). Non-risk-reduction benefits would need to boost GDP by ~0.1–1% to match this.
  • Risk grows with capability (aiding novel bioweapon R&D) and over time (as more dangerous viruses become designable or publicly derivable).

Benefit side: reduction in larger risks

  • The author’s dominant concern is highly catastrophic or existential risk from systems able to fully automate human cognitive work — especially AI takeover, put at roughly a 30% chance, with ~20% expected fatalities from takeover and other AI sources.
  • Open weights reduce these risks by:
    • accelerating external alignment/safety research via arbitrary fine-tuning, helpful-only model access, and weights/activations access for model-internals work;
    • increasing societal awareness of AI capabilities and risks, which also helps against other large risks such as AI-enabled coups.
  • Countervailing effects: open weights may help non-lab researchers advance general capabilities (probably net bad); more speculatively, they might slow capabilities by reducing frontier-lab revenue and hence investment.
  • The benefit shrinks if AI companies provide better model access to safety researchers, or if one believes future risks will be handled competently.

Implications for advocacy

  • The author will not advocate for open-weight release at this capability level: it pays “a large tax in blood” for an uncertain future risk reduction, and other asks are more leveraged.
  • Releasing a model with a reasonable chance of exceeding the threshold is characterised as a unilateral and aggressive action by a company, absent an impartial, unconflicted, legitimate third-party decision (e.g. a representative citizens’ assembly).
  • It would also be a mistake for loss-of-control-focused people to argue such releases are net-harmful — the release is probably net-positive on the author’s view, and picking a fight with open-weight advocates is politically costly.
  • Things the author thinks people should not say or advocate for:
    • that safety policies should include a no-open-weights rule at the “substantially assist amateurs” CBRN threshold;
    • that a given company’s open-weight release was very net-harmful on CBRN grounds;
    • that a company’s CBRN open-weight threshold should trigger earlier.
  • Things the author thinks it is good to say, where true:
    • companies uphold their safety policies poorly, and released a model exceeding their own stated threshold;
    • CBRN evals are shoddy and thresholds poorly justified;
    • a company greatly weakened its safety framework without serious justification;
    • companies should deploy API safeguards, filter datasets, and advocate for better DNA synthesis screening.
  • Pressuring companies toward honesty and faithful adherence to their own policies is endorsed, as is documenting non-compliance to inform the case for regulation.
  • Weakening a CBRN commitment is acceptable if accompanied by a clear public case that discusses the real downsides — though the author thinks that is unlikely to happen.

When the author’s view would change

  • Open weights become probably bad once models can significantly accelerate AI R&D (roughly 1.5x faster algorithmic progress) or are capable of Autonomous Replication and Adaptation (ARA).
  • Bioweapon risk becomes much higher once AIs can significantly accelerate bioweapons R&D (e.g. Anthropic’s CBRN-4); open weights are then more straightforwardly bad, though with less certainty than the AI R&D case.
  • The author still doubts advocacy is the right lever here: large companies and governments are unlikely to want open releases at that capability level, so governance work mostly need not focus on preventing intentional releases above the bar.
  • The precedent-setting argument — oppose early releases to establish a norm against later ones — is rejected as a costly fight over an ask one doesn’t believe in.
  • Views could shift the other way if misalignment risk looked basically resolved.
  • Open release is also bad when it leaks substantial state-of-the-art algorithmic advances not already known to relevant actors; currently most labs are leaky enough that this is not decisive.

Mitigations judged worthwhile

  • For open-weight models: filtering problematic virology data out of training and/or robustly unlearning the relevant capabilities. This does not eliminate risk (fine-tuning can restore it) but might cut it by 2–3x. Downstream developers should not fine-tune those capabilities back in.
  • Strong safeguards on non-open-weight models (e.g. constitutional classifiers), though the author is unsure current deployments suffice.
  • Much better DNA synthesis screening — improved automatic screening software plus serious Know Your Customer policies — which could eliminate the majority of risk at this capability level, but not once AIs can substantially aid frontier bioweapons R&D.
  • General biosecurity improvements.
  • Companies releasing such models should at minimum filter virology data and publicly advocate for improved DNA synthesis screening; high-quality evaluations are a precondition.
  • Failing to evaluate and seriously mitigate would be unacceptable regardless of the net-benefit judgement; mitigations and evals should be scrutinised by unconflicted third-party experts able to comment publicly.

Comment thread points

  • Tao Lin: the post overestimates the impact of large open models on external safety research — the safety community has barely used DeepSeek R1/V3 weights, working instead with R1-distill-8B and QwQ-32B, so what matters is when small models cross the bioterrorism threshold; filtering biology data from small models is also cheaper because it does not cost customers.
  • Greenblatt’s replies: benefits should grow as open-source tooling improves; people do use Llama-70B and will use bigger models; substantial chance further open models don’t matter, some chance they matter a lot.
  • ChristianKl: Bruce Schneier’s movie-plot-threat contest surfaced many low-resource high-damage ideas that terrorists never actually pursue.
  • Greenblatt’s reply: agreed the risk is reduced by the small number of scope-sensitive bioterrorists, but that suffices for the low probabilities in question, especially with meme-spread effects.
  • anaguma: questions the 0.1% figure as too low.
  • Against Moloch: pandemics may make society dumber and more reactive at a critical time; open weights bring near-SOTA capability to actors such as North Korea.
  • Thane Ruthenis: the central extinction route is a dangerous model sandbagging internal evals, getting open-sourced, then going the rogue-replication route.
  • Chris_Leong: value in some people saying in advance that it is a terrible idea, to have credibility later.