Total research transparency would be nice

Core claim

Total research transparency — the mechanism at the heart of the AI Futures Project’s Plan A — would radically simplify both deciding what the rules for AI alignment should be and enforcing them, removing the need for third-party evaluators to justify in advance what information matters. It is a very tall order, and the realistic fallback is a narrower audit-based regime, but anything less makes it harder to know whether these systems can be trusted.

Standfirst: “We could learn what’s scary and stop doing that.”

The context: Plan A

  • The AI Futures Project (authors of the AI 2027 scenario) released Plan A, a positive recommendation for how the US should handle the arrival of superintelligence.
  • At root it is an arms control deal: the US and China jointly regulate frontier AI development, so each can take the time needed to gain confidence in system safety without fearing the other will cut corners and race ahead.
  • An old, simple idea — but AIFP works through implementation detail and narrative plausibility: how public demand might build after further “Mythos moments” and job anxiety, how a crude stopgap deal could be assembled in a hurry, and how to bootstrap from there to a mature regime.
  • Cotra calls it the most comprehensive vision available for things going well even under a fast takeoff toward god-like superintelligence and hard alignment, and recommends reading it.

What “total research transparency” means

  • Plan A’s core mechanism, in AIFP’s words: “We’ll agree to let each other see all the AI research. Then, if we don’t like something someone is doing, we’ll talk about it and perhaps agree to ban it.”
  • The radical version imagined: anything any AI researcher can see (training code, scaling curves, architecture designs) everyone can see; any experiment any researcher can run (including fine-tuning) anyone can run.
  • Carve-outs: model weights are secured — including from AI researchers themselves, so no one directly accesses them. Customer data and a small amount of other sensitive information is likewise protected from both researchers and the public.

Why it would help: the case from experience

Cotra writes as someone who just ran a multi-month, multi-stakeholder process (METR’s Frontier Risk Report, May 2026) to inform the public about loss-of-control risk at AI companies.

The current regime forces question-first investigation

  • The science of AI capabilities is nascent and fast-moving; “nascent” is generous for alignment science.
  • Present norms require third-party evaluators like METR to justify in advance why each scrap of information they want to publish is tightly connected to misalignment risk.
  • Consequence: METR had to know the shape of its risk argument before designing the questionnaire. Every question needed an at-a-glance logical link to misalignment risk.
  • They could probe known concerns — e.g. whether models used architectures allowing reasoning in “neuralese” rather than English (all four participants — OpenAI, Anthropic, Google DeepMind, Meta — said no) — but there was no room to discover unknown unknowns.
  • Open-ended prompts that could seed deeper investigation, like “describe your training process in detail,” were off the table.
  • A friend’s analogy: it’s like doing an operational security audit of a building where, instead of asking for the blueprints, you ask “Are you aware of any doors that might be unlocked?”
  • Answers went up the chain through company points of contact, then lawyers and comms, returning as a “Very Official Response.”
  • Everything shared was subject to redaction at company discretion, so for months only a tiny silo inside METR could see and discuss the underlying information.
  • This is a poor setup for sensemaking on a fraught subject nobody fully understands — once redactions cleared and the team could talk to a wider set of peers, a host of new considerations and framings immediately appeared.
  • Illustrative extreme: METR’s first embedded red-teaming exercise involved a silo of one employee, who kept a poker face through three weeks of one-on-ones with his manager to avoid leaking bits about how it was going.

What transparency would unlock

  1. Discovery without pre-justification — outside scientists wouldn’t need to specify in advance what will turn out to matter, and could think out loud with anyone from day one, dramatically accelerating the scientific conversation and the odds of converging on which practices are good or bad.
  2. Near-trivial verification — once a practice is agreed necessary or dangerous, compliance can simply be observed. Far less need to specify every edge case of every rule.
  3. Real-time peer enforcement — violations of the spirit of a vague principle (e.g. the stated intention not to put “too much training pressure” on chains of thought — which several companies have reported accidentally doing anyway) could be called out by a large body of technically informed outsiders, notably researchers at competing companies, who have both the most relevant expertise and the best incentives.
  4. Public-interest legitimacy — these systems will be societal infrastructure making millions of consequential decisions with growing discretion. Knowing how they were trained makes it much harder to insert secret loyalties into AI that governments and militaries depend on.

The obstacles Cotra concedes

  • It’s a very tall order: it means strong-arming the world’s most powerful companies into effectively wiping out ludicrously valuable IP.
  • It also means the US voluntarily surrendering a large part of its geopolitical lead over China. (Compute is another chunk of that lead — and Plan A regulates competition on that axis too, likely shrinking the US advantage further.)
  • Realistic fallback: a more limited and less flexible transparency regime built on third-party auditing. That’s METR’s job — prototyping the regime that could power “Plan A- or B+.”
  • Cotra is optimistic it can get far, and is especially keen on more embedded assessments.
  • The closing caveat: anything less than total research transparency “would make it harder to know we can trust the AI systems that could soon have tremendous power over us.”

Comment thread

  • Kamila Selig: precedents exist for auditable but not public data — FedRAMP, financial audits. Both are imperfect and purpose-built, but both are far closer to truth than a Gartner-magic-quadrant-style quasi-evaluative/quasi-marketing questionnaire, which is roughly where independent evals sit today.
  • Charlie Sanders: this would be a Fifth Amendment takings of proprietary American IP worth many trillions, for the express purpose of handing it to a declared foreign adversary — “that seems…impractical?”

Notes

  • Footnote in the original: a takeoff steep enough to justify Plan A’s urgency also concentrates power by default, since a small initial lead compounds rapidly into an overwhelming advantage.
  • Cotra works at METR, which produced the Frontier Risk Report (19 May 2026) referenced throughout.