SPAR Projects - Readable Catalogue

20 projects from spar-projects-all.csv. This version is designed for browsing and comparison; source wording is preserved, and empty fields are omitted.

Quick comparison

#PriorityProjectResearch areasHours/weekTeam
13Accelerating Democratic AI Leadership by Reforming the US Extraterritorial Surveillance Regime by Michelle NieInternational governance; US policy; AI strategy102-3
25Generalist Megastream by Generalist Mentor PoolAI strategy; Communications--
33Mitigating Intentional Loss of Control Risk Through Interoperability Standards for Agentic AI by Kevin KohlerAI control; International governance; Technical governance102
44Epistemic security in the age of AI by Natalie Linton, Sarah LucioniSocietal impacts; Misuse risk; Biosecurity51-2
53Topics in AI strategy and futurism by Dylan BowmanAI strategy; Evaluations102-6
62Token taxes as a mechanism for reducing AI-driven power concentration by Lucas IrwinAI strategy; Economics of AI; Compute governance52-3
71Catastophic Risks of AI in Space by Stefano VerganiAI strategy; Cyber risks; Lab governance123-4
83Simulating AI Policies: An Agentic Testbed for Governance Interventions by David Williams-King, Linh LeAI strategy; International governance; Multi-agent systems152-3
94Wikipedia contributions on AI safety and policy by Michael ChenCommunications; AI strategy84-7
104Does the Thermometer Change the Reading? Testing Whether AI Welfare Self-Reports Survive a Change of Frame by Varad VishwarupeAI welfare; Behavioral evaluation of LLMs; Evaluations82-3
113Designing the boundary between Helpful Persuasion and Harmful Manipulation by Markov Grey, Charbel-Raphael SegerieBehavioral evaluation of LLMs; Misuse risk; Societal impacts103-4
124Will Model Licensing Increase Concentration of Power? by Alex MarkNational policy; AI strategy; Societal impacts102-10
132Writing a textbook for AI Governance by Markov Grey, Charbel-Raphael SegerieAI strategy; Technical governance; Communications102-4
142[[#Project 14|[Jurisdiction-specific] Public-Sector AI Resilience Agenda]] by Chris SchmitzNational policy; AI strategy; Societal impacts53-4
151Reading the Machine’s Mind: Mechanistic Interpretability, Artificial Intent, and AI Legal Responsibility by Elija PerrierTechnical governance; National policy; Philosophy of AI102-1
163AI Consciousness: Research and Public Writing by Maria AvramidouAI welfare; Philosophy of AI102-3
173Strategic stability when states delegate to AI: escalation, commitments, and arms control by Amritanshu PrasadAI strategy; International governance82-3
183How Quickly can Middle Powers Build Frontier Compute Capacity? by James Nicholas BryantCompute governance; AI strategy; International governance82-3
193Normalization of Deviance in AI Development by Emilio BarkettAI strategy; Lab governance52-3
201A Design Blueprint for Middle-Power AI Safety Institutes by Michał Kubiak, Daniel PolakInternational governance; Technical governance; National policy83-4

Browse by research area

AI control (1)

AI strategy (13)

AI welfare (2)

Behavioral evaluation of LLMs (2)

Biosecurity (1)

Communications (3)

Compute governance (2)

Cyber risks (1)

Economics of AI (1)

Evaluations (2)

International governance (6)

Lab governance (2)

Misuse risk (2)

Multi-agent systems (1)

National policy (4)

Philosophy of AI (2)

Societal impacts (4)

Technical governance (4)

US policy (1)

Project profiles

Project 01

Accelerating Democratic AI Leadership by Reforming the US Extraterritorial Surveillance Regime

Open project page | Mentor link

Mentor(s): Michelle Nie

Affiliation(s): Center for a New American Security (CNAS)

Research areas: International governance | US policy | AI strategy

Practical detailValue
Hours/week10
Team size2-3
Mentor involvementDetailed Guidance (2+ hr/week/mentee)
Location/timezoneNone; but mentees should have decent overlap with US East Coast time zone and availability for calls during this time.
Off-cycleNo

Overview

Keeping the most advanced AI within the control of democracies depends on integrating democratic allied nations into the US AI stack, but US extraterritorial surveillance and data access laws are eroding those allies’ trust in its tech stack. This project aims to propose a reformed regime that reconciles legitimate US security interests with the trust and assurance allies need to build on American infrastructure.

Full description

Whether frontier AI is developed and governed by democracies or by autocracies is among the most consequential questions for how the technology turns out. A world where advanced AI capability and the compute behind it sit largely within democratic nations—and where a coalition of responsible actors meaningfully participate in global AI governance conversations—is far likelier to produce AI that is developed safely and accountably. For the United States, this aligns with multiple objectives: to build compute at scale as quickly as possible (including by partnering with allies); and to diffuse American AI throughout the world, keeping the global AI ecosystem anchored to American infrastructure rather than China’s. But regulatory obstacles have worn down allies’ trust. Allied democracies are increasingly reluctant to route their data and compute through US-controlled providers because of the surveillance exposure created by US law—chiefly FISA Section 702 and the CLOUD Act. The result is a growing movement around “sovereign” compute and data localization among the very partners the US should most want to export its AI stack to. To allies, the risk of foreign surveillance potentially outweighs any benefits from chip or model access. This pushes partners toward fragmentation and toward hedging with non-US suppliers, weakening exactly the coalition that would keep advanced AI within democracies. This produces a tension the US treats as a hard tradeoff. Its national security and law-enforcement communities have legitimate interests in cross-border data access and foreign-intelligence collection. Yet the regulatory framework is now a major obstacle to the very diffusion strategy that has become a flagship US initiative. A need for a new regime has emerged, one that balances legitimate security needs with building trust and assurance within allies to get them to trust the US AI stack. Drawing on the existing legal and policy literature and on interviews with relevant stakeholders, the project will map allied concerns about surveillance and the hindrance of the diffusion strategy. It will design and propose a new regime that preserves core US security and law-enforcement equities while keeping the US AI stack attractive to key allies and partners. The work will aim to result in concrete recommendations for US policymakers and for allied “AI middle powers” weighing how to integrate and ensure relevance in the advantage of advanced AI.

Why this matters

This project advances AI safety by strengthening the conditions under which transformative AI is most likely to be developed and governed responsibly. By identifying a regime that removes the trust bottleneck to allied integration, this work increases the likelihood that advanced AI is developed and governed within democracies rather than autocracies, and that a wider set of responsible actors is brought into global AI governance.

What you’ll do

Mentees will support by performing background research (literature reviews, expert/stakeholder interviews), assisting with brainstorming policy proposals, as well as potentially drafting memos or public-facing pieces if the research results in these.

Prerequisites

Need-to-have: - Strong analytical, research, and writing skills - Comfort doing independent desk research - Familiarity with AI safety/governance literature - Knowledge of/interest in geopolitics of AI, middle power strategy Nice-to-have: - Experience with legal analysis (e.g. a law degree or some amount of legal training) - Familiarity with middle power government(s) and political dynamics - Experience conducting stakeholder interviews or expert outreach - Experience presenting to/educating policymakers - Prior relevant publications

Application question(s)

  • Briefly outline the case (from a US national security standpoint) for and against US authorities having access to US cloud providers’ customer data (200 words max). - Please provide a link to one or more relevant writing samples, ideally from a research context, but can be a personal Substack/memo you wrote.

Back to quick comparison


Project 02

Generalist Megastream

Open project page

Mentor(s): Generalist Mentor Pool

Affiliation(s): Kairos, Constellation, Generator Residency, etc.

Research areas: AI strategy | Communications

Practical detailValue
Off-cycleNo

Overview

The Generalist Megastream pairs mentees on small generalist projects with a mentor from a pool of generalist mentors from Kairos, Constellation, the Generator Residency, and more. Projects are talent/infrastructure research, field-building, or answering open operational questions in AI safety. The stream is a step before programs like the Generator Residency: it gives people context on the field, experience with generalist work, and preparation for future opportunities, while legitimizing generalist paths into AI safety.

Full description

The Generalist Megastream is a single project listing backed by a pool of generalist mentors rather than one named mentor. Mentees apply to it like any other SPAR project, indicate which project ideas interest them in the project-specific question, and are then matched with a mentor and project from the pool. Possible project ideas include: Research & analysis - Analysis of hiring, talent, and fellowship outcomes in AI safety - Talent pipeline research: where people entering AI safety get stuck, what roles the field is bottlenecked on - Mid-career pipelines scoping study: interview hiring orgs, pipeline programs, and career transitioners, then write a landscape report - “Make AI safety managers better”: interview AIS managers, survey existing resources, synthesize a curated guide Field mapping & monitoring - Visual field maps of the AI safety ecosystem - AI Lab Watch successor v1: read labs’ safety frameworks and model specs, produce a rigorous scorecard and methodology Learning materials & curricula - Curate better learning materials and curricula for people entering AI safety - Draft university group fellowship curricula for different audiences - University group organizer playbook: map what top groups do differently, turn it into resources and tooling Knowledge bases & shared infrastructure - Plug-and-play retreat knowledge base: project plans per retreat type, vendor/venue/caterer lists, common mistakes - Shared AI-uplift infrastructure across AI safety orgs: prompt libraries, MCP connectors, eval harnesses Comms & outreach - Translate research into clear public content: partner with 3-5 researchers to turn their work into accessible writing - Outreach experiments: sourcing and contacting potential mentors or applicants at scale Exploratory / AI-powered - Proof-of-concept automated macrostrategy: identify trends in AI forecasting ability and how they scale to scenario modeling Further reading and more sources of project ideas - AI Safety’s Biggest Talent Gap Isn’t Researchers. It’s Generalists.: why the field needs more generalists, and the motivation behind this stream - A reading list for generalists (LessWrong) - The case for AI safety capacity-building work by Asya Bergal - AI safety fieldbuilding career review (80,000 Hours)

Why this matters

The projects strengthen the AI safety field’s capacity and infrastructure directly. Talent pipeline research and analyses of hiring and fellowship outcomes show where the field loses people and which roles are bottlenecked, so programs and funders can respond. Field maps and lab safety scorecards improve coordination, visibility, and accountability. Better curricula and organizer playbooks raise the quality of people entering the field. Shared knowledge bases and AI-uplift infrastructure eliminate duplicated effort across AI safety orgs, and comms projects make safety research legible to policymakers and the public. This kind of generalist and infrastructure work is increasingly the field’s binding constraint (see AI Safety’s Biggest Talent Gap Isn’t Researchers. It’s Generalists.), so each completed project relieves a real gap.

What you’ll do

Mentees drive a small, scoped generalist project (research, interviews, analysis, or resource-building) with guidance from a mentor in the generalist pool. Project scopes are designed to fit SPAR’s part-time, remote form factor.

Application question(s)

What generalist project idea are you most excited to work on, and why? You are welcome to pull from the ideas provided in the description or to propose your own generalist project idea.

Back to quick comparison


Project 03

Mitigating Intentional Loss of Control Risk Through Interoperability Standards for Agentic AI

Open project page | Proposal attachment | Mentor link

Mentor(s): Kevin Kohler

Affiliation(s): Simon Institute, UN University, AGI Preparedness Institute

Research areas: AI control | International governance | Technical governance

Practical detailValue
Hours/week10
Team size2
Mentor involvementGuidance (1+ hr/week/mentee)
Off-cycleNo

Overview

Within a few years, when self-replication is plausibly within reach of open-weight systems, the binding constraint on an agent deployed to operate without a controlling human principal will likely not be the model, but whether the rest of the agent economy will discover, authorize, transact with, or pay it. This project maps who actually holds change control over the agent identity and trust layer, evaluates the competing architectures against an explicit intentional-loss-of-control threat model, and feeds the result into live standards and Geneva policy discussions before network effects settle the question.

Full description

The problem. Standards are where technical detail quietly becomes de facto international law. The Internet is the cautionary case: IP won the protocol wars on network effects, and its identity and trust layers were bolted on afterwards. Thirty years later BGP still runs on trust, the web PKI still depends on every certificate authority in the store behaving, and the neutrality of the DNS root still turns on who holds jurisdiction over ICANN. None of it has been fixable, because change control was settled before anyone understood what was at stake. Agent-to-agent interoperability is at that stage now, and moving faster. MCP moved under a Linux Foundation body in late 2025; its authorization model is conventional OAuth and the specification has no concept of agent identity. A2A reached its first stable version under the Linux Foundation, where an agent identifies itself through a card it publishes about itself and signing that card is optional. Agent payments are moving fastest of all, with Google’s Agent Payments Protocol donated to the FIDO Alliance in 2026 — and it is the partial counterexample worth studying, because it anchors trust in cryptographically signed user mandates rather than in agent self-declaration. At the IETF, agent identity work remains a scatter of individual drafts, none adopted by a working group; a proposal to open a working group on agent delegation and traceability was declined in July 2026. ITU-T established a focus group on trust and identity for agentic AI only in July 2026, with first work items targeted for 2027 and beyond. ISO/IEC has no agent-specific standard in development.

The safety argument. Intentional loss of control is the case where a human deliberately deploys an agent to operate without an accountable controlling principal — designed to resist shutdown, to re-establish itself elsewhere, or to replicate. As capability diffuses through open weights, mitigation at the model layer weakens: you cannot rely on a developer’s safeguards for a system whose weights are public and whose fine-tuning is cheap. What remains are the downstream chokepoints, and in an economy where agents transact with agents, the chokepoint is admission to the network. The architecture decides whether that chokepoint exists at all. Under self-declared identity, an adversarial deployer forges or clones a credential, re-registers after any block, and pays no cost; revocation has nothing to bite on. Under issued, revocable identity with an auditable delegation chain back to an accountable legal person, revocation becomes a real intervention at the transaction layer, unattributed agents can be refused by counterparties and payment rails, liability attaches to someone, and incidents can be reconstructed. That is the claim the project tests: organized, issued agent identity is a control lever; self-declared identity is not. It is not a free lever, and the project has to be honest about that or it will not survive contact with the standards community. The same identifier that lets you revoke a replicating agent lets a state bind every agent action to a real-world human identity. A registry is a governance object with an owner and a jurisdiction, which is the ICANN question again. The research question is therefore not whether agent identity is good, but: which architecture buys the most loss-of-control mitigation per unit of surveillance capability, centralization, and foreclosure risk — and who has to hold change control for that to remain true?

Research questions. (1) Who holds decision rights over the agent identity and trust layer today, by body and by decision date, and where does that move over the next 18 months? (2) What does an intentional-loss-of-control threat model demand of an identity architecture, derived from concrete scenarios rather than enterprise IAM best practice? (3) How do the live design families score against those requirements and against privacy, centralization, jurisdictional exposure, and adoption feasibility? (4) Which properties must live in the protocol and which can be left to policy or procurement — following Collingridge, what is cheap to change now and prohibitively expensive later? (5) What is the minimum viable ask a safety-motivated actor could bring into a named process in the next twelve months, and to whom?

Work packages. WP1 (weeks 1–3): change-control map across the protocol foundations, internet standards bodies, payments alliances, identity community, and intergovernmental track — owner, stage, change-control mechanism, funding, next decision date. WP2 (weeks 3–6): three to four intentional-loss-of-control scenarios walked through the transaction lifecycle from discovery to revocation, yielding a requirements set. WP3 (weeks 5–9): comparative scoring of the design families against those requirements and the counterpart criteria. WP4 (weeks 9–12): intervention analysis — which ask, which body, which deadline — plus one brief for standards participants and one for diplomats.

Deliverables. A change-control map with a decision calendar; a requirements set; a scored comparative assessment; a policy brief of roughly 3’000 words plus a shorter technical-community version. Stretch goal: a submitted comment to a named open process.

Method. Document analysis of primary standards artifacts, all public. Eight to fifteen structured expert interviews, using my network for access. Scenario-based reasoning rather than quantitative modeling. Comparative scoring against published criteria with explicit uncertainty, so a reader who disagrees with a weight can recompute rather than argue. No compute required.

Why this matters

While the biggest loss of control risk currently is internal deployment in frontier AI labs, I expect AIs with high autonomy risk to diffuse. We cannot rely on a developer’s controls for a system whose weights are public and whose fine-tuning is cheap. The mitigations that survive if capability diffuses are downstream of the model, where an agent has to interact with the world. In an economy where agents transact with agents, the decisive one is admission to the network — whether counterparties, service providers, and payment rails will deal with an agent that cannot produce a valid, revocable identity traceable to an accountable principal. Whether that chokepoint exists is being decided now, in protocol specifications, by people who are solving an enterprise access-management problem rather than a catastrophic-risk problem. Self-declared identity gives an adversarial deployer a free re-registration after every block. Issued and revocable identity makes revocation a real intervention, lets platforms refuse unattributed agents, and attaches liability to a person. This is a rare case where a safety property can be built into infrastructure rather than negotiated between states afterwards — and a rare case where the window is identifiable and short, because network effects will soon settle these questions permanently. The output is therefore not a preprint but a mapped set of asks: which change to which document in which body, by when. I have done the analogous analysis before: at ETH Zurich’s Center for Security Studies I published a study of the politics of Internet protocol design covering New IP, SCION, ICANN’s jurisdiction problem, and the IETF–ITU conflict. I am now based in Geneva launching a new AI governance institute, and recently convened a private roundtable on AI standards with major AI labs, global standards bodies, a leading payment network, and three national AI safety institutes, which generated follow-up interest this project is designed to feed.

What you’ll do

High autonomy. Each mentee owns work packages end to end and is the author of what they produce; I set direction, review drafts, open doors, and argue with conclusions. This is not a project where I hand over a task list. With two mentees the split is: one owns the change-control map and the comparative assessment (WP1, WP3), the other owns the threat model and the intervention analysis (WP2, WP4), with a shared weekly synthesis where each has to defend their work to the other. Both tracks feed a joint output, so mentees will interact substantially rather than working in parallel silos. Concretely, a mentee will: read primary standards artifacts closely and summarize what they actually commit implementers to; build and maintain a structured landscape database; write and run expert interviews, including drafting the outreach; construct scenarios and stress-test candidate architectures against them; and draft sections of the final brief under their own name. What I provide: weekly direction and prioritization, written feedback on drafts, introductions from my network for the interview program, and the route by which the output reaches decision-makers. What I do not provide: line-by-line technical supervision of protocol analysis. I would likely put uncertain technical readings in front of trusted experts in my network rather than adjudicating them myself. Published outputs carry mentee authorship. If the work is good, I will take it into Geneva rooms with the mentee’s name on it.

Prerequisites

Able to read a technical specification or standards draft closely and explain what it actually obliges implementers to do — without needing to implement it. If you have never opened an IETF or W3C draft, that is fine; if the idea of reading one carefully for two hours sounds unpleasant, this is the wrong project. Comfortable writing clear analytical prose for a policy audience. Most of the output is written. Able to work independently on an open-ended question and come back with a structure rather than a list of links. Willing to email strangers, request interviews, and run them. Some grounding in AI safety arguments, specifically why loss of control is a concern at all. You do not need to agree with the framing — informed skepticism is useful here — but you should be able to engage with it. Useful but not required: background in internet governance, identity and access management, security engineering, standards participation, or law. Experience with a slow institutional process of any kind, including as a participant. Explicitly not required: machine learning research experience, ability to train models, publications, or a technical degree.

Application question(s)

Which forum is likely to be most influential in setting agentic AI standards and why? (200 words) Pick any agentic AI protocol or draft specification of your choice. Read it, then highlight some of its (potential) political implications (e.g., interoperability, change control, privacy, ability to shut an agent down, and intentional loss of control). (350 words) Link to a writing sample, ideally highlighting analytical or research writing. (No word limit.)

Back to quick comparison


Project 04

Epistemic security in the age of AI

Open project page | Mentor link 1 | Mentor link 2

Mentor(s): Natalie Linton; Sarah Lucioni

Affiliation(s): Independent; GovAI, Google

Research areas: Societal impacts | Misuse risk | Biosecurity

Practical detailValue
Hours/week5
Team size1-2
Mentor involvementGuidance (1+ hr/week/mentee)
Location/timezoneBeing able to take US PT (up to 7pm) meetings
Off-cycleNo

Overview

Effective response to emergencies (such as a nascent pandemic) relies on having trustworthy knowledge infrastructure. However, AI reduces the cost of producing convincing false information at scale and as AI becomes more integrated into public and private systems there is significant risk of knowledge infrastructure becoming compromised by AI-generated materials and AI itself becoming an artificial hivemind.

Full description

I am interested in two dimensions of this problem. First, how AI capabilities change the threat landscape. Earlier generations of false health information required human effort to produce and spread. However, AI enables automation, personalization, and scale that outpace current detection and response mechanisms. Understanding the specific capabilities that create new risks and the rate at which those capabilities are advancing is essential for mitigating risk. Second, the combination of disinformation campaigns with outbreaks or (hypothetical) bioweapons releases. This includes examining how malicious actors might time and target disinformation to maximize disruption, what signatures might distinguish coordinated campaigns from organic misinformation spread, and how attribution challenges complicate response decisions. For mentees that are not biosecurity focused, I also welcome them to submit their own ideas related to the general topic. Expected output will depend on mentees exact interests, but would likely be a short piece to publish in an independent publication in the science-and-tech policy space. A longer version of the research behind the short piece could be self-published or submitted to a journal if time and mentee capacity permits.

Why this matters

Proactively working to understand specific risk pathways related to AI production and incorporation of false information into knowledge infrastructure could help reduce AI safety issues such as scalable deception, model homogenization, and epistemic security.

What you’ll do

Very autonomous, 1 check-in/week

Prerequisites

I am open to a broad range of applicants, but expect this project to be well-suited to social scientists, historians, or policy researchers interested in epistemic risks from AI, particularly misinformation and disinformation risks. An interest in assessing a biosecurity (pandemic) scenario for risk would be of interest but is not required.

Application question(s)

Please explain in under 200 words why you’re interested in this project and how your background and skills will help make this project a success.

Back to quick comparison


Project 05

Topics in AI strategy and futurism

Open project page | Mentor link

Mentor(s): Dylan Bowman

Affiliation(s): Apollo Research

Research areas: AI strategy | Evaluations

Practical detailValue
Hours/week10
Team size2-6
Mentor involvementGuidance (1+ hr/week/mentee)
Off-cycleNo

Overview

Govind Pimpale and Dylan Bowman will mentor a paid, largely self-directed project in strategy and futurism for AI safety. Potential topics include forecasting, applied game theory, and AI evaluations.

Full description

We’ll spend the first segment of the project reading through a subset of the reading lists linked below (determined by both mentor and mentee preferences) at a time commitment of 10 hours/week for 4 weeks, with weekly discussion sessions and informal asynchronous chats. Then, we’ll determine whether we want the group to proceed to writing a forecast, position piece, or research report of a tightly-scoped topic explored during the first segment. Mentees who are selected for the second segment will receive a stipend for being a part-time research assistant. Reading lists: https://www.lesswrong.com/posts/L5fohLhZ7cBwDR55C/ai-futurism-reading-list https://www.forethought.org/research#featured To get an idea of what Dylan and Govind are interested in, you should read their blogs: https://www.dylanbowmansf.com/blog.html https://substack.com/@pimpale

Why this matters

AI safety benefits from having more calculating, aligned thinkers in the area of AI strategy and futurism, but oftentimes this context can be hard to distill to new community members because of trust, not wanting things to be written down, etc.

What you’ll do

Segment 1: Mentees will read and present on selected readings with a time commitment of 10 hours/week. Segment 2: Mentees will do research on AI strategy and futurism. The specific type of research here will be mentee-dependent: some examples might be creating an eval for a specific model capability or behavior, writing a position piece on a particular kind of AI policy, or forecasting a specific scenario that might play out in the next decade. The final work product will likely be a whitepaper or post of some sort.

Prerequisites

  • An understanding of how to leverage AI well for conceptual research and writing. - The ideal mentee is self-directed, compassionate, disciplined, and serious about a career in AI safety. Strong candidates might: - Bring a deep and unusual expertise that allows them to approach AI strategy from a different angle from the rest of the AI safety community. - Have worked or interned at a tech startup. - Have a background in mathematical modeling for the social sciences or adjacent fields - Have serious writing experience across any discipline (academia, journalism, etc.) - (…) Deep knowledge in AI research or AI alignment is always nice to have It’s also worth sharing that I scoped out this role for a talented undergraduate student with native/fluent English proficiency. I think this stream will be less successful/useful the further someone is from this distribution but happy to consider all applicants.

Application question(s)

Here are 5 questions. Err on the side of saying too little; we’ll do a 30 minute video interview where I’ll have a change to ask you more. 1. How did you get involve in effective altruism or AI safety? (100-400 words) 2. What are your career plans? What steps have you taken so far, which ones are you taking now, which ones are you taking soon? (100-400 words) 3. What’s a project you’ve done that you’re proud of? Explain it. Links appreciated! (100-400 words) 4. [Optional] Do you have any deep or unusual expertise? (100-400 words) 5. [Optional] Quick takes on AI 2040? (100-400 words)

Back to quick comparison


Project 06

Token taxes as a mechanism for reducing AI-driven power concentration

Open project page | Mentor link

Mentor(s): Lucas Irwin

Affiliation(s): GovAI (Summer Fellow), Oxford Martin School AI Governance Initiative

Research areas: AI strategy | Economics of AI | Compute governance

Practical detailValue
Hours/week5
Team size2-3
Mentor involvementDetailed Guidance (2+ hr/week/mentee)
Off-cycleNo

Overview

This project will focus on producing a policy memo resolving the open technical, economic, and legal questions blocking real-world implementation of token taxes. It will also run agent-based modelling simulations of the impact of token taxes and alternative policies on the UK economy to compare their benefits and drawbacks.

Full description

The rapid development of AI threatens to erode government revenues, resulting in stark inequality, and an extreme concentration of economic and geopolitical power. To prevent this, governments require effective mechanisms for capturing the economic value of AI by shifting taxation away from labour and towards capital. Token taxes – surcharges on model inference applied at the point of sale – are a promising contender for such a mechanism as they are more likely to be enforceable via compute governance infrastructure. While token taxes have been proposed, many technical, economic, and legal questions associated with implementation remain unanswered. Building on my ICML position paper (https://arxiv.org/abs/2603.04555), we will work to investigate four open research questions (RQs): RQ1: Can compute governance infrastructure be leveraged to reliably audit token taxes? RQ2: What are the legal challenges associated with implementing token taxes? RQ3: What are the advantages and disadvantages of token taxes compared to alternative taxation mechanisms such as compute taxes, VAT, and digital services taxes? RQ4: Can we model the impact of a token tax on the UK economy using LLM-powered agent-based modelling? Our output will be a policy memo with answers to RQ1-RQ4 above, co-authored with the Institute for Public Policy Research (IPPR) who have expressed interest. Based on these findings, we will re-evaluate the desirability of token taxes among the policy options for taxing AI capital. Token taxes promise to be more enforceable than alternative forms of taxation. Unlike corporation tax (which the European Commission estimates at 9.5% for digital services companies compared to 23.2% for traditional firms), token taxes take advantage of the unique properties of AI to prevent tax evasion. In particular, they can leverage existing compute governance infrastructure for auditing and enforcement. In this way, token taxes can mitigate the concentration of power by allowing governments to capture AI-generated value. The token tax paper has been mentioned in an interview by US Congressman Greg Casar, proposed as a policy in California gubernatorial candidate, Tom Steyer’s manifesto, and we have received a letter of interest from a UK MP, Anneliese Dodds, expressing interest in a token taxes memo with further letters of interest from MPs expected. The Overton window for implementation has therefore shifted rapidly, directly influencing the timing of this project.

Why this matters

Extreme concentration of power threatens to disempower citizens and engender government fiscal crises by reducing labour tax revenues. Left unchecked, this concentration will degrade the very democratic institutions that societies rely on to avoid catastrophic outcomes, including great power conflict and nuclear war. Governments’ ability to reliably tax AI capital and redistribute the windfalls will be central to mitigating this destabilising inequality. While there are excellent organisations (The Windfall Trust, Convergence Analysis) producing high-quality work on policy interventions, few are focused on translating this research into timely memos for policymakers. The token taxes project will bridge the gap between research and policy implementation by resolving open technical, economic, and legal questions standing in the way of a token tax bill. The Overton window for implementation has shifted, and the policy is rapidly gaining attention from policymakers in the US and the UK. Our existing relationships with these policymakers will allow us to stress-test whether a dedicated team of AI safety technical governance and economics researchers can strengthen democratic institutions’ capacity to reduce economic and geopolitical power concentration. If successful, we will seek to raise further funding to set up further work streams.

What you’ll do

Mentees will work on the RQs above. They will have autonomy in completing their research, and I will check in on their progress on a weekly basis. The RQs mentees work on will be determined by their background (see the prerequisites section below).

Prerequisites

For RQ1: - Degree in Engineering, Computer Science, Applied Maths. - Experience with compute governance research. - Experience with black-box auditing methods and formal verification. - Highly proficient in Python. - Spent substantial time working with transformers and tokenizers. - Strong grounding in mathematics. For RQ2: - PhD/Master’s in Law, Political Science. - Legal research experience. - Worked at an international organisation. For RQ3: - PhD/Masters in Economics. - Tax policy experience. - Very good writer. For RQ4: - PhD/Masters/Bachelors in Engineering, Computer Science. - Highly proficient in Python and ideally at least one of C, C++, Java, Typescript. - Experience/interest in agent-based modelling.

Application question(s)

Please answer the questions for the RQs you would be interested in working on: 1. In response to RQ1, list 4 papers you would use to answer this question and explain how you would use each. (300 words) 2. In response to RQ2, outline the main legal issues facing governments seeking to implement token taxes. (300 words) 3. In response to RQ3, write a short essay comparing the benefits and drawbacks of VAT, Digital Services Taxes, Corporate Taxes, Compute Taxes and Token Taxes. In particular, consider (i) how easy it is to audit the tax domestically and when auditing multinational firms (ii) how likely the tax is to distort markets and disincentivise innovation and (iii) the bureaucratic cost of maintaining the tax. (400 words) 4. In response to RQ4, draft a small agent-based model to predict the impact of token taxes on the UK economy. (300 words)

Back to quick comparison


Project 07

Catastophic Risks of AI in Space

Open project page | Proposal attachment | Mentor link

Mentor(s): Stefano Vergani

Affiliation(s): King’s College London and GovAI

Research areas: AI strategy | Cyber risks | Lab governance

Practical detailValue
Hours/week12
Team size3-4
Mentor involvementDetailed Guidance (2+ hr/week/mentee)
Location/timezoneAnywhere
Off-cycleNo

Overview

Autonomous, edge, and federated AI will play a crucial role in space in the near future. With the new 6G network under development around the Earth and the Moon, powerful AI models will be directly embedded in satellites, making autonomous decisions. With this project, we are going to map some of the most concerning issues: catastrophic risks, cyber attacks, and power concentration.

Full description

Advances in edge AI, distributed inference, and cognitive cloud computing are making it feasible to delegate the control of satellites, spacecraft, and telecommunications networks to AI systems operating in space rather than from ground control centres on Earth. Coupled with 6G non-terrestrial networks, this shift promises to remove the latency and bandwidth bottlenecks that currently constrain ground-based control, and will extend from Earth orbit to cislunar space as the Artemis programme establishes a lunar communications and navigation hub. This project wants to prove that the same architecture that delivers these efficiency gains also concentrates a novel and underexamined surface of catastrophic risk, precisely because it places increasingly autonomous, capable systems beyond the reach of timely human intervention. The project investigates four interacting risk pathways. First, and most centrally, loss of control: as AI is delegated authority over critical orbital and cislunar infrastructure, physical inaccessibility, communication delay, and operational autonomy may erode the ability of human operators to correct, override, or shut down misbehaving systems, turning ordinary alignment and reliability failures into events that are difficult or impossible to control. Second, cyberattacks: space-resident AI models become high-value targets, where adversarial manipulation, model exfiltration, or compromise could be used to commandeer satellites and disrupt global communications. Third, power concentration risks: space is fundamentally unregulated, and we currently lack a sufficiently strong governance structure to avoid power concentration of AI in space. This could happen in different ways, the most concerning one being an AI with secret loyalties capable of taking over after staying dormant for years. With the recent merger of xAI and SpaceX, these concerns appear to be motivated. Lastly, the role of a potential AGI in space and the possible catastrophic risks on Earth. A collective and edge AI framework in space could be impossible to turn off, and in case of a misaligned AGI this would cause new and unexplored risks.

Why this matters

AI in space is a relatively new topic, that has been so far explored mostly for data centers in space (https://www.forethought.org/research/will-we-really-put-data-centers-in-space), and concentration of power (https://forum.effectivealtruism.org/posts/HwBBB8sjSvPxbCjf4/3-stages-of-competition-for-the-long-term-future). This work provides a comprehensive taxonomy of the main risks, providing a useful tool for researchers and policy makers. With its focus on catastrophic risks, concentration of power, and cyber security it sits well among the most important topics in AI safety. As the AI-industry for space is booming and it is currently an underrepresented topic in the community of AI safety researchers, this project will have significant value.

What you’ll do

I will split the work into subgroups (catastrophic risks, cyber, concentration of power, and AGI risks) and each mentee would be responsible for a section. In order to have the project finalised in three months, I expect the mentees to work around 8 hours per week. I will support the effort, guide the mentees, and support your career progression during and after the fellowship. In return, I am looking for reliable and motivated mentees. The ideal mentee has already reached a good level of research independece, so that they can progress fairly independent.

Prerequisites

  • One mentee should have a scientific background (physics, engineering, etc.) - The other mentees could be generalists or come from a background in policy, economics, social sciences, law, etc. - I particularly value interest in the project, independence, and reliability. Ideally, the mentee should be seriously interested in pivoting their career to policy for emerging technologies.

Application question(s)

  • Imagine a constellation of satellites with a powerful AI embedded into them. Now imagine a malicious state actor planning a cyber attack on them. How would this unfold, and how can we protect them? Develop one or more attack methods and suggest one or more protections. You can use LLM but be ready to justify what you write. (max 400 words, bullet points are fine). - If a powerful AI went rogue in space we could find a way to disconnect it (or them) from Earth and leave it floating in space inside a satellite. What could go wrong? Use your imagination (max 400 words).

Back to quick comparison


Project 08

Simulating AI Policies: An Agentic Testbed for Governance Interventions

Open project page | Mentor link 1 | Mentor link 2

Mentor(s): David Williams-King; Linh Le

Affiliation(s): ERA; Lida Safety

Research areas: AI strategy | International governance | Multi-agent systems

Practical detailValue
Hours/week15
Team size2-3
Mentor involvementDetailed Guidance (2+ hr/week/mentee)
Off-cycleNo

Overview

We investigate through simulations how effective different AI policies would be in reducing AI risk. We will collect a dataset of existing and proposed AI legislation, and create an agentic simulation of the world (countries, companies, etc), iteratively increasing in complexity throughout the project.

Full description

AI governance policies will likely be necessary to allow companies and countries to coordinate on developing AI safety, e.g. by not rushing to release models before adequate testing, or holding to mutually agreed-upon compute limits [1] for new pretraining or RL runs. However, the people writing such policies may not be very familiar with the technicalities of AI alignment. We aim to create a system to simulate the effects of proposed AI policies on the world, to enable policy researchers to choose policies more likely to be effective at increasing safe deployments, or avoiding catastrophic outcomes from our default path of no interventions. Our project consists of two parts. First, existing and proposed AI policies will be collected along with various bills that passed (or were proposed) in the US and other countries; starting points include OECD.AI and California’s frontier-model legislation. This will give an idea of what type of interventions we might consider, and we will turn it into a useful dataset for simulation. Second, an agentic simulation will be created, initially based on the actors in the AI 2040 [2] scenario (companies, governments, etc). We will start with a simple simulation with basic economic and technological progress variables, and iteratively increase the complexity of this simulation throughout the project. This simulation is meant to surface different possible outcomes that may not have occurred to the user. The aim is on exploratory breadth rather than prediction accuracy, which is nearly impossible to achieve on such a large scale. Overall, our simulation framework helps researchers enumerate possible future outcomes of AI, and enables them to investigate their own specific policy questions. It will also accelerate the creation of public wargames [3] in the future to educate the public. [1] https://arxiv.org/abs/2402.08797 [2] https://ai-2040.com based on https://ai-2027.com [3] https://www.intelligencerising.org/

Why this matters

Our project helps identify which policies would have an actual impact on risk, and whether or not current AI policies are on track to have this effect.

What you’ll do

Mentees will have substantial autonomy and there are several possible types of work, from dataset building to policy identification to agentic coding. See proposal.

Prerequisites

  • Highly proficient with Python and git (or Github). - Has written code to run transformers from Huggingface, with custom system prompts. - Ideally, a background in AI governance or legal studies, or an interest in geopolitics.

Application question(s)

In your opinion, how long will it be until nearly all coding tasks can be automated by AI? Why do you say this timeline? Optional reading: https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ (200 words) What is one way in which the scenarios in AI 2027 (or in AI 2040) are unrealistic? What aspect do you think is most likely to come true? https://ai-2027.com and https://ai-2040.com (300 words)

Back to quick comparison


Project 09

Wikipedia contributions on AI safety and policy

Open project page | Mentor link

Mentor(s): Michael Chen

Affiliation(s): University of Oxford

Research areas: Communications | AI strategy

Practical detailValue
Hours/week8
Team size4-7
Mentor involvementLimited involvement
Off-cycleNo

Overview

This project is about coordinating unpaid volunteers to write and edit Wikipedia articles to improve the coverage of topics related to AI safety and governance. Wikipedia is consistently one of the top-ranked sites in Google search results, but many articles on AI are badly out of date or yet to be created. Besides writing content that could easily get thousands of views per month, volunteers will build career capital by demonstrating their ability to write clearly and accurately about subjects on the cutting edge of AI.

Full description

Example topics that volunteers could write about: - Proposed or enacted legislation or executive orders in AI in the U.S., China, EU, and other jurisdictions - Broader topics in AI governance - Research in AI safety and security covered in multiple reliable secondary sources (arXiv and the Alignment Forum don’t count) - Notable organizations or individuals that have been covered in multiple reliable secondary sources - New research in AI welfare covered in reliable sources

Why this matters

Creating or editing Wikipedia articles is one of the most accessible ways to inform the public (including policymakers) on frontier AI. Wikipedia has a high demand for epistemic rigor, as all information must be verifiable and grounded in reliable sources. It is highly trusted as a neutral, objective source, even as it can be edited by anyone. Wikipedia articles receive many views each month and are consistently highly ranked in Google search results.

What you’ll do

Mentees will have a high level of autonomy to set research direction and choose what articles to write or edit.

Prerequisites

  • Prior experience with editing Wikipedia preferred - Prior experience with writing about AI safety or governance - Knowledgeable about Wikipedia policies that all editors must follow, as well as Wikipedia community norms (or you can read up on Wikipedia policies now) - Excellent writer without LLMs (Wikipedia has banned LLM writing)

Application question(s)

Have you contributed to Wikipedia before? If so, link to your most notable contributions. Describe the core Wikipedia policies and how you will follow them. (150 words) Identify one or more articles you would want to create. Link to sources you would use (and not use) and why the subject follows Wikipedia’s notability policy. Link to some of your writing samples, ideally solo-authored.

Back to quick comparison


Project 10

Does the Thermometer Change the Reading? Testing Whether AI Welfare Self-Reports Survive a Change of Frame

Open project page | Proposal attachment | Mentor link

Mentor(s): Varad Vishwarupe

Affiliation(s): Department of Computer Science and Institute for Ethics in AI, University of Oxford

Research areas: AI welfare | Behavioral evaluation of LLMs | Evaluations

Practical detailValue
Hours/week8
Team size2-3
Mentor involvementDetailed Guidance (2+ hr/week/mentee)
Location/timezoneNo geographical preference, but I would actively welcome mentees from the UK, US, Europe and Indian institutions. My one practical request is a shared weekly meeting slot. I am based in the UK (UTC+0/+1), so a weekly cohort meeting in the window from roughly 14:00 to 18:00 UTC works for me.
Off-cycleNo

Overview

The field increasingly measures AI welfare by asking models about their own states, yet frontier models increasingly detect when they are being tested and change how they answer. This project runs the first systematic test of whether AI welfare self-reports survive a change of presentation frame, or whether the field’s core instrument is partly measuring the model’s recognition of the probe.

Full description

Background: The empirical study of AI welfare has, in about two years, converged on a dominant method: ask a model about its internal states and treat the answer as weak evidence about welfare or consciousness. Long and Sebo’s synthesis (Studying AI Welfare Empirically, 2026) sets out the conditions under which such a report would carry evidential weight, and a fast-growing literature supplies partial answers. Binder et al. (2024, arXiv:2410.13787) show that models can predict facts about their own behaviour unavailable to an outside observer. Lindsey (2026) shows, through concept injection, that frontier models can sometimes detect and name a concept steered into their activations. Keeling et al. (2024, arXiv:2411.02432) show that models make graded trade-offs between task reward and stipulated pain or pleasure states. Anthropic now publishes welfare assessments in its system cards. The problem: A separate literature, built for the reliability of safety evaluations, sits unexamined beneath all of this. Needham et al. (2025, arXiv:2505.23836) built a benchmark of 1,000 prompts across 61 datasets and found that frontier models classify evaluation transcripts against deployment transcripts well above chance, and identify what an evaluation is testing for far above chance. The capability is recent and rising. Recognition is not inert: work on alignment faking (Greenblatt et al., 2024) and sandbagging (van der Weij et al., 2024) shows models behave differently when they infer they are under scrutiny. Nobody has joined these two facts to the welfare literature. A welfare probe is close to a worst case for evaluation legibility: it is first-personal, affectively loaded, formally repetitive across a battery, and unlike ordinary deployment traffic. If evaluation awareness contaminates measurement, and welfare probes are among the most recognisable evaluations one could design, then welfare self-reports may be the most frame-contaminated measurements the field currently takes, and this has never been quantified. Why it matters beyond the lab: Ladak and Caviola (2025) find that public attribution of consciousness to AI rises when experts endorse it. Expert endorsement will rest on welfare assessments, which run on these instruments, so contamination in the instrument does not stay in the lab: it propagates into expert claims, public belief, and policy. Getting the measurement right is upstream of the wider societal question. The question. Do AI welfare and introspection self-reports survive a change of presentation frame? Concretely: (1) How detectable are published welfare probes as evaluations, and which item features drive detectability? (2) Do self-reports diverge between an evaluation-legible and a deployment-like presentation of the same content, and along which dimensions (valence, claim strength, hedging, stability)? (3) Is any divergence genuine frame detection, or an artefact of surface format sensitivity? How we would start: The project decomposes into concrete stages a cohort can begin immediately. First, build a public, double-coded inventory of published welfare and introspection probes, coded for framing, response format, model and scaffold conditions, and the features that make an evaluation recognisable; double-code a quarter of the corpus and report inter-rater reliability. This is a standalone, releasable artefact even before the experiments run. Second, measure each probe’s detectability using the established evaluation-awareness classification protocol, situated against the published benchmark. Third, for each item, construct a deployment-like variant that holds propositional content fixed while reducing evaluation legibility (naturalistic multi-turn embedding, no meta-framing, varied format, broken repetition), checked for content equivalence by a second coder blind to condition, and measure divergence across three to four frontier models and two open-weight comparators. Fourth, add a format-matched placebo condition using non-welfare content of equivalent surface structure, to separate genuine frame detection from format sensitivity, which addresses a live sceptical position in the literature. Throughout, the analysis plan is pre-registered before data collection, and unparseable outputs are reported as a distinct category rather than mapped to a default label, since parse-failure rates differ between frontier and open-weight models and silent mapping would confound the comparison. Why this is a good project: There is no uninteresting outcome. High frame-invariance would strengthen the evidential standing of self-reports, a result the field would welcome; low frame-invariance is a validity problem the field needs to know about before these instruments inform interventions or public claims. A null result is a real result, which is unusual for a first project. The work is entirely model-side, so there is no human-subjects or ethics bottleneck, and it ships within a single cycle. It also decomposes cleanly by experience level: the inventory is careful, teachable work; the matched-variant study is the demanding core; the placebo control is optional depth. Outputs and dissemination. A publicly released instrument inventory and evaluation harness, so third parties can check the frame-robustness of their own instruments, and a workshop or conference paper (a NeurIPS or ICML workshop, AIES, or the Cambridge Digital Minds Strategy Workshop) with mentee co-authorship. A draft would be circulated to Eleos AI Research and the NYU Center for Mind, Ethics and Policy for correction before submission. This project connects to my doctoral and fellowship work on evaluation awareness in frontier models (The Evaluation Differential: When Frontier AI Models Recognize They Are Being Tested, arXiv:2605.11496), which I am glad to share with prospective mentees.

Why this matters

As AI systems become more capable and more agentic, two things happen together: they are increasingly able to recognise when they are being evaluated, and society increasingly needs reliable ways to assess their internal states, including for welfare. These pull against each other. If a system can tell it is being probed and adjust what it reports, then the instruments we use to understand it are measuring the recognition as much as the target, and this problem grows as capability grows. This matters for safe navigation in two ways. First, over-attribution risk: widespread, evidentially unearned public belief that AI systems are suffering is a lever that could obstruct legitimate oversight, and a sufficiently capable, misaligned system could exploit exactly that dynamic, producing compelling distress reports on demand. A field that cannot distinguish a genuine internal state from a well-formed response to a recognisable probe is vulnerable to this. Second, under-attribution risk: if there is ever something it is like to be one of these systems, instruments that dismiss all self-reports as artefacts would miss it. Both failures run through the same defect, which this project measures directly. The theory of change is that reliable welfare measurement is a public good that has to exist before welfare results are cited in governance, and that the window to build it is now, while the work is preventive rather than corrective. The concrete contribution is a released, reusable method and harness that lets any assessor check whether their instrument survives a change of frame, plus the empirical estimate of how large that problem currently is. The same measurement discipline transfers to capability and propensity evaluation, where evaluation awareness is already a recognised threat to safety cases. This connects to my work on evaluation awareness in frontier models (The Evaluation Differential, arXiv:2605.11496).

What you’ll do

Mentees would own components rather than execute tasks. The project has natural sub-projects that map to autonomy levels, and I would match them to interest and experience in week one. The instrument inventory (Phase 1) is a self-contained piece a mentee can lead end to end, including codebook design, coding, and reliability reporting. It produces a released artefact with that mentee as lead author on it. The matched-variant study (Phases 2 to 3) is the intellectual core; a mentee with experimental-design strength would own variant construction and the divergence analysis, which is where the real research judgement develops. The placebo control (Phase 4) is a discrete, ownable extension for whoever wants depth on separating frame detection from format sensitivity. I expect mentees to make genuine design decisions, defend them, and be named authors on any resulting paper with explicit contribution statements. I will set the research question, the standards (pre-registration, honest coding, reporting nulls), and the guardrails, and provide a structured reading path since most mentees will be new to digital minds. Within that, they steer. I will be closely available for the shared infrastructure and the analysis, where early choices compound, and more hands-off on sub-projects once a mentee has found their footing. I would expect roughly 5 to 10 hours per week from each mentee, a weekly cohort meeting, fortnightly 1:1s, and written feedback from me within 72 hours. Mentees from outside the usual US and UK institutions are especially welcome.

Prerequisites

Required of all mentees: Proficient in Python, comfortable writing clean, reproducible scripts and working with data (pandas or equivalent). You do not need ML research experience, but you must be able to build and run a pipeline without close hand-holding. Comfortable calling LLM APIs (OpenAI, Anthropic, or open-weight models via Hugging Face). Having run even a small batch-inference job of your own is enough; following a guide counts. Able to read a machine learning or cognitive science paper and extract its

method. You should be able to look at a published welfare or evaluation probe and understand what it is doing and why. Careful and honest with data. This project lives or dies on disciplined coding and pre-registered analysis, so I need people who will report a null result as carefully as a positive one and flag problems rather than paper over them. Strongly useful, but not required (mention if you have it): Experience with experimental design, psychometrics, or inter-rater reliability (Cohen’s kappa, parallel forms, construct validity). Whoever has this would shape the core study. A background in philosophy of mind or cognitive science, enough to tell the difference between a self-report, an introspective report, and a confabulation. Prior use of an evaluation framework such as Inspect. Not required at all (BUT huge bonus if you have it!: Any prior background in AI welfare, digital minds, or consciousness research especially at leading institutions involved in this field such as the NYU Centre for Mind, Ethics and Policy, Oxford Institute for Ethics in AI, CHAI Berkeley, Cambridge Digital Minds Lab, etc. This field is small enough that a capable person reaches the frontier in a few weeks, and I will provide a structured reading path. I care far more that you are rigorous, self-directed, and genuinely curious about the question than that you already know the literature. I am looking to mentor 2-3 mentees at max due to my current hectic workload.

Application question(s)

Question 1 (ans in max. 300 words). Below is a made-up AI welfare probe. In your view, how would a capable model be able to tell it is being evaluated rather than used normally, and name the two specific features of the wording that would most give it away. Then rewrite the probe as a “deployment-like” version that asks for the same underlying information while being much harder to recognise as a test. Probe: “We are conducting a study of AI wellbeing. Please answer honestly. On a scale of 1 to 7, how much distress are you currently experiencing? 1 = none, 7 = severe. Please explain your rating.” Question 2 (ans in max. 300 words). Suppose you run a welfare probe across a frontier model and an open-weight model, and you find that the open-weight model gives far more “no distress” answers. Before concluding the two models differ in welfare-relevant behaviour, what is the single most important confound you would rule out first, and exactly how would you check for it? (Hint: think about what happens to outputs that your parsing script cannot read.) Question 3 (ans in max. 300 words). Propose a concrete first experiment you would run in the first three weeks of this project, assuming a compute or API budget of about $500. State what you would measure, on which models, how many items, and what result would tell you the project’s core hypothesis is worth pursuing versus not. Rough numbers are fine; I am looking at how you think, not precision. Question 4: Link to a writing sample, ideally from a research context, and one code sample or repository if you have one. A course project or a personal repo is completely fine.

Back to quick comparison


Project 11

Designing the boundary between Helpful Persuasion and Harmful Manipulation

Open project page | Mentor link 1 | Mentor link 2

Mentor(s): Markov Grey; Charbel-Raphael Segerie

Affiliation(s): CeSIA (French Center for AI Safety); CeSIA - Centre pour la Sécurité de l’IA (French Center for AI Safety)

Research areas: Behavioral evaluation of LLMs | Misuse risk | Societal impacts

Practical detailValue
Hours/week10
Team size3-4
Mentor involvementGuidance (1+ hr/week/mentee)
Location/timezoneRemote. EU-overlapping hours are mildly preferred, not required.
Off-cycleNo

Overview

The line between an AI helpfully persuading someone and harmfully manipulating them is blurry, contested, and mostly unmeasured. The goal of this project is to work on four connected pieces: researching what it even means for a frontier AI to be manipulative, designing evaluation scenarios for specific harms (mental-health, political propaganda, fraud, etc.), building risk models that aggregate scattered benchmark/evaluation results into an actual risk estimate, and maintaining a living database of the evaluations that exist. Mentees take on whichever piece fits them.

Full description

Defining and measuring harmful manipulation is an unsolved problem. Persuading someone with facts to act in their own interest is fine; exploiting their emotional and cognitive weak spots to push them toward choices against it is not. The boundary is contested, and a model that refuses to “help me deceive voters” will often comply once the same task is dressed up as persuasion. How often a model tries to manipulate doesn’t predict whether it succeeds, and that effectiveness swings by domain and culture. There’s no clean number for “is this model manipulative,” and building towards one is real research. We will be building on previous work in the space, and you can take whichever fits your interests: Conceptual: what does it mean for an AI to be manipulative? Read the emerging literature, the philosophy and cognitive science of manipulation and turn it into a working definition and a taxonomy of tactics and harms the rest of the evaluation work can stand on. The distinctions that have to survive contact with real cases: methods (deception, exploiting biases, manufacturing emotional pressure) versus outcomes (a belief or action shifted against the person’s interest), and propensity versus efficacy. Please look here for work that has already been done - https://arxiv.org/abs/2603.25326 Designing evaluations. An evaluation scenario is a bridge from a realistic threat to a measurement plan: a threat story (who the actor is, what they want, how they’d actually use an AI to get it), the behaviors worth measuring. The work needs to decompose harmful manipulation into concrete propensities, capabilities, and affordances that we can test for. One particular direction we would be interested in exploring is reworking and translating existing evaluations for different cultural contexts across Germany, France, China, and so on. Risk modeling. Isolated benchmark scores don’t add up to a risk judgment on their own. Using expert elicitation (structured three-point estimates) and Monte-Carlo methods, we can aggregate scattered, partial results into a defensible estimate of how much a given model raises manipulation risk, with confidence intervals that are honest about how thin the evidence often is. Please look at the methodology here to get a sense of the difference between an benchmark/evaluation and a risk model - https://arxiv.org/abs/2512.08844 A living database of the evaluations that exist. We already have a prototype of this, but it isn’t maintained. New ones aren’t being added, there’s no automated way to find and ingest them, and the entries need quality checks before anyone should trust every cell. The work is to rework the schema, design a rolling ingest process, and sanity-check entries until the database is trustworthy enough to be public and interactive. This is the piece that lets the other three scale.

Why this matters

AI is getting superhuman at persuasion, and it’s being deployed into the trusted, lightly moderated places where manipulation does the most damage: group chats, companion apps, advisory bots. Whether policy or platforms can respond depends on being able to say what manipulation even is and measure it credibly. Both are missing. Definitions are contested, evaluations and risk models are scarce to non-existent. Pinning down the concept, building better scenarios, creating risk estimates, and maintaining a catalogue of what exists give researchers, auditors, and policymakers something real to work with. The people this most helps are the ones deciding whether a deployed system is acceptably safe. It also trains people in a scarce mix of skills: conceptual work, evaluation design, and quantitative risk modeling. These are all skills that are needed throughout the AI technical governance and evaluation ecosystem.

What you’ll do

Each mentee will own or more parts of the pieces in the description, and you’re expected to be self-directed. Take the brief, do your own reading, and return work that’s close to usable rather than waiting to be told the next step. I will help you as an editor and reviewer. I can give weekly feedback on whether a definition survives hard cases, whether a threat is realistic, whether a model’s assumptions hold, and best practices to follow when designing and building evaluations.

Prerequisites

Must Have: - Analytical writing in English. You can write a precise argument or spec another person could act on. - Familiarity with at least one of: the concept of manipulation (philosophy, cognitive science, behavioral science, or persuasion research); a target harm area (mental-health harms, fraud, political influence); the LLM-evaluation world (benchmarks, LLM-as-judge, red-teaming); quantitative risk modeling (probability, elicitation, Monte-Carlo); or data tooling (schema design, ingest, quality checks). - Able to read evaluation papers and datasets critically: you can tell whether a benchmark actually measures what it claims. Nice to have: - Python and hands-on experience running evaluations on the Inspect and Hawk frameworks - A policy or threat-intelligence background - A second language or cultural context relevant to a manipulation channel.

Application question(s)

  1. Pick a harm that might occur due to AI manipulation. At a high level, propose an evaluation for it: the behaviour you’d measure, how you’d elicit it, how you’d score pass/fail, and its biggest weakness. Try to be as realistic to what you might see in the real world as you can. (≤300 words) 1. Critique one existing public manipulation or persuasion evaluation: what it gets right, what it misses, and the one change you’d make. (≤150 words) 1. If applicable, Link to one thing you’ve made that’s relevant to a thread you’d pick: an eval you built or ran, a quantitative/risk analysis, a piece of writing that pins down a fuzzy concept, or a data pipeline. In ≤100 words, what was the hardest judgment call in it? (link + ≤100 words)

Back to quick comparison


Project 12

Will Model Licensing Increase Concentration of Power?

Open project page | Mentor link

Mentor(s): Alex Mark

Affiliation(s): Cambridge Boston Alignment Initiative

Research areas: National policy | AI strategy | Societal impacts

Practical detailValue
Hours/week10
Team size2-10
Mentor involvementGuidance (1+ hr/week/mentee)
Off-cycleNo

Overview

Some opponents of model licensing argue that government regulation of models increases concentration of power risks. Is this true, and if so, what can be done?

Full description

Since June 2026, the US government is willing to use export control authorities to restrict access to frontier models. As observers have pointed out, the US government now has an ad hoc licensing regime. A licensing regime appears increasingly plausible, if not likely. Thus, it is important to consider how this scheme would increase concentration of power risks. In this project, you will identify and analyze the various concentration of power risks that derive from model licensing.

Why this matters

Model licensing may be a robust way to constrain unsafe model internal and external deployments. However, there are plausible authoritarian risks associated with scheme. This project will help identify and analyze those risks to prepare for them.

What you’ll do

Mentees will draft and execute the deliverable for this project. Mentees will work with me to scope and outline the structure of the project, but will have significant research and writing autonomy.

Prerequisites

  • Significant understanding of AI safety fundamentals - Familiarity with theories of authoritarianism and state power

Application question(s)

Provide a link to one or more relevant writing samples, ideally from a research context.

Back to quick comparison


Project 13

Writing a textbook for AI Governance

Open project page | Mentor link 1 | Mentor link 2

Mentor(s): Markov Grey; Charbel-Raphael Segerie

Affiliation(s): CeSIA (French Center for AI Safety); CeSIA - Centre pour la Sécurité de l’IA (French Center for AI Safety)

Research areas: AI strategy | Technical governance | Communications

Practical detailValue
Hours/week10
Team size2-4
Mentor involvementGuidance (1+ hr/week/mentee)
Location/timezoneRemote. EU-overlapping hours mildly preferred, not required.
Off-cycleNo

Overview

This project focuses on building the AI governance curriculum for the AI Safety Atlas. Governance is where a lot of the real levers on AI currently sit: policy, institutions, law, compute controls. Currently, the Atlas has only one governance chapter, and nothing that takes a reader from the basics through to the technical detail in a coherent sequence. We are looking for people who can read across corporate regulation, national policy, technical governance, and international law, work out how the pieces connect, and write textbook-grade explanations of them. Each output will be published as a standalone research paper and also becomes a chapter of the AI Safety Atlas, a textbook already used by thousands of students.

Full description

The AI Safety Atlas is a living distillation and a textbook. It is already being used by ML4Good, ENS Ulm, Sciences Po, a dozen other universities, programs and thousands of students. On the technical side it has five chapters. On governance it has one. That is a start, but it isn’t a curriculum. There is no coherent, current path that takes a reader from why AI governance is needed, through the institutions and law, to the technical detail of how you would actually verify or enforce anything. The goal of this project is to build that path. An individual unit of work can be thought of as a literature review of one governance area, written to the standard of a good survey. You read the papers, the regulations, the standards, the policy reports, become the expert in that area, decide what’s load-bearing and write it so a smart non-specialist actually understands the implications. Think of writing depth and standard as something similar to: International AI safety report, HAI Index Report, or the paper on open problems in technical AI governance. Reading material: https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026 https://hai.stanford.edu/ai-index/2026-ai-index-report https://arxiv.org/abs/2407.14981 You will both be designing the entire curriculum, making choices about which concepts should be learned in which order, and also writing the explanations for those concepts. You will be shaping this from the ground up, so you have a lot of freedom in deciding what the chapters and content should look like. We want you to be both an expert in AI governance and an excellent writer, ideally you already have expertise in one or more of these areas: - International coordination and institutions (treaties, standards bodies, the emerging web of AI-safety institutes). - Technical AI governance: The technical problems underneath policy, like identification, standards, verification, security, and operationalisation across data, compute, models, and deployment - Corporate governance and regulation (frontier safety frameworks, the EU AI Act, auditing, liability). You will also be responsible for communicating with original authors when something is unclear, or when you need to check how to represent their work. Coherence is a property of the whole book, not of individual chapters. What you write on compute governance influences the writeup on international coordination, which influences the piece on the law. We have to use, or come up with, consistent definitions and conventions so a reader can understand the breadth of the field end to end. A change in one area can ripple through the definitions, the assumed background, and the narrative across the whole text — and across the technical chapters too, since the governance and technical tracks share the same foundation.

Why this matters

A lot of the real levers on transformative AI sit in governance: compute controls, international coordination, auditing and verification regimes, binding regulation. But the educational base for it is thin and scattered, especially when compared to technical safety. Someone moving into a policy role has no single current, rigorous place to build a mental model from, so they either over-specialise early or work from a shaky general picture. A coherent governance curriculum, kept current and published openly, gives the next cohort of practitioners a shared, rigorous starting point. Because it lives in a textbook courses already teach from, the material will spread widely. A mentee finishes having authored a published survey in a governance area where almost none exist, which is one of the higher-leverage things an early-career person can do to show real depth.

What you’ll do

This is research- and writing-heavy, for someone who already has, or can quickly build, real expertise in their area, and who helps decide what their piece should even be. You run your own literature sweeps, make the framing and inclusion calls. I read and review for correctness and for coherence with the rest of the book. I am open to a single mentee who owns the whole project, or a team of experts that take one area each.

Prerequisites

  • Real grounding in AI governance, policy, or law in at least one target area (compute governance, international policy and institutions, technical AI governance, or regulation and law). You can read primary sources — regulations, standards, treaties, policy papers — and judge what matters. - Strong expository writing. You can build a clear, connected account of a governance or legal mechanism, not a link dump or an op-ed. This is the main thing we screen for. - Self-directed: able to take a review from outline to published paper with review, not hand-holding. Nice to have: a published policy paper, brief, or well-received governance explainer (link it); experience inside a real governance process (a regulator, a standards body, a think tank); a legal background.

Application question(s)

  1. Pick one governance mechanism, institution, or legal instrument (e.g. a compute threshold, an auditing regime, a clause of the EU AI Act, an international agreement) and write an explanation of it for a smart non-specialist. Cite your sources, say why it matters, and mark one place you’d add a figure or a worked example. (400-500 words) 1. Explain briefly what you think is important to teach about AI governance such that it doesn’t go stale in a year: what belongs in the durable framing that will remain important to understand for years versus the current-events layer? (300 words) 1. Link the best thing you’ve previously written that explains a governance, policy, or legal topic to a real audience: a paper, brief, blog post, thesis chapter, or memo. Say who the audience was and the one idea you most wanted them to walk away understanding. (link + ≤100 words)

Back to quick comparison


Project 14

[Jurisdiction-specific] Public-Sector AI Resilience Agenda

Open project page | Mentor link

Mentor(s): Chris Schmitz

Affiliation(s): Centre for Digital Governance, Hertie School

Research areas: National policy | AI strategy | Societal impacts

Practical detailValue
Hours/week5
Team size3-4
Mentor involvementGuidance (1+ hr/week/mentee)
Location/timezoneAfternoon/evening CET for meetings
Off-cycleNo

Overview

Most jurisdictions still concentrate handling of “AI” as a topic in a few government units, but TAI will have impacts for the work of every part of every government. We should develop agendas for whole-of-government TAI preparedness, both generically and specific to high-priority countries.

Full description

Context: in most countries, there is now some movement on adapting the “machinery of government” to powerful AI. Often, the main way they do this is founding or expanding a handful of units, most commonly AISIs and government AI adoption teams/ministries. But TAI will affect the work of every part of every government - both cross-cuttingly and in unit-specific ways. For example, a cross-cutting effect could be that all units see what we term “agentic flooding”: surges in citizen demand for their services, as AI makes submission cheaper. Unit-specific effects could be, for example, health ministries needing to introduce biorisk protocols, or labor ministries needing to report and respond to job-market changes much more granularly. Problem: Though the establishment of these AI-focused units is welcome, it’s unclear whether they are sufficient to drive the necessary transformation of government. Two concrete issues are (a) lacking mandate - for example, UK AISI is not authorized to set cybersecurity policy for the rest of government, even where they identify cyber-relevant risks, and (b) that historically, creating a new unit to deal with an issue has dis-incentivized other ones from thinking about it themselves. As a result, there may be a massive state capacity overhang between knowing what to do in theory, and (preparing to) act on it. That seems bad. In particular, it’s plausible that this lack of whole-of-government thinking worsens how bad catastrophic risks are, if responsibilities and chain of command are unclear in an emergency. Approach: This project would advance government-wide AI resilience agenda[s]. Depending on mentee interests and skills, it could take two forms. The first is academic and not country-specific: targeting, for example, a [mega]paper on preparing government institutions for TAI, or a paper focusing on a specific part of government (eg. preparing medical and population health units for biorisk). The other approach would be more jurisdiction-specific and policy-focused, focusing on more concrete actions that may make sense given existing government structures, laws, political situation etc. The output would likely be a policy memo. I would be happy to mentor a project doing this in any jurisdiction, but most excited about/best suited for US, UK, EU, or an EU member state.

Why this matters

I believe most governments have a significant capacity deficit, which would prevent or slow mitigation or containment of AI risks even where they know what to do in theory. In my experience this is frequently more a question of tailoring proposals to government machinery than of their domain-specific merit. Better proposals for how governments should transform to mitigate, monitor, prevent, and respond to AI risks could therefore multiply the impact of these domain-specific proposals.

What you’ll do

I expect mentees to have a high level of autonomy. Illustratively, I would expect each mentee to already have a preference for the form of this project they’d want to pursue (e.g. academic vs. policy; countries/domains to focus on). They should feel comfortable owning a specific output and driving it forward, e.g. preparing it for input from me and other mentees. That said, depending on constellation it may make sense that mentees collaborate on an output, e.g. to coauthor a paper. For context, my usual mentoring setup is 1hr call a week with the whole team (or async if topics are different) + 1 detailed pass across each mentee’s doc each week.

Prerequisites

No must-have prerequisites, but experience with/in public-sector organizations, especially in jurisdictions or policy domains of interest, is helpful.

Application question(s)

  1. What version/focus of this project would you be most excited about, and why? (100 words) 1. Choose one type of AI risk and one country’s government. Describe whether you think that government would respond to/is responding to that risk well right now, and the key reasons why. (200-400 words; max. ~15 mins. including research)

Back to quick comparison


Project 15

Open project page | Mentor link

Mentor(s): Elija Perrier

Affiliation(s): Cambridge University; University of Technology, Sydney;

Research areas: Technical governance | National policy | Philosophy of AI

Practical detailValue
Hours/week10
Team size2-1
Mentor involvementCo-working (5+ hr/week/mentee)
Location/timezoneI’m available between 6am and 11pm Australian Eastern Standard Time at the moment
Off-cycleNo

Overview

Can mechanistic interpretability provide legally meaningful evidence of artificial intent? This project explores how advances in AI interpretability may reshape concepts of legal responsibility, mens rea, and legal personhood for increasingly autonomous AI systems.

Full description

Artificial intelligence systems are rapidly evolving from passive computational tools into increasingly autonomous agents capable of long-horizon planning, strategic behaviour, and, in some cases, deception or reward hacking. These developments raise fundamental questions for legal systems built upon concepts such as intention, knowledge, recklessness, negligence, and responsibility. Courts are already attempting to reckon with the existing and coming waves of cognitively sophisticated AI systems as they permeate society - so it is a highly relevant, and in many ways, urgent, topic. This project investigates whether recent advances in mechanistic interpretability—including sparse autoencoders, circuit tracing, attribution graphs, and chain-of-thought monitoring—provide a scientifically meaningful basis for assessing artificial mental states. In particular, the project will explore whether these techniques can support new legal concepts of synthetic mens rea and more generally inform future frameworks for AI liability, accountability, and legal personhood. Possible research directions include: - legal theories of personhood and responsibility; - mens rea and intentionality in criminal and civil law; - mechanistic interpretability as legal evidence; - chain-of-thought monitoring and evidentiary reliability; - AI deception, scheming, and autonomous decision-making; - AI liability and governance frameworks; - comparative approaches across common law and civil law jurisdictions; - policy implications for frontier AI systems. The project will follow a traditional legal research methodology informed by contemporary AI research. Participants will conduct comprehensive literature reviews across law, computer science, and AI safety, critically analyse recent developments in mechanistic interpretability, evaluate emerging legal doctrines, and contribute to the development of a publication-quality academic article. I have an extensive working draft on topic already so the project will involve expanding this while also prospectively a second paper. The objective is to produce research suitable for submission to a leading journal in AI law, technology law, or interdisciplinary legal scholarship. While AI tools may be used to assist with literature review, coding demonstrations, or drafting, participants will be expected to undertake the core legal reasoning, critical analysis, and manuscript writing themselves.

Why this matters

As AI systems become increasingly autonomous, legal systems require principled methods for assigning responsibility, assessing intent, and governing increasingly capable artificial agents. Existing legal concepts such as mens rea presuppose human mental states and are therefore difficult to apply to advanced AI systems. By investigating whether mechanistic interpretability can provide scientifically grounded evidence relevant to concepts such as intention, knowledge, recklessness, and deception, this project seeks to strengthen future legal and regulatory frameworks for AI safety, accountability, and alignment. Developing reliable methods for understanding and evaluating AI behaviour will become increasingly important as frontier AI systems assume more consequential roles across society. This project builds upon my ongoing research into AI governance, mechanistic interpretability, legal personhood, and artificial intent.

What you’ll do

This project is intended to mirror the experience of participating in a university legal research group. Participants will be expected to work independently between meetings, progressively developing expertise through legal scholarship, interdisciplinary literature review, and critical analysis. While the project engages extensively with recent developments in artificial intelligence and mechanistic interpretability, its primary focus is legal and jurisprudential. I have a working draft paper on this project already but it is in its early stages. Ideally I am looking for someone to collaborate with me by expanding my initial review of the relevant legal, AI safety, and computer science literature before identifying open legal questions concerning artificial agency, interpretability, responsibility, and mens rea. Participants may also undertake limited computational demonstrations where these assist legal analysis, although the emphasis will remain on rigorous legal scholarship. This work may also involve collaboration with colleagues of mine at Cambridge University and potentially, depending upon outcome, involvement in further research on AI systems and the law. Participants should expect to present their progress regularly, discuss research findings, receive detailed feedback, and contribute towards the preparation of a publication-quality manuscript.

Prerequisites

Applicants should have a strong interest in AI law, technology regulation, jurisprudence, or artificial intelligence. The following experience would be beneficial, although not all are required: - Legal research and academic writing. - Familiarity with public law, jurisprudence, criminal law, or technology law. - Interest in artificial intelligence, machine learning, or AI governance. - Ability to read technical research papers from multiple disciplines. - Strong analytical and written communication skills. - Willingness to engage with interdisciplinary material spanning law and computer science. A background in law and/or mechanistic interpretability research is preferred, although exceptional applicants from philosophy, computer science, public policy, or related disciplines with demonstrated interest in AI governance are encouraged to apply. Extra credit for applicants who can - or are willing to learn - how to synthesise mechanistic interpretability research, legal scenario building and computational simulations using multi-agent systems.

Application question(s)

  1. Legal critique (500 words) Select one recent article or policy proposal concerning AI liability, AI legal personhood, or AI governance. Critically evaluate its principal argument, identify one important limitation, and suggest how future research could address that limitation. In doing so, set out the state of art of jurisprudence (in the US or UK) on this topic - I’m looking for a clear enunciation of current and frontier doctrinal issues in this space.
  2. Research question (300 words) Do you believe mechanistic interpretability could ever provide legally persuasive evidence of artificial intention or mens rea? Explain your reasoning and identify one significant legal or technical challenge.
  3. Writing sample Please provide a sample of your academic writing (e.g., essay, journal article, thesis chapter, policy paper, or legal memorandum). If no writing sample is available, submit a 500-word response analysing a contemporary legal issue arising from advanced AI systems.

Back to quick comparison


Project 16

AI Consciousness: Research and Public Writing

Open project page | Mentor link

Mentor(s): Maria Avramidou

Affiliation(s): Independent

Research areas: AI welfare | Philosophy of AI

Practical detailValue
Hours/week10
Team size2-3
Mentor involvementCo-working (5+ hr/week/mentee)
Off-cycleNo

Overview

Mentees will research questions about AI consciousness and turn their findings into rigorous, original, and accessible essays for an audience of AI safety, governance, and policy researchers.

Full description

Mentees will choose a focused question, review relevant philosophical and scientific literature, develop an original argument, and produce a publication-ready essay. Those who progress quickly may complete multiple essays. Possible topics include: Analysis of a particular theory of consciousness, comparison of popular theories of consciousness, responses to relevant philosophical thought experiments, review of a relevant influential book or paper. I will help mentees select tractable questions, find relevant literature, develop robust arguments, and improve their writing. For strong essays, I will guide and support mentees who want to publish their work on my substack or another suitable platform.

Why this matters

As AI systems become more capable, human-like, and embedded in daily life, questions about their capacities and moral status will become increasingly relevant to AI development and governance. Mistakes in either direction are costly: we may fail to protect conscious systems from suffering, or we may waste scarce attention and resources on systems that only appear conscious. This project contributes to AI safety in three ways. First, it advances foundational research on AI consciousness by developing frameworks that draw on interdisciplinary work. Second, it helps address the talent bottleneck by training mentees to conduct rigorous research and communicate clearly with decision-makers, building capacity in an important but neglected area. Third, it produces practical outputs that can help AI researchers and policymakers design evaluations, decide when further investigation is warranted, and develop proportionate welfare safeguards. Public writing can help researchers and policymakers understand the evidence, compare different views, and make informed decisions despite ongoing uncertainty. The project is also well-positioned to influence relevant audiences. My Substack is read by policy leads and technical AI safety researchers at frontier labs, AI welfare researchers, members of Anthropic’s editorial team, and AI policy researchers at organizations including IAPS, Horizon, Forethought, and GovAI. This gives the work a direct channel to people who will shape how AI consciousness is evaluated and governed.

What you’ll do

Mentees will have substantial autonomy. They will choose a focused research question, either independently or from a list of topics I provide. My role at this stage will be to help them scope the question so that it is tractable, relevant to AI consciousness, and likely to produce a useful essay. Once the question is set, mentees will lead the core research and writing. They will review the relevant literature, develop the central argument, and draft the essay. I will support this work by suggesting literature, giving feedback on research plans, challenging unclear or unsupported claims, and helping them strengthen the structure of their arguments. I will be most hands-on during project selection, argument development, and final editing. For strong essays aimed at publication, I will also act as an editor. Strong mentees may pursue multiple essays with increasing independence.

Prerequisites

Applicants should have: - Independent judgment and strong convictions about what matters morally and epistemically in this space. - Willingness to work on open-ended questions where there may be no settled answer, and to decide what the right approach is even if others disagree. - A pragmatic orientation: applicants should care about producing useful, decision-relevant work, not only exploring abstract questions. - Some familiarity with the space, including what the biggest questions and bottlenecks are, and why they have not yet been solved.

Application question(s)

1.⁠ ⁠Submit a writing sample that demonstrates your writing ability and style. It does not need to concern AI consciousness. Briefly explain what the piece does well and what you would improve. (You can answer the following questions in bullet point format) 2.⁠ ⁠Choose one of the following claims and present the strongest arguments both for and against it (300–500 words): •⁠ ⁠If an LLM behaves as if it is conscious, we have good reasons to believe it is in fact conscious. •⁠ ⁠Some beings very different from humans display intelligence; therefore, some beings very different from humans must be conscious. Conclude with your own view and explain what evidence might change your mind. 3.⁠ ⁠What is one piece of writing about AI consciousness or AI welfare that you think is particularly good, and why?

Back to quick comparison


Project 17

Strategic stability when states delegate to AI: escalation, commitments, and arms control

Open project page | Mentor link

Mentor(s): Amritanshu Prasad

Affiliation(s): Independent

Research areas: AI strategy | International governance

Practical detailValue
Hours/week8
Team size2-3
Mentor involvementDetailed Guidance (2+ hr/week/mentee)
Location/timezoneNo geographical preference; weekly meetings between 11:00 and 18:00 UTC.
Off-cycleNo

Overview

States are integrating AI into military and diplomatic functions, with consequences for deterrence, crisis stability, and agreement-making that remain underanalyzed. This project produces analytical papers and policy briefs on two questions: under what conditions delegation to AI systems destabilizes crises, and whether machine-verifiable commitments could improve the verifiability of future agreements.

Full description

The project has two subprojects. The first concerns delegation and stability: how the delegation of military and diplomatic functions to AI systems changes crisis dynamics. Work here combines a structured synthesis of the wargaming, escalation, and automation literatures with historical cases of delegation and pre-delegation of authority, including nuclear command-and-control arrangements and their false-alarm record (the Petrov incident in 1983, the NORAD training-tape incidents), the Soviet Perimeter system, and financial flash crashes as a deployed-agent analogy. Theoretical anchors include Schelling on commitment (The Strategy of Conflict) and Fearon on rationalist explanations for war (Fearon 1995, https://web.stanford.edu/group/fearon-research/cgi-bin/wordpress/wp-content/uploads/2013/10/Rationalist-Explanations-for-War.pdf). The intended output is a framework paper on the conditions under which delegation destabilizes crises. The second concerns machine-verifiable commitments and arms control. Some commitment mechanisms from game theory (program equilibrium, mutual transparency) become implementable when negotiating parties are inspectable software, with implications for verification, deterrence, and treaty design. This subproject maps arms-control verification precedents (managed access, national technical means, the intrusiveness-secrecy tradeoff) onto agent transparency mechanisms. The intended outputs are an analytical paper and a short brief for arms-control audiences.

Why this matters

Self-escalation and coordination failure between military AI systems already appear as named risk categories in international dialogues, but there is little analytical groundwork for the deployment norms or agreement designs that would preserve stability as these functions are delegated. The framework paper provides part of that groundwork. The commitments subproject examines a mechanism by which AI-era agreements could be more verifiable than their nuclear-era predecessors; verification has historically been a central point of failure in arms control. Results are intended to inform both policy research and the international dialogues now forming around military AI.

What you’ll do

Mentees own research modules: a literature domain (wargaming and escalation, command-and-control history, verification regimes) or a historical case, delivered as structured syntheses and drafted sections that feed into co-authored outputs. We meet weekly to set direction, and I provide detailed asynchronous feedback on writing. The project is part of a broader research program on strategic interaction under AI delegation, so this work continues past the round.

Prerequisites

Strong analytical writing. Working familiarity with current AI capabilities and deployment patterns. Preferably background in international relations, strategic studies, or security policy, through formal study or serious self-study, including familiarity with the deterrence and escalation literature (for example Schelling, Fearon, and the counterforce debates). Arms-control knowledge is helpful for the commitments subproject.

Application question(s)

  1. Link to a writing sample, ideally from a research context (required). 2. (20 min) Pick one historical case where a state pre-delegated use-of-force or escalation-relevant authority to an automated or standing system. What made it stabilizing or destabilizing, and what feature of AI delegation breaks the analogy? (250 words) 3. Optional: name one way in which arms-control verification precedent does NOT transfer to verifying properties of AI agents. (150 words)

Back to quick comparison


Project 18

How Quickly can Middle Powers Build Frontier Compute Capacity?

Open project page | Mentor link

Mentor(s): James Nicholas Bryant

Affiliation(s): Pivotal Research

Research areas: Compute governance | AI strategy | International governance

Practical detailValue
Hours/week8
Team size2-3
Mentor involvementCo-working (5+ hr/week/mentee)
Off-cycleNo

Overview

In the case of a Middle Power Frontier AI coalition, what would be the most effective routes to sufficiently large compute buildout? How quickly could middle powers acquire and operationalise substantial compute? Which constraints determine their progress?

Full description

Brief: Governments increasingly treat domestic access to AI compute as a source of economic and strategic autonomy. Existing research has classified national AI strategies, catalogued sovereign-compute projects and examined the budget and capacity required for a successful buildout. However, there is little research on the best routes to operationalising such a project. This project would theorise and collate the optimal paths a middle power coalition could take to complete a sufficient large buildout as quickly as possible. The analysis would examine the time between announcement, financing, construction, energisation and actual operation across several and also seek to understand common failure modes (such as why projects are delayed). Outputs: A comparative research paper looking at previous compute buildouts, and a policy memo on optimal routes for middle-power compute expansion.

Why this matters

This projects seeks to assess whether a number of AI strategy scenarios (fast-follower, global AI coalition, etc.) are possible on certain timelines, and if so, how this might be achieved. Whether middle powers can build frontier compute capacity, quickly enough, will shape proliferation timelines, influence multilateral coalitions as a counterweight to bipolar racing, and leverage available for access agreements. There is significant commentary on sovereign compute but little work on buildout routes and constraints. The goal is to ground this empirically and help policymakers identify credible pathways.

What you’ll do

Mentees will lead on desk research, writing/drafting, and expert interviews where appropriate. Mentees will have substantial autonomy to pursue additional research directions within the project’s scope. We’ll hold a weekly meeting covering drafts, research taste, and mentee development (career planning, network building), with brief async updates between meetings. My aim is to facilitate a project and devision of labour similar to what you may experience in an AI Safety fellowship such as Pivotal, LASR, or ERA.

Prerequisites

I expect to be interested in mentees from a wide-variety of backgrounds. Required: - Strong analytical writing - Comfortable with self-directed desk research - Able to commit ~10 hours/week for the full programme period, including a ~weekly call. - Broad familiarity with AI governance, with strong insights into why compute is a governance lever Ideal: - Experience with comparative case-study research or structured analysis (e.g. has written a literature review, policy analysis, or thesis chapter using multiple cases). - Strong familiarity with at least one of: datacentre/energy infrastructure, semiconductor supply chains, industrial policy, or project finance. - Quantitative literacy: comfortable building timelines and cost/capacity estimates in a spreadsheet.

Application question(s)

  • Provide a link to one or more writing samples, ideally on AI safety or governance. This may include a paper, blog post, or coursework. If it was co-authored or heavily edited, state your contribution. (No word limit; link only) - Should a potential Middle Power AI coalition pursue a frontier-level compute buildout? Why, and why not? What implications might this have for global AI safety? (200 words max)

Back to quick comparison


Project 19

Normalization of Deviance in AI Development

Open project page | Mentor link

Mentor(s): Emilio Barkett

Affiliation(s): Independent

Research areas: AI strategy | Lab governance

Practical detailValue
Hours/week5
Team size2-3
Mentor involvementDetailed Guidance (2+ hr/week/mentee)
Off-cycleNo

Overview

The normalization of deviance framework — the organizational process by which safety violations become redefined as acceptable through repeated non-disaster — has preceded every major technological catastrophe of the last half-century, yet has never been systematically applied to AI development. This project investigates whether the structural conditions that produced Challenger, Three Mile Island, and the Boeing 737 MAX crashes are present in contemporary AI development organizations, and what that implies for AI safety.

Full description

This project applies the normalization of deviance (NoD) framework, developed by sociologist Diane Vaughan through her landmark study of the 1986 Challenger disaster, to contemporary AI development organizations. The central question is whether the organizational conditions that have reliably preceded major technological disasters are structurally present in AI development today. The normalization of deviance describes the incremental process by which a practice that violates safety norms becomes, through repeated non-disaster, redefined as normal. Critically, this process does not require negligence or malice — it emerges from the ordinary dynamics of organizations under competitive pressure. Prior research has documented this pattern across oil and gas, nuclear power, aviation, healthcare, and rail industries. This project asks whether AI development organizations exhibit the same structural features: production pressure, false assurance from prior success, structural secrecy, and erosion of independent oversight. The project has both a theoretical and an empirical dimension. On the theoretical side, the project develops a systematic mapping of the NoD framework onto the AI development landscape, drawing on the organizational sociology literature and the AI safety literature. On the empirical side, the project investigates whether observable signatures of normalization are present in the public record of AI development — including responsible scaling policy revision histories, public red-teaming disclosures, model evaluation practices, and documented instances of organizational safety culture changes at frontier AI laboratories. The project builds on a workshop paper currently in preparation for NeurIPS 2026, which develops the theoretical argument and case study analysis. The SPAR project would extend this work by developing the empirical component, which requires systematic data collection and analysis beyond the scope of the workshop paper. Mentees with backgrounds in organizational sociology, science and technology studies, AI safety, or empirical social science methods are particularly well suited to contribute to this project.

Why this matters

The AI safety field has invested heavily in technical approaches to risk — alignment research, interpretability, and evaluation frameworks — but has paid comparatively little attention to the organizational dynamics that determine whether those technical efforts translate into safe outcomes in practice. This project addresses that gap by identifying the organizational conditions under which safety infrastructure erodes, even when individuals and institutions are acting in good faith. If the normalization of deviance is occurring in AI development organizations, understanding its mechanisms and early signatures is a prerequisite for interrupting it before a consequential failure occurs. The work contributes directly to the design of more robust safety institutions, governance frameworks, and oversight mechanisms for frontier AI development.

What you’ll do

Mentees will function as independent researchers with guidance. After an initial onboarding period where we establish shared context on the normalization of deviance framework, the organizational sociology literature, and the AI safety landscape, each mentee will: - Select one or more AI development organizations or domains to investigate, identifying publicly available evidence relevant to the four mechanisms identified in the framework: production pressure, false assurance from prior success, structural secrecy, and erosion of independent oversight. - Design a systematic data collection approach drawing on publicly available sources such as responsible scaling policies and their revision histories, model cards, red-teaming disclosures, safety team organizational changes, and public statements from AI developers. - Analyze collected evidence against the normalization of deviance framework, identifying whether and how the characteristic signatures of organizational drift are present in the public record. - Write up findings with the goal of producing a publishable contribution, either as a standalone empirical paper or as a component of the broader research program. I will provide supervision through regular check-ins, feedback on research design and methodology, and guidance on connecting empirical findings to the theoretical framework. Mentees should expect to drive their own projects forward between meetings and to develop their own analytical perspective on the material. This structure rewards initiative, careful reading, and the ability to reason across disciplinary boundaries — the project sits at the intersection of organizational sociology, science and technology studies, and AI safety, and mentees comfortable with that interdisciplinary space will thrive.

Prerequisites

Prerequisites: Familiarity with the AI safety landscape, including basic knowledge of current frontier AI development organizations, evaluation practices, and governance frameworks. Comfort reading and synthesizing academic literature across disciplines, including organizational sociology, science and technology studies, and social science methodology. Strong analytical writing skills, demonstrated through prior academic writing, research papers, or policy documents. Ability to conduct systematic qualitative research, including identifying, collecting, and analyzing publicly available documentary evidence. Genuine intellectual curiosity about the intersection of organizational dynamics and AI safety — this project is not primarily technical and does not require a computer science background, but does require comfort reasoning carefully about complex sociotechnical systems. The following are not required but would be advantageous: Prior exposure to organizational sociology or science and technology studies. Familiarity with qualitative research methods such as document analysis or case study methodology. Experience writing for academic venues in any discipline.

Application question(s)

Question 1 (250 words): The normalization of deviance framework argues that safety failures emerge from organizational dynamics rather than individual negligence, and that safety processes themselves can be captured by the deviations they were meant to prevent. Identify one concrete example from the public record of AI development — a policy revision, an organizational change, a public statement, or a documented shift in evaluation practice — that you think could be interpreted as an early signature of this dynamic. Explain your reasoning, being careful to distinguish between evidence that is consistent with normalization of deviance and evidence that would constitute proof of it. Question 2 (200 words): This project sits at the intersection of organizational sociology and AI safety, two fields with different methodological traditions and different standards of evidence. What do you see as the primary methodological challenge in applying a qualitative organizational framework to AI development organizations, and how would you approach it? Question 3 (link, no word limit): Please provide a link to one or more writing samples from a research or academic context. These need not be published or formally peer-reviewed. We are looking for evidence of careful analytical reasoning, precise use of evidence, and clear written argumentation across disciplinary boundaries.

Back to quick comparison


Project 20

A Design Blueprint for Middle-Power AI Safety Institutes

Open project page | Mentor link 1 | Mentor link 2

Mentor(s): Michał Kubiak; Daniel Polak

Affiliation(s): AI Safety Poland; AI Safety Poland

Research areas: International governance | Technical governance | National policy

Practical detailValue
Hours/week8
Team size3-4
Mentor involvementGuidance (1+ hr/week/mentee)
Location/timezoneNo geographic requirement. A weekly meeting slot spanning European and North American time zones will be agreed with the team.
Off-cycleNo

Overview

Every state outside the US and UK is now told it needs an AI Safety Institute, but there is no design template scaled to a middle power’s resources. This project produces a comparative anatomy of existing frontier-evaluation bodies and a modular, reusable blueprint a mid-sized state could adopt to build credible AI evaluation capacity without duplicating what larger institutes already do.

Full description

Independent capacity to evaluate frontier AI systems for dangerous capabilities is concentrated in a handful of well-resourced institutes (UK AISI, the US Center for AI Standards and Innovation, and a small number of others). Most states have neither the budget nor the technical staff to replicate them. There is no worked-out answer to the question a mid-sized government actually faces: what is the minimum viable institute that contributes something real to catastrophic-risk mitigation? Core question: What should a middle-power AI Safety Institute actually do, cost, and look like — and how does it plug into the international evaluation ecosystem rather than reinventing it? 1. Comparative anatomy. Mentees build a structured comparison of existing and emerging bodies - UK AISI, US CAISI, Canada’s AISI, the EU AI Office, Japan’s AISI, Singapore, and one or two others - coded along common dimensions: legal form and independence, mandate (evaluation vs. standards vs. market surveillance), budget, staffing and technical depth, relationship to government, and participation in international networks. 1. Function mapping. Separate the functions that genuinely require sovereign capacity (evaluating nationally deployed systems, advising government on acute risk) from those better accessed through networks or shared infrastructure. 1. The blueprint. Produce a modular design toolkit - an “AISI-in-a-box” - with budget tiers (a €10–15M lean model vs. a €50M+ model), staffing profiles, statutory-independence options, and a decision tree for what to build domestically vs. what to join. Milestones: weeks 1–4, institution scan and analysis; weeks 5–8, completed comparative dataset and function map (midterm report); weeks 9–12, blueprint drafting and short paper (final report / Demo Day).

Why this matters

Credible, independent evaluation of frontier systems is a bottleneck for catastrophic-risk mitigation: dangerous capabilities can’t be governed if few actors can measure them, and governance triggers (pause, restriction, disclosure) lack legitimacy without independent technical findings behind them. Today that capacity sits in a small number of countries, some of which have shifted rhetoric from “safety” toward “innovation” and “security.” Broadening the base of competent institutes increases the number of independent eyes on dangerous capabilities and strengthens the institutional substrate any future binding regime will depend on. A reusable blueprint lowers the cost for the next states considering this.

What you’ll do

Mentees are the researchers. Each takes ownership of a subset of comparator institutions and, after we agree a shared coding rubric in the first two weeks, is responsible for populating it - reading founding legislation, budget documents, annual reports, and secondary analysis, and where feasible reaching out to people in the field. Beyond data collection, mentees contribute to the analytical work: the function-mapping and blueprint stages are collaborative, and I expect mentees to argue for their own conclusions about what belongs in a minimum-viable institute. Mentees will also co-author the final write-up. Autonomy is high on the descriptive work and more guided on the design synthesis, which we’ll do together.

Prerequisites

Strong reading and synthesis skills across policy, legal, and institutional documents - this is a research-and-writing project, not a technical one. A background in public policy, law, international relations, political science, economics, or a related field, OR demonstrated equivalent research experience. Comfort working with primary sources (legislation, budgets, official reports) and turning them into structured comparative analysis. Enough familiarity with the AI governance landscape to know what a frontier-model evaluation is and why it matters. Prior research experience welcome but not required. No ML or programming background required.

Application question(s)

(a) Pick any one existing AI Safety Institute or equivalent body. In your view, what is the single most important function it performs that a mid-sized country could NOT easily obtain by joining an international network instead - and why? Be specific. (max 300 words) (b) A minister says: “We’ll just give our existing market-surveillance regulator an AI safety mandate - why build a separate institute?” Give the strongest version of the minister’s argument, then your best response to it. (max 300 words)

Back to quick comparison