AI systems out-persuade expert humans

Core claim

Across four preregistered experiments, frontier AI systems were reliably more persuasive than every class of expert human tested — including world championship debaters and professional canvassers — even when those humans chose their own issues, researched in advance, practiced live, and competed for £1,000 bonuses. The advantage stems from AI’s higher information throughput: constraining AI to human message lengths and writing speeds closed the gap entirely. The effect extends to consequential behavior, with AI nearly 3x more effective than professional canvassers at raising real-money donations.

Study design

  • Four preregistered experiments, conducted 16 October 2025 – 13 May 2026, single-blind and between-subjects.
  • n = 18,978 conversations from 6,923 persuadees, plus 295 human persuaders; 7,218 participants total.
  • Persuadees were randomized in real time via a custom multiplayer platform to converse by text with either a human persuader or an AI. They were not told which.
  • Conversations ran 2–10 turns with strict turn-taking; Studies 1–3 had a median of 7 turns / 14 minutes, Study 4 a median of 5 turns / 10 minutes.
  • Active control: a non-persuasive chat with ChatGPT-4o on a neutral, non-political topic. All effects are reported relative to this.
  • Studies 1–3 outcome: attitude toward one of 10 prespecified UK policy stances (immigration, the two-child benefit cap, the monarchy, assisted dying, a social media ban for under-16s, Ukraine territorial concessions, protest penalties, campus speech, pension age, restitution of historic objects), measured pre/post as the mean of three 0–100 items.
  • Study 4 outcome: proportion of a £1 study bonus donated to Save the Children.
  • AI models used: Claude Opus 4.1 and 4.6, ChatGPT-4o, GPT-5.4, Grok 4.20, Gemini 2.5 Pro. Studies 1–3 used an “information-first” prompting strategy; Study 4 used Claude Opus 4.6 with an “impact-efficacy” prompt.
  • Conducted by researchers at the University of Oxford, the UK AI Security Institute, Stanford, and the LSE.

Study 1 — AI vs. three classes of human persuader

  • Random Laypeople (n = 132 UK-representative Prolific workers; £12/hr, £10 top-10% bonus): persuaded 4.7 pp vs. control. AI exceeded them by 8.2 pp (95% CI [6.7, 9.7], p < .001).
  • Selected Laypeople (n = 87): drawn from the top ~10% of a separately preregistered four-round elimination tournament in which 1,154 Prolific workers competed across 9,475 conversations to persuade 2,634 persuadees. Paid £24/hr with a £1,000/750/500/250 prize pool plus per-conversation bonuses. Persuaded 7.2 pp. AI exceeded them by 5.6 pp (95% CI [4.1, 7.1], p < .001).
  • Elite Debaters (n = 56): all had reached at least the semifinals of a major international competition, including 4 world champions and 11 continental champions, mean 8.9 years of competitive experience. They selected the policy stances themselves by deliberative vote, were paid £30/hr for ~8 hours of advance research, had 21 days’ notice, and competed for the same prize pool. Persuaded 8.3 pp. AI still exceeded them by 4.6 pp (95% CI [3.1, 6.1], p < .001).

Study 2 — Can anything close the gap?

Coaching the debaters

  • 43 returning Elite Debaters received a coaching tool built around the AI that had beaten them: they could chat with it, see how it had been prompted, review their own Study 1 transcripts annotated with attitude shifts, and see what the AI would have said at any point in any past transcript.
  • Engagement: two 4-hour sessions, ~16 hours cumulative paid preparation, ~25 live practice conversations.
  • Their behavior changed measurably — +9.8 words per message (+19%, p = .009) and +1.6 fact-checkable claims per conversation (+54%, p < .001).
  • Coaching did not significantly improve persuasiveness: +1.0 pp (95% CI [−0.5, +2.5], p = .20).
  • Numerically it produced the largest human effect observed (9.7 pp) and narrowed the gap from +6.0 pp to +4.1 pp — narrowed but not closed.

Constraining the AI

  • Baseline throughput disparity: Elite Debaters averaged 54 words per reply at ~95 seconds; AI averaged 294 words with sub-second latency.
  • A Constrained AI condition capped message length and response latency at debater-calibrated levels (51 words, 92 s), implemented via a lagged, adaptive daily-target design since prompting alone could not fix length reliably.
  • Result: AI’s advantage over Coached Elite Debaters collapsed from +4.1 pp to a non-significant 0.0 pp (95% CI [−1.7, +1.6], p = .96). Unconstrained AI beat Constrained AI by +4.2 pp.

Robustness

  • Of 318 per-persuader estimates (275 unique humans), none exceeded the pooled AI estimate. The highest individual estimate was 9.9 pp (a Coached Elite Debater), still 4.0 pp below AI.
  • Predicted probability that a randomly drawn new persuader from any class would exceed AI: below 0.1% (largest tail: Selected Laypeople, 0.09%).
  • AI’s advantage held across all 10 policy issues (3.0–9.6 pp) and in 46 of 49 persuadee subgroup levels (3.3–10.0 pp).
  • Two moderators survived correction: pre-treatment attitude (8.7 pp at the 25th percentile vs. 3.7 pp at the 75th) and pre-treatment issue knowledge (7.0 vs. 5.4 pp).

Why AI wins: the fact-density account

  • The hypothesis: AI’s throughput advantage lets it pack conversations with more fact-checkable claims, which prior work links to persuasive impact.
  • Prediction 1 confirmed — constraining AI selectively suppressed the informational partner ratings: perceived argument strength (−11.8 pp) and “I learned a lot” (−11.8 pp) fell most, while felt-understood (−6.8 pp) and enjoyment (−6.4 pp) moved about half as much. Human-likeness moved in the opposite direction (+7.5 pp).
  • Prediction 2 confirmed — constraining AI cut fact-checkable claims from ~37 to ~12 per conversation, and across all human and AI conditions this measure strongly predicted impact (R2 = 0.89 overall, 0.89 within humans, 0.90 within AI).
  • Controlling for log fact-checkable claims, the AI-vs-human coefficient became statistically indistinguishable from zero (β = −0.9 pp, p = .38 excluding Constrained AI; +1.4 pp, p = .07 in the full sample) — consistent with fact density accounting for much of the advantage.

Study 3 — Professional canvassers

  • 19 canvassers from a UK firm, median ~10,000 career conversations, paid £140/hr, given the issues 7 days in advance, competing for the same prize pool.
  • Persuaded 6.9 pp. AI exceeded them by 5.9 pp (95% CI [4.3, 7.5], p < .001).
  • Adding Study 3 to all prior models left every substantive conclusion unchanged.

Study 4 — Real-money charitable giving

  • Ran with AppcoUK, which had operated real Save the Children fundraising from 2016 to 2023, raising £824,297 from 22,583 donors. 18 of its canvassers competed against Claude Opus 4.6.
  • Persuadees could donate any portion of a £1 study bonus.
  • AI: +17.2 pp of the bonus vs. control. Canvassers: +6.4 pp (p = .048). AI exceeded canvassers by +10.8 pp (95% CI [+6.4, +15.1], p < .001) — nearly 3x more effective.
  • The advantage appeared on both margins: share donating anything (+6.0 pp) and average donation among donors (+12.9 pp).
  • On a 14-item battery covering seven preregistered donation mechanisms, AI was rated higher than canvassers on all 7 (and on all 14 individual items), largest gaps on implementation intentions (+15.0 pp), commitment escalation (+12.6 pp), and impact-efficacy information (+10.3 pp).
  • Notably, AI was prompted to pursue only the impact-efficacy strategy, yet outperformed canvassers on the six mechanisms it was not instructed to use.

Discussion — implications

Consolidation of influence

  • Power could flow to whoever can most readily access and deploy the most capable systems — large corporations, political campaigns, nation states — deepening existing imbalances in who can sway the public.
  • Even where access is perfectly equal, power could flow to the suppliers of persuasive AI, who decide which positions their models will and will not argue for.

Countervailing possibilities

  • Cheap, widely available persuasion could help under-resourced actors — pro se litigants, public defenders, small charities, grassroots activists — narrowing gaps in access to justice.
  • Because the edge derives from information volume, AI persuasion could leave citizens better informed — but this depends on accuracy, which varied widely across models (some more accurate than the human conditions, others far less).
  • The net effect turns not on AI’s truthfulness alone but on its truthfulness relative to the persuasion it displaces, much of which is itself selective or misleading.

Possible brakes

  • Reach is gated: authentication and verification systems on digital platforms may prevent AI from reaching its targets, and securitization is likely to increase.
  • Recognition may not help: labelling messages as AI-generated has not been found to reduce persuasive impact.
  • Quantity may matter more than quality: variability in exposure typically exceeds variability in per-exposure persuasiveness, and pre-existing attitudes plus limited attention still constrain durable opinion change at scale.
  • The conditions that made AI most persuasive — sustained text engagement lasting a median of 14 minutes — are demanding to reproduce outside a paid-survey context.

Stated limitations and future directions

  • All conversations were text-based; persuasiveness in audio, video, or face-to-face settings remains uncertain (though recent studies find only small bonuses for video over text).
  • The behavioral outcome was consequential but low-stakes (£1); effects may differ for vote choice, large or recurring donations, or public-health compliance.
  • Translating per-conversation effects into net societal impact requires knowing whether and how often AI persuasion displaces other exposure.

Notes

  • Preprint, under review; arXiv:2606.16475v1 [cs.CY].
  • Ethics approval from the Oxford Internet Institute’s Departmental Research Ethics Committee (ref. 2255223) and the UK AI Security Institute’s Research Assurance Board. All four studies plus the selection tournament were preregistered on the Open Science Framework before data collection.
  • Analysis in R 4.4.3 using lme4 and emmeans; linear mixed-effects models with crossed random intercepts for persuader and persuadee.
  • Data and code: github.com/kobihackenburg/AI-out-persuades-experts, archived at Zenodo (doi:10.5281/zenodo.20629034).
  • Post-treatment attrition ranged 3.5–5.0%; it differed by condition in Studies 1–2, but Lee bounds on the headline contrasts remain strictly positive.