Universities are relying on AI-detection software to catch cheating. How well do the programs work?

The problem in a nutshell

  • Universities face a surge of written material that may be AI-generated or heavily AI-shaped, and are turning to AI-detector tools to cope.
  • Like other AI tech, these tools are far from perfect — and the stakes for wrongly-accused students are high.

A cautionary tale

  • Lauren Jager, a chemistry undergrad at Idaho State University, saw PhD application portals warn that AI-flagged personal statements would get the entire application disregarded.
  • She had used no AI, but ran her essays through online detectors “just for safety” — they came back at almost 100% AI.
  • She suspects her clean, rule-following writing style (she nearly majored in English) triggered the flags.
  • She rewrote her statement deliberately “less perfect” to score ~30% AI, called it “good enough”, and sent it — and was later accepted to a PhD at the University of Utah.

How detection has evolved

  • Educators vs. cheaters is an ancient arms race (going back to Sumerian cuneiform).
  • Turnitin-style tools detect text similarity against a large corpus of published work — good against direct plagiarism.
  • Cheaters shifted to paying ghostwriters; institutions responded with in-person/timed exams (which have their own downsides: disadvantaging some groups, rewarding rote memorization).
  • LLMs (e.g. ChatGPT) made producing written work cheap and fast. Ironically, though LLMs are built on prior work, their output isn’t caught by plagiarism tools that match full sentences to existing text.
  • Cath Ellis (Western Sydney University) calls it a fundamental change — “the volume of stuff coming through has just rocketed.”

How AI detectors work

  • On-market products include Copyleaks, GPTZero, ZeroGPT, and tools from Grammarly, QuillBot and Turnitin.
  • Many rely on perplexity — a measure of how predictable each word in a sequence is.
  • AI text tends to be more statistically predictable (lower perplexity) → flagged as machine-generated; less predictable phrasing is read as human.

Do they actually work? The evidence

  • 2025 paper on GPTZero (described as the most widely used detector): caught most fully-AI papers with high confidence, but had a ~16% false-positive rate on human writing. Reliability at distinguishing human-authored text is “limited”.
  • 2023 study (OpenAI, Writer, Copyleaks, GPTZero, CrossPlag): better at spotting GPT-3.5 than the more advanced GPT-4; inconsistent on human text, with false positives.
  • The Declaration of Independence problem: Reddit users found the 1776 text often flagged as AI-written. Nature ran it through ZeroGPT and got 95–100% AI results.
  • Pangram Labs (New York) claims a near-zero false-positive rate. Its approach trains on human-written text that’s then AI-rewritten, learning how each new chatbot writes — avoiding perplexity as the sole measure. Independent assessments rate it among the most accurate.
  • William Walters (Southern Illinois University): of 16 tools tested, only 3 performed strongly across both AI and human writing. Newer GPT models (latest public: GPT-5.5) are even better at mimicking human writing.

Why detectors shouldn’t drive high-stakes decisions

  • Mike Perkins (British University Vietnam): “they don’t [work reliably]” for sensitive cases; false-positive concerns make them unsuitable for anything high-stakes for a student.
  • A key trust problem: teachers are used to trusting plagiarism scores, but similarity tools showed the matching text as evidence — AI-detection tools have no such evidence.

Evading detection

  • Hybrid texts break detectors: pure AI text is caught well, but manipulated text isn’t.
  • Students can have another AI rewrite passages or use “humanizer” tools designed to lower AI-detection scores.
  • Detection firms are chasing humanizers, but evasion tools emerge fast — “a huge arms race that doesn’t really help anyone.”
  • Marzena Karpinska (Simon Fraser University) et al. used Pangram on 186,000 articles from 1,500 US newspapers (June–Sept 2025): ~9% detected as partially or fully AI-generated.
  • Useful for large-scale trends, but not proof of any individual’s guilt: “We certainly cannot mass-reject people because of it.”

The bias problem

  • 2023 Stanford study: 7 detectors tested on 91 pre-ChatGPT (pre-2020) TOEFL essays by Chinese students — over half wrongly labelled AI-generated (avg 61.3% false-positive rate).
  • The same detectors accurately classified 88 essays by US students aged 13–14 → strong bias against non-native English writers.

Takeaways

  • AI detectors can reveal broad trends but are unreliable for judging individuals.
  • False positives disproportionately harm careful writers and non-native English speakers.
  • They shouldn’t be used as evidence in high-stakes academic decisions.

Nature 655, 535–537 (2026). doi: 10.1038/d41586-026-01358-2