False Alarms, Missed Catches: The Two Ways AI Detectors Get It Wrong
AI detectors don't just let disguised AI text slip through — they also flag plenty of human writing that never touched a language model. Here's why both failure modes happen, and how to read a detector score without treating it as a verdict.
A Score, Not a Verdict
Run a piece of writing through an AI detector and you get back a number: 12% AI, 87% AI, whatever the tool's confidence engine produces. It looks precise. It isn't. Under the hood, most detectors measure two statistical properties of a text — perplexity (how predictable each next word is) and burstiness (how much sentence length and structure vary) — and compare the pattern against what language models tend to produce. That's a reasonable signal. It is not proof of anything.
The trouble is that this signal fails in two opposite directions at once, and most people only brace for one of them. They worry about a slick AI paragraph slipping past the scanner undetected. Fewer people worry about the reverse: a human being who wrote every word themselves getting flagged as a machine. Both happen regularly, for structural reasons that have nothing to do with whether the writer actually used AI.
Two Failure Modes, Side by Side
| Failure type | Who gets caught in it | Why it happens |
|---|---|---|
| False positive | Technical writers, formal report authors, non-native English speakers | Disciplined, predictable prose scores low on perplexity and burstiness — the same pattern detectors are tuned to read as machine-generated |
| False positive | Students following a strict rubric or citation format | Rigid structure and repeated phrasing mimic the templated cadence detectors associate with AI output |
| False negative | AI text that's been lightly edited or run through a paraphraser | A few swapped words and reshaped sentences blur the statistical fingerprint the detector was trained to spot |
| False negative | Output from newer, larger models | These models already vary sentence length and word choice more, so the raw text no longer looks as uniform as what the detector learned from |
One of the most cited findings here comes from a 2023 Stanford analysis, widely reported at the time, which found that several popular detectors flagged more than half of TOEFL essays written by non-native English speakers as AI-generated — while essays from native speakers passed through almost entirely clean, as reported. That's not a rounding error in a beta tool. It's a pattern that quietly punishes the plain, grammatically conservative English that non-native writers are trained to produce, and it can follow a student or job applicant into a real consequence: a failed assignment, a rejected application, an accusation that's hard to disprove after the fact.
What Tilts a Score, Regardless of Who Actually Wrote It
- Formal register — reports, technical documentation, academic writing — reads as flatter and more predictable than casual prose, which nudges perplexity down.
- A fixed template or rubric, the kind students are told to follow, produces the same repeated structure a detector is tuned to flag.
- Non-native phrasing tends to favor simpler, more standard sentence construction, which lowers burstiness even though a person wrote every sentence.
- A single light editing pass over AI-generated text — swapping a few words, breaking up a run-on sentence — is often enough to blur the statistical fingerprint a detector was trained on.
- Newer models already vary their own sentence length and vocabulary more than the models most detectors were trained against, so their raw output can look less 'AI-typical' than it used to.
A Few Numbers Worth Holding Loosely (As Reported)
How to Actually Use a Detector Result
- 1Treat the scan as a starting point, not a ruling
A high score means look closer, not guilty. A low score means nothing was flagged, not that the text is clean.
- 2Read the piece yourself
Look for the tells a percentage can't catch: vague claims standing in for real evidence, generic examples that could belong to any topic, statistics with no source attached.
- 3Check the facts and the citations
AI-generated text is prone to confidently wrong details and citations that don't hold up. Verifying a handful of claims often tells you more than any detector score.
- 4Get a second reading
Run the text past a second tool built with a different approach — AI Humanizer Lab's AI detector is a reasonable second data point alongside your own read, especially when the first score is borderline.
- 5Make the call in context
Weigh the score against what you already know — the writer's usual voice, their track record, what's at stake if you're wrong in either direction — before treating a number as a decision.
One Input, Not the Whole Story
None of this makes detectors useless. They're decent at catching the obvious cases — unedited output from an older model, dropped into a document with no changes made. They're far shakier at the edges, which is exactly where most real disputes live: the technical writer whose prose is naturally even, the AI draft someone spent twenty minutes smoothing over.
The honest way to use a detector is the same way you'd use a spell-checker that occasionally flags a real word: fast, worth running, genuinely useful, and never the final word. Pair the scan with an actual read of the piece, and treat the score as one more data point instead of a verdict. Both the writer's reputation and the reader's trust hold up better that way.
Make your writing sound human
Humanize AI-generated text in one click with AI Humanizer Lab.
Try for free