How AI Detectors Score Perplexity and Burstiness
These two metrics drive most detectors, and neither is as precise as the tools imply. Here is what they actually measure.
Two Numbers Behind Most Detectors
Most AI detectors, from GPTZero to Turnitin to Originality.ai, present a single confidence score that an AI wrote a piece of text. Behind that one number, the engine is almost always weighing two statistical measures of the text: perplexity and burstiness. Understanding what those two terms actually measure is the fastest way to understand why detectors are sometimes wrong.
Neither metric detects AI directly. They detect statistical patterns in how the words are arranged, and then they assume that certain patterns are more likely to come from a machine than a person. That assumption holds often enough to be useful, and fails often enough to cause real problems, especially for students and non-native writers.
Here is what each metric actually measures, how detectors combine them, and where the whole approach breaks down.
The Two Metrics in Plain Terms
| Metric | What it measures | What a high score suggests |
|---|---|---|
| Perplexity | How predictable each word is to a language model | Low perplexity reads as machine-written; high reads as human |
| Burstiness | How much sentence length and structure vary | Low variation reads as machine-written; high reads as human |
| Combined | Both metrics weighted into one score | A percentage chance the text is AI-generated |
What Perplexity Actually Measures
Perplexity is a measure of how surprised a language model is by the next word in a sentence. If the model finds the text highly predictable, word after word, the perplexity is low. Language models generate text by choosing the most likely next word, so their own output tends to be low-perplexity. Human writing, with its odd word choices and unexpected phrasing, tends to be higher.
A detector translates this into a judgment: low perplexity means the text was probably generated by a machine, because it reads the way a machine would predict. The logic is sound in aggregate but weak on individual sentences. A human who writes very plainly and conventionally produces low-perplexity text too.
This is the root of the false-positive problem. Technical writing, legal writing, and the work of careful non-native speakers all tend toward predictable, conventional phrasing. Detectors read that as low perplexity and flag it, even though a person wrote every word.
Where Burstiness Comes In
- Burstiness measures variation across the whole text, not word by word.
- Human writing tends to mix short and long sentences in irregular bursts, which scores high on burstiness.
- Machine writing tends toward uniform sentence length and structure, which scores low.
- Detectors combine low burstiness with low perplexity as a stronger signal that text is generated.
- A human who edits for variety raises burstiness and reduces the chance of a false flag.
A detector reporting 87% AI is not making a measurement the way a thermometer measures temperature. It is reporting a probability from a statistical model, and the same text can score differently across tools and across versions of the same tool. Treat the number as a rough signal, never as proof, and never as grounds for an accusation on its own.
Why Honest Writing Gets Flagged
The most common reason real writing gets flagged is that it is unusually clean. A student who writes carefully, uses conventional sentence structures, and keeps their grammar consistent is producing exactly the low-perplexity, low-burstiness text that detectors associate with machines. The detector is not lying about the metrics. It is misreading the cause.
The same happens with non-native speakers who have learned to write formally, and with professionals in fields that demand a plain, uniform style. These writers are punished for clarity. The statistical patterns that detectors flag as machine-like are, for them, simply signs of careful writing.
This is why no detector should be the sole judge of whether a person used AI. The metrics are real, but the inference from metric to verdict is where the errors live, and those errors fall hardest on writers whose style happens to be conventional.
What the False-Positive Problem Looks Like
What This Means in Practice
If you write plainly and get flagged, the fix is not to game the detector. It is to add the natural variation that human writing tends to have: a mix of sentence lengths, a few specific examples, and the occasional concrete detail that a model would not default to. This raises burstiness and perplexity without changing your meaning, and it also makes the writing better.
Tools like AI Humanizer Lab can help with this surface layer, smoothing a draft so it carries more natural variation. But the deeper point is to understand what the detector is measuring, so you can read its score with the right skepticism. The metrics are genuine. The conclusions drawn from them are fallible, and a single percentage should never be treated as the final word on whether you wrote your own work.
Make your writing sound human
Humanize AI-generated text in one click with AI Humanizer Lab.
Try for free