Why AI Detection Scores Need to Show Their Work
A percentage alone doesn't tell you why a detector flagged your writing. Here's what changes once you can see the perplexity and burstiness readings behind the number.
A Number With Nothing Behind It
Picture getting an email that says your essay scored 87% AI-generated. No breakdown, no highlighted sentences, no explanation of what triggered the number. Just a score, sitting there like a verdict. You wrote the thing yourself, and now you're supposed to prove a negative to someone who won't tell you what the machine actually noticed.
This is how most AI detection tools work today. A document goes in, a percentage comes out, and everything in between stays sealed off as proprietary. That's a strange way to run something with real consequences attached — a grade, a job application, a client relationship. The signals detectors rely on aren't secret sauce. They're measurable patterns in text, and there's no good reason a user shouldn't be able to see them.
The Two Patterns Doing the Actual Work
Strip away the marketing and most AI detectors lean on two underlying signals. The first is perplexity, which is a fancy way of asking: how surprised is a language model by the next word in this sentence? Large language models generate text by repeatedly picking a statistically likely next word. The result tends to read smoothly, but it's smooth in a specific, measurable way — low surprise, word after word.
The second signal is burstiness, and it's about rhythm rather than word choice. Human writing tends to be uneven. A short, blunt sentence follows a long, winding one. Paragraph structure wanders. Machine-generated text, left unedited, tends to even that out — sentences that cluster around a similar length and a similar level of complexity, page after page.
Neither signal is proof by itself. A contract clause or a lab report written by a person can score low on both, because dry, formal writing is naturally more predictable and more even than a blog post or a text message. That's exactly why a single blended score is the wrong level of detail to hand someone — it collapses two different questions into one number and throws away the part that would let you sanity-check the result.
Two Signals, Side by Side
| Signal | What it actually measures | Typical human pattern | Typical AI pattern |
|---|---|---|---|
| Perplexity | How predictable each next word is, given what came before | Occasional unexpected word choices, personal quirks, mild inconsistency | Statistically 'safe' word choices, fewer surprises |
| Burstiness | How much sentence length and structure vary across the text | Short sentences mixed with long ones, uneven rhythm | Sentences clustering around a similar length and shape |
When a detector only outputs a score, there's no way to check whether it got thrown off by something ordinary — a non-native English speaker's phrasing, a technical subject with limited vocabulary, or a short piece of writing that doesn't give the model enough signal either way. Students have been accused with nothing but a percentage as evidence. Freelancers have lost work over a number they had no way to contest. None of that requires bad intent from the detector's maker — it just requires a tool that hides its reasoning, applied to a situation where the reasoning is the only thing that matters.
What the Public Numbers Actually Say
How to Read a Detection Report Instead of Just a Verdict
- 1Look at the confidence range, not the headline number
A tool that says 'likely AI-influenced, moderate confidence' is giving you more usable information than one that says '73%' with no context. Treat a lone percentage as the start of a question, not the end of one.
- 2Check which signal actually drove the flag
A high perplexity flag on a technical report means something different from a low-burstiness flag on a personal essay. If the tool won't say which pattern tripped it, you can't judge whether that pattern makes sense for the kind of writing in front of you.
- 3Weigh what the tool can't see
Genre, subject matter, the writer's first language, and whether the text was edited afterward all affect these signals in ways a score can't capture. A transparent report gives you room to factor that in; a bare number doesn't.
Quick Answers
- Does high perplexity always mean AI wrote it? No. Legal writing, technical documentation, and non-native English often score similarly, because the vocabulary is naturally more predictable word-for-word.
- Can burstiness alone catch AI-written text? Not reliably. Some people naturally write in an even rhythm, and plenty of AI output gets manually varied through editing or careful prompting.
- Is a black-box detector automatically wrong? No — but it can't be checked, appealed, or reasoned about, which matters a great deal when the outcome affects a grade or a paycheck.
- Does 'open methodology' mean a detector publishes its source code? Not necessarily. Usually it means showing which signals contributed to a score and why, without requiring anyone to hand over the whole model.
Start With the Reasoning, Not the Number
If a detector's output is a single score, the only thing you can do with it is accept or reject it on faith. If it shows perplexity and burstiness readings alongside the score, you can actually reason about whether the flag makes sense for the specific piece of writing in front of you.
AI Humanizer Lab's free AI Detector was built around that difference. Instead of handing back a bare percentage, it breaks a score down into the perplexity and burstiness signals behind it, so you can see what actually drove the result rather than just trusting it. Running a piece of text through it is a reasonable next step if you want a second opinion you can actually inspect, not just a number to take on faith.
Make your writing sound human
Humanize AI-generated text in one click with AI Humanizer Lab.
Try for free