5 Things AI Detectors Actually Measure In Your Text
AI detectors don't read for meaning — they run a handful of statistical checks on word choice, rhythm, and punctuation. Here's what those checks actually look for, with real examples.
Five Things A Detector Actually Checks
Feed a paragraph into an AI detector and it doesn't know anything about who wrote it or why. It runs a set of statistical checks against the text and returns a percentage. Some of those checks are more useful than others, and once you know what they are, the whole process stops feeling like a black box.
Most detectors lean on five signals: how predictable each word is, how much your sentence lengths vary, which specific words show up too often, whether your style stays consistent the way one person's writing usually does, and how you punctuate. None of these proves anything on its own. Together, they're what's driving the score you see on screen.
The Five Signals, At A Glance
| Signal | What It Looks At |
|---|---|
| Perplexity | How predictable each next word is, given everything before it |
| Burstiness | How much sentence length and rhythm swing across a passage |
| Token distribution | Whether specific words or phrases appear far more than a human sample would produce |
| Stylometry | Whether sentence structure and word habits stay consistent, the way one person's writing usually does |
| Punctuation patterns | How marks like em dashes, semicolons, and lists get used compared to typical human habits |
Perplexity: Guessing The Next Word
Perplexity measures how surprised a language model is by each word in a sequence. Low perplexity means the next word was easy to guess. High perplexity means it wasn't.
Take "The cat sat on the mat." Every word lands exactly where you'd expect, so a model scores it as low perplexity. Now compare it to "The cat sat on the committee." Same grammar, same rhythm, but "committee" breaks the pattern, and perplexity spikes right there.
AI-generated text tends to sit in that low-perplexity range for an entire passage, because the model is built to pick the statistically likely next word most of the time. Human writing drifts into odd word choices, tangents, and phrasing that no predictive model would rank as the obvious move.
Burstiness: The Rhythm Of Your Sentences
Burstiness looks at variation, not average length. Human writing tends to burst: a long, winding sentence, then a short one. Two medium ones after that. Then a fragment. Left on its own, AI output often settles into a narrower band of sentence lengths, one after another, almost like a metronome.
This is part of why a lightly edited AI draft can still get flagged even after someone swaps in different vocabulary. Changing words doesn't change the sentence-length pattern sitting underneath them.
Token Distribution: The Words AI Reaches For
- Delve — a transition word that shows up constantly in AI drafts and rarely in casual human writing
- Tapestry — a stand-in for "mix" or "combination" that's almost always a tell when it appears
- Boasts — a formal substitute for "has" or "includes" that models default to
- Underscores / highlights — used to introduce a point instead of just stating it plainly
- Testament to — a stock phrase standing in for "shows" or "proves"
Frequency Is the Signal, Not the Word
None of these words is banned, and people use "delve" now and then without help from a model. The signal isn't the word itself — it's frequency. How often it turns up relative to everything else on the page is what a detector is actually counting.
Stylometry: Your Own Fingerprint
Stylometry predates AI detection by decades. It's the same method used to analyze function-word frequency and sentence construction when scholars settle disputes over anonymous or contested authorship. Applied to AI detection, it checks whether a passage reads like one person writing under one consistent set of habits, or like something stitched together from a source with no fixed habits of its own.
That makes stylometry harder to fake than perplexity or word choice. It isn't about dodging a handful of red-flag words — it's about whether the whole piece holds together as one person's writing.
Punctuation Patterns: Small Marks, Big Tells
Detectors also track punctuation habits. AI writing has a well-documented preference for the em dash, for semicolons linking two related clauses, and for tidy three-item lists that resolve into a rule of three. Human punctuation is messier: commas doing the work a semicolon should, dashes used once and then abandoned for a page, fragments with no punctuation flourish at all.
OpenAI retired its own AI-text classifier in 2023 after reporting that it correctly identified only about 26% of AI-written text as AI-written, alongside a real false-positive rate on human writing. Researchers with ties to MIT and other institutions have since published work cautioning against leaning on any single detector for decisions with real stakes attached, like a grade or a hiring call. A detector score is a hint worth investigating, not a verdict.
What To Do With That Hint
If a detector flags something you wrote yourself, it doesn't mean you're lying about authorship. It usually means your prose runs unusually even in rhythm or leans on a few overused phrases without you noticing. AI Humanizer Lab's free AI Detector breaks a scan down signal by signal, so instead of one flat number, you can see whether it's the punctuation, the word choice, or the sentence rhythm doing most of the flagging — and fix the actual thing instead of guessing at it.
Make your writing sound human
Humanize AI-generated text in one click with AI Humanizer Lab.
Try for free