What Actually Moves Perplexity and Burstiness — and What's Just Noise
Perplexity and burstiness are the two numbers behind most AI-detection scores. Here's a mechanics-first look at what genuinely shifts them, what only feels like it should, and why the score is a probability, not a verdict.
Two numbers, one guess
Most AI detectors boil a document down to two measurements before they ever print a score. Perplexity asks how surprised a language model is by your word choices — low surprise, low perplexity, more "AI-like." Burstiness asks how much your sentence lengths and rhythms swing from one line to the next — flat and even reads as machine-made, jagged and uneven reads as human.
Neither number is measuring truth. Both are measuring resemblance to the output of models that were trained to pick the statistically likely next word. Once you know that, you can stop guessing at fixes and start looking at what actually pushes those two numbers around.
What moves the needle vs. what just feels like it should
| Change you make | Effect on perplexity/burstiness | Why |
|---|---|---|
| Swapping words for synonyms | Little to none | The sentence shape and predictability stay the same — a thesaurus doesn't touch structure. |
| Adding a concrete number, name, or detail | Raises perplexity | Specific facts are less predictable than generic claims, so the model assigns them lower probability. |
| Mixing short and long sentences on purpose | Raises burstiness | Detectors measure variance in sentence length directly; uniform length is the signal they're built to catch. |
| Rewriting in a more casual, conversational register | Often lowers scores, inconsistently | Anecdotal — conversational phrasing seems to break up predictable patterns more than plain paraphrasing, but results vary by tool and topic. |
| Passive-to-active voice conversion | Marginal | Changes surface grammar, not the underlying word-probability distribution. |
| Adding transition words (moreover, furthermore) | Can lower the score further | These are exactly the connective tissue high-probability text leans on. |
| Formal, technical, jargon-heavy writing by a human | Can still score low | Technical prose is often genuinely predictable — the model isn't wrong that it's low-surprise, it's wrong that low-surprise means AI. |
In informal side-by-side checks, rewriting a paragraph in a looser, more conversational voice tended to move detector scores more than swapping in synonyms ever did — even when the synonym pass touched more words. This isn't a controlled finding, and it doesn't hold the same way across every detector or every topic. Treat it as a lead worth testing on your own text, not a rule to lean on.
Specificity is the cheapest lever you have
Perplexity climbs when your next word is hard to guess. Generic writing is easy to guess — "the company saw significant growth" is a sentence a model could have written in its sleep. "Revenue went from $2.1M to $3.4M in the back half of the year" is not, because a language model has no way to predict that exact figure from context alone.
This is why detail-heavy writing tends to score as more human than vague writing, independent of who actually wrote it. Names, dates, numbers, brand-specific details, and idiosyncratic examples all raise the surprise factor. It's also why a student's first draft full of hedge words and abstractions can flag even when a human typed every letter — the ideas themselves are predictable, not the typing.
Where specificity tends to hide in a draft
- Swap a general claim for the one data point that supports it
- Replace "experts say" with the actual source or study
- Trade a category ("a popular tool") for the actual name
- Add a concrete example instead of describing a type of example
- Let one sentence carry an opinion or judgment a model wouldn't default to
Building burstiness on purpose
- 1Read your paragraph lengths out loud
If every sentence takes about the same breath to say, that's the flat rhythm burstiness metrics are built to flag.
- 2Cut one sentence down to three or four words
A short, blunt sentence next to longer ones creates the variance detectors measure directly.
- 3Let one sentence run long, with a subordinate clause or two
Not padding — an actual layered thought, the kind that comes from explaining a real exception or caveat.
- 4Re-check the mix
You're not aiming for a formula. You're aiming for the kind of uneven rhythm that shows up when someone writes the way they'd actually talk through an idea.
The limit that matters most
None of this makes a detector score a verdict. It's a probability estimate built on patterns from training data, and formal human writing — a legal brief, a lab report, a dense academic paragraph — can land in the same low-perplexity, low-burstiness range as machine output simply because careful, formal prose is also predictable. The detector isn't lying in that case. It's just measuring a property that human writing and AI writing can both share.
That's the honest way to use these tools: as a directional signal on one piece of text, not a lie-detector test on a person. A single score tells you how one passage compares to typical model output right now, on one detector's model of "typical." It won't tell you who wrote it.
If you want to see whether a specific change actually shifted anything, run a section through AI Humanizer Lab's free AI Detector before you edit it and again after. Comparing the two scores on the same passage tells you more than reading a checklist ever will.
Make your writing sound human
Humanize AI-generated text in one click with AI Humanizer Lab.
Try for free