OpenAI Built an AI-Writing Detector, Then Pulled the Plug Six Months Later
OpenAI launched its own AI Text Classifier in January 2023 and shut it down by July. The reason it failed, and why it still matters, has nothing to do with bad engineering and everything to do with the math behind detection itself.
A Detector From the Company That Made the Problem
In late January 2023, OpenAI released a free tool that promised to tell you whether a piece of writing came from a person or from a machine. It was a plain text box on a webpage: paste in an essay, an email, a cover letter, and get back a rough guess. No login walls, no fine print about accuracy, just a simple classifier trained by the same lab that had just handed the world ChatGPT.
It seemed like the obvious next move. Teachers were panicking, editors were nervous, and OpenAI had more raw material to train a detector on than anyone else alive. If any company could tell its own model's output apart from a human's, surely it would be the one that built the model.
That confidence lasted about six months. In July 2023, OpenAI quietly pulled the tool, citing accuracy too low to be useful. There was no press conference. No successor product. Just a line in a blog post and a dead link where the classifier used to be.
How It Actually Worked
The classifier leaned on two ideas that still show up in detection tools today: perplexity and burstiness. Perplexity measures how predictable each word choice is to a language model — AI writing tends to pick the statistically expected word more often than a person would. Burstiness measures variation in sentence length and rhythm across a passage; human writing tends to swing between short punchy lines and long winding ones, while machine text often settles into a steadier cadence.
Neither idea is unreasonable on its own. The trouble was scale. To get a stable enough reading, the tool needed at least 1,000 characters of text, roughly 150 to 200 words, before it would even attempt a verdict. Anything shorter, including cover letters, short-answer exam responses, or a single paragraph pulled from a longer document, fell outside what the model could reliably judge. That minimum wasn't a quirk of the interface. It was an admission that the underlying signal is faint and only becomes visible once you've collected enough of it.
The Numbers OpenAI Reported Before Shutting It Down
OpenAI had access to its own model's training data, its own output logs, and presumably every internal signal a detector could want. It still couldn't push the catch rate much past one in four, while also mislabeling human writing often enough to matter. If the company that built the text generator couldn't reliably catch its own output, that's not a story about one bad tool. It's a signal about the ceiling on this whole approach.
Why the Math Was Working Against It From the Start
Perplexity and burstiness measure statistical tendencies, not fingerprints. A distracted writer producing flat, repetitive sentences can look 'AI-like' by these metrics. A careful prompt asking for varied sentence length and idiomatic phrasing can push machine output the other way. Neither case involves any actual deception, yet both would confuse a classifier built on averages.
There's also a moving-target problem baked into the whole idea. Every time OpenAI (or anyone else) improved its language models to sound more natural, it simultaneously eroded the very statistical gap its own detector depended on. A detector trained against GPT-3-era text tells you less and less as the models it's checking keep getting better at mimicking human rhythm. OpenAI was, in effect, competing against itself, and losing.
What the Shutdown Actually Signaled
- Detection based on writing statistics has a real, fairly low ceiling — not a bug to be patched, but a limit built into the method.
- A roughly 9% false-positive rate on human writing is not a rounding error. Applied at the scale of a school district or hiring pipeline, that's a lot of real people wrongly accused.
- A 1,000-character minimum means short-form writing was never something this generation of detectors could meaningfully judge.
- No single score, from any vendor, should be treated as a verdict rather than a hint.
What This Means for Anyone Relying on a Detector Today
Detection tools have improved since 2023, and some now combine several signals rather than leaning on perplexity alone. But the core lesson from OpenAI's own experiment hasn't gone anywhere: these tools produce probabilities, not proof, and the gap between AI-typical and human-typical writing keeps narrowing as models get better.
That's why it's worth treating any single detector score as one opinion rather than a final answer, especially when the stakes are a grade, a job application, or a reputation. Running a passage through AI Humanizer Lab's free AI Detector as a second check costs nothing and takes a minute, and it's a reasonable habit before anyone treats one tool's number as the truth.
Make your writing sound human
Humanize AI-generated text in one click with AI Humanizer Lab.
Try for free