How Do AI Detectors Work? A Plain English Guide
How AI detectors score text, what perplexity and burstiness mean, why results differ between tools, and how to read a detector report.
The one sentence version
AI detectors measure how predictable your writing is. The more predictable, the more likely a model wrote it.
What detectors actually measure
Every detector works from the same idea: language models pick the most likely next word, so their output is statistically smoother than human writing. Detectors look for that smoothness.
Two metrics come up the most:
Perplexity measures how surprised a language model is by the text. Low perplexity means the words were easy to predict, which is exactly what another model would produce. High perplexity means surprises, which is more typical of a human.
Burstiness measures variation in sentence complexity. Humans write in bursts: a long, winding sentence followed by a short punchy one. AI keeps things steady. Low burstiness is a flag.
How the scoring works
Most detectors run your text through their own language model and calculate a probability score for each sentence (or the whole passage). That score gets mapped to a label:
- Human: the text looks unpredictable enough to be human written
- Mixed: some sections look human, some look generated
- AI: the text is statistically very smooth and predictable
The thresholds vary between tools, which is why the same paragraph can score differently on different detectors.
Sentence level vs. whole document
Some detectors give you a single number for the whole text. More useful ones, including Prosify’s detector, highlight individual sentences so you can see exactly which parts were flagged. This lets you fix only what needs fixing instead of rewriting everything.
Why results differ between tools
Each detector uses a different model, different training data and different thresholds. A sentence that scores 40% AI on one tool might score 70% on another. That’s normal.
What matters more than any single score is the pattern: if multiple detectors flag the same sentences, those sentences are the ones to rewrite.
What detectors can’t do
- They can’t tell you who wrote something, only how predictable it is.
- They struggle with short text (under ~50 words) because there isn’t enough signal.
- They can produce false positives on formulaic writing: cover letters, legal boilerplate, nonnative English, and technical documentation can all read as “smooth” without any AI involvement.
- They’re a snapshot. As models improve, detector accuracy will shift too.
How to read your report
- Look at the overall score first. Under 20% AI is generally safe. Over 60% means most of the text reads as generated.
- Check the highlighted sentences. Those are the ones pulling the score up.
- Rewrite the flagged sentences using the tips in our guide to making AI text sound human.
- Rerun the check. One or two rounds usually gets you below the threshold.
The bottom line
AI detectors are a useful signal, not a verdict. Treat them as a revision tool: run your draft, see what’s flagged, fix those parts, and move on.