
How AI detectors work is simpler than the marketing suggests. A detector does not see your ChatGPT history, your keystrokes, or your intentions. It reads the text you paste, extracts statistical patterns, and estimates how closely those patterns match machine-generated writing.
Search the phrase and you will mostly find two definitions — perplexity and burstiness — then a warning that scores are imperfect. That is true, but incomplete. Commercial tools wrap those signals inside classifiers with different training data and thresholds. Separately, copy-paste from ChatGPT can leave invisible Unicode that style-only explainers never mention.
This guide walks through the full stack.
Key Takeaways
What a detector actually does
An AI detector is a classifier. Pipeline in plain terms:
- You submit a passage.
- The system tokenizes it and computes features (predictability, rhythm, phrase habits, sometimes watermark checks).
- A model trained on labeled human and AI samples maps those features to a probability, label, or highlighted spans.
- You see a score.
It does not:
- Prove who wrote the draft
- Match text against a plagiarism database (that is a different check)
- Distinguish “brainstormed with AI” from “pasted a full essay”
- Stay reliable under roughly 150–200 words
If you already have a percentage and need to interpret it, read how to read AI detector scores.
Perplexity: how predictable is the next word?
Perplexity measures how surprised a reference language model is by your sequence. Low perplexity means each word was a highly probable continuation — the path models prefer when they generate fluent text.
| Signal | AI drafts often look | Human drafts often look |
|---|---|---|
| Perplexity | Lower — safe, high-probability words | Higher — odd verbs, hedges, local references |
| Feel | Smooth, polite, statistically “unsurprising” | Bumpier word choice |
Limits: short samples are noisy. Formal academic prose and careful ESL writing can score low perplexity with no AI involved. Perplexity is one feature, not a verdict.
Burstiness: how much does rhythm vary?
Burstiness measures variation in sentence length and complexity across a passage. People mix short punches with long explanations. Model output often keeps an even cadence — claim, elaboration, transition, repeat.
Detectors treat low burstiness as machine-like. Identical paragraph templates across an essay raise flags even when vocabulary looks fine.
The same caveat applies: polished student writing can look low-burstiness honestly. See AI detector accuracy and false positives for non-native writers.
Classifiers: why the same essay scores differently
Most products do not stop at hand-tuned perplexity rules. They feed features into a supervised classifier (often a transformer fine-tuned on millions of paired samples). Training data defines “AI-like” for that vendor. One tool may overweight older GPT blogs; another may include Claude, Gemini, and student essays differently.
That is why the same paragraph can be 18% on one site and 67% on another. Different features, different thresholds, different definitions of suspicious.
Other signals vendors may combine:
- N-gram / stock-phrase frequency — overused transitions and AI “tells”
- Stylometry — function-word habits, syntactic complexity
- Zero-shot likelihood probes — DetectGPT-style checks of how text sits on a model’s probability surface
- Statistical watermarks — only when a provider embeds a generation-time signature
None of these is proof alone. Together they produce a probability.
Two kinds of “watermark” people confuse
Statistical watermarks
Some model providers can bias token selection toward a detectable pattern. Specialized checkers look for that signature. Generic third-party detectors may ignore it and rely only on style features. Not every model embeds one.
Invisible Unicode on copy-paste
A different layer shows up when you copy from ChatGPT into Docs or Word:
- Zero-width spaces (U+200B)
- Word joiners and similar control characters
- Odd non-breaking spaces that never appear on screen
You cannot see them. They are not proof of cheating. They are paste residue. They can still affect scanners, word counts, and integrity workflows — and most “how detectors work” articles never mention them.
PassMyEssay is a ChatGPT watermark remover for that paste layer: strip invisible characters, flag stock AI phrases, then rewrite with meaning locked. Cleanup and clearer writing — not a promise to bypass Turnitin or any institutional checker.
What detectors get right and wrong
Useful: highlighting sections that sound generic, evenly paced, or full of template transitions — the same places a careful reader would ask for evidence.
Wrong: treating probability as guilt; over-flagging formal and non-native English; scoring tiny snippets as if they were full essays; implying a single number settles authorship.
For the difference between database plagiarism matching and AI-style detection, see AI plagiarism vs AI detection.
A practical workflow that respects the tech
- Draft under your course AI policy — brainstorming help is not the same as submitting chat prose.
- Clean paste from ChatGPT at PassMyEssay so invisible Unicode does not sit in your master file.
- Revise flagged or flat sections: add named evidence, vary sentence length, cut empty openers.
- Self-check with a detector only as feedback — never as a green light for policy violations.
- Keep process files (notes, outlines, earlier drafts) if a score is questioned.
Section highlights beat whole-document percentages. Mixed drafts are common: your thesis may be original while one body paragraph still carries machine-smoothed cadence. Revise the spans that look like you did not write them. Leave the rest.
Teachers, LMS tools, and policy
Institutions may run Turnitin, GPTZero, or other checkers through the LMS. Policy weight sits on top of probabilistic tech. Know your syllabus before optimizing for any score. If you revise AI-assisted drafts when allowed, the goal is ownership: claims you can defend aloud, sources you opened, and a voice that matches your other work.
Bottom line
AI detectors work by measuring how predictable and uniform your text looks — perplexity, burstiness, classifier features, and sometimes watermark signals. They estimate resemblance to machine output. They do not know the truth about your process.
The gap left by most ranking pages is the clipboard layer: invisible characters and stock ChatGPT phrasing that survive a skim. Clean that first, then revise for specificity.
Paste your draft into PassMyEssay to strip hidden characters and polish machine-sounding lines — free to try, meaning locked, no card required. Then add your examples and submit writing you can defend.
Keep Reading
Related guides
AI Detector False Positives: Why Honest Essays Get Flagged
What AI detector false positives are, why formal and ESL writing trips classifiers, and a process-first plan if Turnitin or GPTZero flags your work.
AI Detector for Essays: What to Expect Before You Paste
SERP for AI detector for essays is paste-box tools. Here is what those checkers actually score, when false positives hit, and how to clean ChatGPT watermarks first.
Why Find-and-Replace Fails on Text You Pasted From AI
Your search finds nothing on a word that is visibly on screen. Here is the invisible character causing it, how to confirm it, and how to fix the whole document at once.
AI Writing Workflow for Research Papers: Claim–Evidence Pipeline
A practical AI writing workflow for research papers—notes matrix, narrow section prompts, citation open-or-cut audit, methods lock, and ChatGPT watermark cleanup on PassMyEssay.
Make your draft clearer
Use PassMyEssay to rewrite AI-assisted text responsibly, check weak sections, and keep your meaning intact.
Try PassMyEssay