How AI Detectors Work: Scores, Signals, Limits, and Responsible Use
Learn how AI detectors classify text, what their scores mean, why errors happen, and how to use results responsibly with human review.
AI detectors do not watch a document being written and they do not uncover a hidden “AI” label in ordinary text. They analyze the submitted words and estimate whether their statistical patterns resemble text produced by language models. The result is a classification signal—not direct proof of authorship, intent, or misconduct.
Quick Answer
Most AI-writing detectors use one or more of four approaches: a classifier trained on human and machine-written examples; probability-based signals that measure how expected each word is; comparisons between a passage and small rewrites of that passage; or an embedded watermark when the generating system supports one. Commercial tools may combine several signals in an ensemble.
Every approach has limits. Accuracy changes with the model, language, subject, text length, editing history, and decision threshold. That is why an AI score should begin a review, not end it.
Key Takeaways
- An AI detector predicts which class a passage resembles; it does not reconstruct the writing process.
- Perplexity and burstiness explain useful ideas, but they are not a complete description of every modern detector.
- A percentage can represent confidence, the share of qualifying text flagged, or a proprietary score. Check the tool’s definition before interpreting it.
- False positives and false negatives are both possible.
- Short, translated, heavily edited, formulaic, or out-of-domain text can be harder to classify.
- High-stakes decisions require corroborating evidence and human review.
What Does an AI Detector Actually Predict?
An AI detector is a classification system. It receives text, extracts features or representations, and returns a label or score such as “likely human,” “likely AI,” “mixed,” or “uncertain.” The precise target differs by product.
Some tools estimate the probability that a passage belongs to an AI-generated class. Others estimate the portion of eligible prose that resembles generated writing. Still others turn several internal measurements into a normalized score. These outputs may look similar on screen, but they are not interchangeable.
That distinction matters. A result of 70% does not automatically mean that exactly 70% of the words were written by AI. Nor does it prove that a particular model created the passage. It means only what that detector’s current documentation says it means under its own model and thresholds.
The Main Ways AI Detectors Work
1. Supervised classifiers
A supervised detector is trained on labeled examples: text known to be human-written and text generated by one or more language models. During training, the classifier learns combinations of patterns that help separate those categories. A modern system may use a deep neural network rather than a simple checklist of visible writing habits.
The advantage is adaptability: developers can retrain a classifier on newer models, languages, and domains. The limitation is distribution shift. Performance measured on the training or test set may not transfer cleanly to a different model, subject, age group, language background, or editing workflow.
2. Token probability and perplexity
A language model assigns probabilities to possible next tokens. Perplexity summarizes how surprising a sequence is to a particular model. When the next words are consistently easy for that model to predict, perplexity is lower; when choices are less expected, it is higher.
Earlier detection systems often used predictability as a direct signal because generated text can cluster around high-probability word choices. GLTR, a research tool introduced in 2019, visualized how frequently a passage used words from a model’s most probable choices. In a human study, its visual cues improved participants’ detection rate for the tested samples.
But low perplexity is not an AI fingerprint. Clear technical prose, template-based writing, simple language, and some non-native English writing can also be predictable. The score also depends on which model calculates it. A passage can look unsurprising to one model and less predictable to another.
3. Variation, sometimes called burstiness
Burstiness describes variation across a document—for example, changes in sentence length, structure, vocabulary, or local perplexity. Human writing may vary more from sentence to sentence, while generated prose can be unusually even.
This is an explanatory concept, not a universal rule. GPTZero’s current support documentation says it stopped directly using perplexity and burstiness in autumn 2023 after moving to a deep-learning architecture. Modern detectors may learn more complex representations that are not reducible to two visible metrics.
4. Probability curvature and perturbation
Some research methods compare a passage with several small rewrites. DetectGPT, presented at ICML 2023, is based on the observation that model-generated passages can sit in characteristic regions of a language model’s probability function. The method perturbs the passage and compares its log probability with the probabilities of those variations.
This can detect a model’s statistical signature without training a separate binary classifier. However, it assumes access to an appropriate model and remains sensitive to domain, source model, text length, and rewriting. It is one method family, not a definitive test for every real-world document.
5. Watermarks and provenance signals
A watermark is different from style-based detection. The generating system deliberately biases token selection according to a hidden rule, and a compatible detector later tests for that pattern. Research by Kirchenbauer and colleagues showed how an invisible statistical watermark could be embedded during generation and detected in a sequence of tokens.
Watermarks can provide stronger provenance when the generator participates and the signal survives. They are not automatically present in ordinary AI text, they do not identify output from systems that never applied the watermark, and rewriting or translation may weaken them. The ACM’s policy statement also cautions that currently proposed watermarks can be affected by editing and manipulation.
How Training Data and Thresholds Change the Result
A detector needs a decision rule. Imagine a model that returns a score from 0 to 1. The developer must choose where “human,” “uncertain,” and “AI” begin. Moving that threshold changes the balance between two errors:
- False positive: human-written text is classified as AI-generated.
- False negative: AI-generated text is classified as human-written.
A stricter threshold may reduce false accusations but miss more generated text. A looser threshold may catch more generated text but flag more human work. There is no threshold that removes both errors across every population and use case.
Training data creates another boundary. A detector trained on long English essays from a small set of models may behave differently on marketing copy, code, poetry, translated prose, short answers, or output from a newer model. Independent evaluation therefore needs representative data, transparent definitions, and separate reporting of false-positive and false-negative rates—not a single headline accuracy number.
What Does an AI Detector Score Mean?
Start with the label beside the number, then read the product’s current documentation. Three common interpretations are:
| Score type | What it may describe | What it does not prove |
|---|---|---|
| Classification confidence | How strongly the model favors one category under its calibration | The exact percentage of AI-written words |
| Qualifying text share | The portion of eligible prose classified as likely generated | Who used a tool, when, or with what intent |
| Composite score | A product-specific combination of several signals | A score that can be compared directly with another detector |
HumanizeAI’s AI Detector documentation uses a color-coded meter to describe how human-like or AI-like submitted text appears. Use the displayed interpretation for that scan, and do not combine it mathematically with a plagiarism similarity score. The two systems answer different questions, as explained in AI Detector vs. Plagiarism Checker.
You can scan a complete passage with HumanizeAI’s AI Detector to see its current classification. Keep the original file and treat the output as one diagnostic signal alongside your writing history and review context.
Why Text Length, Language, and Editing Matter
Text length
Short text contains fewer observations. A greeting, caption, thesis statement, or single paragraph may not provide enough context for stable classification. Minimum-length requirements vary by tool; Turnitin’s current AI Writing Report, for example, requires at least 300 words of qualifying long-form prose and says shorter submissions are likely less accurate.
Language and writing background
Performance can vary across languages and writer populations. A 2023 study in Patterns tested seven detectors on 91 TOEFL essays by non-native English writers and 88 U.S. eighth-grade essays. The detectors misclassified the TOEFL set far more often. This result does not establish one universal error rate for all tools, but it demonstrates why language background and dataset fit matter.
Editing, paraphrasing, and mixed authorship
A document can combine human ideas, generated sentences, grammar correction, translation, and manual revision. The final text may not fit a binary label. Research has also shown that paraphrasing can reduce detector performance. That does not make evasion responsible; it shows that final-text classification has limited visibility into how the document was produced.
Domain and format
Academic abstracts, legal clauses, customer-support templates, code, bullet lists, and creative prose follow different conventions. A system validated on one domain should not be assumed to perform equally in another. Even a vendor’s own report may exclude non-prose sections because its model was not designed for them.
How False Positives and False Negatives Happen
Human and model distributions overlap. Language models learn from human writing, while people often write in predictable, conventional ways. As generative systems improve, a detector must distinguish increasingly similar distributions.
Sadasivan and colleagues formalized this problem and showed how detector performance is constrained as machine and human text distributions become more similar. Their experiments also found that paraphrasing could sharply reduce detection performance. Separately, Weber-Wulff and colleagues tested 14 detection systems on human, generated, translated, manually edited, and paraphrased documents and concluded that the tested tools were not accurate or reliable enough to stand alone in academic misconduct decisions.
The practical lesson is balanced: a detector can still be useful for triage, review, and quality assurance, but its output must be interpreted within the conditions it was designed for.
A Responsible Decision Framework
| Situation | Use the score for | Do not use it for | Next evidence |
|---|---|---|---|
| Writer checking their own draft | Finding passages worth reviewing | Guaranteeing how every other detector will respond | Source check, clarity edit, policy review |
| Editor reviewing published content | One QA signal | Replacing fact-checking or plagiarism review | Sources, originality, author disclosure |
| Teacher reviewing an assignment | Starting a conversation under policy | Automatic accusation or penalty | Drafts, notes, version history, student explanation |
| Employer reviewing work | Flagging a need for process review | Determining intent from a percentage | Brief, revisions, approved-tool policy, subject knowledge |
- Define the question. Are you reviewing possible AI assistance, source overlap, factual quality, or policy compliance? These are different tasks.
- Check input eligibility. Confirm the tool supports the language, format, and minimum length.
- Read the score definition. Determine whether the number means confidence, text share, or a composite score.
- Review highlighted passages. Look for a coherent pattern rather than reacting to one sentence.
- Gather process evidence. Drafts, notes, citations, document history, and disclosure records can speak to authorship more directly.
- Apply the relevant policy. Permitted grammar help, translation, ideation, and generated drafting may be treated differently.
- Use human judgment. Turnitin’s own guidance says its AI report may be wrong and should not be the sole basis for adverse action.
For a practical response to an unexpected flag, continue with AI Detector False Positives: Causes, Evidence, and What to Do. If your concern is specifically about one vendor, use the focused guide explaining why ZeroGPT may say you used AI.
Frequently Asked Questions
Can AI detectors be wrong?
Yes. They can produce false positives and false negatives. Error rates vary with the detector, threshold, data, language, model, domain, length, and editing history.
What does an AI score mean?
It depends on the product. It may describe classification confidence, the share of qualifying prose flagged, or a proprietary composite score. Read the current documentation rather than assuming it is the literal percentage of words written by AI.
How much text is needed?
There is no universal minimum. Longer coherent prose generally provides more evidence, but each detector sets its own requirements. Do not draw a serious conclusion from a sentence or short fragment unless the tool has validated that use case.
Can edited AI text be detected?
Sometimes, but editing can change the signals a detector uses, and mixed human–AI documents are especially difficult to reduce to a binary label. A negative result does not prove human authorship, and a positive result does not prove misconduct.
Are perplexity and burstiness still used?
They remain useful concepts and may appear in some systems, but they do not describe every modern detector. GPTZero, for example, says it no longer directly uses either metric after moving to a deep-learning architecture.
Can a detector identify the exact AI model?
Some systems may estimate likely model families, but confidently attributing text to one exact model is harder than classifying it as AI-like. Similar models, editing, and shared training patterns limit attribution.
Use the Result as a Signal, Not a Verdict
AI detection is a probabilistic measurement problem. Good use begins with the right question, a supported input, and a clear definition of the score. Responsible use ends with context, corroborating evidence, and a human decision.
Try HumanizeAI’s AI Detector with a complete passage. If you need higher usage limits for an ongoing workflow, compare plans and usage options.
Sources and References
- Mitchell et al., “DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature”, ICML 2023.
- Kirchenbauer et al., “A Watermark for Large Language Models”, 2023.
- Gehrmann, Strobelt, and Rush, “GLTR: Statistical Detection and Visualization of Generated Text”, ACL 2019.
- Sadasivan et al., “Can AI-Generated Text be Reliably Detected?”, 2023.
- Liang et al., “GPT detectors are biased against non-native English writers”, Patterns, 2023.
- Weber-Wulff et al., “Testing of detection tools for AI-generated text”, International Journal for Educational Integrity, 2023.
- Turnitin, “Using the AI Writing Report”, accessed October 7, 2026.
- GPTZero, “How do I interpret burstiness or perplexity?”, updated July 24, 2026.
- HumanizeAI AI Detector documentation, accessed October 7, 2026.