Wikipedia's Signs of AI Writing: The Full List, and Why Following It Can Make Your Writing Worse
Wikipedia's full list of AI-writing tells, a copy-paste prompt to check your drafts, and the research on why the list can't be applied mechanically.
Wikipedia's volunteer editors built the most-cited public catalog of AI-writing tells. This post gives you the full list in plain language, a copy-paste prompt that checks your own drafts against it, setup steps to turn that prompt into a reusable skill in Claude, Gemini, Perplexity, or ChatGPT, and the peer-reviewed research on why treating the list as a style guide will quietly wreck your writing.
The TL;DR
Wikipedia's "Signs of AI Writing" catalogs about a dozen patterns that show up disproportionately in unedited AI output. The patterns are real and the guide is accurate. What it is not is a style guide, and the data makes that plain: Mark Twain used 10.13 em dashes per 1,000 words, GPT-4.1 uses 10.62, and Meta's Llama models use zero. Run the list mechanically and you will strip voice out of writing that was never AI-generated in the first place. Use it as a prompt to find candidates, then use judgment on which ones actually matter.
What You'll Learn
- What Wikipedia's guide is, who built it, and the problem it was built to solve
- All twelve patterns, in plain language, with what each one actually looks like
- Why the em dash is a model-specific artifact rather than an AI tell, with the numbers
- What the peer-reviewed research says about detecting AI writing by pattern
- A copy-paste prompt that flags candidates in your drafts without prescribing deletions
- How to turn that prompt into a reusable skill in Claude, Gemini, Perplexity, or ChatGPT
- What to do with each pattern once you have found it, and how long each fix takes
- Why a checklist can only measure what is present, never what is missing
The Breakdown
What is Wikipedia's Signs of AI Writing guide?
It is a public reference maintained by WikiProject AI Cleanup, a group of volunteer editors who needed a shared standard for catching undisclosed AI-generated text before it lands in the encyclopedia. It documents roughly a dozen recurring patterns and it exists for one narrow purpose: flagging content whose authorship was not disclosed.
They had a real problem. A Princeton study presented at the WikiNLP workshop at EMNLP in November 2024 found that 4.36% of the 2,909 English Wikipedia articles created in August 2024 contained significant AI-generated content. Of the articles flagged by two independent detectors, eight were self-promotional and eight pushed a viewpoint on a polarising topic.
That context matters for how you use the guide. It was written to protect an encyclopedia's neutrality from undisclosed automation. It was not written to help a marketer sound more like themselves, and the two jobs pull in opposite directions.
What are the actual signs, in plain language?
- Formulaic em dashes. Appearing more often, and in positions where a comma, colon, or parentheses would read more naturally.
- Rule-of-three lists. Gravitating toward triplets like "innovative, transformative, and groundbreaking."
- Vague attribution. "Studies show" or "experts say" with no named source.
- Overemphasis on significance. Calling routine facts "vital" or a "testament" to something.
- False ranges. "Ranging from X to Y" constructions that sound specific and convey nothing.
- Section summaries that restate. "In summary," "in conclusion," "overall," adding no new information.
- Superficial analysis tacked onto facts. "Highlighting" or "illustrating" without explaining the actual relevance.
- Negative parallelism. The contrast-reframe pattern: "it's not just X, it's Y."
- Editorializing asides. Inserted opinions about importance, like "it's important to note."
- Letter-style phrasing in non-letter content. "I hope this message finds you well" in a blog post.
- Leftover collaborative phrases. "I hope this helps!" left in unreviewed output.
- Formatting habits. Excessive boldface, inconsistent list use, irregular title case in headings.
The guide itself is careful about this, and most people quoting it are not: no single sign proves anything. These work as a combined signal, not as individual tripwires.
Does having one of these signs mean something was written by AI?
No, and the evidence against that reading is stronger than most people realise. Start with the em dash, because it is the tell everyone reaches for first.
The only primary measurement of em dash rates I could find is a March 2026 preprint by E. M. Freeburg, which generated roughly 240,000 words across twelve instruction-tuned models and compared them against a 57,232-word human baseline. Put those numbers next to counts from public-domain literature and the heuristic falls apart.
| Writer or model | Em dashes per 1,000 words |
|---|---|
| GPT-4.1 | 10.62 |
| Mark Twain, Huckleberry Finn | 10.13 |
| Claude Opus 4.6 | 9.09 |
| Herman Melville, Moby-Dick | 8.12 |
| Charles Dickens, A Tale of Two Cities | 5.33 |
| Henry David Thoreau, Walden | 4.27 |
| Jane Austen, Pride and Prejudice | 3.47 |
| Human baseline (8 published essays) | 3.23 |
| GPT-5.4 | 1.43 |
| Gemini 2.5 Flash | 1.28 |
| Llama 3.1 8B and Llama 3.3 70B | 0.00 |
Twain out-dashes every model except one. Two Llama models produce none at all. The em dash is not a property of AI writing, it is a fine-tuning artifact of particular model families, and it is fading as labs tune against it. Note the trajectory inside OpenAI's own lineup: GPT-4.1 at 10.62, GPT-5.4 at 1.43.
The vocabulary tells have the same problem. Kentaro Matsui's December 2025 analysis of 27.5 million PubMed records found that the "AI vocabulary" terms started rising in 2020, two years before ChatGPT existed. And Dmitry Kobak's Science Advances study, which found "delves" spiking 28-fold in 2024 abstracts, benchmarks that against "zika" spiking 40-fold in 2017 for entirely ordinary reasons.
Word frequency is a population-level signal. It tells you something real about a corpus of 15 million abstracts. It tells you almost nothing about the document in front of you.
Why does the false-positive problem make this more than a style debate?
Because the machines that automate this pattern-matching are wrong often enough to damage real people.
A Stanford team led by Weixin Liang tested seven GPT detectors against 91 TOEFL essays written by non-native English speakers and 88 essays by US eighth-graders. The detectors were near-perfect on the native-speaker essays. On the non-native essays the average false positive rate was 61.22%. Every one of the seven detectors flagged 18 of the 91 essays as AI-authored, and 89 of the 91 were flagged by at least one.
A separate multi-institution study led by Debora Weber-Wulff ran 14 public detectors plus Turnitin and PlagiarismCheck across 756 tests. All scored below 80% accuracy. On machine-paraphrased AI text, overall accuracy dropped to 26%.
So the tools are simultaneously too aggressive on human writing that pattern-matches to AI, and too permissive on AI writing that has been lightly rewritten. If you have ever wondered why ZeroGPT flags text you wrote yourself, that is the mechanism.
Why does following the checklist make marketing content worse?
Because Wikipedia is optimising for neutrality and you are not.
Marketing teams picked the guide up, treated every item as a banned-words list, and started scrubbing those patterns out on the theory that removing tells makes writing better. It does not. A team that mechanically strips every em dash, every triplet, and every strong statement is not humanizing the draft, it is sanding off the only parts that had a pulse.
Look at what is actually on the list. Editorializing asides. Confident claims about significance. Contrast constructions. Those are not defects in marketing writing, they are the job. A Wikipedia editor removes an editorializing aside because Wikipedia has a neutral point of view policy. You do not have one. You are supposed to have a point of view, and saying so is the entire reason anyone reads you instead of a spec sheet.
What is left after a full mechanical scrub reads clean and says nothing, which is a worse outcome than sounding slightly AI-generated in the first place.
What can a detection checklist never tell you?
It can only catalog what is present. It has no way to measure what is missing.
A checklist cannot verify a claim. It cannot confirm that an example actually happened. It cannot tell you whether the argument holds, whether a source exists, or whether anyone would cite this over what already ranks. A guide built to detect the absence of disclosure has nothing to say about the presence of value.
That gap has commercial consequences. SE Ranking ran a controlled 16-month experiment publishing 2,000 unedited one-click AI articles across 20 fresh domains. Presence in the top 100 fell from 28% to 3% by month three. Over 16 months the entire set earned 1,092,079 impressions and 1,381 clicks, roughly one click per two articles. Six AI-assisted but human-edited articles on SE Ranking's own blog drew 555,000 impressions and over 2,300 clicks in 13 months.
The difference was never surface polish. It was whether anything had been added.
How can I check my own writing against this list?
Paste the block below into any AI assistant along with your draft. It is written to flag candidates and explain the tradeoff, not to prescribe deletions, which is the difference between a useful check and an automated voice-remover.
You are a writing pattern checker. I will give you a draft.
Check it against these twelve patterns from Wikipedia's Signs of AI
Writing guide:
1. Em dashes used where a comma, colon, or parentheses would read
more naturally
2. Rule-of-three lists ("innovative, transformative, groundbreaking")
3. Vague attribution ("studies show," "experts say," no named source)
4. Overemphasis on significance ("vital," "testament to," "crucial")
5. False ranges ("ranging from X to Y" that conveys nothing specific)
6. Section summaries that restate without adding ("in conclusion")
7. Superficial analysis tacked onto facts ("highlighting,"
"illustrating," with no explanation of relevance)
8. Negative parallelism ("it's not just X, it's Y")
9. Editorializing asides ("it's important to note")
10. Letter-style phrasing in non-letter content
11. Leftover collaborative phrases ("I hope this helps!")
12. Formatting habits: excessive boldface, inconsistent lists,
irregular heading case
Rules:
- Flag only clear matches. Do not flag theoretical applications.
- For each match, give the exact quote and the pattern number.
- For each match, state in one line whether removing it would
improve the sentence or would flatten it. Be willing to say
"keep this one."
- Do not rewrite anything unless I ask.
- At the end, list what the draft is MISSING: unsourced claims,
places where a specific example would help, and any section
that states a conclusion without evidence.
Here is the draft:The last instruction is the one that matters. Without it you get a subtraction tool. With it you get something closer to an editor.
You ran the check and it found tells. Now what?
Here is what to do with each pattern, and roughly what it costs you.
| Pattern found | Why the model produced it | The fix | Time |
|---|---|---|---|
| Formulaic em dashes | Trained on dense editorial prose that uses them for pacing | Replace with a period or comma, or split into two sentences. Keep them where the aside genuinely interrupts. | 30 sec each |
| Rule-of-three lists | Triplets score as balanced in training data | Cut to two items or expand to four. Asymmetry reads human. | 1 min |
| Vague attribution | The model has no source to name, so it hedges | Name the source, date it, link it. If you cannot, cut the claim. | 5 to 15 min |
| Overemphasis on significance | Reward tuning favours enthusiasm | Delete the adjective. The fact carries itself. | 15 sec |
| Negative parallelism | Contrast framing is a high-probability construction | Keep it where the contrast is real. Cut it where it is decoration. | 1 min |
| Section summaries | Filler that restates without adding | Delete the paragraph. | 10 sec |
| Superficial analysis | The model gestures at relevance it cannot establish | Either explain the actual relevance or cut the sentence. | 2 min |
| Editorializing asides | Trained on hedged institutional prose | In marketing writing, usually keep. Just say the thing directly instead of announcing that you are about to. | 30 sec |
Notice the distribution. Every fix on that list is mechanical and takes under two minutes, except one. Sourcing a claim takes five to fifteen minutes and it is the only one that changes whether the piece is worth citing.
That split is the whole argument. The mechanical passes are worth automating precisely because they are mechanical. The sourcing is worth your time precisely because it is not.
Run this automatically instead of by hand. HumanizeAI's AI Humanizer runs the pattern pass, the stat-verification pass, and the experience-signal check in one go, and it is tuned to flag rather than flatten. Free to start. If you want the full manual method first, the complete guide to humanizing AI text walks through it end to end.
How do I turn this into a reusable skill or custom assistant?
Set it up once and it runs on every draft after that.
| Platform | Where the instructions go |
|---|---|
| Claude | Create a Project and paste the block into custom instructions, or build it as a Skill if you have access |
| Gemini | Create a new Gem and paste the block into its setup field |
| Perplexity | Create a Space and paste the block into the system prompt field |
| ChatGPT | Build a Custom GPT with the block as instructions, or add a shortened version under Settings > Custom Instructions |
Setup runs a few minutes on each. The Claude Skill and the Custom GPT versions are the two worth doing if you only do one, because both persist across conversations without you re-pasting anything.
What should I check for besides the tells?
The things a pattern checker structurally cannot see. HumanizeAI's H.E.A.R.T. framework is built as an addition checklist for exactly this reason. We use this when we write content for our clients.
HUMAN FIRST. Written for a specific reader, not a search engine.
EVIDENCE OVER CLAIMS. Every statistic carries a named source, a date, and a live link.
ANSWER FIRST. The core answer lands in the first 150 to 200 words.
REAL VOICE. At least one specific firsthand example only someone who did the work could write.
TRUST SIGNALS THROUGHOUT. Named author with credentials, outbound citations, internal cross-references.
Run it as a second pass:
You are an addition checker. I will give you a draft.
Check it against five principles and report only what is MISSING:
HUMAN FIRST: Is this written for one specific reader, or for a
generic audience? Name who it appears to be written for.
EVIDENCE OVER CLAIMS: List every factual claim with no named,
dated, linked source.
ANSWER FIRST: Does the core answer appear within the first 200
words? Quote where it appears, or say it does not.
REAL VOICE: List any specific firsthand example. If there are
none, say so plainly.
TRUST SIGNALS: Is there a named author, are there outbound
citations, are there internal links?
For each gap, say what would fix it. Do not rewrite.
Here is the draft:Pattern detection tells you what to remove. This tells you what to add. Removing every flagged tell from a piece with no sources and no real examples gets you clean, empty prose. Clean and empty is what makes a page uncitable.
Founder Observation
I ran our own contrarian article on this exact guide through the checklist before we published it. It caught four instances of negative parallelism, the "it's not X, it's Y" pattern, that I had not consciously noticed while writing, because the whole piece was built around a contrast argument.
Here is what I did with them: I kept the ones that were doing real work and cut the one that was pure decoration. The keepers earned their place, because the piece genuinely was arguing that one thing is not another thing. Negative parallelism is a good construction. It is on the list because models overuse it, not because it is bad.
That is the moment the distinction clicked for me. The checklist did its job perfectly and still could not tell me whether the argument underneath was any good. Earlier this year I went looking for the actual mechanism behind our own authenticity scoring, the real internal documents rather than the marketing description, and what I found was a system built almost entirely around addition: a stat-verification step that kills any claim without a live source, an Experience dimension with an explicit no-fabrication rule, and a standard that asks whether a piece would get cited over what already ranks.
Nobody on my team set out to build the opposite of a detection checklist. We built a pipeline to solve our own content problem, and it turned out to be an addition machine sitting in an industry full of subtraction tools.
Research & Supporting Evidence
- Freeburg, E. M. "The Last Fingerprint: How Markdown Training Shapes LLM Prose." arXiv:2603.27006, 27 March 2026. Human baseline of 3.23 em dashes per 1,000 words against GPT-4.1 at 10.62 and both Llama models at 0.00, measured across roughly 240,000 generated words and a 57,232-word human baseline. Single-author preprint, not peer reviewed, and the only primary measurement of em dash rates that currently exists. https://arxiv.org/abs/2603.27006
- Kobak, D., González-Márquez, R., Horvát, E., Lause, J. "Delving into LLM-assisted writing in biomedical publications through excess vocabulary." Science Advances 11(27), 4 July 2025. At least 13.5% of 2024 abstracts showed LLM processing; "delves" carried an excess frequency ratio of 28.0. Based on 15.1 million PubMed abstracts. https://doi.org/10.1126/sciadv.adt3813
- Matsui, K. "Delving Into PubMed Records." Perspectives on Medical Education 14(1), 2 December 2025. Of 135 candidate AI-influenced terms, 103 rose meaningfully in 2024, but the terms "began increasing in 2020, preceding ChatGPT's 2022 release." Based on 27,501,542 records. https://pmejournal.org/articles/10.5334/pme.1929
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., Zou, J. "GPT detectors are biased against non-native English writers." Patterns 4(7), 10 July 2023. Average false positive rate of 61.22% on essays by non-native English speakers, against near-perfect accuracy on native-speaker essays. Seven detectors, 179 essays. https://doi.org/10.1016/j.patter.2023.100779
- Weber-Wulff, D., et al. "Testing of detection tools for AI-generated text." International Journal for Educational Integrity 19(26), 25 December 2023. All 14 tested detectors scored below 80% accuracy; accuracy fell to 26% on machine-paraphrased AI text. 756 tests. https://link.springer.com/article/10.1007/s40979-023-00146-z
- Brooks, C., Eggert, S., Peskoff, D. "The Rise of AI-Generated Content in Wikipedia." WikiNLP workshop, EMNLP, 16 November 2024. 4.36% of 2,909 English Wikipedia articles created in August 2024 contained significant AI-generated content. https://aclanthology.org/2024.wikinlp-1.12.pdf
- Reinhart, A., et al. "Do LLMs write like humans? Variation in grammatical and rhetorical styles." PNAS 122(8), 25 February 2025. GPT-4o uses present participial clauses at 5.3 times the human rate and nominalisations at 2.1 times; "tapestry" appears in 23% of GPT-4o outputs. Corpora of 8,290 human texts and matched LLM output. https://doi.org/10.1073/pnas.2422455122
- Khromova, Y. "How AI-Generated Content Performs: Experiment Results." SE Ranking, 19 March 2026. 2,000 unedited AI articles across 20 domains earned 1,381 clicks in 16 months. Vendor research, and the result is unflattering to the vendor's own product. https://seranking.com/blog/ai-content-experiment/
- Wikipedia:Signs of AI writing. The primary source, maintained by WikiProject AI Cleanup. https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing
Mini Case Study: what a full mechanical scrub costs
Illustrative composite, built from the patterns above rather than a single named client.
Take a 1,200-word product post that a marketing team runs through a strict checklist pass. The checker flags 9 em dashes, 4 rule-of-three lists, 6 editorializing asides, 3 instances of negative parallelism, and 2 section summaries.
A mechanical scrub removes all 24. Elapsed time is about 20 minutes and the piece now reads smooth. It also no longer contains a single sentence where the writer says what they think, because the editorializing asides were the opinions and the negative parallelism was the argument.
The judgment pass looks different. It removes the 2 section summaries, 7 of the 9 em dashes, and 3 of the 4 triplets. It keeps every aside and every contrast construction. Then it spends the remaining 15 minutes on the thing the checker flagged under "missing": two claims with no named source. One gets a citation. The other gets cut, because no source existed.
Same starting draft, roughly the same total time. One version is clean and says nothing. The other is slightly less clean and is now the only page on the topic carrying a verifiable number.
Key Takeaways
- Wikipedia's guide catalogs about a dozen real patterns in unedited AI output, built for editors catching undisclosed AI content, not for teaching anyone to write.
- No single pattern proves AI authorship. The guide says so explicitly and most people quoting it skip that part.
- The em dash is a model-specific fine-tuning artifact, not an AI tell. Mark Twain used 10.13 per 1,000 words, GPT-4.1 uses 10.62, and both Llama models use zero.
- The "AI vocabulary" terms began rising in 2020, two years before ChatGPT launched, so word frequency is a population-level signal and a poor document-level one.
- Automated detectors falsely flag non-native English writers 61.22% of the time and fall to 26% accuracy on lightly paraphrased AI text.
- Wikipedia optimises for neutrality. Marketing optimises for a point of view. Half of what the list flags is what a marketing writer should be doing more of.
- A checklist can only measure what is present in a draft. It cannot verify a claim, confirm an example, or tell you whether the argument holds.
- Every fix on the list is mechanical and takes under two minutes, except sourcing a claim. Sourcing is the only one that changes whether the piece is worth citing.
FAQ
What is Wikipedia's Signs of AI Writing guide? Wikipedia's "Signs of AI Writing" is a public reference maintained by WikiProject AI Cleanup, a group of volunteer Wikipedia editors. It catalogs roughly a dozen patterns that appear disproportionately in unedited AI-generated text, including formulaic em dash use, rule-of-three lists, vague attribution, and stock phrases such as "it's important to note." It was built to help editors flag undisclosed AI-generated content before publication, not to serve as a writing style guide.
Is the em dash actually an AI tell? No. A March 2026 preprint by E. M. Freeburg measured em dash rates across twelve instruction-tuned language models and found GPT-4.1 at 10.62 per 1,000 words while both Llama 3.1 8B and Llama 3.3 70B produced 0.00. For comparison, Mark Twain used 10.13 per 1,000 words in Huckleberry Finn and Herman Melville used 8.12 in Moby-Dick. Em dash frequency reflects how a specific model family was fine-tuned, not whether text was machine-generated.
How do I make AI text sound human without changing what it says? Fix the mechanical patterns and add what the draft is missing. Replace formulaic em dashes with periods or commas, break up rule-of-three lists, and delete summary paragraphs that restate earlier content, all of which preserve meaning. Then add the things a pattern check cannot supply: a named and dated source for every factual claim, at least one specific firsthand example, and a direct answer to the reader's question in the first 200 words.
Does removing AI tells actually help content rank or get cited? Removing tells alone does not. SE Ranking published 2,000 unedited AI articles across 20 domains and tracked them for 16 months; the set earned 1,381 clicks in total, roughly one click per two articles, while six human-edited AI-assisted articles on the same company's blog drew over 2,300 clicks in 13 months. The difference came from added sourcing, examples, and structure rather than from surface polish.
Will editing out these patterns hurt my writing voice? It can, and this is the main risk of applying the list mechanically. Several items on Wikipedia's list, including editorializing asides, confident statements of significance, and negative parallelism, exist because Wikipedia requires a neutral point of view. Marketing content is supposed to have a point of view, so stripping those patterns often removes the only sentences in a draft where the writer says what they actually think.
Can AI content detectors be trusted? Not reliably. A Stanford study published in Patterns in July 2023 found seven GPT detectors produced an average false positive rate of 61.22% on essays written by non-native English speakers. A separate study of 14 detectors published in the International Journal for Educational Integrity in December 2023 found all scored below 80% accuracy overall and dropped to 26% accuracy on machine-paraphrased AI text.
How can you tell when someone writes with AI? You often cannot tell from the text alone, and the research supports that caution. Individual patterns such as em dashes, the word "delve," or rule-of-three lists appear in both human and machine writing, and a 2025 analysis of 27.5 million PubMed records found the so-called AI vocabulary began rising in 2020, before ChatGPT existed. The more reliable signals are structural: claims with no verifiable source, examples that are generic rather than specific, and prose that describes a topic without contributing anything to it.
What words does AI use the most? Peer-reviewed corpus studies identify a consistent set. Kobak et al. in Science Advances (July 2025) found "delves" appearing 28 times more often than expected in 2024 biomedical abstracts, followed by "underscores" at 13.8 times and "showcasing" at 10.7 times. Reinhart et al. in PNAS (February 2025) found "camaraderie" at 162 times the human rate and "tapestry" at 155 times, with "tapestry" appearing in 23% of GPT-4o outputs.
How long does it take to humanize a 1,500-word AI draft by hand? Roughly 30 to 45 minutes, and the split matters more than the total. The mechanical fixes, including em dashes, triplets, filler summaries, and inflated adjectives, take under two minutes each and account for maybe 15 minutes. Sourcing unverified claims takes 5 to 15 minutes per claim and consumes the rest, which is why automating the mechanical passes and spending your own time on sourcing is the efficient division.
Ready to stop doing this by hand?
The pattern pass is mechanical, repeatable, and exactly the kind of work worth handing to a tool. HumanizeAI's AI Humanizer runs the tell check, the stat-verification pass, and the experience-signal check together, tuned to flag patterns rather than strip them blindly. Free to start, no card required.
Additional Resources
- How to Humanize AI Text: The Complete Guide for Marketers. The pillar guide, start here.
- Answer Engine Optimization Playbook. How content gets cited by AI answer engines.
- What Makes a Page Citable to AI? The 7 Structural Markers
- Why Does ZeroGPT Say I Used AI When I Didn't?
- Does Humanizing AI Text Help or Hurt AI Visibility?
About the Author
Steve Palomares has spent 25+ years building software companies. Now owner of HumanizeAI, he writes about AI content strategy for marketing, AEO, GEO and growing software businesses with AI. Based in North Texas.