How to Tell If Something Was Written by AI (Beyond the Word List)
Stop reading the prose and start checking the claims. That is the whole method, and I say it as someone who edits for a living and spent the first year of this trying to do it the other way. Style will tell you a text feels generated. It will not tell you it was, and it will hand you confident wrong answers about careful writers. What holds up is boring and verifiable: does the source exist, does the quote appear where it is attributed, do the numbers reconcile.
This matters when you are holding a specific document, a job application, a student essay, a vendor's proposal, a news article, and you need a defensible read on it. Vibes are not defensible. Here is what is.
Check the citations first
Invented sources are the single strongest signal you have, for the simple reason that someone copying a real reference does not usually conjure one out of nothing. Researchers writing in Scientific Reports went through 636 citations spread across 84 generated papers. Of GPT-3.5's citations, 55 percent were pure invention. GPT-4 got that down to 18 percent, which reads like progress until you see the next number: among the citations that pointed at something real, 24 percent still had substantive errors in them, typically a wrong volume, a wrong page range, a year that was off.
That second failure is the more useful tell. A completely invented source is easy to explain away as sloppiness. A real paper cited with the wrong volume number and a year that is off by two is the signature of something that reconstructed the citation from statistical memory rather than looking at it.
How common is it? One team went looking across more than two million papers and pulled out about 4,000 fabricated references. The trend is the alarming part. Roughly one paper in 2,828 carried a fake citation in 2023. By 2025 it was one in 458. In the opening weeks of 2026, one in 277. Courts ran into this before publishers did, back in 2023, when attorneys in Mata v. Avianca were fined over a brief resting on six decisions that nobody could locate because none of them had ever existed.
One recent case makes the argument of this whole article better than I can. In March 2026, the Sixth Circuit sanctioned two attorneys $15,000 each over more than two dozen fabricated or misrepresented citations. The court never made an express finding that AI was used at all. It did not need to. The citations were checkable, they failed, and that was sufficient. You do not have to prove authorship to establish that a document is unreliable.
Practically: search the exact title. Ask for a DOI, ISBN, or a link. If a source cannot be found in a library database or a plain search, it very likely does not exist. And if the writer did use AI and simply says so, that is a different conversation entirely, since there is a correct way to cite ChatGPT in APA and MLA and using it is the opposite of deception.
Then the quotes and the numbers
Quotations are the second checkable layer, and they fail in a specific way: a real person, correctly named, saying something they never said. In February 2026, Ars Technica pulled an article the same day it published after an engineer pointed out that quotes attributed to him were not in his blog. Paraphrase had been rendered as direct quotation. The editor called it a serious failure of standards, and the reporter was later terminated.
You can put numbers on this too. When the BBC tested AI assistants on news questions, 51 percent of the answers came back with significant problems, and 13 percent of the quotes drawn from BBC articles had been altered or were nowhere in the piece being cited. A bigger follow-up spanning 22 public media organizations across 18 countries found 45 percent of responses carried at least one significant problem, sourcing being the biggest category of failure. So paste the quotation into a search engine and see whether it exists where the writer says it does. That one move catches a surprising amount.
Numbers break differently. They stay fluent and quietly stop adding up. Take CNET's run of AI-assisted finance explainers in 2023, which ended with corrections on 41 of 77 pieces. The compound interest article informed readers that $10,000 at 3 percent earns $10,300 in a year. It earns $300. Read that sentence at normal speed and nothing catches. The error only surfaces if you stop and work it out.
Book recommendations break the same way. There was a syndicated summer reading supplement in the Chicago Sun-Times in 2025, listing novels that did not exist, every one of them pinned to a real and famous name. One credited to Min Jin Lee. Another to Andy Weir. Genuine author, believable title, believable premise, and no such book anywhere on earth. Same trick as the fake legal citation, different outfit.
Look for what only someone present could know
Generated text is assembled from what has been written about a subject, so it is fluent about the general and empty about the particular. Google's own guidance on people-first content asks whether writing demonstrates expertise "that comes from having actually used a product or service, or visiting a place," which is a useful test to borrow.
What that looks like in practice: named tools with version numbers, actual costs, the setting that was wrong by default, the step that failed the first time, the thing the writer expected and did not get. First-hand accounts are full of loose ends and small irritations because reality has them. A generated account resolves everything tidily and names nothing checkable.
This is judgment rather than proof, and it is where careful human writing can be misread, so treat a smooth surface as a reason to check the citations rather than a verdict on its own.
Check for machine debris
Sometimes the evidence is simply left in the document. A researcher at the University of Louisville built a dataset of 500 published academic papers showing signs of undeclared AI use, and catalogued what was still visible in the final text. Nearly half contained a reference to the model's own knowledge cutoff, phrases in the family of "as of my last knowledge update." Almost a third contained "Certainly, here." About 12 percent contained the words "Regenerate response," and nearly 9 percent had the model identifying itself as a language model. Only 3 percent of those papers were ever corrected.
Newsrooms leak the same way. Press Gazette's running tracker of AI errors in journalism logged a Bristol Live post that opened with the chatbot's own preamble, "Since you didn't include specific instructions," and a Telegraph article that shipped with an AI editing suggestion still in the text. Wikipedia's volunteer cleanup project, which processes this at volume, treats leftover machine markup and fabricated citations as its strongest indicators while warning that stylistic patterns are "only potential signs of a problem, not the problem itself."
Two non-textual checks belong here. Confirm the author exists: Sports Illustrated was caught in 2023 publishing reviews under invented bylines with AI-generated headshots that were on sale as stock images. And watch for anachronism, content confidently describing a state of the world that no longer holds or missing an obvious recent development, which is a reasonable inference about a training boundary rather than a precise measurement, since research shows a model's effective cutoff often differs from its stated one.
The tells that prove nothing
Four popular signals are worth actively discarding, because acting on them gets innocent people in trouble.
Em dashes. There is no evidence behind this one. It spread through social media, and the punctuation mark it accuses has been a staple of English prose for centuries. Writers who have used em dashes for thirty years are now being accused of using AI, and some have started removing them from their own work, which is a real cost for a claim with nothing under it.
Detector scores. The strongest argument against them comes from OpenAI, which withdrew its own detector in 2023 after it correctly identified only 26 percent of AI text while flagging 9 percent of human text as machine-written. Bloomberg later ran two detectors against 500 college application essays submitted before ChatGPT existed and still saw 1 to 2 percent falsely flagged, sometimes with near-total confidence. When Vanderbilt turned Turnitin's AI detector off, it told instructors to compare against the student's earlier work and to look for factual inaccuracies and hallucinated sources instead. That is the method in this article, recommended by a university that tried the alternative. Run a free AI detector if you want the number, by all means. Just file it where it belongs, as one weak input next to the citation check, not as the answer.
Polished, formal, error-free prose. This is the one that does real damage. In a Stanford study, seven detectors falsely flagged more than 60 percent of human-written TOEFL essays, because consistent, careful word choice is statistically predictable and predictability is what these systems actually score. Judging by polish penalizes non-native speakers and disabled writers for writing well.
A single vocabulary hit. One "delve" tells you nothing. The words ChatGPT leans on and the structural rhythm behind same-y AI prose are real patterns, and they earn their keep by telling you where to look harder. That is all they do. Corroboration, never the finding.
Aim for a probability, not a verdict
The Global Investigative Journalism Network puts it well in its guide for reporters: the goal has shifted from definitive identification to a probability assessment and informed editorial judgment. That is not a retreat. It is what the evidence supports, and it points you at the useful work.
So run the checks in order. Do the sources exist and are they cited correctly. Do the quotes appear where they are claimed. Do the numbers survive arithmetic. Is there anything here that required a person to be present. Is there machine debris in the file. A document that fails those is unreliable whether or not a model touched it, which is the only conclusion you actually need, and the only one you can defend.
