Back to Blog
AI detectionneurodiversityacademic integrityaccessibility

Why AI Detectors Flag Neurodivergent Writing (and What Helps)

By Priya Raman8 min read

Why AI Detectors Flag Neurodivergent Writing (and What Helps)

You wrote every word. The detector returned a number anyway, and when you try to work out what it objected to, the answer turns out to be the way you write: the consistent structure, the precise term reused because it is the correct term, the sentences that come out at a similar length because that is how your thinking arrives on the page. There is no revision that fixes that. It is not a habit you picked up. That is the particular unfairness autistic and ADHD students describe to me, and it deserves a straight answer rather than a reassuring one.

So here is the straight answer, including the part most articles on this topic skip: the mechanism is real and well documented, the documented cases are real, and the evidence base specifically about neurodivergent writers is much thinner than the internet suggests. Knowing exactly where that line falls is what lets you defend yourself without making a claim someone can knock down.

What a detector is actually measuring

AI detectors do not recognize machine writing. They score two statistical properties. Perplexity measures how predictable your word choices are to a language model: if the next word is usually the obvious one, perplexity is low. Burstiness measures variation, mainly in sentence length and structure, where human prose tends to lurch between long and short and generated prose tends to even out.

Language models produce low-perplexity, low-burstiness text. So detectors treat those two properties as the signature of a machine. Read that again, though, because the flaw is sitting in plain sight: they are not measuring who wrote something. They are measuring how predictable and how varied it is. Any person whose natural writing is predictable and even lands in the same bucket as the machine, having done nothing wrong.

Stanford researchers showed exactly how this works, and the experiment is worth walking through. They took 91 TOEFL essays, all written by people, and ran them past seven detectors. More than 60 percent came back flagged as AI. Then they went back to those same essays and used AI to widen the vocabulary. Suddenly under 12 percent were flagged. Running it the other direction was just as stark: simplify the word choice in essays by American eighth-graders and misclassification jumps from about 5 percent to nearly 57 percent. Nobody's authorship changed in either direction. The predictability of the words did. That is also the reason detectors flag non-native English writing so heavily: constrained, consistent word choice reads as low perplexity.

That is the mechanism people are pointing at when they say detectors are biased against neurodivergent writers. If precise repeated terminology, learned sentence templates, or a rigorously consistent structure describe how you write, you are producing exactly the pattern the software scores as suspicious.

What the research does and does not show

Now the part I would want to know if I were the one accused. The Stanford study is about non-native speakers. It contains no neurodivergent participants at all, and it gets misattributed constantly, including by law firms advertising to accused students. It gives you the mechanism, not evidence about you.

Only one peer-reviewed study has looked directly at this. Chambers and Kelley, published in the 2025 Artificial Intelligence in Education proceedings, compared roughly 60,000 Reddit posts from likely-autistic and general communities using an OpenAI detector. Significantly more of the likely-autistic texts were flagged. But under 2 percent of either group was flagged in absolute terms, the texts were Reddit posts rather than academic essays, the detector was an older one, and autism was inferred from which communities people posted in rather than diagnosed. The authors themselves noted the connection between autistic writing features and AI-text features was not straightforward.

So: a real directional effect, small in absolute size, measured on the wrong kind of text. And for ADHD and dyslexia specifically, I could not find a single credible study, university finding, or expert measurement linking them to detector false positives. The claim circulates widely. The evidence for it is blog posts, many published by companies selling something.

Nobody has measured the false-positive rate for neurodivergent students on real coursework. That gap has practical consequences, not just academic ones. Walk into a hearing saying "studies show detectors discriminate against autistic students" and someone will go looking, and what they find will be thinner than what you claimed. There are better arguments available, and they hold up under exactly the scrutiny that one collapses under.

The cases where this became real

In January 2026, the New York Supreme Court in Nassau County annulled an academic integrity finding against Orion Newby, an Adelphi University student with Level 2 autism. His professor's report rested on a Turnitin AI score of 100 percent for an essay he had written with a tutor from Adelphi's disability support program. The court granted his petition, annulled the violation, and ordered the record expunged.

Be careful what you claim that ruling said, because it gets repeated wrong all the time. No finding that detectors are unreliable. No finding that they discriminate against autistic writers. What the judge found was that Adelphi broke its own rules. Newby never got the advisor he was owed. The officer who issued the determination then turned around and decided the appeal against it. His counter-evidence went unweighed. So he won on procedure, and I would not call that a consolation prize, because it is the part every student in this situation can actually use.

An earlier case ran differently. Bloomberg reported that Moira Olmsted, an autistic student at Central Methodist University, received a zero after a detector flagged a reading summary. Her work was re-reviewed and regraded, but she was warned that a second flag would not get a second look. She had done nothing wrong either time and left carrying the risk anyway.

A federal case is pending in Michigan where a student with anxiety and OCD argues under the ADA and Section 504 that disability-linked traits, formal tone and meticulous structure, were read as AI markers. Those are allegations, untested so far, and the only ruling to date went against her on an unrelated procedural point. It is worth watching. It is not yet something to cite.

When your accommodations are the trigger

The cruelest version of this happened to Marley Stevens at the University of North Georgia. Turnitin flagged her two-page paper. She had used the free Grammarly browser extension, and that was the extent of it. No chatbot. Nothing paid for. The zero pushed her GPA below 3.0, and the rest cascaded from there: HOPE Scholarship gone, academic probation, plus $105 she had to pay for a seminar about cheating.

Grammarly happens to be a standard accommodation for dyslexia and other language-based disabilities. Universities put it in their assistive technology guides. Your own disability office may have handed you the link. And what does it do? Smooths grammar, regularizes phrasing, which is the exact thing that drives perplexity down. So put Grammarly's own words in front of whoever is accusing you: the company states that no AI detector can conclusively determine whether AI was used, and admits detectors have been found biased against non-native English writers.

One caution so you do not overreach. The Grammarly pathway is documented. The equivalent claim about text-to-speech, dictation, or word prediction triggering a flag is plausible from the same mechanism, but I found no documented case of it. If assistive tech is part of your defense, name the tool you actually used.

And note the trap in the Stanford result: the fix that cleared those essays was using AI to vary the vocabulary. Students may need the tools they are accused of using in order to stop being accused of using them.

What actually helps

Almost everything that works is procedural rather than scientific. Ask what the accusation rests on, in writing. In the Newby case the report offered a percentage with no documentation of how it was produced, and that thinness was central to the outcome. A number is not a method.

Quote the vendor against itself, because these admissions are strong and rarely used. Turnitin's claimed false-positive rate of under 1 percent applies only to documents already scoring above 20 percent AI; its sentence-level false-positive rate is around 4 percent; and it deliberately shows no score at all between 1 and 19 percent because low scores are too unreliable to publish. Its own chief product officer has said the company accepts missing roughly 15 percent of real AI writing in order to hold false positives down. A tool tuned that conservatively is not built to be proof, and its maker says so.

Check your institution's own rules, then hold it to them. Newby won there. Many universities have already conceded the point in policy: Vanderbilt disabled Turnitin's AI detector in 2023, noting that at 75,000 papers a year even a 1 percent error rate means about 750 papers wrongly flagged. If your school kept the tool but its policy says detection is not sole evidence, that sentence is doing more work for you than any study.

Bring your disability office in early rather than after a determination, and have them document your accommodations as part of the record. Offer your process: drafts, outlines, notes, and especially version history showing the essay being built over days. Ask for comparison against your own earlier work, which is what Vanderbilt tells instructors to do instead of running a detector. And insist on human judgment rather than a score, since a detector score cannot prove you used ChatGPT and the better-run integrity offices already know it.

If you want to see what a grader might see before anyone else does, running a finished draft through a free AI detector takes a minute. It will not make an unfair flag fair. It does mean the number is not a surprise, and it tells you whether to have your drafts ready.

Consistency is not evidence

The honest summary is narrower than the headline and more useful. Detectors reward unpredictable prose and penalize consistent prose, which means writing carefully and writing the same way every time can read as machine-made. Whether that lands harder on neurodivergent students is genuinely under-researched, and pretending otherwise weakens the case. What is settled is that the number proves nothing on its own, the companies selling these tools say so themselves, and students win these disputes on process, documentation, and procedure. Consistency is a trait. It is not evidence.

Priya Raman

Priya Raman

Academic writing coach

Coaches grad students through theses and application essays. Writes about AI detection in the classroom and using AI help without crossing disclosure lines.

Check your text with the free AI Detector

See instantly whether your writing reads as AI-generated — free, no signup.

Run the AI Detector →
Why AI Detectors Flag Neurodivergent Writing