Skip to content

AI Detection

Why AI Detectors Flag Non-Native English Writers More Often (and What to Do About It)

7 min read

The Stanford study that exposed a 61% false positive rate, what it means for international students, and how to defend your work.

A 61% misclassification rate, and the students it costs

In July 2023, a team of Stanford researchers led by Weixin Liang and James Zou published a paper in the journal Patterns with a finding that should have made every university administrator in America stop and reconsider their AI detection policies.

The team tested seven of the most widely used AI writing detectors against a corpus of essays. The detectors performed almost flawlessly when scoring writing by U.S.-born eighth graders — false positive rates close to zero. Then the researchers fed the same detectors essays written by non-native English speakers, taken from the TOEFL exam.

The result: a 61.22% misclassification rate. More than three out of every five essays written by international students were flagged as AI-generated, even though every single one was demonstrably human-written.

Of 91 TOEFL essays in the test, 89 were flagged by at least one detector. Eighteen were flagged unanimously by all seven.

The published paper has a DOI you can verify yourself: 10.1016/j.patter.2023.100779. It is one of the most cited studies on AI detection in education, and its findings have been replicated in subsequent research.

Inline chart for article 02
Source data visualization.

How a tool meant to detect cheating became a tool that flags accents

The mechanism behind this bias is not a programming bug. It is the core design of how most AI detectors work.

Detectors like GPTZero, Originality.AI, and ZeroGPT score text based on a metric called perplexity — essentially, how surprising the next word is, given the words that came before it. Highly perplexing text (varied vocabulary, unusual phrasing, unexpected sentence structures) is interpreted as human-written. Low-perplexity text (predictable word choices, common grammatical patterns) is flagged as AI-generated.

The problem, as Stanford’s James Zou explained to the Markup in 2023, is that AI models like ChatGPT learned to write by ingesting an enormous corpus of polished, professional English — academic papers, journalism, technical writing. Their default style is competent, grammatical, and bounded by a relatively narrow vocabulary range.

Non-native English writers, by necessity, often produce writing that sits in the same statistical neighborhood. They tend to use more common words. Their sentence structures hew closer to grammatical norms learned in formal instruction. Their phrasing is often more direct, less idiomatic.

To a perplexity-based detector, that pattern looks like AI. To a human reader, it looks like a careful international student doing their best work.

A bibliography is the most defensible evidence you have. Papyra’s Document Correction tool generates academic writing where every citation is traced back to a real, verified source through CrossRef and OpenAlex — the kind of paper trail that ends an academic integrity meeting in five minutes.

Who pays the cost

The bias is not abstract. It produces real consequences for real students, disproportionately affecting populations that already navigate higher institutional barriers.

In February 2025, an executive MBA student at Yale — a French entrepreneur learning in his second language — was suspended for a year after the school’s Honor Committee accused him of using AI on a final exam. The primary evidence was a GPTZero scan. According to the lawsuit he subsequently filed against the university, his lawyers submitted GPTZero scans of academic papers written by Yale faculty, including former University President Peter Salovey. The detector flagged those as probably AI-generated too.

The student is alleging in federal court that Yale violated the Civil Rights Act on the basis of his national origin. The case is ongoing.

In October 2025, the Australian Broadcasting Corporation reported that Australian Catholic University had referred nearly 6,000 cases of alleged AI cheating during the 2024 academic year — about 90% of all academic integrity referrals at the institution. Internal review found that roughly one in four referrals were dismissed after investigation. The students involved had spent weeks producing search histories, handwritten drafts, and document version logs to clear their names.

A senior university official later acknowledged to the ABC that “any case where Turnitin’s AI detection tool was the sole evidence was dismissed immediately.” The students whose cases ended in dismissal had still spent weeks defending themselves against allegations that should never have been filed.

Why the bias is hard to fix

Some detector companies have pushed back on the Stanford finding. Originality.AI, in a public response, claimed that its specific detector achieved a 5.04% false positive rate on non-native English samples — significantly lower than the Stanford average. Educational Testing Service, the organization that administers the GRE, ran a 2024 study using simpler detection methods and reported finding no bias against non-native writers.

The methodology questions are legitimate. The Stanford study tested seven detectors, not all detectors. Different tools use different scoring approaches.

But the underlying issue — that perplexity-based detection systematically penalizes restricted vocabulary and conventional grammar — is mathematically real. Any detector relying on those metrics will, at scale, produce more false positives for writers whose style happens to match those characteristics. Non-native English writers are the most prominent affected group, but they are not the only one. Researchers have documented similar elevated rates for students with autism, ADHD, and dyslexia. A Purdue professor named Rua Mae Williams, who is autistic, was flagged by an AI detector while submitting their own academic writing.

The fundamental tension is that detectors trained to catch a particular style of machine-generated writing will misfire whenever that style overlaps with the natural style of human writers. Until detection moves beyond statistical surface features, the bias will persist.

Inline visualization for article 02
Practical framework summary.

What students can do, practically

If you are a non-native English speaker — and the data suggests you are facing detection rates several multiples higher than your native-speaking classmates — there are concrete steps that protect your work.

Maintain a verifiable writing process. Use Google Docs, Microsoft 365, or a similar platform with automatic version history. Write incrementally over days or weeks. A document log with hundreds of small edits is virtually impossible to confuse with single-session AI output. If your institution requires a defense, this log is the strongest evidence you can produce.

Verify every citation before you submit. AI tools that produce hallucinated references — invented authors, fake DOIs, made-up journal names — leave students with bibliographies they cannot defend. Every paper you cite should be one you can locate in CrossRef, OpenAlex, or your university library catalog. If a source cannot be verified, it should not be in your paper.

Read your institution’s academic integrity policy. Most universities require human judgment, not algorithmic flags, before any sanction can be imposed. Knowing that distinction in advance often resolves cases at the first meeting. If your institution treats a detector score as proof, that is grounds for appeal under most academic integrity frameworks.

Document your sources. Save the actual PDF of every paper you cite. Keep them in a folder organized by your bibliography order. If you are accused, the ability to immediately produce the original sources for every citation is the cleanest possible refutation.

The broader shift

Universities are beginning to adjust. Vanderbilt disabled Turnitin’s AI detector in August 2023, citing the unacceptable false positive rate. Johns Hopkins, the University of Pittsburgh, and the University of British Columbia followed. The University of Michigan now formally advises its faculty against using AI detection scores as evidence of misconduct. The trend is clear.

In the meantime, the most valuable thing a student — and especially a non-native English-speaking student — can do is to write in a way that produces evidence. A bibliography of verified sources. A document with version history. A process that can be defended. The detectors will keep being wrong. The students who are prepared will be fine.

—

Want a bibliography that defends itself?

Papyra writes academic essays where every citation is traced to a real, published paper — verified through CrossRef, OpenAlex, and Semantic Scholar before it touches your draft. No invented authors. No fabricated DOIs. The kind of paper trail that ends an integrity meeting in five minutes.

See how Papyra works →

papyra

Writing for students who want integrity without guesswork.