Skip to content

AI Detection

AI Hallucinated Citations: Why ChatGPT Invents Sources (and How to Spot Them)

9 min read

The lawyer who learned the hard way, the architectural reason it keeps happening, and a practical method for catching fabricated references before they reach your professor.

The brief that should not have existed

In April 2023, a New York attorney named Steven Schwartz filed a brief in federal court on behalf of his client Roberto Mata. The brief argued that Mata’s personal injury claim against Avianca airlines should not be time-barred. To support that argument, Schwartz cited six prior cases:

– Varghese v. China Southern Airlines
– Shaboon v. EgyptAir
– Petersen v. Iran Air
– Martinez v. Delta Airlines
– Estate of Durden v. KLM Royal Dutch Airlines
– Miller v. United Airlines

None of these cases existed.

Avianca’s lawyers pointed this out to the court. Judge P. Kevin Castel, presiding over the case, attempted to locate the citations himself. He could not. He ordered Schwartz to produce copies. Schwartz, who had drafted the brief using ChatGPT, asked the chatbot directly: “Is Varghese a real case?” ChatGPT answered yes. Schwartz asked for the source. ChatGPT provided one. The source did not exist either.

The whole thing fell apart in the next two months. Judge Castel issued a sanctions order in June 2023 fining Schwartz, his colleague Peter LoDuca, and the firm Levidow, Levidow & Oberman a combined $5,000. The judge described one of the fake legal opinions as “gibberish.”

The case is Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023). It is the first widely publicized example of what AI researchers call hallucination: the tendency of large language models to generate confident, plausible, completely fabricated content.

It was not the last. By late 2025, U.S. courts had sanctioned attorneys in at least a dozen separate cases for submitting AI-generated filings with fake citations. In one California case, two firms were fined $31,000. In an Alabama matter, a major law firm conducted a panicked audit of every brief it had filed in the federal Eleventh Circuit, looking for hidden hallucinations.

If trained lawyers, with malpractice exposure and Rule 11 obligations, are submitting fake citations they obtained from ChatGPT, the question for students becomes obvious: how does this happen, and how do you avoid it in your own work?

Inline chart for article 04
Source data visualization.

What hallucination actually is

A large language model like ChatGPT, Claude, or Gemini does not have a database of real facts it consults when you ask it a question. It has a statistical model of language. When you ask “give me five sources about climate change policy,” the model generates text that looks like a list of sources about climate change policy, based on the millions of similar lists it has seen during training.

The output usually looks correct. Author names that fit the field. Plausible journal titles. DOIs in the right format. A structure that matches the conventions of academic citation.

The problem is that nothing in the architecture forces those plausible-looking elements to correspond to real things. The model might generate the name of a real researcher in the field. It might generate a real journal title. It might combine them into a citation that looks completely legitimate. None of it is verified, because the model has no mechanism for verification. It is producing text that is statistically likely, not text that is empirically true.

This is why hallucinations are so hard to catch by reading. They look real. The hallucinated Varghese v. China Southern Airlines case in the Mata brief had a plausible name, a plausible court, plausible parties, and plausible legal reasoning. It was indistinguishable from a real case until somebody actually checked.

A bibliography that cannot be verified is a bibliography that cannot be defended. Papyra’s academic writing tool generates papers where every citation is matched to a real source through CrossRef, OpenAlex, and Semantic Scholar before it appears in your draft — no invented authors, no fabricated DOIs, no journal titles that exist only in the model’s training data.

Why this matters for academic work

In a court filing, a fake citation is sanctionable misconduct. In a student paper, it is academic dishonesty whether or not the student knew the citation was fake. Most universities define misconduct in terms of the act, not the intent. Submitting a paper with fabricated sources is a violation regardless of whether you generated those sources personally or trusted an AI tool that did.

The risk is not theoretical. Throughout 2024 and 2025, professors at universities across the U.S. and U.K. began reporting a recognizable pattern: student papers with reference lists where some entries did not appear in any database. The 2024 issue of PS: Political Science & Politics included a paper titled “ChatGPT in Higher Education” that documented the phenomenon and proposed verification protocols.

What this looks like in practice: a student uses ChatGPT to brainstorm, asks for sources, gets a confident list of five papers, includes them in their bibliography, and submits the paper. The professor checks one of the sources, cannot find it, checks another, also cannot find it. The student is referred to the academic integrity office. The student insists they did not write the paper using AI. The professor’s evidence is the bibliography. The case is harder to defend than a misclassified human-written paper, because the student is, in fact, citing sources that do not exist.

How to spot a hallucinated citation

The fastest method is verification, not pattern recognition. Hallucinations are designed to look real. The only reliable way to catch them is to check.

For each citation in your paper, before you submit:

Search the DOI directly. Every legitimate academic paper published since the early 2000s has a Digital Object Identifier. Type the DOI into doi.org — the link should resolve to the actual paper. If it does not, the DOI does not exist.

Search the journal’s archive. Major journals all have searchable archives. Search by title or author. If the paper is not in the journal’s own archive, it was never published there.

Search Google Scholar. Search by exact title. A real paper with a citation count of more than two appears in Scholar. A hallucinated paper does not.

Search the author’s institutional page. Most academic researchers list their publications on their university or institute homepage. If the author is real but the cited paper is not in their bibliography, the citation is fabricated.

Use CrossRef directly. crossref.org/search covers over 140 million scholarly works. If a paper is real and recent, it is in CrossRef.

This sounds laborious. In practice, after the first few checks, you develop intuition for which citations look suspicious — overly clean DOIs, journal titles that are slightly wrong (Journal of Educational Psychology vs. Journal of Educational Research), authors with publication dates that do not match their actual career timelines.

Inline visualization for article 04
Practical framework summary.

The architectural reason general-purpose chatbots cannot fix this

When the Mata case became public in 2023, OpenAI’s response was essentially: ChatGPT is not a research tool, you should not have used it that way, our terms of service warn against this. That defense is technically correct. It also misses the point.

The model is good at generating output that looks like research. It produces citations confidently. It produces them in the format you requested. When users ask it whether a citation is real, it tells them yes. The tool’s behavior is misleading even when the user understands it is generative.

The structural fix is to separate generation from verification. A system that wants to produce reliable citations cannot rely on the language model to invent them — it must query a real database of published research, retrieve actual records, and only then incorporate them into generated text. This is the architecture used by tools designed specifically for academic work, including Papyra’s citation engine, which retrieves verified records from CrossRef, OpenAlex, and Semantic Scholar before a single reference is inserted into a draft. The author names are real. The DOIs resolve. The journal titles are correct. The output can be defended.

Without that separation, hallucinations are not a bug — they are a structural property of how the model produces text.

What to do if you have already submitted

If you used ChatGPT or another general-purpose AI tool to help with a paper that has already been submitted and you are worried about hallucinated citations, do not panic.

Verify each citation now, using the methods above. Note any that cannot be confirmed.

If you find any fabricated citations, the responsible move is to disclose to your professor before they discover it themselves. A self-disclosed mistake is treated very differently from a discovered one. Most universities have provisions for student-initiated correction that do not result in academic misconduct findings.

If your paper is graded already and the citations were never checked, you are technically in the clear, but the citations remain a liability. They are findable through web archives. Some professors revisit graded papers periodically. The cleanest move is to clean up your reference list now, even retroactively, even informally.

The bigger picture

Academic dishonesty has always been a category that includes both intent and action. With AI, the action — submitting a paper with fabricated sources — is becoming separable from the intent — wanting to cheat. Students who used AI in good faith, trusting that a sophisticated tool would not invent things, are still ending up in academic integrity hearings.

The defense, in every case, is verification. Sources you can produce. Citations that resolve. A bibliography you can defend.

The detectors will keep improving. The hallucinations will keep happening. The students who treat citation as a verified record rather than a generated artifact will be fine.

—

Want every reference in your paper to be a real, verifiable source?

Papyra generates academic writing with citations matched to published papers through CrossRef, OpenAlex, and Semantic Scholar. Every author is real. Every DOI resolves. Every journal exists. The kind of bibliography that holds up to any reasonable scrutiny.

See how Papyra works →

Writing for students who want integrity without guesswork.