ChatGPT Fake Citations: Why and What to Do
ChatGPT invents citations that look real. Why it happens, how often, how to spot the fakes, and how to fix your reference list before you submit.
The short answer
Yes, ChatGPT invents citations, and yes, some of yours are probably fake if you have not checked them.
A Scientific Reports study collected 636 citations from ChatGPT across 42 topics and found that 55% of GPT-3.5's citations and 18% of GPT-4's were fabricated. Many of the real ones carried errors too. (Scientific Reports)
It happens because the model generates plausible text, and a citation is just very structured text. Nothing in that process checks whether the reference exists.
The fix takes one evening, not one crisis:
- Separate the references an AI tool touched from the ones you found yourself.
- Verify every AI-touched reference against real databases.
- Delete the fakes and replace them with sources you actually read.
- Spot-check the rest.
Here is why it happens, how often, and the full cleanup workflow.
Why does ChatGPT make up citations?
ChatGPT predicts one token at a time based on patterns in its training data. It learned what citations look like: author names that fit a field, journal titles that sound established, years that match a topic's research wave, DOIs with the right shape. When you ask for sources, it produces text with exactly those properties.
Nothing in that process checks whether the produced reference exists. There is no built-in lookup against Crossref or PubMed. Existence was never part of the objective. Plausibility was.
So the model does not lie, and it does not retrieve. It composes. Sometimes the composition lands on a real paper, because real papers dominated the training data. Often it lands on a reference that never existed, assembled from real fragments.
That also explains why asking ChatGPT "are these sources real?" does not help. The model has no internal flag separating retrieved facts from composed text, so it often confirms its own inventions with full confidence.
How often it happens
Fabrication rates are measured, not anecdotal. The peer-reviewed numbers:
| Study | Model tested | Fabricated citations |
|---|---|---|
| Walters & Wilder, Scientific Reports, 2023 | GPT-3.5 / GPT-4 | 55% / 18% |
| Bhattacharyya et al., Cureus, 2023 | GPT-3.5 | 47% (only 7% were real and fully accurate) |
| Chelli et al., JMIR, 2024 | GPT-4 / Bard | 28.6% / 91.4% |
| Linardon et al., JMIR Mental Health, 2025 | GPT-4o | 19.9% |
Walters and Wilder collected 636 citations across 42 topics and found 55% of GPT-3.5's citations and 18% of GPT-4's were fabricated, with substantive errors in many of the real ones. (Scientific Reports)
In medicine, Bhattacharyya and colleagues found that of 115 references GPT-3.5 generated, 47% were fabricated, 46% were real but inaccurate, and only 7% were both authentic and accurate. (Cureus)
For systematic-review prompts, Chelli and colleagues measured hallucination rates of 28.6% for GPT-4 and 91.4% for Google's Bard. (JMIR, 2024)
In mid-2025, Linardon and colleagues tested GPT-4o: 19.9% of 176 generated citations were fabricated, and 45.4% of the real ones contained errors, most often broken DOIs. (JMIR Mental Health, 2025)
Newer models improved. None got close to safe. Roughly one in five citations from a current model is invented, before counting the merely wrong ones.
Niche topics make it worse, and your thesis is a niche topic
The 2025 study found something that matters for thesis writers specifically. Fabrication tracked topic familiarity. On a heavily published topic, major depressive disorder, GPT-4o fabricated 6% of citations. On less published topics, binge eating disorder and body dysmorphic disorder, fabrication jumped to 28% and 29%. (JMIR Mental Health, 2025)
The mechanism explains why. Dense literature gives the model strong memorized patterns, so its plausible-text generator reproduces real references more often. Thin literature forces it to compose, and composition invents.
A thesis lives in thin literature by design. You write about the narrow intersection nobody covered, which is exactly where fabrication rates triple. The general fabrication statistics understate your risk.
What a fake citation looks like
ChatGPT does not usually invent obviously fake references. It blends.
The common patterns:
- Real authors, invented title. The researchers exist and work in the field. The paper does not.
- Real title, wrong everything else. The paper exists, but the year, journal, or author order is wrong.
- Working DOI, wrong paper. The DOI resolves, just not to the source your citation describes.
- Frankenstein reference. Author from one paper, title from a second, venue from a third.
- Too perfect a fit. The title restates your claim almost word for word. Real literature rarely matches your sentence that exactly.
Every fragment passes a plausibility check, and your eye checks fragments. A real researcher's name, because that name appeared near your topic in training data. A real journal, because that journal publishes this field. A title assembled from the field's phrasing. The tell sits in the combination: that author never wrote that paper, that journal never published that title.
That is why skimming your reference list catches nothing. Only a database lookup or an exact-title search exposes the fake, which is why our guide on how to check if a citation is real starts with the DOI and the exact-title search rather than with reading the reference list harder.
Triage your reference list
Sort your references into three buckets. Be honest, nobody is watching.
Bucket 1: an AI tool suggested it. Highest risk. Verify every single one before it stays in your bibliography.
Bucket 2: you copied it from another paper's reference list without opening it. Medium risk. Fabrication is unlikely, but copied metadata errors are everywhere, and the source may not say what the citing paper claimed.
Bucket 3: you read the source yourself. Low risk. Spot-check the metadata.
This ordering saves you from the two bad extremes: trusting everything, or panic-rechecking 200 references the night before submission.
Verify the risky ones
For each bucket 1 and 2 reference, run the standard checks. Resolve the DOI at doi.org and confirm the page matches. Search the exact title in quotes on Google Scholar. Compare authors, year, venue, and pages field by field.
With a long list, batch the first pass. The free Check Your Draft citation checker cross-checks every reference against Semantic Scholar, OpenAlex, arXiv, PubMed, CrossRef, Google Books, DBLP, and Open Library, then flags each entry as verified, hallucinated, or outdated. The hallucination detector does the same job for a single suspicious source: it extracts the citation details, searches the databases, and shows the strongest match with the evidence behind it. Our guide to the best citation verification tools compares the alternatives.
Then read what the tool flags. One warning applies to every checker, including ours: "verified" means the source exists and the metadata matches. Only reading proves the source supports your sentence.
Replace, do not patch
When a reference turns out fake, delete it. Then fix the hole properly.
Do not ask ChatGPT for a replacement citation. You would be drawing from the same well that produced the fake. Search the literature instead: Google Scholar, Semantic Scholar, or your library's databases. Find a real paper, read at least the abstract and the relevant section, and cite it because it supports your point.
Sometimes no real source supports the claim. That is information. The claim came from the model, not the literature, so rewrite or cut the sentence. A fabricated citation usually props up a fabricated certainty.
While you are in the reference list, fix the survivable errors too: wrong years, mangled author names, preprints that now have published versions. The citation updater finds published versions of arXiv preprints automatically, and LaTeX users can clean duplicate entries with the BibTeX cleaner.
Will you get caught if you skip this?
The odds are worse than they feel.
Nothing in the standard submission pipeline verifies references automatically. Turnitin checks text similarity, not source existence, as our post on whether Turnitin checks citations explains. A fabricated reference is original text, so similarity software waves it through.
What catches fabricated citations is a person: a supervisor who knows the field, notices an unfamiliar title, and searches it. Markers do exactly that, and a reference that does not exist is not a debatable accusation. A detector score invites argument. A dead citation ends one. Universities treat fabricated sources as misconduct regardless of whether the fabrication was yours or a chatbot's, because you signed the submission. Our post on can professors tell if you used ChatGPT makes the same point from the other side: a fabricated reference is the one AI trace a marker can prove.
The consequences stopped being hypothetical in 2023. A federal judge fined two lawyers and their firm $5,000 after they filed a brief citing six court cases ChatGPT had invented, and stood by the fakes when challenged. (CBS News) Springer Nature retracted a $169 machine-learning textbook in 2025 after Retraction Watch checked 18 of its citations and found two-thirds either did not exist or were badly wrong. (Retraction Watch) Senior Australian academics apologised to a parliamentary inquiry after AI-generated case studies in their submission, produced with Google Bard, turned out to be fiction. (Information Age)
Universities responded. Libraries now publish dedicated guides on the problem. UNC Charlotte's guide warns that a hallucinated citation "may look real and mix together a combination of real and made-up elements." (UNC Charlotte Library) When librarians write warning pages, markers read them too.
Students sit right in the blast radius. In the UK, the 2026 HEPI student survey found 94% of full-time undergraduates use generative AI to help with assessed work, and 12% submit AI-generated text directly. (HEPI, 2026) Combine the numbers: most students use these tools, current models invent roughly one citation in five, and the rate climbs on exactly the niche topics theses cover. Unchecked reference lists now carry fabricated sources as a statistical default, not as bad luck.
You do not need luck here. You need an evening with a checker.
Using ChatGPT for sources without the mess
AI tools can help with literature work when you flip the workflow.
Ask for search directions, not citations: topics, keywords, author names, subfields to explore. Then find the actual papers yourself in Scholar or Semantic Scholar. When a chatbot does hand you a specific reference, treat it as a lead to verify, never as a bibliography entry.
Keep the boundary clean: the model suggests, the databases confirm, you read. Students who follow that split get the speed of AI assistance without inheriting its fabrications. For which tools fit which stage of the work, see our guide to the best AI for academic writing.
FAQ
Does ChatGPT know when it invents a citation?
No. The model has no internal flag separating retrieved facts from composed text. Confidence in the output tells you nothing, which also means asking it "are these sources real?" often returns a confident yes for invented references.
Are the newer models safe now?
Safer, not safe. GPT-4 fabricated 18% of citations in the 2023 Scientific Reports study, and GPT-4o still fabricated 19.9% in the 2025 JMIR Mental Health study, with rates near 29% on niche topics. One reference in five is a career risk in a thesis. Verify regardless of model.
Do browsing or research modes solve this?
Retrieval-backed modes ground answers in fetched pages, which lowers fabrication when they actually retrieve. You still verify, because models mix retrieved and composed content, and metadata errors survive retrieval. The checking workflow stays the same either way.
Can ChatGPT check its own citations?
Do not trust it to. Models often confirm their own inventions with full confidence. Verification needs real database lookups, either manual or through a citation checker.
Is one fabricated citation really a big deal in a thesis?
Yes. A marker who finds one invented source rereads everything else with new eyes. Universities treat fabricated references as an integrity issue because you attested to sources you never read. The cost of checking is one evening. The cost of one fake is your credibility.
I already submitted with unchecked references. What now?
Check them now anyway. If you find a fabricated source, talk to your supervisor before someone else finds it. Owning the error early, with a corrected reference list in hand, goes over materially better than being confronted with it.
Practical takeaway
ChatGPT composes citations the way it composes everything: by plausibility, with no existence check. The measured fabrication rate on current models sits near one in five and climbs on niche topics like yours.
So sort references by who found them. Verify everything an AI touched. Delete fakes, replace them with papers you read, and fix the metadata errors while you are in there. Run the free citation checker for the batch pass, and spend your own time on the flagged entries. One evening, and this whole risk class is gone from your draft.
Related reading