Note · sourced · on AI and verification
A real citation that does not support the claim

The failure that survives grounding.

An invented citation announces itself: the link dies, the identifier returns nothing. The harder case is the citation that resolves to a real paper by real authors and does not contain the finding attached to it. Retrieval does not catch this, and neither does a quick look at the reference list.

Two things can go wrong when a system attaches a source to a sentence. The source can fail to exist, or the source can exist and fail to say what the sentence says. Only the first one is easy to catch, and it is the one that gets the attention.

A fabricated citation is a mechanical error with a mechanical test. Paste the identifier into PubMed. If nothing comes back, the citation is void, and no medical knowledge was required to find that out. This is why fabricated citations are the failure most people can name and most tools now screen for.

The second failure has no mechanical test. Everything about it is in order: the identifier resolves, the authors are real, the journal is indexed, the paper is about the right disease. Whether the paper contains the specific finding claimed is a question that can only be answered by reading it, and by reading it closely enough to notice that the study measured a different outcome, or enrolled a different population, or reported the association as conditional where the sentence reports it flatly.

The measurements that exist put this second failure well above the level anyone would tolerate if they saw it counted. A Stanford team ran a human evaluation of four generative search engines and found that, on average, 51.5 percent of generated sentences were fully supported by their citations, and 74.5 percent of citations supported the sentence they were attached to (Liu, Zhang and Liang, Findings of EMNLP, December 2023). The second figure is the one worth sitting with. Roughly a quarter of the citations these systems produced were attached to sentences those citations did not support.

In medicine specifically

What an audit of medical answers found

A later group built an automated framework, SourceCheckup, to put the same question to medical queries: given a model's answer and the sources it offered, does each statement actually appear in one of them? They ran 800 questions, half generated from Mayo Clinic reference pages and half taken from real queries posted to Reddit's r/AskDocs, producing roughly 58,000 statement and source pairs across seven models, and validated the automated judge against three licensed physicians, reaching 88.7 percent agreement with the doctors' consensus, higher than the doctors' 86.1 percent agreement with each other (Wu et al., Nature Communications, April 2025).

The best-performing configuration was GPT-4o with live web retrieval, which produced valid URLs reliably. Its response-level support was 55 percent, meaning that in 45 percent of answers at least one statement was not supported by any source the model had provided. On a separate consumer-health question set, a clinician reading the answers end to end judged 40.4 percent of responses fully supported, close to the framework's own 42.4 percent on the same questions.

Two details rule out the obvious objections. The authors sampled 110 statement and source pairs their system had marked unsupported and had physicians re-read them; the physicians confirmed 105. And when they merged all of a response's sources together rather than checking statements against them one at a time, 95.1 percent of the previously unsupported statements remained unsupported. The gap is in the writing, not the bookkeeping.

On what this means for retrieval as a fix, the authors report that a substantial fraction of the references produced by retrieval-augmented models did not fully support the claims made in the responses, and offer two candidate explanations without settling between them: the model extrapolating retrieved information with what it already carries from training, or hallucination.

The mechanism

Retrieval scores topical similarity; support is a different question

A retrieval system is built to return documents that are about the query. That is what its scoring rewards and what it is evaluated on. Whether a returned document contains a passage that establishes the specific sentence a model is about to write is a separate property, and nothing in the retrieval step tests for it. Grounding narrows the candidate pool. It does not verify entailment.

The break happens in three recognizable places. The retrieved paper can be adjacent: the right condition, a different population, outcome, or timeframe. The generated sentence can overreach the passage it was drawn from, dropping a qualifier, a subgroup, or a confidence interval that the source treated as essential. Or the retrieved page can be reporting rather than establishing, repeating a claim it took from somewhere else, so the citation lands on a repetition and the primary evidence is never inspected.

Law offers the cleanest documented case, because the systems there were sold on the claim that this problem had been solved. A preregistered evaluation of the leading commercial legal research tools tested products marketed as producing "hallucination-free" citations and found that Lexis+ AI and Thomson Reuters' Westlaw AI-Assisted Research and Ask Practical Law AI each hallucinated between 17 and 33 percent of the time (Magesh et al., Journal of Empirical Legal Studies, 2025).

What makes that study useful here is its definition. The authors count a response as hallucinated if it is incorrect or misgrounded, where a misgrounded response is one whose key propositions are cited but where the citation misinterprets the source or points to an inapplicable one. Their worked example is a tool citing Planned Parenthood v. Reynolds, a real case that had not been overturned, while relying on that case's description of Casey, which had been. Every link resolves. The reasoning rests on law that no longer stands.

"These errors are potentially more dangerous than fabricating a case outright, because they are subtler and more difficult to spot."
Magesh, Surani, Dahl, Suzgun, Manning and Ho, on misgrounded citations

The same structure transfers to medicine without modification. A real trial, correctly cited, attached to a claim about a population it did not enroll, is harder to catch than an invented trial and lands with more authority.

The base rate underneath

The published literature has the same failure, at a measured rate

It would be convenient to treat this as a machine problem. The published record does not allow that. A systematic review and meta-analysis of quotation accuracy in medicine pooled 46 studies covering roughly 32,000 quotations and references and found that 16.9 percent of quotations were incorrect, with about half of those classed as major errors, at 8.0 percent (Baethge and Jergas, Research Integrity and Peer Review, July 2025). Their definition of a quotation error is the failure this note is about: a reference not supporting the authors' claim. Meta-regression across the period covered found no significant improvement over time.

What that rate produces at scale was documented in a single disease. A neurologist reconstructed the complete citation network for one claim, that beta amyloid precursor protein or beta amyloid is abnormally and specifically present in the muscle fibres of patients with inclusion body myositis, across all English-language papers indexed in PubMed that addressed it. The network held 242 papers and 675 citations, generating 220,553 citation paths supporting the claim (Greenberg, BMJ, July 2009).

Inside that network, the supportive primary-data papers received 94 percent of the 214 citations made to primary data, while the six papers whose data weakened or refuted the claim received 6 percent. The paper also names a practice it calls citation diversion: citing a source while altering what it says. One primary-data paper had reported no beta amyloid precursor protein or beta amyloid in three of five patients studied, and its presence in only a "few fibres" in the remaining two. Three later papers cited that result as having confirmed the claim. Over the following decade, those three citations developed into 7,848 supporting citation paths.

Three sentences, each attached to a real and correctly identified paper, each describing as confirmation a result that paper had reported as largely negative. This is the material a retrieval system indexes. A model searching that corpus and reproducing its citation practices is not malfunctioning. It is doing what the literature taught it.

The checks

Six questions that test whether a citation carries its sentence

These apply to any sourced claim, whether a model, a review article, or a clinician's handout produced it. They take longer than checking that a link resolves, which is the point: the cheap check and the useful check are different checks.

What to checkWhat it tells you
Find the sentence in the source Open the paper and locate the specific passage that carries the claim, not the paper's general topic. If the claim is real, a passage exists that says it. If no passage can be found after looking, the citation is a pointer to a subject rather than evidence for a statement.
Match the population Compare who was enrolled against who the claim is about. Age range, disease subtype, severity, and case definition are where adjacent sources hide. A finding in one subtype cited for the condition as a whole is a real result attached to the wrong people.
Match the strength Check whether the source says "associated with" where the sentence says "causes," or reports a result under conditions the sentence has dropped. Overreach usually shows up as a missing qualifier rather than a wrong number.
Establishing or repeating Ask whether the cited source generated the evidence or is quoting someone else. If it is quoting, the citation belongs to the paper that ran the study. Chains of repetition are how a weak result acquires the appearance of a settled one.
Read what was left out Read the sentences around the quoted one, and the limitations section. Sources routinely qualify their own findings in ways that citations discard. A source that says "in a small, unreplicated sample" is not supporting a categorical claim.
Look for the disconfirming source Search for the result that would weaken the claim and see whether it exists and went uncited. Citation bias is invisible from inside a reference list, because the missing papers leave no mark on the page. This is the check that catches the failure the other five cannot.

Applied to a single claim, these take a few minutes and settle most cases. Applied to every claim in a long answer, they cost more than the answer saved, which is the arithmetic that decides where verification is worth spending. That question has its own method.

The standing practice

Record the passage, not just the paper

A citation to a document is a promise that something inside it supports the claim. The promise is only checkable by whoever is willing to go and read the document, which in practice means it is checkable by almost no one. Recording which passage was relied on converts the promise into evidence that travels with the claim.

VictorOS is built on that distinction. When a claim enters a condition map as confirmed, the record carries the source identifier, the date a person confirmed it, and the literal span of text from that source that supports it. The unit of custody is a claim tied to a passage rather than a claim tied to a paper. Where no passage can be found, the claim does not enter the map as established, and where the evidence is contested, the map records the dispute instead of picking a side.

This is deliberately more expensive than citing. It is also the only version of the work that can be audited by the person the claim is about, which is the standard the project is trying to meet.

The principle underneath

A citation is a promise most readers cannot collect on

The reason misattribution persists at a stable 17 percent in the published literature, and higher in generated text, is that verifying it requires access, time, and training. Someone has to retrieve the paper, often past a paywall, read a methods section, and hold the claim next to what the study measured. Clinicians have that capacity and rarely have the minutes. Patients usually have neither.

So the citation functions as a signal rather than as evidence. It says that a claim belongs to the kind of statement that has support somewhere, and readers reasonably treat the presence of a reference as the end of the inquiry. Generated text inherits that trust and produces references at a volume no reader can audit, which widens the gap between what is cited and what is supported without changing how the page looks.

Closing that gap does not require patients to evaluate methodology. It requires the passage to travel with the claim, so that "does the source say this" can be answered by looking rather than by trusting. The six checks are how a reader recovers that when it is missing. Carrying it by default is how a source stops making the reader do it.

The invented citation is the failure that gets caught, because catching it is free. The genuine citation attached to a sentence it does not support is the failure that propagates, and in Greenberg's network it propagated for a decade into 7,848 chains resting on three papers that had called a largely negative result a confirmation. Nothing on any of those pages looked wrong.

Carrie Schluter, BCPA
Board Certified Patient Advocate
Author of VictorOS
PatientLead Health
Current as of August 2026

Carrie Schluter is a Board Certified Patient Advocate and the author of VictorOS, a rare disease knowledge system built on a single rule: every clinical claim traces to a primary source with a human confirmer.

The rule is written that way on purpose. Tracing to a source is the part that can be automated and the part that fails quietly. The confirmer is the part that reads the passage and decides whether it says what the claim says.

Sources
  1. Liu N, Zhang T, Liang P. Evaluating Verifiability in Generative Search Engines. Findings of the Association for Computational Linguistics: EMNLP 2023. December 2023;7001–7025. aclanthology.org/2023.findings-emnlp.467
  2. Wu K, Wu E, Wei K, Zhang A, Casasola A, Nguyen T, et al. An automated framework for assessing how well LLMs cite relevant medical references. Nature Communications. 2025;16:3615. nature.com/articles/s41467-025-58551-6
  3. Magesh V, Surani F, Dahl M, Suzgun M, Manning CD, Ho DE. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. Journal of Empirical Legal Studies. 2025;22(2):216–242. Published online 23 April 2025. doi.org/10.1111/jels.12413
  4. Magesh V, Surani F, Dahl M, Suzgun M, Manning CD, Ho DE. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. Preprint, 30 May 2024 (definitions of groundedness and the Reynolds example are quoted from this version). arxiv.org/abs/2405.20362
  5. Baethge C, Jergas H. Systematic review and meta-analysis of quotation inaccuracy in medicine. Research Integrity and Peer Review. 2025;10:13. Published 23 July 2025. link.springer.com/article/10.1186/s41073-025-00173-z
  6. Greenberg SA. How citation distortions create unfounded authority: analysis of a citation network. BMJ. 2009;339:b2680. ncbi.nlm.nih.gov/pmc/articles/PMC2714656
  7. Jergas H, Baethge C. Quotation accuracy in medical journal articles: a systematic review and meta-analysis. PeerJ. 2015;3:e1364 (the earlier review, reporting a total quotation error rate of 25.4 percent across 28 studies; superseded by the 2025 analysis above). peerj.com/articles/1364

Verification note: every figure in this note was read from the primary source rather than from a summary of it, and several were narrowed after checking. The 51.5 and 74.5 percent figures are averages across four generative search engines (Bing Chat, NeevaAI, perplexity.ai, YouChat) as evaluated by human raters in Liu et al. (2023); they describe citation recall and citation precision on those systems as they existed in 2023, and are not a measurement of any current product. In Wu et al. (2025), the 800 questions are 400 generated by a model from Mayo Clinic reference documents and 400 real user queries from Reddit's r/AskDocs, and the text says "generated from" rather than "drawn from" for that reason. The 55 percent response-level support figure is for GPT-4o with retrieval across the full 800-question set; the 40.4 percent end-to-end clinician figure and the 42.4 percent framework figure are from a 100-question subset of HealthSearchQA, and the paper separately reports 31.0 percent response-level support on its Reddit-derived questions, so the rate varies substantially with question type and the text says which set each figure comes from. The 88.7 and 86.1 percent agreement figures are the automated judge against the doctors' consensus and the average agreement between doctors, respectively; the paper reports no statistically significant difference between the judge's annotations and the doctors' consensus (p = 0.21), so the comparison establishes that the judge is usable, not that it outperforms clinicians. The paper's two candidate explanations for unsupported statements under retrieval are offered as conjecture in its discussion ("this might be due to... or hallucination"), and the text presents them that way rather than as a finding. Magesh et al. is cited in its peer-reviewed 2025 form for the 17 to 33 percent range, which appears in both the published and preprint abstracts; the definitions of grounded, ungrounded, and misgrounded, the Planned Parenthood v. Reynolds example, and the quoted sentence about subtlety were read in the May 2024 preprint and are attributed to it, because the published article is paywalled and the wording could not be re-checked against it. Baethge and Jergas (2025) report 16.9 percent (95 percent CI 14.1 to 20.0) of quotations incorrect and 8.0 percent (95 percent CI 6.4 to 10.0) major, pooled across 46 studies and roughly 32,000 quotations, with a meta-regression slope of −0.002 (p = 0.85) indicating no significant change over the period covered; the "no improvement over time" statement is theirs and applies to the studies included, not to any single journal. The earlier 2015 review by the same authors reported 25.4 percent, and the gap is a change in how the denominator was counted rather than a change in the literature: the 2025 paper re-ran the 2015 dataset under its current method and obtained 17.0 percent (95 percent CI 12.9 to 22.0), which is why the note treats roughly 17 percent as the stable figure and lists the 2015 review as superseded. The Greenberg (2009) network figures are 242 papers, 675 citations, and 220,553 supporting citation paths as reported in the abstract; the figure legend gives 218 nodes because 24 papers make and receive no citations about the claim, and the body text separately gives 220,609 paths in total, so the text here follows the abstract's count of supporting paths. The claim is stated with "specifically" because the specificity is what the refuting papers weakened: they reported these molecules in muscle fibres across many other diseases, which is a different objection from their being absent. The 94 and 6 percent split is of the 214 citations made to primary-data papers (p = 0.01), not of all citations in the network. The 7,848 figure is the number of supporting citation paths that developed over ten years from the three citations that described one primary-data paper as having confirmed the claim; Greenberg writes that whether those data confirm the claim is "perhaps open to interpretation" before concluding that they are "at the least... exaggerated and generalised," so the text describes what the cited paper reported (a largely negative result) rather than asserting that the three citing papers were flatly wrong, and does not adjudicate the biology. The internal description of VictorOS records was checked against the stored graph provenance, where a confirmed claim carries the source identifier, the confirmation date, and the supporting span of source text.

VictorOS organizes evidence; it does not practice medicine. This note describes a method for testing whether a cited source supports the claim attached to it. It does not give medical advice, diagnose, or recommend or discourage any treatment for any person. The studies named here are cited with their dates and identifiers so you can check them yourself, which is the entire argument. Disease facts in VictorOS guides are based on articles retrieved from PubMed and cited with stable identifiers. These materials support your medical team; your clinicians remain the ones who diagnose and treat.