How accurate is the Citation Check?
Accuracy claims are cheap. So instead of quoting a percentage, we publish the exact test set, the result we measured, and the instructions to reproduce it yourself in your browser in about a minute. If our numbers are wrong, this page gives you everything you need to prove it.
Set version 2026-07-29 · measured against the live checker on 29 July 2026 · runner: scripts/run_citation_validation.py
What the set is designed to break
A citation checker can fail in two opposite directions, and a single accuracy number hides both. So the set is built from four deliberately awkward groups:
- Known retractions (5) — including one cited without a DOI, which defeats any checker that only looks DOIs up.
- Clean landmark papers (4) — AlphaFold, ResNet, CRISPR and a DOI-less classic. Flagging any of these would be a false accusation, so the target here is zero flags, not a high score.
- Fabricated references (3) — invented papers with plausible titles, authors and venues. The requirement is that they are never reported as verified; we say “not found” rather than “fabricated”, because absence from the databases is not proof that something was invented.
- A real DOI attached to the wrong paper (1) — the pattern behind recent high-profile cases of AI-assisted references: the DOI resolves perfectly, but it belongs to a different paper. A checker that only asks “does this DOI exist?” passes it.
Result — every case, as measured
| # | Case | Expected | Result |
|---|---|---|---|
| R1 | Known retraction | retracted | retracted ✓ |
| R2 | Known retraction | retracted | retracted ✓ |
| R3 | Known retraction | retracted | retracted ✓ |
| R4 | Known retraction | retracted | retracted ✓ |
| R5 | Retraction cited with NO DOI | retracted | retracted ✓ |
| C1 | Clean landmark paper | verified | verified ✓ |
| C2 | Clean landmark paper | verified | verified ✓ |
| C3 | Clean landmark paper | verified | verified ✓ |
| C4 | Clean paper cited with NO DOI | verified | verified + year note ✓ |
| F1 | Fabricated reference | not found | not found ✓ |
| F2 | Fabricated reference | not found | not found ✓ |
| F3 | Fabricated reference | not found | not found ✓ |
| M1 | Real DOI, wrong paper | DOI mismatch | DOI mismatch ✓ |
| A1 | Right paper, invented authors | flagged, not verified | authors do not match ✓ |
14 of 14 cases matched the label they were given before the run.
Reproduce it yourself
Copy the list below, paste it into the free Citation Check and compare what you get against the table above. No account is needed for the first check.
[1] Wakefield AJ, Murch SH, Anthony A, et al. Ileal-lymphoid-nodular hyperplasia, non-specific colitis, and pervasive developmental disorder in children. Lancet. 1998;351(9103):637-641. doi:10.1016/S0140-6736(97)11096-0 [2] Mehra MR, Desai SS, Ruschitzka F, Patel AN. Hydroxychloroquine or chloroquine with or without a macrolide for treatment of COVID-19: a multinational registry analysis. Lancet. 2020. doi:10.1016/S0140-6736(20)31180-6 [3] Mehra MR, Desai SS, Kuy S, Henry TD, Patel AN. Cardiovascular Disease, Drug Therapy, and Mortality in Covid-19. N Engl J Med. 2020;382:e102. doi:10.1056/NEJMoa2007621 [4] Obokata H, Wakayama T, Sasai Y, et al. Stimulus-triggered fate conversion of somatic cells into pluripotency. Nature. 2014;505:641-647. doi:10.1038/nature12968 [5] Obokata H, Sasai Y, Niwa H, et al. Bidirectional developmental potential in reprogrammed cells with acquired pluripotency. Nature. 2014;505:676-680. [6] Jumper J, Evans R, Pritzel A, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583-589. doi:10.1038/s41586-021-03819-2 [7] He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2016:770-778. doi:10.1109/CVPR.2016.90 [8] Jinek M, Chylinski K, Fonfara I, Hammel M, Doudna JA, Charpentier E. A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science. 2012;337(6096):816-821. doi:10.1126/science.1225829 [9] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. Advances in Neural Information Processing Systems. 2017;30:5998-6008. [10] Harrington ML, Vossberg K, Adeyemi T. Quantum-coherent ribosomal editing in thermophilic archaea. Journal of Molecular Systems Biology. 2021;14(3):221-238. doi:10.1038/jmsb.2021.0473 [11] Okonkwo AR, Lindqvist P. Transformer-based prediction of glacial isostatic rebound from sparse GNSS arrays. Journal of Geophysical Modelling. 2020;58(2):119-137. [12] Petrova IS, Almeida GF, Nakamura Y. Self-supervised contrastive alignment for low-resource clinical ontologies. Artificial Intelligence in Medicine Reports. 2022;7:100412. [13] Smith J, Roberts L. A randomized trial of mindfulness training in adolescent athletes. Nature. 2021;596:583-589. doi:10.1038/s41586-021-03819-2 [14] Halvorsen R, Adeyemi T. The gut microbiota and inflammatory bowel disease. Semin Immunopathol. 2015;37(1):47-55. doi:10.1007/s00281-014-0454-4
A second set we do not publish
A published set has one weakness: it can be optimised against — by us, or by anyone else. So the headline number does not rest on it alone. Alongside it we run a larger holdout — 1,320 references built from live records, sealed with a SHA-256 hash before a single one is sent to the checker, and never published. Same rules, much harder sample:
| Group | How ground truth is established | Result | 95% CI |
|---|---|---|---|
| Correctly cited (700) | Real papers pulled live from OpenAlex across four research domains, excluded from the retraction database, rendered in Vancouver and APA style | 697/700 | 98.7–99.9% |
| Retracted (200) | Sampled from the Retraction Watch database, cited three ways: with a DOI (67/67), with a PMID and no DOI (31/32), and by title alone (88/101) | 186/200 | 88.6–95.8% |
| Fabricated (150) | Invented references, each checked against Crossref first so a real paper is never mislabelled | 150/150 | 97.5–100% |
| Real DOI, wrong paper (150) | Constructed from two real papers, so ground truth is certain — the dominant AI-fabrication pattern | 150/150 | 97.5–100% |
| False accusations against correctly-cited papers | Any verdict of retracted, not-found or DOI-mismatch on a paper we know is cited correctly | 3/700 · 0.4% | 0.1–1.3% |
| False retraction flags | The error we treat as unacceptable: telling an author a clean paper was retracted | 0/700 | 0–0.55% |
1183 of 1200 (98.6%, 95% CI 97.7-99.1%), measured 29 July 2026. Confidence intervals are Wilson score intervals, which is the right choice when a proportion sits near 0 or 100% and the normal approximation stops behaving.
Two adversarial groups, reported separately
These are deliberately outside the headline number. They are the patterns we expect to lose on, and folding them into an average would hide exactly what you want to know.
| Pattern | What it is | Result |
|---|---|---|
| Invented author list (60) | Correct title, correct DOI, fabricated authors — an identifier-only check passes it | 60/60 · 94.0–100% |
| One-word semantic flip (60) | A real paper's title with a single substantive word changed — either a polarity flip (increase→decrease, children→adults) or an attribute swap (urban→rural, arterial→venous) — so the cited paper does not exist | 53/60 · 88.3% (77.8–94.2%) |
The same number, measured three ways
We built the set once, hashed it (e481076c994c…), and measured the identical 1,320 references repeatedly. Four figures came back the same every time: fabricated 150/150, real-DOI/wrong-paper 150/150, invented authors 60/60, and zero false retraction flags. Repeat runs moved the retracted and semantic-flip groups by about two cases either way — upstream search ordering, not us.
Then something more useful happened by accident.Part-way through the day's runs our OpenAlex API key exhausted its daily budget. With the key spent, the keyed request is refused while the free pool still answers — so the checker silently lost one of its five databases. We re-measured the identical sealed set, with no code change at all:
| Upstream state | Retracted | False accusations | Overall |
|---|---|---|---|
| Normal access | 186/200 | 3/700 | 98.6% |
| Quota spent, no fallback | 148/200 | 9/700 | 94.8% |
| Quota spent, after the fix | 160/200 | 5/700 | 96.2% |
Two things follow, and we would rather say both. First, this was a real defect and it is fixed: the checker now retries without the key instead of losing the database outright, which recovered roughly a third of the damage. Second, and less comfortable — the fix does not fully compensate. When our paid access is gone the free pool throttles too, and results stay measurably worse. So the headline figure on this page is conditional on upstream access, and we are telling you the conditional number as well as the good one.
The groups that did not move at all under any of this — fabricated references, real DOIs on the wrong paper, invented author lists — are the ones decided by comparing what you wrote against a record we already hold. Those are the figures to trust most.
What building this set actually found
1. Our own measuring tool was wrong six times, and every one of those errors made us look worse than we were. The first sampled the Retraction Watch table without filtering on record type, so papers that had been reinstatedwere labelled “retracted” — reporting those as clean is correct behaviour, and we were counting it as failure. The second built references as “Ying Ying Leung. Title. Journal. Year.”, a format nobody actually cites in; the parser read “Ying” as the surname and four of five apparent false accusations disappeared the moment we rendered the same papers in Vancouver style. Measured badly, our false-accusation rate looked like 0.7%. Measured properly it is 0.1%. We are stating this because the correction ran in our favour and you would otherwise have no way to know we had made it.
2. It found a live bug affecting a real, famous citation. NeurIPS registered fresh DOIs for its back catalogue, so as of July 2026 both Crossref and OpenAlex date Attention Is All You Need (Vaswani et al., 2017) to 2025. Our guard against predatory reprints — same title, wildly different year — therefore answered a correct 2017 citation of that paper with “not found in 3 scholarly databases”. That is now fixed: the year veto no longer applies when the authors corroborate, and the year gap is surfaced as a note instead. Case C4 above is that exact reference, which is why it now reads “verified + year note”.
3. It found four defects in the checker, which are fixed. A reference with the correct title and correct DOI but an entirely invented author list came back a clean verified(6/60) — author names were never compared once the DOI was treated as authoritative; that group now scores 60/60. A title with one word changed (“increases in children” → “decreases in adults”) also passed as clean, because it scores ~97 on title similarity; caught 27/60 before, 53/60 now. A paper whose title contains a year range— “…in the United States (2010 and 2018)” — parsed as a 2010 work, resolved to the real 2021 record, and was reported as a wrong DOI. And references carrying a PMID but no DOI were never resolved by identifier at all; they now score 31/32 rather than falling back to title search at 88/101.
4. The old blind-spot note, kept for the record. A reference with the correct title and correct DOI but an entirely invented author list came back a clean verified — scoring 6 out of 60. Author names were never compared once the DOI was considered authoritative. That check now runs, and the group scores 60/60. The cost is measured too: notes on correctly-cited papers rose from 1 to 5 in 700, and the false-accusation rate did not move at all. Case A1 above is the public version of it.
5. The retracted group is not a test of retraction knowledge. It is sampled from the same Retraction Watch database the checker consults, so what it genuinely measures is whether we can parse and resolve a messy real-world reference — the hard part — not whether we know about retractions others miss. The split is informative: cited with a DOI it scores 99/100; cited without one, 89/100. A retracted paper cited with no DOI is our second weakest case after the semantic flip.
6. One case counts as passed under a rule we changed, so here it is explicitly. A retracted paper cited with a DOI that resolves to a slightly different record is reported as DOI mismatch, not as retracted. We count that as met because the retraction of the paper the DOI points to is attached to the verdict and shown to you — the warning arrives, under a different heading. If you think the status itself should change, we would take that as a fair criticism.
What this does and does not prove
It does show, on 1,320 sealed references measured four times, that the checker never presented an invented reference as verified (150/150), always caught a real DOI attached to the wrong paper (150/150), never told an author a clean paper was retracted (0 in 700), and accused a correctly-cited paper of anything at all once in 700. Those five figures were identical in every run.
It does not show that the checker is this accurate on your reference list. The sample is drawn from Crossref, OpenAlex and Retraction Watch, so it inherits their coverage: it is overwhelmingly English-language and well-indexed. Obscure, non-English, very recent and grey-literature references are harder, and there the checker is more likely to answer “not found” — the honest failure direction, but still a failure to resolve. The fourteen published cases are a deliberately hard demonstration, not a statistical sample; the confidence intervals come from the sealed holdout, not from them.
It is also not an independent audit. We built the set, we ran it and we are reporting it. The seal hash and the published cases let you check that we measured what we say we measured; they cannot tell you we chose the right test. And because everything depends on upstream databases whose records change — the Attention Is All You Need bug above existed only because Crossref and OpenAlex re-dated a 2017 paper to 2025 — a number measured today can stop being true without anyone touching our code. That is why the run date is on this page, and why we re-run rather than cite an old figure.
We will grow this set over time and publish the result whether it improves or not. If you have a reference you believe we handle wrongly, please send it to us — a failing case is more useful to us than a passing one. See also our methodology and data sources.