How accurate is the Citation Check?
Accuracy claims are cheap. So instead of quoting a percentage, we publish the exact test set, the result we measured, and the instructions to reproduce it yourself in your browser in about a minute. If our numbers are wrong, this page gives you everything you need to prove it.
Set version 2026-07-28 · measured against the live checker on 28 July 2026 · runner: scripts/run_citation_validation.py
What the set is designed to break
A citation checker can fail in two opposite directions, and a single accuracy number hides both. So the set is built from four deliberately awkward groups:
- Known retractions (5) — including one cited without a DOI, which defeats any checker that only looks DOIs up.
- Clean landmark papers (4) — AlphaFold, ResNet, CRISPR and a DOI-less classic. Flagging any of these would be a false accusation, so the target here is zero flags, not a high score.
- Fabricated references (3) — invented papers with plausible titles, authors and venues. The requirement is that they are never reported as verified; we say “not found” rather than “fabricated”, because absence from the databases is not proof that something was invented.
- A real DOI attached to the wrong paper (1) — the pattern behind recent high-profile cases of AI-assisted references: the DOI resolves perfectly, but it belongs to a different paper. A checker that only asks “does this DOI exist?” passes it.
Result — every case, as measured
| # | Case | Expected | Result |
|---|---|---|---|
| R1 | Known retraction | retracted | retracted ✓ |
| R2 | Known retraction | retracted | retracted ✓ |
| R3 | Known retraction | retracted | retracted ✓ |
| R4 | Known retraction | retracted | retracted ✓ |
| R5 | Retraction cited with NO DOI | retracted | retracted ✓ |
| C1 | Clean landmark paper | verified | verified ✓ |
| C2 | Clean landmark paper | verified | verified ✓ |
| C3 | Clean landmark paper | verified | verified ✓ |
| C4 | Clean paper cited with NO DOI | verified | verified ✓ |
| F1 | Fabricated reference | not found | not found ✓ |
| F2 | Fabricated reference | not found | not found ✓ |
| F3 | Fabricated reference | not found | not found ✓ |
| M1 | Real DOI, wrong paper | DOI mismatch | DOI mismatch ✓ |
13 of 13 cases matched the label they were given before the run.
Reproduce it yourself
Copy the list below, paste it into the free Citation Check and compare what you get against the table above. No account is needed for the first check.
[1] Wakefield AJ, Murch SH, Anthony A, et al. Ileal-lymphoid-nodular hyperplasia, non-specific colitis, and pervasive developmental disorder in children. Lancet. 1998;351(9103):637-641. doi:10.1016/S0140-6736(97)11096-0 [2] Mehra MR, Desai SS, Ruschitzka F, Patel AN. Hydroxychloroquine or chloroquine with or without a macrolide for treatment of COVID-19: a multinational registry analysis. Lancet. 2020. doi:10.1016/S0140-6736(20)31180-6 [3] Mehra MR, Desai SS, Kuy S, Henry TD, Patel AN. Cardiovascular Disease, Drug Therapy, and Mortality in Covid-19. N Engl J Med. 2020;382:e102. doi:10.1056/NEJMoa2007621 [4] Obokata H, Wakayama T, Sasai Y, et al. Stimulus-triggered fate conversion of somatic cells into pluripotency. Nature. 2014;505:641-647. doi:10.1038/nature12968 [5] Obokata H, Sasai Y, Niwa H, et al. Bidirectional developmental potential in reprogrammed cells with acquired pluripotency. Nature. 2014;505:676-680. [6] Jumper J, Evans R, Pritzel A, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583-589. doi:10.1038/s41586-021-03819-2 [7] He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2016:770-778. doi:10.1109/CVPR.2016.90 [8] Jinek M, Chylinski K, Fonfara I, Hammel M, Doudna JA, Charpentier E. A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science. 2012;337(6096):816-821. doi:10.1126/science.1225829 [9] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. Advances in Neural Information Processing Systems. 2017;30:5998-6008. [10] Harrington ML, Vossberg K, Adeyemi T. Quantum-coherent ribosomal editing in thermophilic archaea. Journal of Molecular Systems Biology. 2021;14(3):221-238. doi:10.1038/jmsb.2021.0473 [11] Okonkwo AR, Lindqvist P. Transformer-based prediction of glacial isostatic rebound from sparse GNSS arrays. Journal of Geophysical Modelling. 2020;58(2):119-137. [12] Petrova IS, Almeida GF, Nakamura Y. Self-supervised contrastive alignment for low-resource clinical ontologies. Artificial Intelligence in Medicine Reports. 2022;7:100412. [13] Smith J, Roberts L. A randomized trial of mindfulness training in adolescent athletes. Nature. 2021;596:583-589. doi:10.1038/s41586-021-03819-2
A second set we do not publish
A published set has one weakness: it can be optimised against — by us, or by anyone else. So the headline number does not rest on it alone. Alongside it we run a larger holdout that is rebuilt from live records each time and whose references are never published. Same rules, harder sample:
| Group | How ground truth is established | Result |
|---|---|---|
| Retracted (20) | Sampled from the Retraction Watch database; half cited without a DOI | 19/20 |
| Clean (20) | Real papers pulled live from OpenAlex, excluded from the retraction database | 20/20 · 0 false flags |
| Fabricated (12) | Invented references, each checked against Crossref first so a real paper is never mislabelled | 12/12 |
| Real DOI, wrong paper (8) | Constructed from two real papers, so ground truth is certain | 8/8 |
59 of 60 (98%), measured 28 July 2026. Two things are worth saying plainly. First, the retracted group is sampled from the same database the checker consults, so what it really measures is whether we can parse and resolvea messy real-world reference — the hard part — not whether we know about retractions others don't. Second, building this set found a real bug: references whose venue our parser failed to read were being classified as books and quietly skipped, so two retracted papers came back clean. That is now fixed, which is what moved the score from 57/60 to 59/60.
The one remaining miss is honest and still open: when a DOI points to a different paper, we report the mismatch but do not additionally check whether the paper it points to was retracted. The user is warned about the reference either way, but that compound signal is missing today.
What this does and does not prove
It does show that on a set built specifically to trip it up, the checker caught every retraction — including a DOI-less one — raised no false retraction flag against a clean paper, never presented an invented reference as verified, and caught a real DOI attached to the wrong paper.
It does not show a general accuracy rate. Thirteen references are a deliberately hard sample, not a statistical one: they are mostly high-profile, English-language and well-indexed, and the retractions are famous ones. Obscure, non-English, very recent or grey-literature references are harder, and there the checker is more likely to answer “not found” — which is the honest failure direction we chose, but it is still a failure to resolve. Results also depend on upstream databases (Crossref, OpenAlex, DataCite, Semantic Scholar, Retraction Watch), so they can change as those records change. That is exactly why the run date is on this page.
We will grow this set over time and publish the result whether it improves or not. If you have a reference you believe we handle wrongly, please send it to us — a failing case is more useful to us than a passing one. See also our methodology and data sources.