How accurate is the Citation Check?

Accuracy claims are cheap. So instead of quoting a percentage, we publish the exact test set, the result we measured, and the instructions to reproduce it yourself in your browser in about a minute. If our numbers are wrong, this page gives you everything you need to prove it.

Set version 2026-07-28 · measured against the live checker on 28 July 2026 · runner: scripts/run_citation_validation.py

5/5
Retractions caught
incl. one cited without a DOI
0/4
False retraction flags
on clean landmark papers — must be zero
3/3
Fabricated refs never “verified”
reported as “not found”, never invented
1/1
Real DOI / wrong paper caught
the pattern a DOI-exists check misses

What the set is designed to break

A citation checker can fail in two opposite directions, and a single accuracy number hides both. So the set is built from four deliberately awkward groups:

Result — every case, as measured

#CaseExpectedResult
R1Known retractionretractedretracted
R2Known retractionretractedretracted
R3Known retractionretractedretracted
R4Known retractionretractedretracted
R5Retraction cited with NO DOIretractedretracted
C1Clean landmark paperverifiedverified
C2Clean landmark paperverifiedverified
C3Clean landmark paperverifiedverified
C4Clean paper cited with NO DOIverifiedverified
F1Fabricated referencenot foundnot found
F2Fabricated referencenot foundnot found
F3Fabricated referencenot foundnot found
M1Real DOI, wrong paperDOI mismatchDOI mismatch

13 of 13 cases matched the label they were given before the run.

Reproduce it yourself

Copy the list below, paste it into the free Citation Check and compare what you get against the table above. No account is needed for the first check.

[1] Wakefield AJ, Murch SH, Anthony A, et al. Ileal-lymphoid-nodular hyperplasia, non-specific colitis, and pervasive developmental disorder in children. Lancet. 1998;351(9103):637-641. doi:10.1016/S0140-6736(97)11096-0
[2] Mehra MR, Desai SS, Ruschitzka F, Patel AN. Hydroxychloroquine or chloroquine with or without a macrolide for treatment of COVID-19: a multinational registry analysis. Lancet. 2020. doi:10.1016/S0140-6736(20)31180-6
[3] Mehra MR, Desai SS, Kuy S, Henry TD, Patel AN. Cardiovascular Disease, Drug Therapy, and Mortality in Covid-19. N Engl J Med. 2020;382:e102. doi:10.1056/NEJMoa2007621
[4] Obokata H, Wakayama T, Sasai Y, et al. Stimulus-triggered fate conversion of somatic cells into pluripotency. Nature. 2014;505:641-647. doi:10.1038/nature12968
[5] Obokata H, Sasai Y, Niwa H, et al. Bidirectional developmental potential in reprogrammed cells with acquired pluripotency. Nature. 2014;505:676-680.
[6] Jumper J, Evans R, Pritzel A, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583-589. doi:10.1038/s41586-021-03819-2
[7] He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2016:770-778. doi:10.1109/CVPR.2016.90
[8] Jinek M, Chylinski K, Fonfara I, Hammel M, Doudna JA, Charpentier E. A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science. 2012;337(6096):816-821. doi:10.1126/science.1225829
[9] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. Advances in Neural Information Processing Systems. 2017;30:5998-6008.
[10] Harrington ML, Vossberg K, Adeyemi T. Quantum-coherent ribosomal editing in thermophilic archaea. Journal of Molecular Systems Biology. 2021;14(3):221-238. doi:10.1038/jmsb.2021.0473
[11] Okonkwo AR, Lindqvist P. Transformer-based prediction of glacial isostatic rebound from sparse GNSS arrays. Journal of Geophysical Modelling. 2020;58(2):119-137.
[12] Petrova IS, Almeida GF, Nakamura Y. Self-supervised contrastive alignment for low-resource clinical ontologies. Artificial Intelligence in Medicine Reports. 2022;7:100412.
[13] Smith J, Roberts L. A randomized trial of mindfulness training in adolescent athletes. Nature. 2021;596:583-589. doi:10.1038/s41586-021-03819-2

A second set we do not publish

A published set has one weakness: it can be optimised against — by us, or by anyone else. So the headline number does not rest on it alone. Alongside it we run a larger holdout that is rebuilt from live records each time and whose references are never published. Same rules, harder sample:

GroupHow ground truth is establishedResult
Retracted (20)Sampled from the Retraction Watch database; half cited without a DOI19/20
Clean (20)Real papers pulled live from OpenAlex, excluded from the retraction database20/20 · 0 false flags
Fabricated (12)Invented references, each checked against Crossref first so a real paper is never mislabelled12/12
Real DOI, wrong paper (8)Constructed from two real papers, so ground truth is certain8/8

59 of 60 (98%), measured 28 July 2026. Two things are worth saying plainly. First, the retracted group is sampled from the same database the checker consults, so what it really measures is whether we can parse and resolvea messy real-world reference — the hard part — not whether we know about retractions others don't. Second, building this set found a real bug: references whose venue our parser failed to read were being classified as books and quietly skipped, so two retracted papers came back clean. That is now fixed, which is what moved the score from 57/60 to 59/60.

The one remaining miss is honest and still open: when a DOI points to a different paper, we report the mismatch but do not additionally check whether the paper it points to was retracted. The user is warned about the reference either way, but that compound signal is missing today.

What this does and does not prove

It does show that on a set built specifically to trip it up, the checker caught every retraction — including a DOI-less one — raised no false retraction flag against a clean paper, never presented an invented reference as verified, and caught a real DOI attached to the wrong paper.

It does not show a general accuracy rate. Thirteen references are a deliberately hard sample, not a statistical one: they are mostly high-profile, English-language and well-indexed, and the retractions are famous ones. Obscure, non-English, very recent or grey-literature references are harder, and there the checker is more likely to answer “not found” — which is the honest failure direction we chose, but it is still a failure to resolve. Results also depend on upstream databases (Crossref, OpenAlex, DataCite, Semantic Scholar, Retraction Watch), so they can change as those records change. That is exactly why the run date is on this page.

We will grow this set over time and publish the result whether it improves or not. If you have a reference you believe we handle wrongly, please send it to us — a failing case is more useful to us than a passing one. See also our methodology and data sources.