Guides · 7 min read
PubScope vs Scholar Sidekick: Two Different Answers to Fake Citations
Scholar Sidekick is the most serious citation verifier we have looked at, and this page is not going to pretend otherwise. It is free, it needs no account, it publishes a larger accuracy benchmark than we do, and it ships a public API and a hosted MCP server. If all you need is to check whether references are real, it is an excellent tool and you should use it.
We build something different, and the useful comparison is about shape, not about who scores higher. Every Scholar Sidekick figure below is quoted from their own pages on 29 July 2026.
Where Scholar Sidekick is genuinely ahead
- A larger published benchmark.They report a “blind holdout of 1,395 freshly-drawn citations — 1,185 correctly-cited and 210 fabricated or wrong”, with “150 / 150 = 100%” on dominant fabrication patterns and a 0.8% high-confidence false-accusation rate (Wilson 95% CI 0.4–1.4%). Ours is 1,320 references. Theirs is bigger, and they got there first.
- More identifier types. Eight — DOI, PMID, PMCID, arXiv ID, ADS bibcode, ISBN, ISSN, eISSN. We resolve DOIs and arXiv IDs and otherwise work from the reference text. If you need ADS bibcodes or ISBN lookups, we do not do that.
- A free, public developer surface.“Every web tool is free with no account. The REST API and MCP server are free for light, anonymous use”, then $9/$49/$199 per month for 10k/100k/500k requests. Our MCP server has 4 tools and our public REST API is deliberately closed.
- They publish their own blind spots.Their page states plainly that semantic near-miss flips are “caught only 4 / 30”, and lists what the product does not do. That is the standard we are trying to hold ourselves to.
Why you cannot compare the two scores
It would be easy to put their number next to ours and declare a winner. It would also be meaningless, and we are not going to do it. The two benchmarks use different sets, built differently, scored against different definitions of what counts as a catch. Their 0.8% counts “all high-confidence flags” against 1,185 correctly-cited references, and they report a stricter figure of 0.17% alongside it; our 0.1% counts any verdict of retracted, not-found or DOI-mismatch against 700. Those are not the same measurement.
The same applies to the semantic-flip pattern, where our set and theirs are built by different rules — so the fact that the two numbers differ tells you about the two test sets at least as much as about the two products. The only way to compare properly is a single set run against both, which neither of us has done.
One further honesty note: we verified that Scholar Sidekick publishes these figures. We have no way to audit whether they are correct, and the same limitation applies to ours — we built our set, ran it and reported it. Neither benchmark is an independent audit.
Side by side
| PubScope | Scholar Sidekick | |
|---|---|---|
| Price | Currently free — everything is open while we are in beta. Listed plan when we switch it on: $15/month or $99/year, halved for 14 lower-income countries. | Web tools free, no account. API/MCP free for light anonymous use; $9 / $49 / $199 per month for 10k / 100k / 500k requests. |
| Input | Paste a whole reference list in ordinary citation style, BibTeX or RIS. No PDF in the checker either — but a manuscript PDF uploaded to Paper Review has its bibliography checked by the same engine. | Paste-only: “a DOI, PMID, PMCID, arXiv ID, ADS bibcode, ISBN, ISSN, or a link to a paper”. Their page states “No PDF upload”. |
| Retracted papers | Yes — 70,742 Retraction Watch records held locally, plus the compound case: the DOI belongs to a different paper and that paper is retracted. | Yes — “Retraction Watch + Crossref”. |
| Venue trust signals | Flags references to venues our own 46,797-journal database marks DOAJ-withdrawn (1,208) or very low trust (52), with a 0–100 Trust Score per journal. | Their page states it “does not currently flag predatory publishers as a separate signal”. |
| Public API / MCP | MCP server with 4 tools. Public REST API deliberately closed. | REST API and hosted MCP, free tier. |
| Identifier types | DOI and arXiv ID; otherwise resolves from the reference text. | Eight, incl. ADS bibcode, ISBN, ISSN, eISSN. |
| Published benchmark | 1,320 sealed references, hash published before the run, Wilson CIs, measured four times, unstable groups reported as a range. Plus a 14-case public set you can reproduce yourself. | 1,395-citation blind holdout, 150/150 on dominant patterns, 0.8% high-confidence false accusations (CI 0.4–1.4%), 4/30 on semantic flips, 99.9%-stable repeat run. |
| Beyond citations | Journal finder, Trust Score, journal comparison, pre-submission review, advisor, publishing radar. | Citation formatting, verification, open-access and retraction checking. |
The actual difference: a checker versus a submission workflow
Scholar Sidekick answers “is this reference real?” extremely well, from an identifier, through an API, at scale. That is a clean, well-executed product and it is the right tool if that is your question.
PubScope is built around the question that comes before and after that one: you have a manuscript, and you need to know where to send it, whether the venue is sound, whether the paper is ready, and whether its references hold up. Reference checking is one stage of that, which is why it sits next to a journal finder, a Trust Score over 46,797 journals, and pre-submission review. If you only ever needed the one question answered, that surrounding machinery is overhead rather than value.
The venue signal is the one place where the difference bites inside citation checking itself. When a reference points at a journal that has been withdrawn from DOAJ, we say so, because we already hold that journal database. Their page says they do not flag publisher quality as a separate signal. Whether that matters depends entirely on whether your field has a problem with it.
What we are not better at, stated plainly
A bigger benchmark, more identifier types, a free public API and a more mature MCP surface — all theirs. We are not going to claim a lead we cannot demonstrate on a shared test set, and there is no shared test set.
PubScope brings the public signals together for tens of thousands of journals: Web of Science / Scopus / DOAJ indexing, SJR quartile, APC, a 0–100 Trust Score and predatory-risk flags — each linking out so you can confirm it at the source.
Sources
- scholar-sidekick.com ↗ — pricing, identifier types, retraction sourcing and input method, read 29 July 2026.
- scholar-sidekick.com/citation-integrity ↗ — the 1,395-citation holdout, 150/150 figure, false-accusation rate and confidence interval, and the 4/30 semantic-flip result, quoted verbatim, read 29 July 2026.
- scholar-sidekick.com/compare/best-ai-citation-verifier ↗ — their “where Scholar Sidekick does not win” section, source for the no-PDF-upload and no-predatory-flag statements.
- PubScope validation results — our public set, sealed holdout, confidence intervals and measured blind spots.
- PubScope catalogue figures read from the live system on 29 July 2026 (46,797 journals; 70,742 Retraction Watch records; 1,208 DOAJ-withdrawn; 52 below a trust score of 20).