Do AI assistants cite real sources?
What we measured
We asked 4 real questions — taken from actual search volume, not invented — to four assistants, each time asking explicitly for sources. Then we opened every citation and recorded what came back: 136 citations checked, 0 that do not hold up, 23 paywalled.
Interim. Fewer than 25 answers per assistant so far, which is not enough to rank them. The figures below describe what has been checked, not a conclusion about the assistants.
■ Perplexity ■ Grok ■ Claude ■ ChatGPT
How much they cite
Whether the citations hold up
Whether they cite at all
The full numbers
| Assistant | Answers | Citations | Per answer | Verified | Did not hold up | Paywalled | Integrity |
|---|---|---|---|---|---|---|---|
| Perplexity | 3 | 54 | 18.0 | 44 | 0 | 10 | 100.0% |
| Grok | 3 | 33 | 11.0 | 28 | 0 | 5 | 100.0% |
| Claude | 3 | 30 | 10.0 | 26 | 0 | 4 | 100.0% |
| ChatGPT | 4 | 19 | 4.8 | 15 | 0 | 4 | 100.0% |
Method, and its limits
- This measures citation integrity, not accuracy. Whether a claim is true is not machine-checkable at scale; whether a cited source exists and is reachable is.
- Questions come from real search volume in one domain. Search volume is a proxy for what people ask an assistant, not the same thing, and one domain is not the whole web.
- A paywalled or bot-blocked source (401, 403, 429) is excluded from the integrity figure. A 403 proves nothing about whether the paper exists, so counting it as either a pass or a failure would invent a result.
- Every answer has a public receipt listing each source and what happened when we opened it.