Knowledge hub

The Co-Citation Problem: Most "Pages Driving Your Mentions" Don't Mention You

Paprik fetched every page its AI answers cited and checked the text. Up to 90% of the pages co-occurrence credited with driving brand mentions never mention the brand - and the newer the brand, the bigger the error.

ABAbhilashFounder7 min read
An AI answer naming a brand beside twelve source citations, next to two bars: 183 pages credited with mentions, 19 that actually mention the brand.

Author

AB
Abhilash

Founder

Abhilash is the Founder of Paprik AI. He writes about AEO, AI search visibility, and how brands can win in AI-driven discovery.

Every AI visibility tool has a version of the same report: the pages driving your brand's mentions. It is the report that decides where outreach budget goes, which publishers get pitched, and which "influential" listicles the team celebrates being on. And in most tools, the credit behind it is assigned by co-occurrence: your brand was named somewhere in an answer, that answer cited some pages, so those pages get credit for the mention.

Paprik runs the same kind of attribution - and then does something most tools skip: it fetches every credited page and reads it. That check has now covered 4,697 pages, and the result is uncomfortable for the whole category, Paprik included. For our own brand, roughly nine out of ten pages the co-occurrence method credited with driving Paprik mentions do not mention Paprik anywhere in their text. This is the study behind that number: how the error happens, how it scales with brand size, and what attribution has to require before you act on it.

How co-occurrence attribution works, and why it breaks

When an AI engine answers a question like "best tools for X", it does not write one sentence from one source. Across the answers Paprik tracks, an answer cites 12 distinct URLs on average. The engine names brands in one part of the text and attaches citations to another, and nothing in the output reliably connects a specific brand name to a specific source.

Co-occurrence attribution bridges that gap with an assumption: if the brand and the citation appear in the same answer - or even the same paragraph - the cited page probably drove the mention. For a comparison or listicle prompt that pulls a dozen sources into one response, the assumption fails structurally. A brand named in paragraph two and a citation attached to paragraph five sit side by side with nothing connecting them. Proximity is not attribution.

The test: fetch every page and look

The check is as blunt as it sounds. Paprik's verification system takes every URL that its attribution layer credited to any tracked brand, fetches the page, and looks for the brand in the page's actual text. A page that mentions the brand is Confirmed. A fetched page that does not is Co-cited only. A page that could not be fetched stays Unverified rather than being counted either way.

We first ran this on our own domain in August 2026: of the 49 pages co-occurrence credited with driving Paprik mentions at that point, 3 actually mentioned Paprik. The system has kept fetching daily since, and the picture below covers 18 April to 31 August 2026: 4,742 credited URLs across six brands tracked on our domain - Paprik and five competitors in the AI-visibility category - of which 4,697 (99%) were successfully fetched and checked.

The verified results, brand by brand: the most established brand in the set genuinely appears on 59% of the pages credited with its mentions. The next three land at 53%, 44% and 39%. Then the floor drops: 19% for the fifth brand, and 10% for Paprik itself - 19 of 183 verified pages. The competitor rows are anonymized; the pattern, not the names, is the finding. Every brand's attribution is inflated - the question is only by how much.

Bar chart: share of credited pages that actually mention each brand - five anonymized competitors from 59% down to 19%, and Paprik at 10%.
Verified confirmation rates, 18 April to 31 August 2026, competitor rows anonymized. Every brand's attribution is inflated; the newest brand's most of all.

The newer the brand, the bigger the error

That ordering is not an accident, and it is the finding that matters most if your brand is new to AI answers. Established brands really are on the listicles, the review roundups and the comparison tables that AI answers cite, so when co-occurrence credits those pages, it is often right. A brand that is newer to the category - as Paprik is in this set - appears in answers through a handful of pages: its own site, a founder's LinkedIn, one or two directories. The answers citing it also cite TechRadar, YouTube and Reddit for everything else in the response.

Co-occurrence then hands the newer brand credit for all of it. In our data, TechRadar pages collected 20 Paprik-credited citation events without mentioning Paprik once; YouTube collected 20 more across nine URLs; Reddit threads another 10. A team taking that report at face value would conclude TechRadar is driving their AI presence and go optimize for it. Meanwhile the pages that actually mention Paprik - our own site, a partner's blog, LinkedIn - are exactly the ones the verified report confirms. The attribution error hands a brand the appearance of broad coverage precisely when its real coverage is narrowest, which is when the truth is most valuable.

Where the phantom credit hides

One more cut of the data shows where the false credit concentrates. Counting citation events instead of distinct pages - so a page cited fifty times counts fifty times - confirmation rates improve for every brand, to between 44% and 80%. Pages that get cited again and again for a brand tend to genuinely mention it. The phantom credit lives in the long tail: pages that co-occurred with the brand once, in one answer, and never again.

That gives a practical filter even without full verification: a page that repeatedly co-occurs with your brand is worth a look; a page that appeared once in a twelve-source answer is, most of the time, noise.

What attribution should require

Whether you approach this as AEO or GEO, the operating rule is the same: co-occurrence generates hypotheses, and only the page's own text confirms them. Concretely:

  • Attribution needs evidence tiers. Confirmed (fetched, mentions the brand), co-cited (fetched, does not), unverified (fetch failed). Anything presented as one undifferentiated "sources" list is mixing the three.

  • Fetch failure is not absence. A page that could not be checked is unknown, not innocent and not guilty. Our 99% fetch rate is what makes the percentages above meaningful.

  • Headline metrics should ignore unconfirmed credit. Paprik's mention-share figures exclude co-cited-only pages for this reason.

  • Act only on confirmed pages. Outreach, updates and monitoring belong on pages verified to mention you; the co-cited list is a lead list, not a scoreboard.

This is the same discipline behind our finding that API-based tracking sees an 84% smaller brand universe than the product interface: measure where reality is, and verify before you report. If you are building a measurement practice from scratch, how to track AI visibility covers the full method.

A note on scope, honestly stated: this is one domain's data - six brands, one category, four and a half months - and the five competitor rows are anonymized. The mechanism, though, is not specific to us: any brand tracked by co-occurrence in multi-source answers inherits the same inflation, and the newer the brand is to AI answers, the more of its report is phantom.

Frequently asked questions

What is co-citation in AI visibility tracking?

Co-citation (or co-occurrence) attribution credits a cited page with causing a brand mention because both appeared in the same AI answer. It is how most visibility tools populate "pages driving your mentions" reports, because engines do not reliably link a specific brand name to a specific source. The assumption holds poorly in multi-source answers, which average 12 cited URLs in our data.

Doesn't the answer itself show which source goes with which sentence?

Only loosely. Inline citation markers attach to passages, not to brand names, and a passage naming four tools may carry a citation supporting only one claim in it. Some engines cite at the answer level with no positional information at all. Paragraph-level attribution narrows the pool but still cannot say that this page caused that mention - only the page's own text can confirm the connection.

Does this make "influential pages" reports useless?

No - it splits them in two. The verified subset is the most actionable report in AI visibility: pages proven to mention you that engines actually cite. The unverified remainder is a hypothesis list. The failure mode is treating the whole list as the first kind, which for a brand new to AI answers means acting on a report that is mostly noise.

How do I check whether my tool has this problem?

Open its influential-pages or sources report, take the top ten pages credited with driving your mentions, open each one and search the page for your brand name. If most of them fail that search and the tool gave you no warning, its attribution is co-occurrence without verification. Ask the vendor whether pages are fetched and text-checked, and whether headline metrics exclude unconfirmed pages.

Win AI Search

Increase brand visibility across AI search, from insights to action.

Start free trial