Platform · research agent
The sources that decide your visibility are not on your site.
An answer engine describing your category quotes review sites, community threads, industry publications and comparison roundups. Your own domain is one source among them and rarely the loudest. Everything downstream — which gap to close, which page to write, which placement to earn — depends on being able to read those sources reliably, including the ones that would rather not be read.
How the reader escalates against a resistant page is the part that took longest to build, and it stays private. What belongs on a public page is the standard it holds itself to: a date on every source, typed fields instead of prose, and an honest failure instead of a plausible summary.
Read 22 July 2026
Six capabilities live in the workspace · every source dated · a page that cannot be read is reported as unread.
The number
"Most citations are external" is true and imprecise.
The figure quoted around this claim ranges from roughly 39% to 95% depending on the study, and that spread is not sloppiness. The studies count different things. Brand mentions during early commercial discovery are one measure; the share of all citations going to earned and news media is another; a per-engine number is a third. A single tight percentage applied to every engine and every query type is one study's definition presented as a law.
The primary we read for this page is the AirOps 2026 State of AI Search, read on 22 July 2026: about 85% of brand mentions in early commercial discovery come from external domains, brands are roughly five times more likely to be cited through third-party sources than their own, and about 48% of citations come from community and user-generated sources.
Our own baseline says the same thing from the other side. When we scanned our category across engines on the day we published our pricing research, our domains appeared in zero of 63 answer cells. That is not a comfortable number to print, and it is the honest starting point that makes the rest of this measurable.
Why this works
Where research agents quietly fail.
A block gets read as an absence. A page returns nothing, the tool reports no data, and the analysis proceeds as though the competitor has no pricing page. The conclusion is confident and wrong, and nothing in the output shows which source was missing.
A summary replaces a reading. The page was never opened, but the model knows roughly what such a page says, so a plausible paragraph appears. This is the failure mode that produces invented numbers with real-looking sources attached.
The source has no date. A price read in May is quoted in September as current. In this category that gap is enough to be wrong twice over, which is why we found five errors in our own competitor sheet the first time we re-read it with dates attached.
None of these look like failures when they happen. All three produce output that reads well, which is exactly why the standard has to sit in the tooling rather than in the care of whoever is reviewing.
What it does
Six capabilities, one of them is refusing.
Read
Fetch a page that does not want to be read
A page that returns nothing to a simple request is not an absent page. The reader escalates until it either has the content or can say honestly that it could not get it, which is a different answer from silence.
Survey
Search across engines, not one
A single search engine is one opinion about what exists. Queries run across several and the results are merged, because the source that matters for a category is often the one your usual engine ranks tenth.
Extract
Typed fields, not a wall of text
Pricing, contacts, article bodies, reviews and events come back as declared fields against a schema. A number that arrives in a named field can be checked; a number pulled out of prose by eye cannot.
Research
A question, then cited sources
A research question fans out into searches, reads the pages that answer it, and returns claims attached to the source and the date it was read. The date is part of the answer, not a footnote.
Refuse
An honest failure beats a plausible one
When a page cannot be read, the answer is that it could not be read. The failure mode this exists to prevent is a confident summary of a page nobody actually opened.
Budget
The expensive path is opt-in
The cheap route is tried first and the costly one is a deliberate choice for a target that genuinely blocks. Otherwise a routine question quietly bills like a hard one.
The standard
What a source has to carry.
These are the rules a returned source is checked against before anything is built on it.
- The date it was read. Not the year, the day. A figure without one is a rumour, and this category re-prices faster than the pages quoting it get updated.
- The page it came from. A claim points at the page that makes it, not at a homepage or a category hub that happens to be on the right domain.
- How it was read. Some pages give a figure in their source and some render it only in a browser. Which one applies changes how much a number should be trusted, so it travels with the number.
- An explicit gap when there is one. A source that could not be read is recorded as unread. Nothing downstream is allowed to quietly treat it as empty.
Honestly
What it cannot do.
It reads public pages. Anything behind a login, a paywall or a private dataset stays outside, and no amount of escalation changes that. Where a number lives only in a report somebody bought, the honest output says so.
It is not free at the hard end. Reading a page that actively resists costs real money per attempt, which is why that path is a decision rather than a default. A research plan that needs it on every source is usually the wrong plan.
It does not judge for you. It returns sources with dates and the method used; deciding that a source is credible, or that two studies disagree because they measure different things, is analysis. The section above about the 39-to-95 spread is exactly that, and no tool produced it.
Questions this raises
About the research agent, answered directly.
Is this just a web scraper?
A scraper fetches a page. This decides which sources are worth reading for a question, escalates when a page resists, returns fields rather than prose, and attaches a date to everything. The fetching is the least interesting part, which is why it is the part we do not describe here.
Why does the date on a source matter so much?
Because a citation study from 2024 and one from 2026 disagree, and pricing pages move faster than the pages that quote them. Everything this stack returns carries the day it was read. This page was verified against the live module on 22 July 2026.
What happens when a site blocks you?
It escalates, and if it still cannot read the page it says so. What it does not do is quietly return a plausible summary of a page it never opened. A blocked source is a known unknown; an invented one is a defect that ships.
Do you publish how the escalation works?
No. The discipline is on this page and the engine is not. That is a deliberate line: describing that a block is never treated as "no data" is useful to a buyer, and publishing the routing that makes it work would be handing over the part that took the longest to build.
So what is the real share of AI citations from third-party sites?
It depends on what you count, which is why the published figures for 2026 run from roughly 39% to 95%. Brand mentions in early commercial discovery sit near 85% external by the AirOps 2026 report; earned and news media alone are a much smaller share of all citations. Anyone quoting a single tight number across all engines and all query types is quoting one study's definition.
Can I use this on my own sources?
It runs inside the workspace rather than as a standalone product, so it is used on your category, your competitors and the sources that cite them. It is not sold as a scraping subscription and there is no price list for it.
Want to see which sources your category runs on?
Thirty minutes that opens with the pages the engines actually quote when somebody asks about your category, and an honest read on how many of them you could realistically appear on this quarter.