Method
How we measure recommendation share and citation share, method v1.1
The full method behind every number we publish. What the two metrics are and why they are two, query construction, engine coverage, run counts, what we store, how we parse it, and the four places this method is weakest.
Published
Every figure we publish carries a stamp naming the method version that produced it. This is that method. When it changes, the version changes, and figures published under an old version keep their old stamp rather than being quietly restated.
What the two metrics mean here
Recommendation share is the proportion of answers, across a defined query set, in which a company is named as a vendor to consider.
Citation share is the proportion of answers in which a URL on a domain the company owns appears as a cited source.
Recommendation share is the headline metric. Each retainer tier contracts on the one its work moves, citation share on the first and recommendation share on the second, and both are reported on both, never merged.
They are two outcomes produced by two mechanisms. Citation comes from owning a fact on a page you control. Recommendation comes from sources that discuss options in a category, which are pages you do not control. A buyer paying for pipeline is buying the first one, and work that moves one does not automatically move the other.
Both are deliberately narrow. Neither is share of voice, neither is sentiment, and neither is a composite index with weights we chose. Composite scores are easier to sell and impossible to audit, and the first time a client asks why their number moved, you want an answer that is a fact rather than a weighting decision.
We record three separate things per answer, because they move independently:
- Named. The company appears in the text of the answer as a vendor to consider. This is what recommendation share counts.
- Cited. A URL on the company’s domain appears in the answer’s sources. This is what citation share counts.
- Absent. Neither.
A company can be named constantly and cited never, which usually means the model knows of it from training and is not currently retrieving it. That distinction changes what you do about it, so we refuse to collapse the two into one number.
Method v1.0 recorded all three and then reported them under one name. The recording was right and the reporting was wrong. See the changes section at the end.
Query construction
A diagnosis uses 50 queries. An index category edition uses 30. Where a client’s category is narrower than the diagnosis set, engagement scope varies between 20 and 50, and the number is agreed before anything runs. Both retainer tiers track the same set at the same depth: the higher tier adds work, not queries, engines or competitors.
Three intent types, tracked separately:
Category queries. “We are replacing our underwriting tool, what should we be looking at?” The ones where a buyer is choosing and has not decided. These carry the most weight because they are the shortlist question.
The example is written the way a buyer prompts rather than the way a buyer searches, and that is not a cosmetic choice. A keyword string and a described situation retrieve differently, and measuring the first while your buyers type the second produces a number about a query nobody runs.
Category queries are written across a span of specificity rather than at one level. The open form above is one end of it. The other end is the same buyer with their constraints in the prompt: portfolio size, the system of record they already run, the integrations they cannot give up. Both ends are category queries, both are tracked as such, and holding them in one set is the only reason the results can tell a company that never appears from a company that appears for the open question and is gone by the qualified one. Those are different failures and they are fixed by different work, so a set written entirely at the open end cannot report the second one at all.
Problem queries. “How do I check development feasibility for a site.” Buyer has a problem and no vendor in mind.
Brand queries. “What is [company].” Used to test how a model describes a company, and to catch hallucinations. These are the only type that scales per company.
Queries are written in the language buyers use, not the language vendors use. They are approved by the client before anything runs, because the query set is the contract. Changing the query set mid-engagement invalidates the comparison, so we do not.
Engines and runs
Five engines: ChatGPT, Claude, Gemini, Perplexity, and Google AI Overviews.
Three runs per query per engine. Every query is run logged out, without personalisation, from a consistent location.
Three runs is a floor, not a comfort. Generated answers vary between runs, and a single run is an anecdote. Three tells you what a model tends to say. It does not give you a confidence interval, and we do not report one, because reporting statistical confidence from three samples would be dressing up a small number.
Where the three runs disagree substantially, that disagreement is itself a finding and gets reported rather than averaged away.
What we store
One row per run, per query, per engine, per company, carrying: run id, query, engine, run index, company, mention type, cited URL, position, and timestamp.
We store the raw response, not only the parsed result.
This is the design decision that matters most and it is the one most commonly skipped. Parsers are wrong sometimes. Brand names collide with ordinary words, a competitor gets missed, a citation is attributed to the wrong clause. If only the parsed output was kept, a parser fix cannot be applied backwards and the history is permanently built on the old mistake. Because the raw answer is stored, a correction re-reads history instead of re-running and re-paying for it.
We also store a hash of each response. When an engine’s answer to a query has not changed between runs, that is worth knowing, and it is occasionally worth publishing.
Parsing
Citations are extracted from the structured source lists the engines return where they return them, which is most of them. Company mentions are matched against a name list including known aliases, legal entity names, and product names, because “Archer” and “Archer Technologies” and the product name are the same company and a naive match gets this wrong in both directions.
Mentions inside a cited page title, rather than in the model’s own prose, are recorded separately and not counted as named.
Where this method is weak
Four places. All of them are known, none of them are solved, and we would rather publish them than have a client find them.
The API and consumer app gap. We query most engines through APIs. Buyers use the consumer applications, which have different system prompts, different retrieval behaviour, personalisation, and sometimes a different model version. We spot-check roughly ten percent of results through a real browser session to gauge the size of the gap, and we report when the gap is large. We cannot eliminate it.
Google AI Overviews fidelity. AI Overviews requires a search results provider rather than a model API. Vendor claims about fidelity conflict with each other. We trial providers against a set of known queries and inspect the parsing manually before trusting a provider, and we treat this engine’s numbers as the least reliable of the five.
Three runs is a small sample. Enough to distinguish “usually” from “once”. Not enough to detect a small change. We do not report movements we cannot distinguish from run-to-run variance, which means some real small movements go unreported.
Model updates confound everything. An engine can change retrieval behaviour overnight with no announcement. A change in a client’s numbers may be our work, may be a model update, and frequently we cannot separate them. Where we cannot, we say so rather than claiming the credit.
What this measurement cannot tell you
It cannot tell you what a citation is worth.
Attribution from generated answers is unreliable. Many answers produce no click. The buyer who reads your name in an answer and searches for you directly a week later arrives as brand search. We do not model this and we do not quote pipeline figures derived from it.
Recommendation share is a leading indicator of whether you are in the consideration set when a buyer asks a model. That is a real thing to care about, and it is a different thing from revenue. Treating it as the second is where this whole category loses its credibility.
Citation share tells you something narrower: whether the engine reached your pages to answer at all. It is the metric most responsive to work on your own site, which is also why it is the easier of the two to move and the weaker of the two to sell on.
Who runs this method, and the conflict in that
We publish measurements of a category and sell work to companies in that category. That is a conflict, and the method is built so that the part a client could be quietly favoured in, the query set, is frozen and dated before collection. We do not appear in any table we publish, because we are not a vendor in these categories.
We do not claim to be independent. We claim that the measurement reproduces. The conflict of interest policy states the rules, how client rows are marked, and what we do not claim.
Changes from previous versions
v1.1. The headline metric was renamed and split. What v1.0 reported as one number, citation share, is now reported as two: recommendation share and citation share.
Nothing about collection changed. Both were already recorded per answer, as Named and Cited, and v1.0 already said it refused to collapse them. It then collapsed them everywhere else. The clearest case was the section above this one: under v1.0 it called citation share the leading indicator for being in the consideration set. Being in the consideration set is Named. So v1.0 defined the metric one way and used it another, and the definition was the part nobody read. That section now names recommendation share, which is the metric it was describing all along, so the text you read above is the corrected version rather than the one this paragraph is about.
Engagements contracted on recommendation share from v1.1. Since 13 September 2026 each retainer tier contracts on the metric its work moves: citation share on the first tier, recommendation share on the second. Collection did not change, so that is a pricing change rather than a method version. Figures published under v1.0 keep their v1.0 stamp rather than being restated. Nothing on the site currently carries one: no client measurement has been published, the home page block is illustrative and says so, and the REN.PH evidence page reports Bing’s own figures under a stamp naming the console export rather than a method version of ours. Bing does not report recommendation at all, so that page is citation share in the strict sense above and needed no rewording.
v1.0. First published version.