Research methodology · August 2026
How the AI Recommendation Index is measured
This page documents exactly how the data behind the August 2026 study was gathered, in enough detail that you could rerun the protocol yourself. The full run, all 375 question texts and every answer verbatim, is retained and available on request.
The categories
The index runs on a fixed list of 20 UK market categories, chosen to span consumer retail, software, services and trades, each with one real seed brand. The seed brand anchors the category research and appears in the minority of questions that are branded; those questions are excluded from every published statistic, though a seed brand can still show up in the published list when an engine names it unprompted. The list was locked on 3 August 2026 and is versioned: it will not change between monthly runs, because a list that shifts under the numbers would let us cherry-pick. The first two categories ran as a pilot to check cost and output quality before the rest were committed; the pilot's data is part of the run, not discarded. The pilot's measured cost is also why five categories (robot vacuums, mattresses, returns management software, employment law solicitors, private dentists) were deferred to the September run: an API budget decision, made before the remaining categories ran, not after seeing their results.
The questions
For each category we generated 25 buyer questions, 375 in total, using the same pipeline as Skavora's product: live research on the category first, then questions written the way a real buyer asks, with budgets, constraints and situations, never keyword strings. 303 of the 375 questions (81%) are non-branded, meaning they never name the seed brand or any competitor the research identified; the rest are anchored on the category's seed brand. Questions never name an AI platform, and the engine answering is never told any brand is being measured. Every published statistic uses the non-branded questions only. We adopted that rule after catching a real artefact in our own first pass: seed brands were reaching the named-by-all-five list through the questions that name them, which inflates consensus. Excluding branded questions removed that inflation, and we recomputed every figure before publishing.
Auditing the generated questions against those definitions found edges, and we would rather list them than round them off. A handful of non-branded questions mention third-party products as buyer context (an existing Shopify store, a Gmail integration, a NatWest business account); one non-branded question in wireless headphones names a rival product outright (AirPods Pro); and one branded question names two comparator brands the seed retailer stocks. The branded slip cannot touch the figures at all. The non-branded mentions affect only the single category each sits in, most plausibly by nudging the named product into that category's list. They are generation defects, they are disclosed here, and question generation is being tightened for later runs.
The engines, exactly as asked
Each category's identical 25 questions went to five engines on 4 August 2026, every one called with live web search enabled: Gemini (gemini-2.5-flash with Google Search grounding), ChatGPT (gpt-4.1-mini with OpenAI web search), Claude (claude-haiku-4-5 with Anthropic web search), Google AI Overviews and Google AI Mode (both retrieved through SerpApi). Enabled is not the same as used: 32% of Claude's answers and 17% of ChatGPT's cite no source at all, while the other three engines almost always cite, and an uncited answer may draw on training data rather than the live web. Per-answer citations are in the raw files. Being exact about model tiers matters too: API models with web search are close proxies for the consumer apps, not the consumer apps themselves, and the consumer products may answer differently. We name the models so you can judge that limitation rather than take our word it does not exist.
One structural fact deserves its own sentence: three of the five engines are Google surfaces. Gemini, AI Overviews and AI Mode all run on Google models with Google Search underneath, so a five-engine consensus figure partly measures Google agreeing with itself. We therefore also computed consensus at the provider level, a brand recommended by all three of Google, OpenAI and Anthropic: 53 of the 1,448 brands, 4%, against 2% for all five engines. The disagreement finding survives either way.
Two engines were sampled rather than asked everything. Google AI Overviews and Google AI Mode are billed per lookup, so each answered a deterministic, evenly spaced 15 of the 25 questions. Google also sometimes shows no AI Overview for a query at all; those lookups are recorded as no-result, not as an answer, and a further 12 of the 450 sampled lookups failed at the API and were excluded the same way. Actual answered counts per category ranged from 12 to 15 for these two engines and are in the raw files, so their brand lists come from fewer questions and slightly understate their breadth.
What was recorded
For every answer we stored the full text, the sources it cited, and every brand it named, extracted by a structured scoring call. Where a scoring call failed, the affected answers were excluded from the figures entirely rather than estimated; for example, ChatGPT's answers in two categories were scored on 20 of 25 questions for this reason, and Claude answered 20 of 25 in one category after retries. Every per-engine answered and scored count is preserved in the raw run files.
Two honest caveats about that extraction step. First, the question generation and the brand extraction both run on Gemini, which is itself one of the five engines measured. The umpire being a player is fair to poke at, so here is precisely what it could and could not bias: extraction is mechanical name-listing from the answer text in front of it, not a ranking; it runs identically over every engine's answers; and the answers themselves come from each engine independently. Second, extraction is performed by a model and carries some noise: a small number of entries are product lines or organisations rather than companies. What is exact is the matching after extraction: normalised names either match or they do not, with no fuzzy merging.
How the statistics were computed
Brand names were normalised before any counting: lowercased, with everything except letters and digits stripped, so “De'Longhi” and “DeLonghi”, both of which appear in the coffee-machine answers, count once. The spelling shown in the published brand list is the most common form the engines actually wrote. A brand counts as recommended by an engine if it appeared in at least one of that engine's non-branded answers in the category; the per-brand number is how many of the five engines that is true for. Every figure in the article was recomputed mechanically from the raw run files before publication, by script rather than by hand. When we tightened normalisation during checking, the headline 76% figure moved by less than one percentage point, which is the scale of error we believe applies to these numbers.
Normalisation has a limit worth stating plainly: it merges spellings, not identities. Genuinely different name forms of one company stay separate entries, so an agency named once as “Impression” and once as “Impression Digital Agency” appears twice. That bias cuts in one direction: fragmentation understates consensus and inflates the long tail, which is the direction of our own headline finding. We disclose it rather than correct it, because merging name forms by judgement would put our thumb on the same scale.
One consequence of measuring engines rather than judging companies: a brand is in the list because an engine recommended it, so the list includes recommendations a human would dispute, such as a software platform offered to a buyer who asked for an agency. We publish those unedited, because the gap between what was asked and what was recommended is engine behaviour, which is the thing being measured.
What this study cannot see
- One run, one month, one country. Engine answers vary between runs, so per-category numbers are estimates with real variance, not rankings. The monthly re-runs on the identical list are what will separate signal from churn.
- No personalisation. Every question was asked fresh, with no account history or memory. A logged-in user's answers may differ.
- UK phrasing and UK settings throughout. The same categories in another market would produce different lists.
- API models stand in for consumer apps, as described above.
- Five of the twenty categories are not yet measured, and nothing here should be extrapolated to them until they are.
The complete brand list is public at /research/ai-index-brands, and the findings are in the article. If you spot an error, tell us and we will correct it in public.