
The Brand Index: What It Measures and Why Nobody Can Pay to Rank
A public measurement of how often AI assistants name DTC brands when someone asks what to buy. What gets asked, what gets refused, why the roster is a file in git, and the five ways we have already got it wrong.
The Brand Index measures one thing: how often AI assistants name a DTC brand when somebody asks what to buy in that brand's category. Not sentiment, not share of voice, not a quality judgment. Whether the name comes up.
This post is about the instrument rather than the readings. The findings are here, and they move: the index refreshes daily and the registry is still filling, so any number quoted in prose is stale the week after it is written. What does not move is the method, and the method is the part you should interrogate before you believe any of it. If you are looking for what to change rather than how the measurement works, the agent visibility playbook is the other half of this.
The questions are unbranded, and that is the whole design
A question that names the brand guarantees a mention and measures nothing. So the questions never name the brand. They are derived from the store's own catalogue signals, its product types, tags and titles, and they ask what a shopper asks: what are the best merino socks, what is the best olive oil for finishing.
That derivation is the hard part, and it is where most of our mistakes have happened. Ask a canned-water company about t-shirts and it will score zero for not being named among products it does not sell. That is a measurement of our question generator, not of the brand.
A score is only published above a stated floor
Four assistants get asked: ChatGPT, Claude, Perplexity and Gemini. They do not all answer every time. So the index publishes a number only when at least two assistants answered at least four questions. Below that, the brand is withheld rather than shown at zero.
Why nobody can pay to be in it, or to rank in it
The index page says it plainly: no brand pays to be measured or to rank. That sentence is worth as much as the mechanism behind it, so here is the mechanism.
- The roster is a file in version control, not a row in a database somebody can add at two in the morning. Changing who appears is a pull request with a diff and a reviewer, and the history is permanent.
- There is no advertiser relationship with anyone in it. None of the brands measured are customers. Most do not know they are in it until somebody sends them the link.
- Inclusion is free and editorial. There is a request form on the index. It is not a purchase, and paying would not be an option we could offer.
- The score is computed, not curated. It comes out of the same collector for every brand, and the rank is a sort on that number. There is no editorial adjustment step where a thumb could rest.
The honest limit on all of that: we choose who is in the registry, and choosing the roster is an editorial act. We do not claim otherwise. What we claim is that nobody can buy their way into it or up it, and that the roster's history is public so the choosing can be inspected.
The rubric is versioned, and a change retires the old rows
Every score carries the version of the rubric it was computed under. When the method changes materially, rows computed under the old one are retired rather than ranked against the new ones.
This is expensive and we do it anyway. When Perplexity came online as a fourth assistant on 2 August, every score taken on the previous three assistant panel was retired, and the index shrank overnight. Two brands measured on different panels are not two samples of one quantity, and a table that ranks them against each other is quietly wrong in a way no reader could detect.
The failure modes are published, with the defects that produced them
The methodology page documents five classes of error, each with the specific instance that caused it, each a real defect this engine shipped. Two examples, both ours:
- Engine failure attributed to the merchant. Allbirds was told it had no Product schema. We had scanned the homepage, and homepages do not carry Product schema; it was on the product pages the whole time. Another store was told it refuses automated browsers, when what happened was our screenshot vendor being rate limited on a shared IP pool.
- Unverified success credited to the merchant, which is the direction nobody complains about. `og:image` scored on the presence of the tag, so a page whose card renders blank scored the same as one that renders. `llms.txt` was awarded on an HTTP 200, and a catch-all rewrite answers 200 for a file that does not exist.
A measurement nobody can check is a claim. The fastest way to make a method checkable is to say where it has been wrong and what now stops it happening again.
Brands that block us are listed, not dropped
Six stores in the registry cannot be read at all: they return 403 to every request, from a crawler user agent and a desktop Chrome user agent alike. They stay on the index with no score, because dropping them would hide the most on-thesis observation the whole exercise has produced. That is its own post.
None of this makes the index authoritative. It is a partial measurement of a moving target, taken from one vantage point, under a method that has already been wrong five documented ways. What it is, is checkable. The roster, the rubric version, the eligibility floor and the failure log are all published, so a number you disagree with is a number you can go and argue with.
See where your brand sits The free growth audit runs the same unbranded question set against your own catalogue, whether or not you are in the registry.