0. Two methods, and which one ranks you
This engine asks questions in two different ways, for two different jobs. They are described separately below and they are not interchangeable, so it is worth being exact about which one produces the ranking on the index.
The category pool, which is what the index ranks
Each category has ONE fixed question set, written once and unbranded, and every brand in that category is measured against exactly the same questions. The set is put to every assistant on the panel, the answers are pooled, and a brand is credited when an answer names it. Nobody is asked a question chosen for them, and no brand can be ranked above another because it was asked easier questions, because there is only one set of questions in the category.
Today that is 7 categories and 594 questions in total, put to 4 assistants. Those figures are counted from the question set itself at the moment this page renders, so they cannot drift from what actually ran.
Share of voice is therefore a share OF SOMETHING: of the answers given to one category’s shared questions, by the whole panel, in that run. That is the number the index orders on.
Per-brand questions, which are the customer method
A connected brand’s own report derives its questions from its own catalogue, which is what section 1 describes and argues for. That method produces a report about one store and never produces a ranking: two brands asked different questions cannot be placed above or below each other, and doing it anyway is the defect that made the pool necessary.
So when you read section 1’s case against fixed templates, read it as an argument about the per-brand report. The index uses a fixed set on purpose, for the opposite reason: comparability between brands is the whole point of a ranking and is worth the loss of per-store fit. Section 2c, on a fixed core and a growing pool, is likewise about per-brand tracking only.
1. Per-brand questions come from the catalogue, not from a template
This section is about the per-brand report. The index ranking uses the category pool described in section 0.
Every question is derived from the store's own public signals: the merchant's product-type taxonomy first, then tags, then metafield values, then product titles, each weighted by how much it reveals about what is actually sold. Harvested terms are matched against curated lexicons of product nouns, attributes and use cases, ranked by accumulated weight, and used to fill question templates. No model is involved. The same code produces the question set for every brand.
The questions are never branded. A question that names the brand guarantees a mention and makes the score meaningless. What is being measured is whether an assistant surfaces you to somebody shopping for what you sell, so the question has to be about what you sell and nothing else.
Why the fixed templates were abandoned
The previous method had nine category buckets. Every store classified as fashion was asked about sustainable clothing, DTC denim, basic t-shirts and affordable jewelry. Allbirds sells shoes, socks, apparel and underwear, so three of those four questions were about categories it does not stock, and not one was about footwear. It scored 0 of 8, and the report said AI assistants do not recognise or recommend the brand. That claim was disproved in under a minute by asking about sneakers.
The score was not miscalibrated. The questions were wrong, and a wrong question produces a confident zero that looks exactly like a real one. Stores whose signals are too thin to rank a product noun still fall back to generic templates, and the run is marked as having used them so the report can say the questions were generic rather than pretending otherwise.
2. An unconfirmed absence never scores
"We asked enough assistants and you were not named" and "we could not ask" are different findings. Only the first is a fact about the brand. The second is a fact about our run, and reporting it as a low score would be a fabrication with a number attached.
So a visibility score is only produced when at least 2 assistants answered at least 4 questions. Below that bar the run reports no score at all. A zero on the index therefore always means the questions were asked, on topic, enough times, and the brand did not come back. A brand we could not measure is omitted from the ranking and counted underneath it, never ranked last.
The same rule governs everything else the engine reports. A check that could not run is not-assessed, which is removed from the denominator rather than scored zero, so a store is never marked down for our network, our timeout, or our expired credential.
2b. What a zero can and cannot tell you
A zero means the questions were asked and the brand did not come back. It does not mean the brand never appears. Every measurement has a floor below which a real rate is indistinguishable from absence, and we publish that floor per brand rather than leaving a reader to assume it is zero.
The same question goes to every assistant, so their answers are not independent observations. We measured how much they cluster across 4,329 answered questions and 1,196 question groups: the intra-cluster correlation is 0.594, which discounts a run’s observations by a factor of 2.556. A brand asked four questions of four assistants therefore has about six effective observations, and the smallest rate that could be told apart from zero at 95% confidence is 38%. Each brand page states its own figure, because the number depends on how many questions that catalogue supported.
This is the least flattering number on the site and it is the one that makes the rest of it usable. A ranking that does not say what it could not have seen is asking to be read as more precise than it is.
2c. A fixed core for the trend, a growing pool for the measurement
Per-brand tracking only. The index’s own question set is fixed per category and does not work this way.
Tracking a brand over time and measuring it well are different jobs, and they want opposite things. A trend line needs a constant: if the questions change between runs, a movement in the number could be the brand or could be us. A measurement wants to stay current: catalogues change, and a question set frozen forever slowly stops describing what a store sells.
So we do both. Four questions per brand are fixed at its first run and never change; the change column is computed on those alone, so it is stable by construction. Every other question sits in a pool that grows over time. Adding to the pool retires nothing and moves no history, because the core it would have to disturb is untouched and a run stays comparable to an earlier one on the questions both asked.
Roughly a third of measured brands have catalogues too thin to support the full four, and those brands say so on their own page rather than presenting a two-question trend as though it were a four-question one.
2d. What this method gets wrong, listed
These are not edge cases we intend to fix. They are accepted limits of a question set built from a store’s own published taxonomy, and we would rather a reader know them than discover them.
- A small line can outrank the main one. We read what a catalogue declares, not how much of it sells. A skincare brand with a few fragrances can be asked about colognes. Every check we have passes, because cologne really is in the catalogue; what we cannot see is that it is a rounding error of the business.
- Price anchors follow the catalogue, including its bundles. A question can read “under $725” when a product line is full of multi-item kits. The figure is the true median of that line and still reads as absurd.
- An attribute can be real and still be odd. We require an attribute to appear on a product of the type being asked about, which removed questions like “the best fleece bags” from a laptop-bag company. It does not stop a true but unrepresentative pairing, such as asking a key-organiser brand about insulated bags.
- Eleven brands cannot be measured at all. Their published taxonomy names a flavour, a container or a merchandise line rather than a product, or the store is too small to read. We report them as unmeasured rather than scoring them, and for most of them the fact itself is the finding: a catalogue that does not name what it sells is hard for an assistant to enumerate too.
3. Coverage is stated on every run
Every measurement carries the conditions it was taken under: which assistants answered, which did not, how many questions were scored. Not as a disclaimer, as part of the result. Two scores taken under different coverage are not the same kind of number, and a reader who cannot see the difference has no way to know that.
This matters most when it is inconvenient. A vendor outage on our side narrows coverage, and the honest response is to say so on the affected runs rather than publish a thinner measurement that looks identical to a full one.
What actually happens when an assistant is unreachable is this. The run is NOT abandoned: it proceeds with the assistants that answered, and the run record is stamped partial rather than ok, carrying how many of the 4 were reached. Anything computed from a partial run is shown as partial on the index, and the condition raises an operational alert (category_pool_degraded) so a narrowed panel is something a person is told about rather than something a reader has to notice.
The difference between refusing and recording matters to a reader. A refusal would mean the index silently stops moving during an outage, which is invisible. Recording the run as partial means the index keeps moving and says, on the affected rows, that it was measured under a narrower panel than usual.
3b. How a brand gets in, and what gets published before it has a score
Nobody pays to be here and nobody is here by invitation. There are three ways a store enters, and all three end at the same checks.
- Systematic sourcing. We take domains from named public lists of direct-to-consumer stores. The list a domain came from is recorded against it, so a question about whether the population is skewed by its sources is answerable rather than a matter of trust.
- Submission. Anybody can submit a store on the index page. The form states that submitting lists the store publicly once it passes the checks, and records the answer; a submission without that consent is kept and never published.
- Extension suggestion. The Cresva browser extension has a button offering the store being looked at. It sends the host and the id of the scan, carries the same consent sentence, and never sends a name or an email address, because it never asks for one.
The three checks
A store is admitted when it is REACHABLE, when it SELLS ONLINE, and when we can READ A PRODUCT CATALOGUE from it. That is the whole gate. Each check records what was observed and each failure is a statement about what could be read on the day it was read, which is why a failure is re-run rather than appealed: unreadable stores are asked again every week.
Note what is NOT a condition: how big the store is, how much it sells, whether it is a customer, and how many questions its own catalogue could support. The last one used to be a gate and was wrong, because index scoring is pooled and a brand’s own catalogue yield never enters into it.
Admitted, then scored at the category’s next run
Every admitted brand is pooled and scored against its category’s shared question set at that category’s next run. Since the cron takes one category a night, that is at most 7 nights away and the row says which date. A brand admitted into a category whose catalogue we cannot confidently place is listed WITHOUT a category and is not pooled until an operator assigns one, which is a smaller claim than putting it in the wrong category and ranking it there.
Until a brand has a score it is still listed, with its checks shown, and the row says it has not been scored yet. A public index of stores that silently omits the stores it has admitted and not yet measured is describing its own throughput and calling it a market.
Two populations, and every percentile names its own
Two different sets of stores are counted on this site and they answer different questions. The crawled host population is drawn from the Tranco ranking of the web, which is a real measurement of large sites and a poor comparison for a merchant: measured over one crawl of 100, 97 were large multi-brand commerce. The direct-to-consumer registry is this index: stores of the kind a reader here actually runs.
So no percentile on this site is published without naming which population it is a percentile of, and the size of that population is interpolated into the sentence rather than pinned, because a pinned number becomes false the first time the registry grows and the sentence still reads correctly.
4. Tracking observability cannot rank brands
The engine reports what tracking is visible in a page's source. That is a useful thing to know and a disqualifying thing to rank on, because it measures visibility of tracking rather than quality of tracking, and those two run in opposite directions as a team matures.
Moving tags into a server-side container is a step up in sophistication: better data quality, less client weight, more control over what is shared. It also removes those tags from the page source. A brand that does the more advanced thing scores lower. On the first brands measured, the lowest tracking score in the set belonged to a store whose tracking is almost certainly better than the highest.
So the column is reported and never ranked. A metric that falls as competence rises is not a metric, and publishing a league table ordered by it would be publishing the inversion as a finding.
5. Why there is no overall score
There used to be one. It read 41 for Allbirds and 60 for a 46-product indie store, and that was not a calibration problem to be tuned. It was measuring the wrong thing and being framed as another.
The components share no unit. Schema completeness, pixel hygiene, homepage conversion structure and share of assistant recommendation are four different questions, and a weighted mean asserts both that they are commensurable and that ten points of one trade against ten points of another. They are not and it does not.
Worse, it inverted on sophisticated brands by construction. Maturity moves work out of static HTML: tags into server-side containers, rendering into a SPA, content behind a CDN edge. Every one of those lowers what a crawler observes while raising actual capability. Coverage also varied per run, so averaging over whichever collectors happened to succeed made two runs of the same store non-comparable, never mind two different stores.
And it invited exactly the misreading the reports shipped. "Overall audit score of 41 out of 100" reads as "this business is failing", when what was measured was "this homepage exposes little to a document fetch". A reader can act on "0 of 12 schema types present". Nobody can act on 41. The index ranks one named column for the same reason.
6. Failure modes
Every class below is a real defect this engine shipped. They are published because a measurement nobody can check is a claim, and because the four instances of the first class were found separately and looked like four unrelated bugs until they were put side by side, at which point they were obviously one mistake made four times.
The last item in each entry is the point of this document. An error that has been fixed will recur; an error with an invariant behind it has to get past the invariant first.
a. Engine failure attributed to the merchant
Something on our side did not work, and the report described that as a fact about the store.
What was published
- Allbirds was told it had no Product schema, so AI agents could not parse its product data. We had scanned the homepage, and homepages do not carry Product schema. The schema was on the product pages the whole time.
- A store was told it refuses automated browsers. What actually happened was that our screenshot vendor got rate-limited by the store's CDN, on the vendor's shared IP pool. The same store answered our own crawler with a 200 on every attempt.
- Reports said the Meta Ad Library returned no usable response, which reads as a fact about the merchant's ad account. Our own access token had been dead for 182 days.
- A store was told ClaudeBot was blocked when it was not. The check fell through from ClaudeBot to the deprecated anthropic-ai token, so blocking a string ClaudeBot does not honour was reported as blocking ClaudeBot.
Root cause
A failed or mis-aimed observation was allowed to produce a finding. None of these sentences said 'you' and none of them said 'us' either, so a reader supplied the subject, and the subject a reader supplies is always themselves.
Invariant that now prevents it
A collector may emit a claim about a store only from an observation that succeeded, on the page that observation belongs on. When an observation fails or was never attempted, it emits a coverage statement that affirmatively locates the failure with us. Not-assessed is a third state, distinct from pass and from fail, and it is removed from the denominator rather than scored zero.
Enforced by lib/growth/merchant-claim.ts, validated over every failure-path reason each collector can emit by the compliance suite in __tests__/growth-engine.test.ts.
b. Unverified success credited to the merchant
The opposite direction, and the one that is easier to miss because nobody complains about a score that is too high.
What was published
- og:image scored on the presence of the tag. The tag is on virtually every modern site, so the points were close to free, and a page whose card renders blank in Slack scored the same as one that renders.
- canonical scored the same way: the tag exists, so the URL inside it was assumed to resolve.
- llms.txt was worth more than any other single item and was awarded on an HTTP 200. A catch-all rewrite answers 200 for a file that does not exist.
- sitemap presence was inferred from a 200 rather than from parsing what came back.
Root cause
Presence of a pointer was treated as proof of the thing pointed at. Every one of these is a claim we never actually checked, and each inflates the score of a site that would fail in the real world.
Invariant that now prevents it
A check may return ok only when it observed the thing it claims, and not-ok only when it observed the absence. Following a pointer is part of the check, not an optimisation. Anything else is not-assessed. Concretely, two questions have to be answered no when a check is added: if our network is broken, does this blame the merchant, and if their asset is broken, does this still award points.
Enforced by The three-state LinkedAssetCheck shape in lib/growth/collectors/ai-readiness.ts, which replaced the booleans, plus the linked-asset verification pass that actually fetches og:image and canonical targets.
c. Aggregation error
The arithmetic itself was wrong, in a way that was visible on the page.
What was published
- A report stated '104 of 102 product titles' had an issue. A count larger than the population it is drawn from.
- The same store was told '61 of 102 products have an image_link problem: Missing product image'. The ground truth from its own products.json was six missing images and about fifty-five reused ones.
Root cause
Issues were grouped by field rather than by field and issue type, and counted as issues rather than as distinct products. A product with two title problems counted twice, which is where a number larger than the population came from. Collapsing four different image checks into one bucket kept the total and described it with whichever issue happened to be seen first.
Invariant that now prevents it
Any count presented as 'N of M' counts distinct members of M, and the label describes the specific issue type being counted rather than the field it belongs to. A total that can exceed its own population is a bug by construction, not a rounding question.
Enforced by Grouping by field and issue type with a distinct-product Set in lib/growth/collectors/catalogue-quality.ts.
d. Prose leak
The invariant was enforced where it was written and not where the text was actually generated.
What was published
- Collector reasons were validated against the merchant-claim rule. The synthesised report was not, so a model writing prose from those collectors could reintroduce exactly the claim the validator existed to catch, in its own words.
- Vendor detail reached public copy through the same gap: HTTP statuses, JSON error bodies, request ids and billing text are all things a collector sanitises at source and a generated paragraph can reproduce.
Root cause
The validated surface and the published surface were different surfaces. Structured fields were checked; the paragraph a visitor actually reads was not.
Invariant that now prevents it
Generated prose is validated against the same invariant as the structured reasons it is written from, using the set of collectors whose measurement did not complete, including any collector that returned ok without clearing its coverage bar. Vendor detail is replaced wholesale at a single chokepoint rather than patched, because a partial scrub is a leak waiting to be re-found.
Enforced by validateSynthesisProse in lib/growth/merchant-claim.ts, and the single publicReason chokepoint in lib/growth/public-reason.ts shared by the report builder and the PDF builder.
e. Matching error
The measurement was taken correctly and then compared wrongly, which produces a confident zero.
What was published
- Ten Thousand scored 0 on a run where assistants had in fact recommended it. The brand name is derived from a domain label and domains have no spaces, so we were looking for 'Tenthousand' while every assistant wrote 'Ten Thousand'.
- Catalogue text arrives HTML-encoded. A literal apostrophe and ' are different strings, so terms harvested from product data failed to match the same words written normally.
Root cause
One canonical form was assumed on both sides of a comparison whose two sides come from different systems: a domain label against natural prose, encoded markup against decoded text.
Invariant that now prevents it
Any comparison between our derived form and third-party text normalises both sides first, and a name match permits the separators a human would insert while still requiring every character in order and a non-alphanumeric boundary at each end. A zero is only reported when the matcher has been shown to be capable of a hit, and the answers that produced it are stored so the comparison can be re-run when the matcher improves.
Enforced by brandPatterns in lib/growth/collectors/ai-visibility.ts, entity decoding in lib/growth/query-generator.ts, and the full answer text retained in BrandIndexSnapshot.probeDetail so history can be re-scored rather than trusted.
f. Degraded coverage reported as a successful run
The panel narrowed, the runs kept completing, and every signal a person looks at said the same thing it says on a good day.
What was published
- Two of the panel's credentials were out of funds from 2026-09-02. The scorer ran on the assistants that still answered and recorded eight consecutive runs across nine days, every one of them stamped ok, while measuring a smaller panel than the one the page describes.
- The condition was detected correctly the whole time. credential_unhealthy fired critical for both, and the credential probe recorded the same reduced count on all 135 of its runs. What did not happen was anybody being told: the alert address was on a suppression list, so detection was perfect and delivery was the failure.
- Cost fell while coverage fell, because fewer assistants answering is cheaper. So the one number somebody was watching moved in the direction that reads as good news.
Root cause
A run's status was derived from whether the handler returned rather than from how much of the panel it reached, so a narrower measurement and a full one produced the same record. Nothing compared the panel actually reached against the panel the published method claims.
Invariant that now prevents it
A run that reaches fewer than the full panel is recorded with status partial, never ok, and anything derived from it is shown as partial wherever it appears. The condition raises category_pool_degraded, and that alert has a surface in /admin rather than only an address, because an alert is only as good as the inbox it reaches and this one proved it.
Enforced by The status derivation in lib/growth/run-category-pool.ts, which compares platformsAnswered against PANEL_PLATFORMS.length, the alert in app/api/cron/category-pool/route.ts, and the scorer's last runs with their panel size on /admin/connectors.
g. A placeholder sent to the assistant as if it were the brand
The question was well formed, the assistant answered it, the answer was scored correctly, and the brand named in the question did not exist.
What was published
- 705 scans across four question templates were sent with a literal {brand} token where the brand name belongs: is {brand} worth it, best {brand} products, {brand} alternatives, {brand} reviews. Measured over AcVisibilityScan on 2026-09-11, spanning 2026-03-27 to 2026-08-29.
- Every one of the 705 recorded the brand as not mentioned. Not most of them: all 705, and zero recorded a mention, which is what a template variable guarantees rather than what a market measures.
- Nothing looked wrong at any layer. The rows are well formed, the rate is a real number, and a brand that is never mentioned is an ordinary and believable result.
Root cause
The brand name was substituted into the template by a caller rather than carried by the thing being asked, so a caller that did not substitute produced a question that was still structurally valid. An unfilled placeholder is indistinguishable from a real name to everything downstream of it.
Invariant that now prevents it
A query carries its brand at construction and cannot be built without one, so an unfilled placeholder is a type error rather than a run. A guard additionally refuses any probe whose text still contains a template token, because a type is only as good as the one code path that respects it.
Enforced by The query type in lib/growth/query-generator.ts, which takes the brand as a required argument, and the no-template-token guard over every probe the scorer sends.
7. The record
Every index run appends a row that is never updated and never deleted, enforced by the database rather than by convention. It carries each collector score, the coverage conditions, the raw finding counts, and the complete text of every assistant answer.
The full answers are kept rather than excerpts because the matcher has already improved once and will again. An excerpt cut around today's match freezes the record at today's matching quality and makes a future correction impossible. Keeping the whole answer means a past zero can be re-examined against a better method instead of being taken on trust.
The pool’s runs, and the version each was asked under
Every pooled question put to every assistant appends its own row, holding the category, the assistant, the question, the model that answered, the brands it named in the order it named them, and whether it answered, refused or errored. A refusal and an error are kept apart from each other and from an answer naming nobody, because those are three different facts and only one of them is about the market.
Each row also carries the QUESTION-SET VERSION it was asked under, currently 2026-08-31-a, and the field is required rather than optional. The reason is specific: a brand that appears this week and not last week may be a fact about the questions rather than about the brand, and after the fact those two are indistinguishable unless every row says which set produced it. An optional field with a default would let a run claim a version it never used, which is worse than not recording one.
So when the question set changes, the change is in the record rather than in a changelog somebody has to believe: rows from before and after carry different versions, and a comparison across that boundary is visibly a comparison across two sets.
The refresh runs once a day and takes the stalest brands each time. No brand is re-measured inside 6 days, so any one brand is re-measured at least 6 days apart. The floor is a minimum: the true gap depends on how many brands are stale at once. Re-running the same brand daily would render assistant run-to-run variance as brand movement, which is the same class of false signal as every failure mode above: a chart of noise is still a fabricated finding. Sampling different brands each time is not that, and it is what keeps the interval from stretching as the registry grows.
9. What changed in this document
Version 1.2, 2026-09-11. Every change below is a change to what this page SAYS. Where it corrects something, the previous wording is quoted, because a correction nobody can see the before of is indistinguishable from an edit.
- Section 0 is new: the index ranks on the category pool, and the per-brand catalogue-derived method is the customer report. Sections 1 and 2c are now marked as describing the per-brand method only.
- The cadence block now states both: one category a night for the index, and the unchanged per-brand floor. The cycle length is counted from the categories rather than written down.
- Section 3 is CORRECTED. It said “the index refuses to refresh at all when fewer assistants are reachable than the bar requires”. No such refusal exists in the code. A narrowed run is recorded as partial, shown as partial, and raises an alert.
- Failure mode f is new: degraded coverage reported as a successful run, eight runs across nine days, with cost falling as coverage fell.
- Failure mode g is new: 705 scans sent with a literal placeholder in place of the brand name, every one recording a non-mention by construction.
- Section 3b is new: the three intake lanes, the three admission checks, what a listed-but-unscored row shows, and the two benchmark populations.
- Section 7 now describes the pool’s own run records and the question-set version each row carries.
Version 1.1, 2026-08-06: the cadence became derived rather than restated, the panel was reconciled against what actually answered, and a landing-page score whose unit is unknown is withheld rather than printed.
8. Corrections
If a measurement here is wrong about your brand, the useful thing to send is the question we asked and what your own check returned. Both are on your report page, which lists every question put to every assistant. A correction that changes the method bumps the version at the top of this document, and a correction that changes what a score means bumps the rubric instead.