Across the real engines
We ask neutral buying questions — many phrasings, no brand names — across ChatGPT, Perplexity, Gemini and Claude, and record which products each names, and how often.
AI assistants already shape what shoppers buy. We measure what they recommend across ChatGPT, Perplexity, Gemini and Claude, verify it against auditable evidence, and issue it as a signal brands earn — and show.
AI is becoming a discovery channel — a third-party, verifiable signal is trust you can show on the product page.
When a shopper asks an assistant "what should I buy?", a handful of products make the list — and the rest disappear.
That list is a real distribution channel now. But it is invisible, it changes with every query, and no one has made it verifiable.
The signal that lasts isn't the ranking. It's how often a product is recommended — and that is what we measure, and prove.
We ask neutral buying questions — many phrasings, no brand names — across ChatGPT, Perplexity, Gemini and Claude, and record which products each names, and how often.
Every observation is captured with its sources and a timestamp. Frequency is tested against a per-intent null baseline with a conservative bootstrap interval, and source diversity is checked. Reproducible, not asserted.
When a product is consistently recommended for a real shopping need, it earns a category record — shown on the brand's own profile, and, through the Shopify app, on its product pages. Every surface links back to the evidence.
The brand seal is ours. What lands in a store is quieter: a small, neutral badge that adopts your theme and always carries its category, the engines behind it, the date, and a link to the full evidence.
Claims are measured across four independent, web-grounded engines, tested against a null baseline, and re-verified quarterly. Every record links to the raw, timestamped observations behind it. Anyone can audit them.
| Category / intent | Engines | Null baseline | Top frequency | Trend | Status | Updated |
|---|---|---|---|---|---|---|
| office / dressy example: Carets | 4 / 4 | 24% | 77% | ◉ Observed | 09·02 | |
| walking | 3 / 3 | 35% | — | Measuring | 08·17 | |
| hiking | 3 / 3 | 40% | — | Measuring | 08·17 | |
| wide feet | 3 / 3 | 42% | — | Measuring | 08·17 | |
| running | 3 / 3 | 44% | — | Measuring | 08·17 | |
| first barefoot | 3 / 3 | 52% | — | In review | 08·17 |
Each category has its own honest null baseline (chance level for that intent). A record reads Verified only when a brand clears its baseline across engines (with a bootstrap interval) and meets the full C1–C7 bar — enough observations and stability across rounds. The worked example (Carets · office) is Observed at 91% across 4/4 engines: it clears the null (64% vs 24%) but is not yet Verified — its frequency varies with phrasing (50–100% across ten questions), so it fails the stability check (C4). No brand appears as verified before its evidence does.
We're building the verification layer between AI recommendations and commerce — one auditable category at a time. Start with a free check of your own.