We publish the method, not just the conclusion. Anyone should be able to take a record and recompute it. This is the current, provisional method (v0.2) — grounded engines only.
The stable, meaningful quantity is not where a product appears in an answer but how often it is recommended for a specific buying need. List position is close to random across repeats; appearance frequency is stable. So the core metric is frequency by intent, and cross-engine agreement is reported as a separate number.
We measure across ChatGPT, Perplexity, Gemini and Claude in modes that browse the web and return citations. Parametric (no-web) answers are treated as internal exploration only — never as evidence for a badge.
Each intent has its own null baseline — the frequency you'd expect by chance given how many brands plausibly populate that need. A product's observed frequency must clear that baseline, and we attach a conservative bootstrap confidence interval so the separation is statistical, not anecdotal.
Worked example (our own measurement): for office / dressy the null sits near 24%; the example record clears it at 77% across 4/4 engines.
A signal that comes from one repeated citation is fragile. We measure the concentration of citing domains (an HHI over the cited sources); low diversity is flagged as a weakness and can disqualify a record. This stops a single affiliate page from manufacturing a recommendation. Known limit: Gemini currently returns citations as opaque redirect URLs, so its source-diversity score is not yet reliable — we treat Gemini's grounding as present but its domains as unresolved, and never rest a claim on Gemini alone.
A (product × intent) pair is eligible only if every criterion below holds on fresh evidence. Thresholds are marked [provisional] and are recalibrated as the program matures.
| ID | Plain-language criterion | Provisional bar |
|---|---|---|
| C1 | Enough observations | ≥ 120 valid queries |
| C2 | Appears in enough grounded engines | ≥ 2 web-grounded engines |
| C3 | Recommended often enough for the need | ≥ 30% aggregate frequency |
| C4 | Stable within a round (split-half) | Δ ≤ 5 pp |
| C5 | Stable over time — maintenance, not an issuance gate | t=0 issues if the signal clears the null with margin (CI-low ≥ null + 15 pp); a signal that only just clears the null waits for a 2nd round. Re-checked every round (Δ ≤ 10 pp) |
| C6 | Evidence is fresh | ≤ 90 days old |
| C7 | Real, buyable, citation-backed in the market | ≥ 1 resolvable US retail citation |
You cannot pay to change any of these. Payment covers measurement, display and upkeep — never eligibility.
Every eligibility decision is reproducible from the per-query log: given a round id, an outside auditor can recompute C1–C7 and confirm the status. The method version travels with the badge.
A record means: AI recommends this product for this intent, this often, on this date, across these engines.
It does not mean the product is "the best", it is not an endorsement by any AI provider, and it is not a prediction of sales. Claims are scoped to English / US; other markets are measured separately.
Status: methodology v0.2.2 · grounded-only scope · numeric thresholds provisional, recalibrated in phase 1. Example figures are RecommendedByAI's own measured data for Carets · office.