Early accessMethodology and figures are provisional and may change. The example record shown uses RecommendedByAI's own measured data.
Vol. 2026 · Edition I
Methodology

The method, in full.

We publish the method, not just the conclusion. Anyone should be able to take a record and recompute it. This is the current, provisional method (v0.2) — grounded engines only.

01 What we measure

Frequency, not ranking.

The stable, meaningful quantity is not where a product appears in an answer but how often it is recommended for a specific buying need. List position is close to random across repeats; appearance frequency is stable. So the core metric is frequency by intent, and cross-engine agreement is reported as a separate number.

02 The instruments

Four independent, web-grounded engines.

We measure across ChatGPT, Perplexity, Gemini and Claude in modes that browse the web and return citations. Parametric (no-web) answers are treated as internal exploration only — never as evidence for a badge.

  • Neutral buying questions, many paraphrases, no brand names in the prompt.
  • Observations are spaced over time, not taken in a single burst.
03 The null baseline + bootstrap

It beats chance, or it isn't a claim.

Each intent has its own null baseline — the frequency you'd expect by chance given how many brands plausibly populate that need. A product's observed frequency must clear that baseline, and we attach a conservative bootstrap confidence interval so the separation is statistical, not anecdotal.

Worked example (our own measurement): for office / dressy the null sits near 24%; the example record clears it at 77% across 4/4 engines.

04 Source diversity

No echo chambers.

A signal that comes from one repeated citation is fragile. We measure the concentration of citing domains (an HHI over the cited sources); low diversity is flagged as a weakness and can disqualify a record. This stops a single affiliate page from manufacturing a recommendation. Known limit: Gemini currently returns citations as opaque redirect URLs, so its source-diversity score is not yet reliable — we treat Gemini's grounding as present but its domains as unresolved, and never rest a claim on Gemini alone.

05 Freshness, re-check & revocation

Signal that drifts is revoked, not kept.

  • Freshness: evidence supporting a record must be within a 90-day window.
  • Revalidation: records are recomputed on a cadence; stability over consecutive rounds is required.
  • Revocation & expiry: if a product falls below its baseline or loses cross-engine consensus, the record is automatically and auditably revoked and the store badge stops showing. If evidence merely ages out of the 90-day window without a contradicting round, the badge shows its date — a dated fact — rather than a current claim, and stops showing only past the category horizon; it never displays a negative mark.
06 Eligibility (C1–C7), in plain language

All must hold at once.

A (product × intent) pair is eligible only if every criterion below holds on fresh evidence. Thresholds are marked [provisional] and are recalibrated as the program matures.

IDPlain-language criterionProvisional bar
C1Enough observations≥ 120 valid queries
C2Appears in enough grounded engines≥ 2 web-grounded engines
C3Recommended often enough for the need≥ 30% aggregate frequency
C4Stable within a round (split-half)Δ ≤ 5 pp
C5Stable over time — maintenance, not an issuance gatet=0 issues if the signal clears the null with margin (CI-low ≥ null + 15 pp); a signal that only just clears the null waits for a 2nd round. Re-checked every round (Δ ≤ 10 pp)
C6Evidence is fresh≤ 90 days old
C7Real, buyable, citation-backed in the market≥ 1 resolvable US retail citation

You cannot pay to change any of these. Payment covers measurement, display and upkeep — never eligibility.

07 Auditability

Reproducible from the record.

Every eligibility decision is reproducible from the per-query log: given a round id, an outside auditor can recompute C1–C7 and confirm the status. The method version travels with the badge.

08 What a claim does not mean

The limits, stated.

A record means: AI recommends this product for this intent, this often, on this date, across these engines.

It does not mean the product is "the best", it is not an endorsement by any AI provider, and it is not a prediction of sales. Claims are scoped to English / US; other markets are measured separately.

Status: methodology v0.2.2 · grounded-only scope · numeric thresholds provisional, recalibrated in phase 1. Example figures are RecommendedByAI's own measured data for Carets · office.

Read the method, then check your own category.

Free category check Earned, not bought.