How to measure AI visibility properly: a working method
Measuring AI visibility properly requires five things: a prompt set spanning four buyer intents, a declared competitive set, logged-out collection in the target market and language, at least ten samples per prompt per engine, and retention of the raw answers so any number can be traced back to its source.

Step 1 — Build a prompt set that can embarrass you
Cover four intents: category discovery ("best CRM for B2B sales"), comparison ("HubSpot vs Pipedrive"), problem-led ("how do B2B teams choose a CRM"), and brand-specific ("is Pipedrive any good"). Twenty to sixty prompts is the usual working range.
A set weighted to brand-name queries will always flatter the client and will never surface category risk. If the prompt set can't produce bad news, it isn't a measurement instrument.
Step 2 — Declare the competitive set up front
Three to five competitors, named before collection starts. Every run then reports each brand's share on identical prompts and samples. Comparing your number to a competitor's number from a different tool or month is not a comparison.
Step 3 — Collect logged-out, in-market
Set the country and language deliberately. On SERP-based engines that's parameters plus residential egress; on chat engines the prompt language is the locale control, so measuring Japan means writing the prompt in Japanese.
Do not collect from a signed-in account. A personalised session is not reproducible by anyone else, which makes the resulting number unusable in an argument.
Step 4 — Sample enough to have an interval
At n=1 you have no interval. At n=10 the 95% interval around a 30% share is roughly ±25 points. At n=25 it's about ±17. At n=81 it's about ±10.
Pick the sample size from the change you need to detect, not from your budget, and if the budget won't cover it, reduce the number of prompts rather than the samples per prompt.
| Samples per prompt | Interval around 30% share |
|---|---|
| 10 | ±25 points |
| 25 | ±17 points |
| 50 | ±12 points |
| 100 | ±9 points |
Step 5 — Keep the answers
Store every sampled answer with its timestamp, engine, market and parser version. When the client says "that's not what I see", you show them the ten answers rather than relitigating the methodology. This is the single highest-leverage thing you can do for the durability of a retainer.
The failure mode to avoid
Step 6 — Push movement, don't wait for someone to check
A dashboard is there when you want to explore, but few people open one every week unprompted. So alert on share moving beyond its interval, on a new competitor entering the answer set, and on extraction health dropping for an engine. Everything else can wait for the monthly.
See this for your brand
Nostimates shows you how your brand shows up across twelve AI engines and search — in a dashboard, or as data in your own tools. Tell us what you want to measure.