Nostimates
← All posts
Methodology·2 min·Nostimates

How to measure AI visibility properly: a working method

Measuring AI visibility properly requires five things: a prompt set spanning four buyer intents, a declared competitive set, logged-out collection in the target market and language, at least ten samples per prompt per engine, and retention of the raw answers so any number can be traced back to its source.

How to measure AI visibility properly: a working method

Step 1 — Build a prompt set that can embarrass you

Cover four intents: category discovery ("best CRM for B2B sales"), comparison ("HubSpot vs Pipedrive"), problem-led ("how do B2B teams choose a CRM"), and brand-specific ("is Pipedrive any good"). Twenty to sixty prompts is the usual working range.

A set weighted to brand-name queries will always flatter the client and will never surface category risk. If the prompt set can't produce bad news, it isn't a measurement instrument.

Step 2 — Declare the competitive set up front

Three to five competitors, named before collection starts. Every run then reports each brand's share on identical prompts and samples. Comparing your number to a competitor's number from a different tool or month is not a comparison.

Step 3 — Collect logged-out, in-market

Set the country and language deliberately. On SERP-based engines that's parameters plus residential egress; on chat engines the prompt language is the locale control, so measuring Japan means writing the prompt in Japanese.

Do not collect from a signed-in account. A personalised session is not reproducible by anyone else, which makes the resulting number unusable in an argument.

Step 4 — Sample enough to have an interval

At n=1 you have no interval. At n=10 the 95% interval around a 30% share is roughly ±25 points. At n=25 it's about ±17. At n=81 it's about ±10.

Pick the sample size from the change you need to detect, not from your budget, and if the budget won't cover it, reduce the number of prompts rather than the samples per prompt.

Samples per prompt Interval around 30% share
10 ±25 points
25 ±17 points
50 ±12 points
100 ±9 points

Step 5 — Keep the answers

Store every sampled answer with its timestamp, engine, market and parser version. When the client says "that's not what I see", you show them the ten answers rather than relitigating the methodology. This is the single highest-leverage thing you can do for the durability of a retainer.

The failure mode to avoid

A screenshot of one AI answer, taken once, presented as evidence of visibility. It proves the answer existed at that moment. It says nothing about the typical case, and a client who repeats the query and sees something different will stop trusting the whole report.

Step 6 — Push movement, don't wait for someone to check

A dashboard is there when you want to explore, but few people open one every week unprompted. So alert on share moving beyond its interval, on a new competitor entering the answer set, and on extraction health dropping for an engine. Everything else can wait for the monthly.

See this for your brand

Nostimates shows you how your brand shows up across twelve AI engines and search — in a dashboard, or as data in your own tools. Tell us what you want to measure.

No newsletter. One reply from a human.