Nostimates
← All posts
Methodology·5 min·Nostimates

One reading is not a measurement

AI answers vary between identical requests, so a single reading cannot establish a brand's visibility. Sampling the same prompt repeatedly and reporting a rate with its 95% Wilson interval — 34% ±17 points at n=25 rather than a bare 34% — is the minimum standard for a number a client can act on.

One reading is not a measurement

Run the same prompt ten times

You will not get ten identical answers. Brands appear and disappear, citation order shuffles, and a source that led the answer at 9am is absent at 9:05. This is not a bug in the surface. It is what a sampled generative system does.

What a single sample actually tells you

At n=1, a brand either appears or it doesn't, so every score you can report is 0% or 100%. The confidence interval around a single Bernoulli trial spans nearly the whole range. You have not measured visibility; you have observed one coin flip.

The reporting difference

What you say What happens next quarter
"You're at 34%" It reads 28% next week and you look wrong
"You're at 34% ±17 points at n=25" 28% is inside the range and your method holds

The second sentence is the one that survives a QBR. It also lets you say something genuinely useful: at n=25 a move from 34% to 62% clears both intervals and is real, while a move to 37% is not. Narrower claims need more samples, not more confidence: the 95% Wilson half-width around a 30% share is ±25 points at n=10, ±17 at n=25, ±12 at n=50 and ±9 at n=100.

Practical floor

n=10 is the floor for standard reporting, n=25 for a competitive claim, n=50 for before-and-after tests you intend to defend, and n=100 for anything we publish.

See this for your brand

Nostimates shows you how your brand shows up across twelve AI engines and search — in a dashboard, or as data in your own tools. Tell us what you want to measure.

No newsletter. One reply from a human.