What can actually be detected

Tell us how big a change you need to see. We will tell you what it costs to measure — and, when your budget cannot buy it, that you should not buy it. That answer is arithmetic, and it applies to every tool in this category, not only to ours.

Your situation

What it takes

To detect a change of 10.0 points around a 25.0% baseline you need 295 independent observations. With an assumed design effect of 2.50 (measured runs in this system range from about 2.2 to 10, so this plan is if anything optimistic), that is 738 sampled answers — about 21 questions at 12 repetitions across 3 engines, costing about $6.51 per round in model spend.

Independent observations
295
Answers to sample
738
Questions × reps
21 × 12
Model spend per round
$6.51

Why the answer is bigger than you expected

Halving the change you want to detect costs four times the sample. Detecting 20.0 points is cheap; detecting 5.0 points is sixteen times the work. This is the fact that makes most weekly AI-visibility reporting decorative.

Asking the same question repeatedly does not buy independent answers. Responses to one question are correlated, so 738 sampled answers are worth about 295 independent observations. The figures above assume a design effect of 2.5; measured runs in this system range from about 2.2 to 10, so this estimate is on the optimistic side.

A tool quoting its raw response count is quoting the bigger, wrong number. “One million prompts a month” describes spend, not sensitivity.

What we sell, against those numbers

The ai visibility audit at $500 samples 144 answers. Weekly monitoring at $249/month repeats that weekly on a fixed question set.

If the number above is far larger than 144, the honest reading is that a one-off audit tells you where you stand but will not resolve the change you are chasing, and you should either widen the change you are willing to call a win or commit to more sampling.