How the measurement works
Methodology
Anyone can print a percentage. This page is how you check ours — the questions we ask, how many times we ask them, the estimator by name, the point at which we refuse to report a number, and the things we cannot yet see. If your agency wants to pull this apart, everything they need is on this page.
- Questions per run
- 40 (min 20)
- Repetitions each
- 8 (floor 8)
- Confidence
- 95% (z = 1.96)
1. The questions are the instrument
A visibility number is only as meaningful as the questions behind it. Asking one question measures that question, not your category — a brand can lead for “best {category} brands” and be absent from “which is best on a budget”, and quoting either alone is a cherry-pick.
Every run asks 40 questions, never fewer than 20, balanced across five intents: informational, comparative, transactional, local, and problem-led. No single intent may exceed 40% of the set and each required intent must have at least 3 questions. That ceiling is the rule that matters — without it, a set can meet every minimum and still be twenty-five rewordings of one question.
Your questions never contain your brand name. If they did, the engine would simply repeat it back and “visibility” would measure nothing. This is enforced in code: a run whose questions contain any of your brand’s aliases is refused rather than reported.
2. The questions are frozen before we measure
A prompt set is locked before the first run executes against it, and the lock timestamp appears in your report. A locked set cannot be edited — changing a question forks a new version with a new fingerprint, and the two versions are different instruments whose numbers may not be charted together.
This exists to make one specific abuse impossible rather than merely detectable: run a set, notice the client never appears for three questions, quietly drop those three, re-run, and report an improvement.
3. Why we ask the same question many times
An answer engine is not a database. The same question, on the same model, returns a different set of brands from one minute to the next. A single check can report “not mentioned” and “mentioned” for the same brand minutes apart, which is why a one-shot answer is not a measurement.
Each question is asked at least 8 times per engine, 8 by default. Engines are queried at their default temperature and never pinned to greedy decoding: a deterministic answer is not what a real user sees, and pinning it would manufacture false precision. We do not cache responses, we do not set seeds, and we never de-duplicate identical answers — identical answers are real data. Every run checks for this automatically and flags any engine returning byte-identical text across repetitions.
4. The estimator, by name
Rates pooled across questions use a cluster-robust interval with the question as the cluster unit. Rates for a single question on a single engine use a Wilson score interval. At a rate of exactly 0% or 100% the cluster-robust variance is undefined, so we use a Wilson interval on the effective sample size — zero mentions bounds a rate, it does not prove the rate is zero.
Why not simply pool every answer? Because repeated answers to the same question are not independent of each other. If a brand is named in every answer to one question and none to another, asking both ten times does not give twenty independent observations. Treating them as independent produces intervals that are too narrow — on our own real runs, by a factor of roughly two to three.
So every report carries an effective sample size beside the raw answer count: what was collected, and what it is actually worth. Any report quoting a raw response count without one is making the optimistic assumption silently.
5. When we refuse to give you a number
A rate is marked provisional when its 95% interval is wider than ±10%, and the report says how many more answers would be needed — scaled by the clustering penalty, because reaching a target precision on clustered data genuinely takes more data than the naive formula implies.
A run that loses more than 5% of its planned calls is marked degraded, its intervals are widened, and it is not comparable to a clean one. And a run that would exceed its budget is refused before it starts rather than quietly trimmed — a truncated run is a biased sample reported as though it were the run you asked for.
Refusing to answer is a feature. Where the data cannot support a claim, the report says what was measured, why the interval is too wide, and what would be needed — instead of printing a confident number you could act on and be wrong.
6. Before we tell you something changed
A rate that moves between weeks has two possible explanations: your brand moved, or the engine did. Most tools report the first without checking the second.
We publish an engine volatility index — the variance observed across repeat runs divided by the variance the sampling design already predicts. Around 1 means the engine is steady and a week-over-week change is meaningful. Around 4 means it moves roughly twice as much as sampling explains, and small changes mean nothing.
A change is only reported as real when it clears both hurdles: beyond sampling noise at the effective sample size, and beyond that engine’s ordinary run-to-run range. It takes 3 runs before we will compute the index at all, and 5 before we treat it as more than indicative. Below that, the honest answer is that a change is unverified — not that it passed.
7. Comparability over time
Every report carries a prompt-set fingerprint, a content hash of the exact questions asked. Two reports may only be compared if their fingerprints match. Change a single word of a single question and the fingerprint changes, which makes a trend line across a wording change structurally impossible rather than merely discouraged.
When a client moves to a larger question set, we run both instruments in the same window and report the difference between them explicitly, so an instrument change is never mistaken for a real one.
8. Which engines we actually cover
Verified in production and used in client runs: anthropic, anthropic_search, openai. The grounded engine performs a live web search before answering; the ungrounded ones answer from model memory alone. We report them separately, because a brand can be prominent in one and invisible in the other, and blending them hides exactly the difference that matters.
Adapters exist but are not verified for perplexity, google, google_search. They are refused in client runs rather than included on trust, and they are not counted when we describe our coverage.
9. What we cannot do, and what we got wrong
This section exists because a methodology without one is marketing. These are the real limits, including the ones that do not flatter us.
- No Google AI Overviews or AI Mode yet. For a local business this is a material gap — it is where a large share of intent actually resolves. There is no first-party API for it, and we will not run a scraper. This is the next engine we are adding.
- The grounded engine does not always search. Measured across our stored runs, about one grounded answer in eight was produced without a web search — answered from model memory while being recorded as grounded. We are recording this per answer so it can be reported rather than assumed.
- We retracted our own headline claim. We previously believed we had evidence that answer engines drift substantially day to day, based on two runs whose intervals did not overlap. Once we corrected the estimator to account for clustering, those intervals overlapped comfortably and the volatility index came out at or below 1 for every engine we measured. There was no drift; there was an error in our own maths. The claim is gone.
- Effective sample sizes can be much smaller than they look. On a real run, 144 answers were worth about 14 independent observations for one brand, because its visibility was concentrated in a few questions. We print that number rather than hiding it.
- US English only. No multi-country or multi-language measurement.
- We measure; we do not publish on your behalf. We do not write or ship content to your site. A party that produces the content cannot credibly audit whether it worked.
Check it yourself
Every figure in a Perceply report links to the stored answer it came from — the raw response, with the provider’s own response ID. If a number cannot be traced to evidence, it does not get printed.
Run a free scan