Skip to main content
Methodology · How we measure

Measurement you can defend.

Most AI-visibility tools report one number from one run. Answer engines return a distribution, not a rank — so a single sample tells you almost nothing. This is the whole protocol: how we sample, what we count separately, how we test whether your work moved anything, and what we refuse to claim.

The protocol at a glance
Sampling
Repeated runs × phrasings
Evidence
Intervals, not points
Impact verdicts
Probability, with a range
Control group
Untouched prompts
01Visibility is a distribution

Ask the same question twice. Expect two answers.

Generative answers move run to run and day to day. A tool that scans once and prints a rank is reporting noise with a decimal point attached to it.

Repeat the run.

Every measured prompt runs multiple times per phrasing. A citation rate comes from a real sample per platform, not from one lucky draw on a Tuesday.

Freeze the phrasings.

Alternate phrasings of each measured question are generated once and then frozen for the life of the prompt. Rewriting the question between scans would make every week-over-week comparison meaningless.

Report the spread.

Each prompt carries a stability reading next to its rate, and thin evidence is shown as an interval rather than a point. Thirty percent from a handful of checks is not the same claim as thirty percent from nine hundred.

Name the denominator. A headline visibility rate counts only questions that never mention the brand. A question that already contains your name will hand it back to you, so folding those answers into the rate measures recognition and reports it as discovery. Questions that name the brand, and questions that name a rival, are still tracked — they are reported as their own readings, never averaged into the headline.

This is not a stylistic preference. A systematic comparison of five generative search systems against Google organic search found that engines differ substantially in how much they retrieve versus recall from memory, and that their outputs vary across time and across repeated executions of the same query — Kirsten et al.. Any protocol that samples once is measuring the wrong object.

02The visibility funnel

“Not cited” is two problems wearing one label.

Before an engine can cite you, it has to search at all. We record whether retrieval actually happened on every single run, and split the number accordingly.

Pr(cited) = Pr(AI searched) × Pr(cited | searched)

Two terms, two different fixes. Collapsing them into one score is how a content problem gets misdiagnosed as an authority problem, and vice versa.

The AI never searched.

The model answered from what it already carried. Nothing on your site was fetched, so nothing you published that week could have been read.

This is an entity problem, not a content problem. You need to exist in what the model already knows, in the sources it learned from.

The AI searched and skipped you.

Retrieval happened, your pages were reachable, and the engine chose someone else as the source it would cite.

This is the one your next article can actually fix — structure, coverage, and authority against the specific pages that won.

Mentioned is not cited.

Every run records two separate facts: whether your brand was named in the answer text, and whether your site was linked as a source underneath it. They are different levels of visibility with different causes, so we never merge them into a single number. Being named without being linked usually means the model already knows you and did not need your page. Being linked without being named means your page won the retrieval and lost the sentence.

03Proving impact, not correlating it

A citation after your fix is not proof of your fix.

AI referral traffic is growing on its own. Any tool that shows you a rising line after you publish is showing you the tide. Six steps separate the two.

  1. 01

    Log the change

    Every published edit is recorded with its timestamp and the prompts it was meant to move. No log, no verdict — we will not reverse-engineer a cause from a chart after the fact.

  2. 02

    Fix the windows

    Pre and post windows are cut around the change date before any result is read, so the comparison cannot be tuned until it flatters the work.

  3. 03

    Fit your own prior

    A statistical baseline is fitted to your site’s own citation history. Your baseline noise is the floor a change has to clear — not an industry average, not a vendor benchmark.

  4. 04

    Read the posterior

    The verdict is a probability with a credible interval — the report says “8 in 10 chance this is a real improvement” and shows the range the lift most likely sits in. Not a green arrow.

  5. 05

    Hold out the untouched

    Prompts you did not target form a control group. If they moved by the same amount, the platform moved — not your content.

  6. 06

    Drop the shock days

    Days flagged as platform-wide volatility are excluded from both groups, so a model update never gets credited to your content work.

The treated-minus-control difference is then tested against chance itself: we ask how often random assignment alone would produce a gap this large, using inference that assumes nothing about the shape of the noise. That matters at the prompt counts real sites actually have — where the textbook tests quietly stop being valid.

04Evidence grading

Grade the study, not the headline.

The strength of a claim is a property of the design that produced it, not of the number it reports. Evidence-based fields have graded designs this way for decades. GEO has not started.

Evidence grades A through E, by study design and what each design can tell you
GradeStudy designWhat it can actually tell you
ARandomized or quasi-experiment on live engines, with logged changes, a holdout group, and pre/post inferenceWhether a specific change moved a specific outcome
BObservational panel with a control group and time controls, but no randomizationA direction, with the obvious confounders ruled out
CSingle group, pre/post only, no controlThat something changed. Not what changed it
DCross-sectional correlation or a one-shot auditWhich signals tend to co-occur with being cited
EFixed-context simulation — sources supplied to the model, no live retrievalHow a rewrite performs once you are already in the answer

Most of what circulates as GEO advice sits at D or E — including the field's most-quoted percentage. That is not a slur on the work. Fixed-context studies are cheap, repeatable, and genuinely useful for learning how to write. They simply cannot tell you whether an engine would have retrieved you at all, because the sources were handed to the model before the test began.

Prominara's impact verdicts are built to sit at A or B for your own site: live engines, logged changes, an untouched holdout, and inference that survives being questioned. When the evidence only supports a weaker grade — too few prompts, no clean holdout, a window contaminated by platform volatility — the report says so instead of rounding up.

05What we will never claim

Five things we will never put in writing.

  • A guaranteed placement in any AI answer, on any platform, at any price.
  • A timeline. “Results in two to four weeks” is a sales line wearing a measurement costume.
  • An AI SEO hack. There is no ranking API to game, and anyone selling one is selling a guess.
  • A cross-platform ranking. Answer engines do not publish ranks, so every “your rank in ChatGPT” number is invented.
  • An ROI figure without a control group. Compared to what is not an optional question.

We do not control generative ranking, and neither does anyone selling you a subscription. What we control is the measurement — the sampling, the denominators, the controls, and whether a verdict is stated at a confidence the evidence actually supports. If a number cannot survive the question compared to what, it does not ship.

06References

The primary record. Read it yourself.

Four papers this protocol is built on, linked to the originals. Where a headline number gets quoted out of context, the note says what the number actually measured.

Now run it on your prompts.

Every measurement on this page runs on your own questions, on a fixed schedule, with the evidence shown underneath the number. 14 days free, no credit card.