
Loading...
Most AI-visibility tools report one number from one run. Answer engines return a distribution, not a rank — so a single sample tells you almost nothing. This is the whole protocol: how we sample, what we count separately, how we test whether your work moved anything, and what we refuse to claim.
Generative answers move run to run and day to day. A tool that scans once and prints a rank is reporting noise with a decimal point attached to it.
Every measured prompt runs multiple times per phrasing. A citation rate comes from a real sample per platform, not from one lucky draw on a Tuesday.
Alternate phrasings of each measured question are generated once and then frozen for the life of the prompt. Rewriting the question between scans would make every week-over-week comparison meaningless.
Each prompt carries a stability reading next to its rate, and thin evidence is shown as an interval rather than a point. Thirty percent from a handful of checks is not the same claim as thirty percent from nine hundred.
This is not a stylistic preference. A systematic comparison of five generative search systems against Google organic search found that engines differ substantially in how much they retrieve versus recall from memory, and that their outputs vary across time and across repeated executions of the same query — Kirsten et al.. Any protocol that samples once is measuring the wrong object.
Before an engine can cite you, it has to search at all. We record whether retrieval actually happened on every single run, and split the number accordingly.
Pr(cited) = Pr(AI searched) × Pr(cited | searched)
Two terms, two different fixes. Collapsing them into one score is how a content problem gets misdiagnosed as an authority problem, and vice versa.
The model answered from what it already carried. Nothing on your site was fetched, so nothing you published that week could have been read.
This is an entity problem, not a content problem. You need to exist in what the model already knows, in the sources it learned from.
Retrieval happened, your pages were reachable, and the engine chose someone else as the source it would cite.
This is the one your next article can actually fix — structure, coverage, and authority against the specific pages that won.
Every run records two separate facts: whether your brand was named in the answer text, and whether your site was linked as a source underneath it. They are different levels of visibility with different causes, so we never merge them into a single number. Being named without being linked usually means the model already knows you and did not need your page. Being linked without being named means your page won the retrieval and lost the sentence.
AI referral traffic is growing on its own. Any tool that shows you a rising line after you publish is showing you the tide. Six steps separate the two.
Every published edit is recorded with its timestamp and the prompts it was meant to move. No log, no verdict — we will not reverse-engineer a cause from a chart after the fact.
Pre and post windows are cut around the change date before any result is read, so the comparison cannot be tuned until it flatters the work.
A statistical baseline is fitted to your site’s own citation history. Your baseline noise is the floor a change has to clear — not an industry average, not a vendor benchmark.
The verdict is a probability with a credible interval — the report says “8 in 10 chance this is a real improvement” and shows the range the lift most likely sits in. Not a green arrow.
Prompts you did not target form a control group. If they moved by the same amount, the platform moved — not your content.
Days flagged as platform-wide volatility are excluded from both groups, so a model update never gets credited to your content work.
The treated-minus-control difference is then tested against chance itself: we ask how often random assignment alone would produce a gap this large, using inference that assumes nothing about the shape of the noise. That matters at the prompt counts real sites actually have — where the textbook tests quietly stop being valid.
The strength of a claim is a property of the design that produced it, not of the number it reports. Evidence-based fields have graded designs this way for decades. GEO has not started.
| Grade | Study design | What it can actually tell you |
|---|---|---|
| A | Randomized or quasi-experiment on live engines, with logged changes, a holdout group, and pre/post inference | Whether a specific change moved a specific outcome |
| B | Observational panel with a control group and time controls, but no randomization | A direction, with the obvious confounders ruled out |
| C | Single group, pre/post only, no control | That something changed. Not what changed it |
| D | Cross-sectional correlation or a one-shot audit | Which signals tend to co-occur with being cited |
| E | Fixed-context simulation — sources supplied to the model, no live retrieval | How a rewrite performs once you are already in the answer |
Most of what circulates as GEO advice sits at D or E — including the field's most-quoted percentage. That is not a slur on the work. Fixed-context studies are cheap, repeatable, and genuinely useful for learning how to write. They simply cannot tell you whether an engine would have retrieved you at all, because the sources were handed to the model before the test began.
Prominara's impact verdicts are built to sit at A or B for your own site: live engines, logged changes, an untouched holdout, and inference that survives being questioned. When the evidence only supports a weaker grade — too few prompts, no clean holdout, a window contaminated by platform volatility — the report says so instead of rounding up.
We do not control generative ranking, and neither does anyone selling you a subscription. What we control is the measurement — the sampling, the denominators, the controls, and whether a verdict is stated at a confidence the evidence actually supports. If a number cannot survive the question compared to what, it does not ship.
Four papers this protocol is built on, linked to the originals. Where a headline number gets quoted out of context, the note says what the number actually measured.
Aggarwal et al. · KDD 2024
The foundational controlled benchmark for the field. Its widely quoted “up to 40%” is a relative gain measured on sources already retrieved into the answer context — a result about how to write, not a promise that an engine will find you.
Martinez · arXiv:2607.14035
A critical review of how GEO results are produced and reported. The discipline it argues for — repetitions, paraphrases, explicit denominators, controls — is what the protocol on this page implements in production.
Liu, Zhang & Liang · arXiv:2304.09848
Citation fidelity, audited by hand across four generative engines: on average only 51.5% of generated sentences were fully supported by the sources cited beside them. Being cited and being represented correctly are separate measurements.
Kirsten et al. · arXiv:2510.11560
A systematic comparison of five generative search systems against Google organic search. It documents exactly what our sampling is built around: engines vary substantially in how much they retrieve versus recall, and their outputs shift across time and across repeated executions.
Every measurement on this page runs on your own questions, on a fixed schedule, with the evidence shown underneath the number. 14 days free, no credit card.