How we measure
SEOAIO measures your visibility across traditional search and the generative engines, the discipline now called GEO (Generative Engine Optimization) or AEO (Answer Engine Optimization). We run on one doctrine: measure, never manufacture. Every number in the product is something we actually observed, and every number tells you how it was observed. This page explains the method.
The doctrine
A visibility score is only useful if it reflects reality. So we do not estimate where we can measure, we do not extrapolate a single engine into a market-wide claim, and we never present a modeled figure as an observed one. If a pillar has not been measured, the dashboard shows it as not present. It is never filled in with a plausible-looking placeholder.
Sampling AI answer engines
For each site we run a fixed set of prompts against each AI answer engine we support, several samples per prompt, because these engines are not deterministic. We count a citation only when the engine actually returns your domain as a source in its answer. A passing mention of your brand name is not a citation, and we do not count it as one. Results are reported per engine: a Perplexity number is labeled as a Perplexity number, never generalized to engines we have not sampled.
Confidence bands and sample sizes
Small samples produce uncertain percentages, and we say so. Every sampled rate ships with a Wilson score confidence interval and the exact sample size behind it. When you see "33%, CI 6-79%, 3 prompts / 9 samples," that width is honest: it tells you the measurement is early, and it narrows as more samples accumulate. We would rather show you a wide interval than a precise-looking fiction.
The thresholds, with their numbers
Five numbers decide whether a result is allowed to be shown to you at all. Each one is read from the constant that enforces it in the running code rather than typed onto this page, so what you read here is what the code does, not what we remembered to update.
- Before a score exists at all5 of 7 pillars
- The Visibility Score is held until at least 5 of its 7 pillars have been measured, and only measured ones count. An estimate is a model's opinion and a derived value is arithmetic over the others, so neither can buy a score its way out. Below the line you get every pillar we do have, listed separately, and no headline number.
- Before any verdict3 runs
- A prompt needs at least 3 comparable runs before we will call it anything. Under that it reads as not measured. It does not read as stable, and it does not read as a zero.
- Written off as noise25% floor
- We call a prompt noise once the 95% Wilson lower bound on its flip rate reaches 25%. The lower bound, not the observed rate: declaring that a prompt can never prove anything is the destructive call, and one flip across three young runs already looks like 33% while proving close to nothing.
- Cleared as stable20% or below
- A prompt is stable when its observed flip rate is at most 20% and the noise bar is not met. This one is judged on the observed rate on purpose. Stable only means an experiment here could show something, and that experiment still has to clear the change gates on its own.
- On every share we publish95% interval
- Every rate ships with a 95% Wilson score interval and the sample size behind it. A wide band on early data is the honest reading of early data, so we show the width rather than a tidier number.
How we decide where a question sits in a buyer's journey
Every tracked question is placed by its own text, by the rules below, in this order, first match wins. No model is involved and nothing is hand-assigned, so the same question always lands in the same stage and you can check the answer yourself. A question whose text supports no stage stays unclassified and is shown that way.
- Discovery
- The buyer is looking for a category and has nobody in mind. Being absent here means never entering the shortlist.
- Evaluation
- The buyer has names and is comparing them. This is where a competitor's page answers a question about you.
- Decision
- The buyer is choosing and asking about price, trials and whether it is worth it. The shortest path to revenue and the smallest number of questions.
- Retention
- The asker already pays you and is wondering whether to stop. Nothing else in this product measures the answers they are given.
- Unclassified
- The text does not say where the asker is. A question that names a brand and asks what it is fits no stage: it could be a stranger who saw an ad or a customer checking what they bought, and we will not pick one.
The rules, in the order they are applied
- 1.Retention Asks how to cancel, unsubscribe, downgrade or get a refund.
- 2.Retention Asks about leaving, switching away, or why customers stop using it.
- 3.Retention Asks whether to keep paying: renewal, or whether it is still worth it.
- 4.Decision Asks directly whether to buy, choose, or sign up.
- 5.Decision Asks whether a specific product is worth its price.
- 6.Decision Asks about a free trial, a discount, a coupon, or how to start.
- 7.Evaluation Asks for alternatives to, or a comparison against, something named.
- 8.Evaluation Asks what existing customers say: reviews, complaints, testimonials.
- 9.Evaluation Asks whether it can be trusted: legit, safe, reliable, a scam.
- 10.Evaluation Asks what something costs: how much, pricing, the cost of.
- 11.Discovery Asks for the best, top, or a good option in a category.
- 12.Discovery Asks which companies or tools do a thing.
- 13.Discovery Asks how to accomplish the job the category exists for.
Order matters and is part of the rule. "Is it still worth paying for" contains a price question and a renewal question; the renewal is what tells you the person asking already pays, so Retention is tried first.
Provenance on every number
Every metric carries one of three labels. Measured: we observed it directly, and you can see when and how. Derived: computed arithmetically from measured values, with the inputs disclosed. Modeled: an estimate, clearly marked as one, and never blended silently into measured figures. If a collection run fails, it writes no aggregate at all; a crashed measurement is a gap, not a zero and not a guess.
How often we measure
We re-check your pages every 24 hours and re-sample the AI engines every 7 days for each site you monitor. Checking your pages is cheap, so we do it daily. Asking the AI engines costs real money on every question, so that runs weekly and the sample size is set by your plan.
The cadence matters more than it sounds. Anything we tell you about movement is only as good as the number of times we looked, which is why we publish the interval and the confidence range beside the figure rather than just the figure.
When we do the work ourselves
We will help you fix what we find. That creates an obvious problem: a company that does the work and then grades it is marking its own homework, which is exactly what most tools in this category do. So the separation is a rule, not a preference, and these are its four clauses.
- We never sit in the path your site serves from. No proxy, no DNS takeover, no runtime injection. Your server sends your bytes.
- The test that decides whether something moved does not know who moved it. The same statistical gate is applied whether you made the change, your agency did, another tool did, or we did. We do not get an easier test for our own work. This is the clause that does the real work.
- Every applied change is logged with who applied it. Including when the answer is us. Your change log tells your work apart from your agency's and from ours.
- Nothing is ever applied without you approving it. We do not run an agent that edits your site while nobody is watching, and we are not planning to.
What we do NOT do
- No manufactured citations. We never seed, plant, or fake mentions of your brand anywhere.
- No auto-published spam backlinks. Link schemes carry penalty risk, and we will not take it with your domain.
- No black-box single number. Every score decomposes into the measurements that produced it, and you can inspect each one.
- No silent extrapolation. One engine's result is never presented as an all-engine result.
- No unattended edits to your site. Other tools run an agent that rewrites your pages on a schedule. We do not, and the approval step is not a setting you can switch off.
What honest measurement looks like, week by week
Day 1
First real samples
Your first sweep runs real prompts against real engines. Numbers appear with wide 95% bands, because a handful of samples deserves a wide band, and we show it instead of hiding it.
Day 7
Intervals narrow
Repeated sweeps accumulate. The same share from five times the samples is a different measurement, and you watch the band tighten around it.
Day 30
Change detection activates
With enough history, run-over-run movement is tested against noise (Fisher's exact). From here, “your visibility moved” means it moved, not that noise pointed upward.
Tools that show a confident score in minute one are showing you a model, not a measurement. Ours starts honest and gets sharp.
