Methodology

Why two AI visibility tools give you different numbers

G2's answer-engine-optimization category grew from 7 products to 248 listings in thirteen months. Run two of them against the same brand and you will get two answers. Both can be computed correctly and still disagree, and neither vendor will usually be able to tell you where the gap came from, because a number can only be reconciled if the inputs behind it were published.

Here is a worked one. Every figure attributed to the other instrument is quoted from a rival's own public sample audit, read on 13 August 2026, and the arithmetic below is recomputed from those figures rather than asserted. We have picked a report that is unusually transparent for this category: that is what makes it reconcilable at all, and it is the reason nothing here is a complaint about them.

The published figure

Their sample audit reports a share of answer of 2 out of 8 unbranded queries. 9 unbranded prompts were run (alongside 1 branded control kept out of the rate), 1 was excluded as a legitimate non-fit, and the business appeared in 2. That is 25%.

On our instrument the same 2 of 8 carries a 95% interval of 7% to 59%, because 8 questions is a small denominator and the interval is what says so. Both numbers describe the same run. Only one of them tells you how much to trust it.

The four choices that produce the gap

None of these is a bug in either tool. They are decisions, and a buyer comparing outputs without knowing them is comparing two different measurements labelled as one thing.

  1. Choice 1 of 4

    What counts as a question

    Their instrument
    9 unbranded prompts asked, 1 dropped as a legitimate non-fit, 8 scored. The exclusion and its reason are printed in the report, which is rarer than it should be.
    Ours
    Same rule, and the same publication requirement: an excluded prompt leaves the denominator only with a stated reason, and that reason prints on the coverage receipt beside every share it changed.
    What it does to the number
    Dropping one question moved this figure from 22% to 25%. One prompt, 3 points, on a denominator this small. A tool that excludes without printing can move a number that far and still call every figure measured.
  2. Choice 2 of 4

    Whether the unit is a question or an answer

    Their instrument
    Questions. 2 of 8 unbranded queries where the business appeared.
    Ours
    Questions, for the same figure, and this agreement matters more than it looks. We sample each question three times, and reporting 2 of 24 ANSWERS would be a different number describing the same run.
    What it does to the number
    Two tools can both be right and disagree by a factor of three here. Samples of one question are correlated rather than independent, so counting answers as trials narrows the interval below what the data supports.
  3. Choice 3 of 4

    Whether the brand's own name is in the question

    Their instrument
    Unbranded queries only, stated explicitly, with one branded query kept aside as a control.
    Ours
    Both, reported separately and never averaged. A branded question is close to a foregone win: an engine handed a name usually returns it.
    What it does to the number
    This is the single largest source of inflation in the category. A vendor quoting one blended figure over a prompt set containing branded questions is quoting a number that rises whenever they add another one.
  4. Choice 4 of 4

    Whether a sample size travels with the number

    Their instrument
    No interval, and the report says plainly that its score is not calibrated against an external benchmark.
    Ours
    Wilson interval on every share, always. On this run, 25% is genuinely 7% to 59% at 95%.
    What it does to the number
    The honest reading of 2 of 8 is not "25%", it is "somewhere between 7% and 59%". Two tools reporting 25% and 59% for the same brand are not necessarily disagreeing at all: both sit inside the same interval, and neither has enough questions yet.

How to reconcile any two tools yourself

Four questions, in this order. Ask them of us too.

  1. 1. What is the denominator? How many questions, and were any dropped? If any were, which ones and why?
  2. 2. Question or answer? If each question is asked more than once, is the rate over questions or over individual answers?
  3. 3. Is your name in the question? Ask for the branded and unbranded figures separately. If a vendor can only give you one blended number, that is the answer.
  4. 4. Where is the interval? A share with no sample size behind it cannot be compared to anything, including its own value last month.

Our answers are published rather than promised: the instrument disclosure states all four for every metric we ship, including our own detector's known error cases, and every run carries a coverage receipt naming what actually ran. We hold a per-question floor of 3 samples, below which an engine is withheld rather than published thin.

The uncomfortable part, said plainly: on 8 questions, neither instrument can tell this business much. 7% to 59% is most of the range. The fix is more questions, not more confident phrasing, and a vendor who responds to a wide interval by hiding it has made the number worse rather than better.