The refusals behind every number

What we will never tell you

An instrument is defined by what it refuses to output. These are the sentences this product is built never to say, each with where the refusal is enforced, because a promise that lives only in marketing copy is a promise until the first hard quarter. Conduct has its own page: what we refuse to do. The binding legal version is the Disclaimer.

  • Never 1 of 7

    “We can get you into the AI answers.”

    Nobody controls what ChatGPT, Perplexity, Claude or Google AI Mode say, including the people who sell placement in them. We measure what they say and help you change the inputs you actually own. Anyone guaranteeing a position in a system they do not control is describing their hopes, billed monthly.

    Where this is enforced: A test suite fails the build if guarantee language appears in our marketing (no-guarantees.test.ts). The legal version lives in the Disclaimer.

    What it costs us: We lose every buyer who wants to be promised the outcome rather than sold the instrument. Competitors happily take them.

  • Never 2 of 7

    “You scored 0” when the truth is “we did not look.”

    A zero and an absence are opposite findings. “No engine recommended you” is a result about your business; “we never asked” is a fact about our coverage. Rendering the second as the first is the cheapest way to look thorough, and it is manufactured data.

    Where this is enforced: The provenance grammar: anything unmeasured renders NOT MEASURED, dashed and with the reason, never as a number. The coverage receipt on every run says which engines were sampled and how many times.

    What it costs us: Our dashboards show fewer numbers than any rival demo, because a rival demo fills the gaps in.

  • Never 3 of 7

    A single headline score built on pillars we did not measure.

    A composite is only as honest as its weakest input. When fewer than five of the seven pillars are actually measured, the composite HOLDS rather than renders, because averaging measurements with guesses produces a number that looks like the first and is really the second.

    Where this is enforced: The vscore HOLD gate (composite withheld below 5 of 7 measured pillars), locked in the scoring doctrine since vscore-2.0.0.

    What it costs us: New accounts often see “HOLD” where a competitor would show them a big friendly number on day one.

  • Never 4 of 7

    A sentiment score.

    “Your brand sentiment is 95/100 positive” is a model’s opinion of a model’s prose, unverifiable by the person reading it. We publish framing instead: named states (recommended, listed, hedged, criticised, inaccurate) where every label carries the verbatim sentence that earned it, so you can check us against the evidence.

    Where this is enforced: The answer-framing pipeline requires a verbatim evidence span on every label, and the framing rules’ own measured error cases are published on the detector accuracy page.

    What it costs us: We forfeit the prettiest chart in the category. Sentiment donuts demo brilliantly right up until someone reads the answers.

  • Never 5 of 7

    A percentage without its denominator and its interval.

    “Recommended in 25% of answers” hides everything that matters: 2 of 8 and 200 of 800 both read 25%, and only one is evidence. Every share we publish carries the question count, the sample count, and a 95% interval. On 8 questions, “25%” honestly means “somewhere between roughly 7% and 59%”, and we print that even though it sounds worse.

    Where this is enforced: Wilson intervals are attached at the scorer, stored with every share, and rendered wherever the share is. The interval cannot be dropped without dropping the number.

    What it costs us: Our headline numbers come with error bars that competitors’ identical numbers pretend not to have. Wide bands look weak next to false precision.

  • Never 6 of 7

    “Our fix caused that movement.”

    Engines change their models, their sources and their moods on their own schedule. A number that moved after a change is a correlation; calling it a result requires ruling out everything else that moved, which one brand’s data usually cannot. We report the movement, the window, and what else changed, and we let you draw the causal line.

    Where this is enforced: Measurement runs are versioned by prompt set, and the comparability gate refuses to compare runs whose questions differ, so a change in what we asked can never be reported as a change in what engines think.

    What it costs us: We cannot publish the “+340% AI visibility after 30 days” case study format that sells this category, because we cannot prove it and neither can they.

  • Never 7 of 7

    A delta smaller than the noise it floats in.

    Ask the same engine the same question three times and the answers differ. A one-point move inside that churn is weather, not progress. Score changes smaller than the noise floor are not announced as changes, and prompts too volatile to ever host a verdict are flagged before you spend a window waiting on one.

    Where this is enforced: The score noise floor gates change announcements, and the Detectable Effect Floor computes per-prompt volatility from retained runs before an experiment is allowed to conclude anything.

    What it costs us: Our trend lines are quieter than rivals’. A tool that celebrates every wiggle sends more exciting emails.

What we promise instead

The work. Every number measured, dated, carrying its sample size and interval, or labelled NOT MEASURED with the reason. Every framing label backed by the verbatim sentence that earned it. Every known defect in our own rules published with counts rather than waited out. How each metric is produced is on how we measure, and our instrument's own measured error cases are on detector accuracy.