Instrument accuracy
The measured noise floor
Within one measurement run we ask each question of each engine several times, minutes apart. The site cannot have changed between those samples, so when the samples disagree about whether you were cited, that disagreement is our instrument wobbling, not your visibility moving. Most tools have this number and none publish it. Here is ours, recomputed from every stored run.
anthropic
99 question-runs judged6 of 99 repeated question-runs disagreed with themselves (6%, 95% interval 3% to 13%).
Plain arithmetic on those counts: two back-to-back identical runs over these questions could differ by up to 6 answer-share points with nothing having changed. That is the movement we must refuse to call a change.
14 of 173 repeated question-runs disagreed with themselves (8%, 95% interval 5% to 13%).
Plain arithmetic on those counts: two back-to-back identical runs over these questions could differ by up to 8 answer-share points with nothing having changed. That is the movement we must refuse to call a change.
openai
104 question-runs judged8 of 104 repeated question-runs disagreed with themselves (8%, 95% interval 4% to 14%).
Plain arithmetic on those counts: two back-to-back identical runs over these questions could differ by up to 8 answer-share points with nothing having changed. That is the movement we must refuse to call a change.
perplexity
144 question-runs judged3 of 144 repeated question-runs disagreed with themselves (2%, 95% interval 1% to 6%).
Plain arithmetic on those counts: two back-to-back identical runs over these questions could differ by up to 2 answer-share points with nothing having changed. That is the movement we must refuse to call a change.
What this does to the score-change floor
We refuse to report a Visibility Score change smaller than 3 points, a constant our methodology corpus marks as provisional. The figures above are the measured evidence that floor answers to: when an engine's within-run wobble is larger than a reported change, the change is not a finding. As runs accrue, the constant will be held against these measured figures on this page, in public, and corrected in the direction the measurement says.
Method: a question-run is all same-batch samples of one question against one engine, counted from stored runs. A disagreement is a run where some samples produced a citation of the measured site and others did not. Aggregated across all measured sites; no question text or site identity leaves the computation.
