Instrument accuracy
The measured noise floor
Within one measurement run we ask each question of each engine several times, minutes apart. The site cannot have changed between those samples, so when the samples disagree about whether you were cited, that disagreement is our instrument wobbling, not your visibility moving. Most tools have this number and none publish it. Here is ours, recomputed from every stored run.
anthropic
129 question-runs judged6 of 129 repeated question-runs disagreed with themselves (5%, 95% interval 2% to 10%).
Plain arithmetic on those counts: two back-to-back identical runs over these questions could differ by up to 5 answer-share points with nothing having changed. That is the movement we must refuse to call a change.
17 of 203 repeated question-runs disagreed with themselves (8%, 95% interval 5% to 13%).
Plain arithmetic on those counts: two back-to-back identical runs over these questions could differ by up to 8 answer-share points with nothing having changed. That is the movement we must refuse to call a change.
openai
134 question-runs judged8 of 134 repeated question-runs disagreed with themselves (6%, 95% interval 3% to 11%).
Plain arithmetic on those counts: two back-to-back identical runs over these questions could differ by up to 6 answer-share points with nothing having changed. That is the movement we must refuse to call a change.
perplexity
176 question-runs judged5 of 176 repeated question-runs disagreed with themselves (3%, 95% interval 1% to 6%).
Plain arithmetic on those counts: two back-to-back identical runs over these questions could differ by up to 3 answer-share points with nothing having changed. That is the movement we must refuse to call a change.
What this does to the score-change floor
We refuse to report a Visibility Score change smaller than 3 points, a constant our methodology corpus marks as provisional. The figures above are the measured evidence that floor answers to: when an engine's within-run wobble is larger than a reported change, the change is not a finding. As runs accrue, the constant will be held against these measured figures on this page, in public, and corrected in the direction the measurement says.
Method: a question-run is all same-batch samples of one question against one engine, counted from stored runs. A disagreement is a run where some samples produced a citation of the measured site and others did not. Aggregated across all measured sites; no question text or site identity leaves the computation.
