The practical brief

The finding

A stopping rule combining ranking stability with measurement sufficiency converged in 27 of 30 tested platform-topic combinations. [c1] [c4] [c7] [c10] [c13] [c14] [c20]

Why it matters

A steady leaderboard does not necessarily make a competitor’s lead meaningful enough to guide spending.

Stable does not mean precisely ranked
Reporting scopeMedian 95% rank-interval width
Top ten domains per combination5.0 rank positions across 300 pooled domain entries [c8]
Established domains: share interval excludes zero69.1 rank positions across 1,693 pooled domain entries [c8]
All observed domains131.0 rank positions across 7,067 pooled domain entries [c8]
Measured at the full collection window from 125 submitted queries per platform-topic combination, with varying citation-bearing response counts. Entries are pooled across 30 combinations, not necessarily unique domains. Study source

What to try

BLURSOR’s practical interpretation

Our interpretation: ask for uncertainty checks, the chosen sufficiency threshold and separate assessments of the segments you use for decisions.

Study boundary

Validation was retrospective; the study measures domain citations, not marketing returns, and assumes citation distributions remain unchanged during collection.

A settled leaderboard may not justify shifting resources

Before moving content resources because a competitor ranks above you in AI citations, ask whether the apparent lead exceeds measurement uncertainty. This study separates two questions: has the ordering stopped changing substantially, and are citation-share differences sufficiently resolved for comparative reporting?

Even near the top, those questions differ. At the full collection window, the median 95% uncertainty interval among top-ten domain entries spanned five rank positions. Stable ordering did not establish that every neighboring competitor could be distinguished.

Evidence [c4] [c8]

The study tested when citation measurement could stop

Ronald Sielinski used 125 ChatGPT-generated queries per topic across ten topics, submitted to Gemini, SearchGPT and Perplexity. Cited URLs were grouped by registered domain. The outcome was citation share—not answer accuracy, brand recommendation or purchases.

The stopping rule requires both a plateau in ranking agreement and “structural sufficiency”: the spread of citation shares among established domains must be large relative to their average uncertainty. Established domains have confidence intervals—ranges expressing estimation uncertainty—that exclude zero.

Uncertainty was estimated by repeatedly resampling whole queries, preserving citations clustered within responses. Structural signal-to-noise ratio, or SNR, compares share variation with uncertainty. Its default threshold of 1.0 is an operating choice, not a uniquely validated optimum or proof that adjacent ranks differ.

Evidence [c1] [c2] [c3] [c4] [c17]

Collection requirements varied, and certification remained imperfect

Both criteria were met in 27 of 30 platform-topic combinations. Median stopping orders were 42 citation-bearing responses for Gemini, 37 for SearchGPT and 51 for Perplexity. These figures cover only converging combinations; SearchGPT’s median covers seven of its ten topics.

They are not universal query budgets. Responses without citations do not count toward stopping orders, so more submitted queries may be needed.

Stopped rankings had a median top-ten overlap of 80% with full-window rankings across the 27 converged combinations. This is internal consistency, not external accuracy: validation reused observations and full-sample information unavailable at stopping.

In retrospective threshold sensitivity analysis, stricter thresholds generally required more collection and improved ranking agreement among certified combinations, but could leave more combinations uncertified within the collection window. Better agreement therefore comes with a collection and coverage tradeoff.

Evidence [c7] [c9] [c10] [c12] [c18]

Ask what makes your provider’s report decision-ready

Our interpretation is to use this research to scrutinize reporting, not as a proven visibility-improvement tactic. A fixed query package should not automatically be described as sufficient everywhere.

  • Request stability and uncertainty checks, plus citation-bearing response counts alongside submitted-query counts.
  • Ask which sufficiency threshold is used and why it fits your precision needs.
  • Require separate checks for geography, query-type or other filtered comparisons; whole-set convergence does not certify smaller subgroups.
  • For combined-platform reports, ask for within-platform share normalization and separately estimated uncertainty before averaging. Raw-count pooling overweights citation-dense platforms; every constituent platform should converge.
  • Check whether sampled questions represent your customers’ needs. Convergence cannot establish that.

Evidence [c11] [c12] [c14] [c15] [c20]

What the study cannot tell your business

Results cover ten topics, three platforms and one collection period. They do not establish transferable stopping counts or show that improved citation rank increases sales. Sparse, clustered citations can also compromise uncertainty estimates.

The framework assumes citation distributions remain unchanged during collection. Platform updates or genuine drift can look like sampling instability; it cannot distinguish those causes or assess changes between collection periods.

Evidence [c1] [c6] [c11] [c13]

The source and evidence

From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement

Ronald Sielinski · 2026-07-11

Study limitations and disclosures

  • Ronald Sielinski lists an affiliation with IQRush. The supplied source evidence does not establish funding details or a commercial-interest declaration; independence should not be assumed.
  • Scope is limited to generated queries, ten topics, three platforms and one collection period. Customer-query representativeness, marketing returns and transferable stopping counts are not established.
  • Sufficiency is a set-level heuristic, not proof that adjacent ranks differ. Sparse, clustered citations can compromise uncertainty estimates.
  • The default SNR threshold of 1.0 is not a uniquely validated optimum. Retrospectively, stricter thresholds generally increased collection and ranking agreement among certified combinations, but could leave more uncertified.
  • Validation shares observations with the full-window comparison and uses full-sample domain membership unavailable at stopping; independent validation remains future work.
  • The framework cannot separate distribution drift from sampling variability. Filtered subgroups need their own checks; combined-platform reporting requires normalization, separate uncertainty estimates and convergence of each platform.

Evidence behind this briefing

[c1] The dataset uses 125 ChatGPT-generated queries per topic across ten topics and three platforms. Queries were conditioned on randomly selected types from twenty categories; analyses concern cited registered domains, not answer accuracy, brand preference, or causal visibility improvements.

Section 3.1, Data · Read the study

[c2] Uncertainty is estimated by resampling whole queries to preserve within-response citation clustering, using 1,000 bootstrap replicates. Share intervals use BCa correction, but derived rank intervals and trajectory bands use bias-corrected percentile intervals; intervals are conditional on citation-bearing responses.

Section 3.2, Bootstrap Resampling and Citation Share Estimation · Read the study

[c3] Rank stability is detected through model selection for a plateau in weighted Spearman correlation, rather than crossing a fixed correlation target. Certification requires a flat tail of at least 15 responses, latches once declared, and cannot occur before response 30.

Section 3.5, Rank Stability During Collection · Read the study

[c4] Structural sufficiency compares the standard deviation of shares among established domains with their mean 95% CI half-width. SNR ≥ 1 is explicitly a practical, set-level heuristic and operational default—not a canonical statistical threshold or proof that every adjacent rank is distinguishable.

Section 3.7, Structural Signal-to-Noise Ratio · Read the study

[c5] Across the tested topics, citation density ranged from 19.9–50.1 citations per citation-bearing response for Gemini, 14.5–42.2 for Perplexity, and 4.2–6.3 for SearchGPT. SearchGPT convergence counts exclude zero-citation responses and therefore must not be interpreted as submitted-query budgets.

Section 4.1, Dataset and Distributional Structure; Table 1 and accompanying SearchGPT yield discussion · Read the study

[c6] Bootstrap share intervals, and ranks derived from them, can misrepresent uncertainty for sparse, clustered citations. BCa only partially addresses this: the paper flags roughly 5–20 citations as a danger zone and reports that SearchGPT did not reach distributional symmetry within the observed range.

Section 3.3, Reliability of Bootstrap Confidence Intervals · Read the study

[c7] The stopping rule requires both a BIC-iso rank-correlation plateau and structural SNR ≥ 1. It converged for 27 of 30 platform-topic combinations within 125 responses. Platform median stopping orders are conditional on convergence, particularly SearchGPT's seven converging combinations; they are not universal query budgets.

§4.5 Conjunctive Convergence Results, opening paragraph; Table 6 · Read the study

[c8] Stable ordering does not resolve every rank. At 125 responses, pooled 95% rank-CI widths have a median of 5.0 positions among the 300 top-ten domain entries, versus 69.1 among 1,693 established-domain entries and 131.0 among 7,067 full-vocabulary entries. Established means a citation-share CI excluding zero; these are pooled combination-level domain entries, not necessarily unique domains across platforms and topics.

§4.3 Rank Uncertainty at Full Sample Size, Table 5 · Read the study

[c9] In the retrospective validation, certified stopping orders across 27 converged combinations had weighted Spearman agreement of 0.717–0.986 with the full-window ordering, median 0.844, and median top-ten overlap of 80%. Agreement improved in all 11 cases where structural sufficiency delayed stopping. These are ranking-agreement results, not accuracy against external truth or causal evidence of marketing effectiveness.

§4.6 Validation: Do Stopped Orderings Persist?, results paragraph · Read the study

[c10] Validation reuses the pre-stop observations and fixes established-set membership using the full sample. It is an internal, retrospective consistency audit, not independent validation or a metric available at the stopping point. Independent repeated collections or held-out responses are left to future work.

§4.6 Validation: Do Stopped Orderings Persist?, opening paragraph · Read the study

[c11] The study's empirical generalizability is limited to ten topics, three platforms and one collection period. Convergence applies to the sampled query population; it does not establish that the query design represents a brand's actual questions. The authors' expectation that qualitative findings generalize is not an empirical test of that generalization.

§5.1 Limitations, first two paragraphs · Read the study

[c12] Reported stopping orders count citation-bearing responses, not all submitted queries. SearchGPT's topic-dependent zero-citation yield means the operational query budget must be larger than the reported response stopping order.

§4.5 Conjunctive Convergence Results, final paragraph · Read the study

[c13] The framework assumes citation distributions remain stationary during collection. Platform updates or genuine drift can appear as continued sampling instability or failure to converge: the criterion cannot distinguish drift from sampling variability and does not investigate changes between collection periods.

§5.1 Limitations, paragraph beginning “The framework assumes the citation distribution is stationary” · Read the study

[c14] Convergence for a complete platform-topic query set does not establish sufficiency for a filtered geography, query type, or other facet. Filtering reduces effective response counts, widens bootstrap confidence intervals, and may require more responses for rank stability. Subgroup comparisons require their own convergence checks and separately reported sufficiency ratios.

§5.1 Limitations, paragraph beginning “The analysis treats all queries within a platform-topic combination” · Read the study

[c15] Pooling raw citations disproportionately weights citation-dense platforms, and bootstrapping pooled counts repeats that imbalance in uncertainty estimates. Combined reporting should normalize shares within each platform, compute confidence intervals and related uncertainty measures separately before averaging, and be treated as ready only after every constituent platform has individually converged.

§5.1 Limitations, paragraph beginning “Similarly, the analysis treats each platform-topic combination” · Read the study

[c16] The structural SNR threshold of 1.0 is an operational default, not a statistically canonical or uniquely validated optimum. Readers should ask which sufficiency threshold their measurement provider uses. This chunk does not establish the comparative collection and ranking-agreement consequences of stricter thresholds.

§3.7 Structural Signal-to-Noise Ratio · Read the study

[c17] The structural-SNR threshold of 1.0 is an operational default, not a uniquely validated optimum. With the rank-stability test fixed, the study swept thresholds from 0.5 to 1.5 across 30 platform-topic combinations.

§4.7 Sensitivity to the Sufficiency Threshold, opening paragraph · Read the study

[c18] In this retrospective sensitivity comparison, stricter thresholds generally increased collection requirements and weighted Spearman agreement with the full-window ranking among certified combinations, while potentially leaving more combinations uncertified within the 125-response window. Median agreement rose from approximately 0.84 at T = 1.0 to 0.89 at T = 1.2, with median collection increasing from 43 to 53 responses; 27 combinations converged at T = 1.3, 25 at 1.35, and 23 at 1.4. These are conditional ranking-agreement results, not external accuracy or causal evidence.

§4.7 Sensitivity to the Sufficiency Threshold; Figure 11 · Read the study

[c19] Ranking agreement is a retrospective internal consistency audit against the full-window estimate: stopped and final rankings share data, and established-set membership is fixed using full-sample information unavailable at stopping. It is not validation against independent ground truth.

§4.6 Validation: Do Stopped Orderings Persist?, opening paragraph · Read the study

[c20] Threshold selection is a material operating choice balancing collection cost and ranking fidelity. Readers should ask which sufficiency threshold their provider uses and why it fits their precision needs; the study does not establish 1.0 as the sole optimum.

§4.7 Sensitivity to the Sufficiency Threshold, closing sentences · Read the study