The practical brief
The finding
Projected reliability reached 0.81 at fifteen repetitions in the pooled brand-count design, exceeding the paper’s group-level target of 0.80 but not its individual-prompt target of 0.90. [c3] [c8] [c9] [c11]
Why it matters
A visibility report can confuse answer-to-answer variation with a meaningful change, especially when judging a particular brand-query result.
| Repetitions per model-by-prompt cell | Projected reliability in pooled S1+S5 design |
|---|---|
| 5 repetitions | G = 0.58 (unitless; pooled S1+S5 brand-count design) [c3] |
| 10 repetitions | G = 0.74 (unitless; pooled S1+S5 brand-count design) [c3] |
| 15 repetitions | G = 0.81 (unitless; pooled S1+S5 brand-count design) [c3] |
| 20 repetitions | G = 0.85 (unitless; pooled S1+S5 brand-count design) [c3] |
What to try
BLURSOR’s practical interpretation
Our interpretation: ask providers to justify repetition counts with a pilot and a decision-specific reliability target, and distinguish visibility tracking from tests of improvement.
Study boundary
The targets and repetition estimates are source-specific, not universal business standards. Independent validation tested the statistical machinery outside the brand domain.
Check the measurement before redirecting marketing spend
Before using an AI visibility report to redirect marketing spend, check whether it captures a repeatable pattern or a handful of variable answers. This paper examines that measurement risk—not a method for making AI recommend your business.
The study reanalyses five existing brand-recommendation datasets rather than collecting new API responses. Shared conditions included temperature 0.3, a setting controlling sampling randomness; a 1,024-token response limit; and dictionary-based brand-name extraction. Results therefore do not describe every live AI search or assistant experience.
The central question is how often an auditor should repeat the same question before treating brand counts as dependable.
Fifteen repetitions cleared only the group-level target
For pooled data from two source studies, the reliability calculation used 450 iteration-level observations across 75 model-by-prompt combinations. Projected reliability was 0.58 at five repetitions, 0.74 at ten, 0.81 at fifteen and 0.85 at twenty.
The coefficient, G, measures how consistently the design distinguishes those combinations despite answer-level noise. It does not mean that 81% of recommendations were accurate or that fifteen answers prove a marketing intervention worked.
Crucially, the paper distinguishes a G=0.80 group-level decision target from a G=0.90 individual-prompt decision target. Fifteen repetitions exceeded the former; even twenty did not reach the latter. At the study’s observed variance-component ratio, reaching 0.90 would require more than forty repetitions. These are source-specific targets and estimates, not universal business standards.
Detecting a change is also separate from reliability. Statistical power—the chance of detecting an effect under the simulated design—depended on model and effect size. In the Valentine’s Day gender-contrast analysis, the observed Grok effect needed about twenty repetitions for 80% power; the smaller GPT effect needed about forty.
Set the audit budget from a pilot and decision
Independent validation on political-orientation and benchmark-accuracy data reproduced reliability predictions in 37 of 39 prediction cells, with two partial replications and no failures. That supports the calculation method, not fixed brand-audit thresholds.
The paper’s practical prescription is “pilot, then solve”: begin with ten repetitions, estimate variation between model-by-prompt combinations and noise within them, then calculate the repetition count for the chosen target. Fifteen was the solution for its group-level target on its own data—not a guarantee for an individual-prompt decision.
Our interpretation is to specify the intended decision before buying an audit. Describing a portfolio of queries, judging one brand-query result and testing a marketing change require different justifications.
- Request exact prompts, model versions, sampling settings and repetition counts.
- Ask which reliability target matches your decision and how the pilot supports it.
- For claims of improvement, ask for a separate power analysis rather than treating repeatability as proof.
Stable counts do not establish a visibility strategy
This study cannot identify which content, citations or marketing investments increase recommendations. It evaluates measurement reliability, not interventions or commercial outcomes. Repeated agreement also does not establish factual correctness.
Long-term comparability remains qualified. The original no-drift finding covered one dataset’s roughly two-week collection window; drift in several other datasets remained untested. Independent data exposed weaknesses in small-sample drift checks, prompting amended guidance.
The realistic takeaway is narrower than “run fifteen queries”: require an audit design that matches the decision, while keeping measurement confidence separate from evidence that marketing worked.
The source and evidence
Study limitations and disclosures
- All five brand datasets came from the author’s prior studies. Independent validation used non-brand corpora, so brand-specific effects and repetition thresholds still lack independent replication.
- The paper’s G=0.80 group-level and G=0.90 individual-prompt targets are not universal business standards. At its observed variance ratio, fifteen repetitions projected G=0.81, twenty projected G=0.85, and reaching 0.90 would require more than forty repetitions.
- The pooled power analysis is an upper bound for single-model, single-prompt audits. A calibrated simulation for that narrower design remains future work; reliability alone does not establish adequate power to detect a change.
- Shared sampling settings and model versions limit transfer to other systems. The original no-drift finding covered roughly two weeks in one dataset, not long-term stability across all datasets.
- Author Dmitrij Żatuchin is CEO of Rankfor.AI, which commercially implements the protocol. The paper reports no external funding and states that the company did not influence the research. Original datasets are available on reasonable request.
Evidence behind this briefing
[c1] The study reanalyses existing brand-recommendation data rather than collecting new API responses. Its shared test conditions include temperature 0.3, a 1,024-token response limit, specified model families, and dictionary-based brand extraction; results should not be treated as estimates for arbitrary live AI-answer systems.
Sections 4.1 and 4.3, Shared Experimental Parameters · Read the study
[c2] All five reanalysed studies come from the same research group, and the authors disclose a commercial affiliation with Rankfor.AI, which implements the protocol. The introduction restricts the evidence to brand-recommendation auditing and describes independent external validation as a next step.
Section 1, final two paragraphs · Read the study
[c3] For pooled S1+S5 log-transformed brand counts, 450 iteration-level observations across 75 model-by-prompt cells produced estimated between-cell variance of 0.202 and residual variance of 0.724. The resulting D-study projects reliability G=0.58 at five iterations, 0.74 at ten, and 0.81 at fifteen. These are design-specific reliability projections, not recommendation accuracy or proof of sufficient power for every effect.
Section 5.5, Tables 7–8 · Read the study
[c4] The main power table pools eight prompts and three models, so its early high power is explicitly an upper bound rather than a single-prompt audit guarantee. The paper leaves calibrated single-model, single-prompt simulation as future work and uses the D-study as the operational constraint.
Section 5.1, paragraph following Table 2 · Read the study
[c5] Observed S1 Valentine’s Day gender-contrast simulations show that required repetitions depend on model and effect magnitude: Grok’s δ=0.26 reaches approximately 80% power at twenty iterations, while GPT’s δ=0.18 needs approximately forty. Ten iterations support descriptive effect reporting but are underpowered for these confirmatory comparisons.
Section 5.2, Observed-Effect Validation, paragraph following Table 4 · Read the study
[c6] In the small S1 Valentine’s subset—15 prompt-model combinations with ten iterations each—the count-based stability metrics are largely redundant. Their descriptive correlations do not establish the empirical complementarity of the full metric battery; semantic consistency and fairness are outside this matrix.
Section 5.4, Table 6 and concluding paragraph · Read the study
[c7] On S5 category-ownership data, prompt-adjusted brand visibility was more concentrated than raw brand counts: PASOR-Gini ranged from 0.621 to 0.691 versus unadjusted Gini from 0.198 to 0.312. Table 13 describes 3,750 responses across 250 query-iteration cells per model and a shared 1,267-brand universe extracted using Title Case regex with a corpus-wide frequency floor of at least 20. These are descriptive concentration measures, not accuracy or causal marketing effects; no uncertainty intervals are reported.
Section 5.8, Table 13 and following paragraph · Read the study
[c8] External validation supports choosing iteration counts from pilot-estimated variance components rather than treating 15 iterations as a universal constant. The prescribed pilot uses 10 iterations and solves for G(n) at least 0.80; n=15 is the solution on this paper’s own data. The local recommendations concern brand-recommendation auditing at temperature 0.3 with three models, and pooled-prompt power is an upper bound for single-cell audits.
Section 6.2, EV1b discussion; Section 7.1 · Read the study
[c9] In preregistered external validation on independent political-orientation and benchmark-accuracy corpora, D-study reliability predictions replicated in 37 of 39 prediction cells, with two partial replications and no failures. Replication required agreement within 0.05 with empirical reliability calculated using 200 disjoint-subset splits. This validates statistical machinery outside the brand domain, not brand-specific visibility effects or fixed iteration thresholds.
Section 6.2, Table 16 and results paragraph; Section 6.1 · Read the study
[c10] External data exposed defects in the original drift battery: within-cell regression sandwich variance can degenerate with a single cluster, and PSI is unreliable with five observations per half. The amended guidance gates PSI at 20 iterations per half, replaces the within-cell regression test or uses proper cross-cell clustering, and requires practical significance. S1’s no-drift result covers only its approximately two-week collection window; S2, S4 and S5 drift remains untested.
Section 6.3, Operating bounds for the drift battery; Section 7.6, Provider-side drift · Read the study
[c11] Independent validation remains out of domain: brand-specific effect sizes and thresholds still rely on the research group’s own datasets. The authors also disclose that their protocol is implemented commercially by an affiliated company.
Section 7.6, Evidential circularity, partially resolved · Read the study
[c12] The author is CEO of Rankfor.AI, which commercially implements the protocol, and all five reanalysed datasets are from his prior studies. The declaration states that the company did not influence the research. The study reports no external funding; it is a secondary analysis of previously collected data, and original datasets are available on reasonable request rather than directly supplied in this text.
Declarations, Competing interests; Funding; Dual publication; Data Availability · Read the study
[c13] The footnote qualifies H1 under the revised Cliff’s delta formulation: n=5 fails to attain 0.80 power even for a large effect (delta=0.474), and the 80% power threshold is first met at n=15. It confirms insufficient power at low iteration counts.
Footnotes, aH1 · Read the study