The practical brief
The finding
The detector flagged 898 of 10,095 usable retrieved webpages as optimized for generative search—an estimated 8.90%—without testing whether optimization increased visibility. [c1] [c7] [c12] [c16]
Why it matters
Businesses evaluating AI-search services need to distinguish detectable edits, discovery outcomes and supporting-source accountability.
| Measure | Gemini | Google Search |
|---|---|---|
| Detector-positive prevalence | 9.09% of usable channel pages | 8.14% of usable channel pages [c7] |
| Equal-query-weighted prevalence | 9.34% mean within-query share across 965 paired queries | 7.91% mean within-query share across 965 paired queries [c7] |
| Embedded citation density | 8.80 citation occurrences per flagged page, including pages without parsed citations | 4.90 citation occurrences per flagged page, including pages without parsed citations [c9] [c10] |
| LOW citation verifiability | 74.15% of embedded citation occurrences on flagged channel pages | 45.88% of embedded citation occurrences on flagged channel pages [c10] [c16] |
What to try
BLURSOR’s practical interpretation
Our interpretation: request measured visibility outcomes and inspect sources behind important claims rather than treating optimization or citation volume as proof of credibility.
Study boundary
Live pages lacked verified optimization labels, and citation-accountability labels do not establish factual accuracy or claim support.
Separate discovery promises from evidence quality
Before paying for AI-search optimization, separate two promises: being discovered and being well supported. This study measures detectable content modifications and source accountability—not recommendations, leads or sales.
Generative engine optimization, or GEO, means modifying content to encourage selection or citation by AI search. A flagged page does not prove publisher intent, deception or successful optimization.
Detection improved, but aggregate scores hid weaknesses
Researchers constructed 3,200 content instances across 400 informational queries in health, finance, technology and travel: 2,000 optimized instances from eight optimizer families and 1,200 human-written, AI-polished or AI-generated controls. Recorded interventions established positive labels.
Ordinary AI-writing cues may be mistaken for optimization signals. Word TF-IDF, a word-frequency detector, achieved F1 of 0.880 but worst-group accuracy of 0.375 and an absolute false-positive-rate gap of 0.558 between authorship groups. F1 combines precision and recall; the other measures expose uneven performance. These observational diagnostics are consistent with shortcut reliance, not proof of internal decision rules.
Training designed to distinguish optimization from ordinary AI polishing improved ModernBERT’s F1 from 0.862 to 0.944 and worst-group accuracy from 0.725 to 0.883. These are detection results, not visibility gains. Sparse and human-executed edits were harder to detect; strong benchmark performance does not guarantee performance on unseen strategies.
Gemini’s retrieved pages were more often flagged
The researchers analyzed released Google Search and Gemini-grounded retrieval links for 1,000 real-user queries. Pages fetched July 28–31, 2026 yielded usable content for 10,095 of 13,985 unique URLs. The detector flagged 898 pages: 8.90%, with a 95% confidence interval of 8.36%–9.47%.
Giving equal weight to 965 queries with usable pages in both channels produced flagged-page rates of 9.34% for Gemini versus 7.91% for Google Search. The difference was 1.43 percentage points, with a paired-bootstrap 95% confidence interval of 0.46–2.41 percentage points. This observational comparison does not show that optimization caused greater AI visibility.
More references did not establish stronger evidence
The audit examined citations inside flagged webpages—not citations generated in AI answers. Only 489 of 898 flagged pages contained parsed citations, yielding 6,663 citation occurrences, including repeated links.
Gemini-returned pages averaged 8.80 citation occurrences per flagged page versus 4.90 for Google Search, including pages without parsed citations. Yet LOW verifiability shares were 74.15% versus 45.88% of citation occurrences. LOW means limited source accountability or inspectability, not falsehood or failure to support a claim.
- Our interpretation: ask vendors whether their results measure edits, retrieved pages or actual answer citations.
- Check whether important references support the specific claim; this audit did not test that relationship.
- Evaluate discovery and evidence quality separately. This is a purchasing safeguard, not a proven growth tactic.
Use the findings as an audit, not an ROI forecast
The findings concern available sampled pages, not the whole web. Live pages lacked verified optimization labels, and detector or citation-auditor errors can affect estimates.
Only 19.57% of usable pages supplied parseable modification metadata. These publisher-declared dates do not identify when optimization occurred, so the recency pattern does not establish increasing adoption. The study cannot forecast traffic, revenue or returns on an optimization contract.
The source and evidence
GEO-Flag: Detecting and Measuring GEO-Optimized Web Content
Study limitations and disclosures
- Live pages lacked verified optimization labels. Estimates describe detector classifications, not proven publisher intent; detector and citation-auditor errors can affect results.
- Aggregate detection scores can mask authorship-group errors consistent with ordinary AI-writing cues being mistaken for optimization. This observational evidence does not establish detectors’ internal decision rules.
- Sparse and human-executed optimization can be harder to detect. The four-domain, eight-family benchmark tests unseen queries, not guaranteed performance on unseen optimization strategies. Historical controls were considered unlikely, not proven, to have undergone intentional optimization.
- Usable content covered 10,095 of 13,985 unique URLs (72.18%), excluding 3,890 unavailable URLs. Findings concern current versions fetched July 28–31, 2026.
- Verifiability measures accountability and accessibility, not truth or claim support. Repeated embedded citations are counted separately; accessibility can change over time.
- Citation-labeling evaluation covered 1,653 of 2,160 citation occurrences in gold-labeled optimized instances that passed detection. Missed instances and three detector-positive non-optimized instances were excluded; accuracy is conditional, not end-to-end.
- Only 19.57% of usable pages had parseable modification dates. These declared dates do not date optimization interventions or establish growing adoption.
Evidence behind this briefing
[c1] GEOFlagBench is a constructed intervention-detection benchmark, not a measurement of search visibility gains: 3,200 instances across 400 queries and four domains, comprising 1,200 non-GEO controls and 2,000 GEO instances from eight optimizer families. AI-polished controls are not GEO seeds.
GEOFlagBench — Overview · Read the study
[c2] Strong aggregate detection scores can conceal authorship-conditioned errors. Word TF-IDF achieved F1 0.880 but worst-group accuracy 0.375 and an absolute false-positive-rate gap of 0.558. These diagnostics are observational evidence consistent with shortcuts, not proof of detectors' internal decision rules.
Evaluation — Human–AI Original Authorship as a Potential Shortcut; Tables 1–2 · Read the study
[c3] On the fixed query-disjoint split of 2,238 training and 962 test instances, using a 0.5 detection threshold, ModernBERT-IPT improved accuracy from 0.839 to 0.931, GEO-positive F1 from 0.862 to 0.944, and worst-group accuracy from 0.725 to 0.883 relative to ModernBERT SFT. These are benchmark detection results, not measured improvements in discoverability.
GEO Detection Evaluation — Experimental Settings and Results; Table 3 · Read the study
[c4] GLM 5.2 achieved 84.75% source-tier accuracy and 83.00% citation-URL-verifiability accuracy in a conditional evaluation of 1,653 citation occurrences from gold GEO instances correctly detected by the gate. These are not end-to-end accuracies over all citations. Verifiability operationalizes publisher accountability/editorial control and accessibility, rather than factual truth.
Audit Scope and Procedure — Phase 2; Evaluation — Citation URL Labeling Performance; Table 5 · Read the study
[c5] The introduction reports an estimated GEO detection prevalence of 898/10,095 usable pages (8.90%), with 8.14% for Google Search and 9.09% for Gemini-grounded search. Among pages with parseable declared dateModified metadata, the reported union rates are 7.02% in 2024, 12.80% in 2025 and 16.36% in 2026. Of 6,663 citation occurrences on detected GEO pages, 69.34% received LOW labels; LOW describes limited accountability or inspectability, not demonstrated falsehood. These estimates concern sampled released retrieval links and subsequently fetched pages, not the entire web or causal GEO effectiveness.
Introduction — Empirical GEO Prevalence Estimation; real-world Experimental Settings · Read the study
[c6] The real-world audit fetched current page versions on 28–31 July 2026 and obtained usable content for only 72.18% of the 13,985 unique released URLs. The unavailable 3,890 URLs are outside the usable-page prevalence denominator. Recovery did not bypass access or policy restrictions.
Empirically Estimating GEO Prevalence in Real-World Search Results — Experimental Settings — Page Collection and Recovery · Read the study
[c7] The pipeline classified 898/10,095 usable unique pages as GEO: 8.90% (95% Wilson CI 8.36%–9.47%). Page-level rates were 8.14% for Google Search and 9.09% for Gemini. Giving equal weight to each of 965 queries with usable pages in both channels produced rates of 7.91% and 9.34%, respectively, a 1.43-percentage-point difference (paired-bootstrap 95% CI 0.46–2.41 pp). These are detector-based observational estimates, not measurements of publisher intent or causal effects.
Estimated GEO Prevalence — Retrieval Channels; Table 6 · Read the study
[c8] Among pages declaring parseable modification dates, estimated GEO prevalence increased with recency: the unique-union rates were 7.02% in 2024, 12.80% in 2025, and 16.36% in 2026. Only 1,976/10,095 pages had parseable metadata, and the 2023–2026 analysis covered 1,684 pages. This describes declared modification dates, not when GEO was applied or evidence of increasing adoption.
Estimated GEO Prevalence — Modification Time; Figure 5; Appendix E, Table 21 · Read the study
[c9] Citation auditing concerns citations embedded in detector-positive webpages, not search-result URLs or citations generated in AI answers. Only 489 of 898 flagged pages contained parsed citations; 409 were excluded. The analysis counts 6,663 citation occurrences, including repetitions, rather than 3,551 globally deduplicated citation URLs.
Estimated Verifiability of Citation URLs — opening analysis scope; Appendix D, Figure 7 and Analysis Units and Deduplication · Read the study
[c10] Of the 6,663 citation occurrences on flagged pages in the unique union, 68.84% were assigned C3 and 69.34% LOW verifiability. Gemini-returned flagged pages had 8.80 occurrences per flagged page versus 4.90 for Google Search; their LOW shares were 74.15% versus 45.88%. These labels principally reflect publisher accountability under the study rubric, not factual accuracy or whether a citation supports a claim.
Estimated Verifiability of Citation URLs — URL Source Tier and Verifiability; Table 8; Appendix C, Tables 16–17 · Read the study
[c11] GEOFlagBench contains 3,200 web-content instances: 2,000 GEO instances across eight optimizer families and 1,200 non-GEO controls. GEO labels derive from recorded interventions, not textual style or detector predictions. The benchmark balances 400 informational queries across Health, Finance, Technology, and Travel; 164 were retained from GEO-Bench and 236 generated using Sonnet 4.6. Human GEO construction used two people, and historical non-GEO seeds were selected as unlikely—not proven—to have undergone intentional GEO.
Appendix A — Table 9; Domain Selection and Query Collection; Non-GEO Data Construction; Dataset Summary · Read the study
[c12] Real-world pages lack gold GEO labels, so detector and citation-audit errors can affect estimates. Modification metadata is sparse and does not date GEO interventions. Citation URL verifiability measures accountability and accessibility, not correctness or evidential support, and accessibility can change over time.
Discussion and Limitations — Limits of the Real-World Estimates · Read the study
[c13] For the example query “picot examples nursing,” Google Search has 2/5 usable pages classified as GEO (40%), compared with Gemini’s 4/5 (80%). This is a single paired-query example, not an aggregate result; positions denote released URL-list order rather than conventional Gemini rank.
Example of a Paired-Query Comparison, Figure 8 · Read the study
[c14] The paired-query example calculates GEO rates only over successfully retrieved usable pages.
Figure 8 caption · Read the study
[c15] Live-page GEO classifications are empirical estimates: actual GEO intervention cannot be established with certainty, ground-truth GEO labels are absent, and both detection and citation-audit pipelines can err.
Footnote 5; related disclosure in footnote 1 · Read the study
[c16] Citation URL verifiability measures inspectability and source accountability, not factual accuracy or whether the cited source supports the claim.
Footnote 3; reiterated in footnote 6 · Read the study
[c17] The conditional citation URL audit covers 106 of 129 gold-labeled GEO instances containing citations (82.17%) and 1,653 of their 2,160 citation URL occurrences (76.53%). Three non-GEO instances that passed the GEO Flag Tool are excluded.
Footnote 4 · Read the study
[c18] An additional Gradient Reversal Layer control suppresses original-authorship information in the learned representation; this chunk points to Appendix B.1 but supplies no control results.
Footnote 2 · Read the study
[c19] Strong aggregate benchmark performance conceals detection weaknesses for sparse and human-executed GEO interventions. The authors suggest that sparse changes can dilute the detectable signal and that poor Human GEO recall may reflect reliance on human–AI authorship cues; these explanations are possibilities, not established internal detector mechanisms.
GEO Flagging with Existing Detectors — Evaluation — Method-Level Results · Read the study
[c20] The benchmark split evaluates generalization to unseen queries, not explicitly to held-out optimization strategies: all instances associated with a query stay together, with 70 training and 30 testing queries per domain, producing 2,238 training and 962 test instances. Strong performance on this evaluation should not be presented as a guarantee of performance on unseen optimization strategies.
GEO Flagging with Existing Detectors — Experiment Settings — Train and Test Split · Read the study
[c21] Sparse and human-executed GEO interventions can be harder to detect. Although the benchmark covers eight optimizer families, strong performance on it does not guarantee similar performance on unseen GEO strategies, including future methods using smaller changes or detection-avoidance strategies.
Discussion and Limitations — Generalization to Future GEO Strategies · Read the study