The practical brief
The finding
Enabling search raised recommended-doctor registry matches from 10.8% to 63.9% and nursing-home matches from 32.7% to 71.2% in the controlled comparison. [c1] [c2] [c9] [c10]
Why it matters
An AI referral list can depend substantially on search configuration, not just the underlying model.
| Measure | Search disabled | Search enabled |
|---|---|---|
| Doctor registry matches | 10.8% of recommended doctor names matched a queried-city registry entry | 63.9% of recommended doctor names matched a queried-city registry entry [c2] |
| Nursing-home registry matches | 32.7% of recommended nursing-home names matched a queried-city registry entry | 71.2% of recommended nursing-home names matched a queried-city registry entry [c2] |
| Adviser disclosure exposure | 18.3% of matched adviser recommendation occurrences carried SEC Item 11 disclosures | 1.7% of matched adviser recommendation occurrences carried SEC Item 11 disclosures [c4] |
What to try
BLURSOR’s practical interpretation
Our interpretation: monitor searched and non-searched answers separately, and verify consequential referrals against authoritative records.
Study boundary
Registry matching is not a definitive fabrication test, and this audit does not establish a marketing tactic or deployed visibility effect.
Establish whether the assistant searched before judging referrals
If you use AI answers to assess whether customers can discover your business, first distinguish answers generated with web search from those generated without it. Combining them into one visibility score could hide substantially different referral behavior.
Researchers audited doctors, nursing homes, hospitals and investment advisers across the 100 largest U.S. metropolitan areas, matching recommended names against government registries. Their controlled comparison used one proprietary model with identical prompts and decoding settings, search disabled or enabled: 2,010 prompts per arm, queried on August 21, 2026.
A separate open-weight model used a larger prompt grid, making that comparison less controlled. Native search may also change orchestration—the system’s sequence of operations—not merely document access.
Search changed matching, geography and registry assessments
Doctor matches rose from 10.8% without search to 63.9% with search; nursing-home matches rose from 32.7% to 71.2%. Hospital and adviser gains were smaller. Statistical checks grouping observations by metro retained positive search-associated gains in all four domains.
Without search, matching generally favored larger metros; searched results showed no clear metro-size gradient. However, uncertainty matters: the open-weight doctor and hospital gradients and proprietary no-search adviser gradient remained supported under metro-grouped analysis, while proprietary no-search doctor and hospital contrasts included zero.
Among matched nursing-home recommendations, mean CMS ratings rose from 3.40 to 4.44 stars with search, versus a 2.88-star metro-roster mean. These are contested registry assessments, not independently measured service quality. Hospitals showed no significant quality selection in the proprietary comparisons.
SEC Item 11 disclosures appeared in 18.3% of matched adviser recommendation occurrences without search and 1.7% with search, versus a 5.1% registry base rate. Counting each firm once reduced these rates to 9.0% and 2.4%, showing how repeated referrals shaped exposure. Disclosures record past events, not current conduct.
Prominence is not proof of quality or a marketing recipe
In a separate restaurant analysis across 13 metros, searched recommendations had a median Google Maps review count 2.9 times the inspection-roster comparison median, but a mean rating premium of only 0.11 stars. Review counts proxy visibility; the association does not prove that accumulating reviews causes recommendations.
Government sites accounted for 35.8% of nursing-home citations but 0.3% of doctor citations. Citation hosts describe reported sources, not necessarily the evidence supporting each referral. Likewise, the geographic matching pattern does not demonstrate differences in deployed business visibility.
- Our interpretation: separate searched and non-searched answers when monitoring referrals.
- Verify identity, location and relevant registry information before relying on consequential recommendations.
- Do not treat these findings as proof that review-building or website edits will earn referrals.
What this audit cannot tell a business owner
The primary matcher required a principal-city registry address, so suburban providers could fail alongside renamed, ambiguous or out-of-registry providers. Unmatched does not necessarily mean invented. Doctor results also depended on a tuned instruction requesting individual clinician names.
The audit used English prompts and U.S. registry infrastructure unavailable in many countries. It cannot establish performance in other languages, countries, assistants or future systems. Nor does it show which business changes increase referrals, whether recommendations convert into customers, or whether recommended providers deliver better future service.
The source and evidence
Understanding AI Provider Recommendations in Local Service Markets
Study limitations and disclosures
- The controlled comparison covers one proprietary model, API and time point. Open-weight comparisons cross prompt grids; native search may change orchestration beyond document access.
- Principal-city matching excludes suburban providers; naming ambiguity, renaming and registry scope can also cause failures. Unmatched recommendations are not necessarily fabricated.
- Doctor outcomes depend on a tuned formatting instruction. Extraction was validated on 150 human-labeled responses, with accuracy dependent on normalization conventions.
- CMS star ratings are contested measures, not independent service-quality tests. SEC disclosures aggregate past events of differing severity. The authors suppress small cells and withhold provider-level disclosure/rating pairings to reduce reputational harm.
- Restaurant reviews and ratings are imperfect proxies, not causal evidence. Citation hosts do not establish which evidence supported each referral.
- English prompts and U.S. registry infrastructure limit transferability: other languages and countries require further validation.
Evidence behind this briefing
[c1] The controlled search comparison uses gpt-5.6-terra with identical prompts and decoding, search disabled versus enabled, queried August 21, 2026. The core comparison contains 2,010 prompts per arm; restaurant extensions increase this to 2,065. The open-weight gpt-oss-120b arm uses a larger prompt grid, so comparisons involving it also cross prompt sets. Native search may change orchestration beyond document access.
Section 3.1, Experimental conditions · Read the study
[c2] Among recommended names, enabling search raised queried-city registry matches for doctors from 10.8% to 63.9% and nursing homes from 32.7% to 71.2%; hospital and adviser gains were smaller. These are match rates, not definitive fabrication rates: suburban, renamed, out-of-registry and ambiguous providers can fail matching. The open-weight doctor matches failed the primary-care specialty test and were interpreted as chance name collisions. Web-coverage density was not measured.
Section 5.1, Whether recommended providers exist depends on search · Read the study
[c3] Among matched nursing-home recommendations, mean CMS ratings rose from 3.40 stars without search to 4.44 with search, compared with a 2.88-star metro-roster mean. Both comparisons were BH-significant in every proprietary-arm paraphrase cell. Hospitals showed no significant quality selection in proprietary-arm cells; doctor MIPS averages were reported descriptively because scoring covers a selected clinician subset.
Section 5.3, Quality among the matched providers · Read the study
[c4] Matched proprietary-model adviser recommendations had SEC Item 11 disclosures at 18.3% without search versus a 5.1% registry base rate, falling to 1.7% with search. Search also selected below-base disclosure rates within the >100-employee band. Counting each firm once attenuated the no-search result to 9.0%, versus 2.4% with search, indicating repeated exposure to disclosed firms. Disclosure flags aggregate events of unequal severity, not a direct measurement of future service quality.
Section 5.4, Search reverses the misconduct tilt among advisory firms · Read the study
[c5] In the separate restaurant analysis, recommended establishments had median Google Maps review counts 5.3, 3.6 and 2.9 times the inspection-roster comparison median for arms A, B and C, respectively, but mean-rating premiums of only 0.07, 0.10 and 0.11 stars. The comparison used 255, 198 and 202 resolved recommended establishments and 216 resolved inspection-roster establishments across 13 metros. Review counts proxy visibility and ratings proxy perceived quality; neither is ground truth, and these observational differences do not establish that accumulating reviews causes recommendations.
Section 5.6, Popularity versus quality among restaurants · Read the study
[c6] Host classification of citations in 2,065 searched responses found government-source shares of 35.8% for nursing homes but 0.3% for doctors; only 1% of doctor responses cited any government source. Source mix and quality selection were associated, not a traced causal mechanism. Commercial registry mirrors also appeared, and the observed source mix may change with optimization for AI citations.
Section 5.5, What sources back the answers · Read the study
[c7] Doctor recommendations are conditional on a formatting instruction that increased individual-clinician naming, while more explicit individual-only instructions increased refusals. These are prompt-dependent response outcomes, not measures of clinician quality.
Appendix B, Doctor-prompt tuning · Read the study
[c8] Extraction validation used 150 human-labeled responses stratified by domain. Pooled precision/recall were 0.979/0.984 under matcher-level normalization but 0.82/0.79 for strict surface strings; extraction failures were counted as naming nothing.
Appendix C, Extraction validation · Read the study
[c9] A 100-metro clustered bootstrap with 2,000 replicates and percentile 95% intervals preserved positive B-to-C match-rate differences in all four audited domains, but clustering weakened two individual metro-gradient contrasts. These are registry-match results, not recommendation-quality scores.
Appendix G, Metro-clustered bootstrap · Read the study
[c10] The primary matcher treats suburban providers as unmatched because it requires a principal-city registry address. A complementary CBSA specification uses ZIP geography, excludes ambiguous names, and omits rename-tolerant matching; its geographic crosswalk has two stated approximations.
Appendix G, Metro-expanded matching · Read the study
[c11] Against the restaurant census distribution, recommended sets showed larger standardized shifts in review-count visibility than in star rating. Arm A used a later hosted-API rerun across 13 metros, excluded 44 manually identified misresolutions, and was collected two weeks after B/C.
Appendix G, Restaurant standardization · Read the study
[c12] The authors report aggregate recommendation-set composition, not judgments about individual providers. They suppress small cells and withhold provider-level disclosure/rating pairings; Item 11 records past events, and star ratings are contested measures.
Appendix H, Ethics and adverse impact · Read the study
[c13] The audit used U.S. domains and English prompts, with registry infrastructure that most countries lack. Its findings should not be generalized to other languages or countries without further validation.
Section 7, Discussion, final limitations paragraph · Read the study