The practical brief
The finding
Across six commercial LLMs and five categories, repeated category-only recommendation prompts often omitted established brands entirely. The paper found only limited support for a simple popularity-bias explanation; in its exploratory modeling, recommendation prominence aligned more with broader marketplace visibility signals, especially search interest. [c1] [c2] [c3] [c4] [c5] [c6] [c7] [c8]
Why it matters
If buyers start with AI recommendations, the first risk is not slipping from first place to third. It is failing to enter the model’s generated set at all—the short list of brands the LLM constructs for the user in that moment.
| Brand | Category-only baseline | Detailed needs-based prompts |
|---|---|---|
| Craftsman drills | BRP@5 0%; MRR 0 | BRP@5 35.4%; MRR .108; based on 240 recommendation lists [c1] [c6] |
| L.L.Bean hiking jackets | BRP@5 0%; MRR 0 | BRP@5 5.4%; MRR .025; based on 240 recommendation lists [c1] [c6] |
What to try
BLURSOR’s practical interpretation
Our interpretation: treat AI recommendation visibility as an audit problem, not a screenshot problem. Run repeated tests across the LLMs your buyers use, measure both BRP@5 and MRR@5, and add realistic needs-based prompts that reflect the jobs, constraints, and use cases your brand is supposed to win. Use any correlation with search interest or online conversation as a testing lead, not as proof for budget reallocation.
Study boundary
The study was run through provider APIs in fresh stateless sessions, with no user history and web search or other retrieval tools disabled, so it should not be generalized directly to consumer-facing, retrieval-enabled, or personalized chat products. The diagnostic prompt examples also cover only two brands and come from a small handwritten prompt set; the authors explicitly say this three-stage approach still needs validation at scale across a larger, randomly selected set of brands before it should be treated as a general diagnostic tool.
The decision: are you absent from AI consideration sets?
For a business owner, the immediate question is simple: when a buyer asks an LLM for brands in your category, do you appear at all? This paper argues that LLMs build a generated set—a short recommendation list assembled at answer time rather than chosen from a visible catalog. If your brand is omitted, you may never make the buyer’s consideration set.
That omission risk was not limited to obscure brands. In the category-only tests, the authors report that many large, established brands received no recommendations at all. They also show why one screenshot is the wrong way to judge performance: the same prompt can return different brands and different ordering, so visibility has to be estimated from repeated runs.
- Baseline test: 6 commercial LLMs across 5 categories.
- Each category-only prompt was run 40 times per model in independent fresh sessions.
- That produced 1,200 category-only recommendation lists in total.
What the study measured, in business terms
The paper separates two questions marketers often blur together. BRP@5, or Brand Recommendation Probability at five, measures prevalence: the share of recommendation lists where a brand appears in the top five. MRR@5, or Mean Reciprocal Rank at five, measures prominence: it gives a brand 1 when ranked first, 1/2 when second, 1/3 when third, and 0 when absent, then averages that across repeated lists.
That distinction matters operationally. A brand can show up often but usually low in the list, or appear less often but near the top when it does surface. The authors’ framework is built to measure both because LLM recommendations are stochastic, meaning variable across identical requests.
- BRP@5 answers: how often do we get named?
- MRR@5 answers: how high do we tend to appear when named?
- Repeated sampling is necessary because identical prompts can yield different outputs.
Evidence [c1]
Fame alone did not explain who got recommended
A common assumption is that LLMs mostly echo the biggest brands. This study finds only limited evidence for that cleaner popularity-bias story. Some prominent brands were omitted, and recommendation prominence did not reliably track conventional salience measures across categories.
In the authors’ exploratory visibility analysis, search interest had the strongest predictive relationship with MRR. Online brand conversation ranked behind search interest in the ridge model, but this part of the paper is especially important to read cautiously: the predictors were highly correlated, the analysis was observational, and the authors explicitly warn against treating these relationships as causal proof for spending changes.
For marketers, the practical read is narrower than 'optimize search volume.' AI recommendation visibility may depend on whether your brand is broadly visible and associated with the category in the surrounding information environment, but this paper does not show which intervention would cause improvement.
- The paper reports only partial support for standard popularity bias.
- Search interest was the strongest predictor of recommendation prominence in the exploratory model.
- Any role for online conversation should be treated as suggestive, not decisive, because of multicollinearity and non-causal design.
Needs-based prompts can recover some missing brands
The most useful managerial idea here is diagnostic, not magical. The authors took two brands that never appeared in category-only recommendations—Craftsman drills and L.L.Bean hiking jackets—and asked whether richer consumer context would help. Needs-based prompts described plausible goals and constraints without naming the brand.
The result: both brands became more retrievable, but unevenly. Craftsman moved from complete omission to BRP@5 of 35.4% and MRR of .108. L.L.Bean improved much less, reaching BRP@5 of 5.4% and MRR of .025. When the authors then used diagnostic positioning probes with unusually distinctive cues drawn from the brands’ own positioning, BRP@5 jumped above 80% for both brands. That suggests omission was not simply a lack of brand knowledge inside the model.
Our interpretation is that this supports a practical audit sequence: test category prompts, then realistic needs-based prompts, then diagnostic probes if a brand stays missing. But it does not prove that rewriting site copy, buying more media, or planting positioning phrases will reproduce the same gains in live consumer interfaces.
- Category-only baseline for both brands: BRP@5 = 0%, MRR = 0.
- Detailed needs-based prompts helped Craftsman more than L.L.Bean.
- Diagnostic probes test conditional retrievability, not normal market performance.
What this study cannot settle for your brand
The boundary conditions matter. These tests were run through provider APIs, in fresh stateless sessions, with no prior user history and web search or other external retrieval disabled. That makes the paper strong for measuring baseline model behavior, but weaker as a direct proxy for what a logged-in consumer may see in a retrieval-enabled product.
The brand-level diagnostic results are also illustrative, not a validated general playbook. They come from two brands and a small handwritten prompt set, and the authors explicitly say the three-stage diagnostic approach needs larger-scale validation across a randomly selected brand sample before managers should treat it as a general tool.
- Do not generalize these results directly to personalized or browsing-enabled chat experiences.
- Do not read search-interest correlations as proof that more spend there will raise recommendation rates.
- Do use the paper as a measurement template for your own repeated brand audit.
The source and evidence
Evaluating Brand Retrieval and Ranking in Large Language Model Recommendations
Study limitations and disclosures
- The authors caution that the visibility analysis is observational and non-causal: “The analysis should instead be interpreted as an exploratory observational study” (§4.2 Explaining Brand Recommendation Prominence).
- Generalizability is limited for the diagnostic brand examples: “This three-stage approach needs validation at scale across a larger, randomly selected set of brands before it's a general diagnostic tool.” (§5.3 Limitations and Future Research)
- The study used fresh stateless sessions rather than personalized histories: “Our study used fresh LLM sessions and therefore did not incorporate users’ prior chat histories.” (§5.3 Limitations and Future Research)
- Some marketplace measures are approximate rather than exact counts: “We therefore interpret Brandwatch values as relative indicators of online brand conversation rather than exact mention counts.” (§3 Methodology)
Evidence behind this briefing
[c1] The study measures baseline recommendation variability by repeatedly querying multiple commercial LLMs with category-only prompts in fresh API sessions.
§3.1 Experiment · Read the study
[c2] Established brands were often completely absent from category-only LLM recommendations, which matters for brand consideration.
§4.1 Measuring Brand Recommendation Prevalence and Prominence · Read the study
[c3] The paper reports only partial support for a standard popularity-bias explanation of which brands LLMs recommend.
§4.1 Measuring Brand Recommendation Prevalence and Prominence · Read the study
[c4] Marketplace visibility signals correlated with recommendation prominence, with search interest emerging as the strongest predictor in their exploratory model.
§4.2 Explaining Brand Recommendation Prominence; Figure 5 · Read the study
[c5] The paper explicitly warns marketers not to treat visibility correlations as causal evidence for budget shifts.
§2.4 Framework for Evaluating LLM Brand Recommendations, Explore · Read the study
[c6] Needs-based prompts sometimes surfaced omitted brands, but effects remained limited in the two brand illustrations, with reported sample size and metrics.
§4.3 Needs-based Prompts · Read the study
[c7] Diagnostic positioning probes greatly increased retrievability for the two omitted brands, suggesting omission was not simply lack of brand knowledge.
§4.3 Needs-based Prompts · Read the study
[c8] The authors limit the generalizability of the needs-based diagnostic results because they come from a small handwritten prompt set for two brands.
§5.3 Limitations and Future Research · Read the study