The practical brief
The finding
In a 7-day online A/B test on Meituan’s search advertising system, UniPolicy improved click-through rate by 0.71%, revenue per search by 1.58%, and advertising revenue by 1.32%, with a 2.5% increase in P99 latency versus the current production model. Offline, it also achieved the best NDCG@10 among the reported baselines while keeping strong retrieval quality. [c1] [c2] [c4] [c6] [c7] [c8]
Why it matters
For a business running ads, retail media, or marketplace discovery, the retrieval stage hard-codes trade-offs early. This paper suggests that keeping explicit control over commercial value, click propensity, and relevance in one system may outperform collapsing them into a single reward—at least in this tested environment.
| Metric | Reported change | Interpretation |
|---|---|---|
| CTR | +0.71% | Higher click-through in this tested system [c1] |
| Revenue per search | +1.58% | Higher monetization per search in this tested system [c1] |
| Advertising revenue | +1.32% | Positive revenue lift in the 7-day test [c1] |
| P99 latency | +2.5% | Small serving-speed cost rather than zero overhead [c1] |
What to try
BLURSOR’s practical interpretation
Our interpretation: if you are evaluating AI-generated ad or product retrieval, test whether you need separate objective controls rather than one blended optimization target. In practice, that could mean comparing objective-specific candidate generation and quota allocation under a fixed retrieval budget against your current single-policy setup.
Study boundary
The evidence comes from one in-house Meituan platform, with online results compared only against the current production model. Offline business metrics are normalized rather than raw values, and the method is tested on only three objectives: eCPM, CTR, and Relevance. Training cost also grows roughly with the number of policies.
The business decision sits upstream of ranking
If your discovery system uses generative retrieval, the key decision is not just which ad or listing ranks first. It is whether the retrieval model itself should optimize one blended business goal or keep separate controls for different goals. That matters because early retrieval filters the candidate set, and later stages cannot recover options that were never brought forward.
This paper studies that choice in search advertising. The authors define three objectives inside one model: eCPM for commercial value, CTR for click propensity, and Relevance for user experience. Instead of compressing them into one reward, they train objective-specific policies that can be invoked separately and combined later.
- eCPM = expected cost per mille, used here as a revenue-oriented objective.
- CTR = click-through rate.
- Relevance = how well the ad matches the query and user need.
Evidence [c4]
What the study actually built and tested
UniPolicy is a generative retrieval system built on a shared model backbone. It adds objective tokens—special input markers that tell the model which business goal to emphasize—and separate policy paths so the three objectives do not fully compete inside one parameter space.
The training setup also uses richer behavioral feedback than clicks alone. The paper models user signals as Click > Exposure > Miss, meaning a clicked ad is preferred to an exposed-but-unclicked ad, which is preferred to an unexposed candidate. In this generative retrieval setting, that is meant to extract information from impressions and misses that click-only supervision would leave unused.
Offline, the authors compare UniPolicy with single-objective reinforcement learning and several multi-objective baselines. They report the best NDCG@10 for UniPolicy while maintaining strong Hit performance. NDCG is a ranking metric that gives more credit when the desired result appears nearer the top.
- Online test: UniPolicy versus the current production model.
- Offline tests: UniPolicy versus single-objective and naive multi-objective baselines.
- Offline business-value tables use normalized scores, not raw CTR, revenue, or relevance levels.
What improved in production, and what stayed limited
The strongest evidence for operators is the live result. In Meituan’s 7-day A/B test, UniPolicy improved CTR by 0.71%, revenue per search by 1.58%, and advertising revenue by 1.32%, while P99 latency increased by 2.5% relative to the production system. That is a modest but commercially relevant pattern if your main concern is whether multi-objective control can survive deployment overhead.
The offline analysis helps explain why. The paper reports that different objective tokens steer the model toward different business preferences: the eCPM token performs best on normalized eCPM, the CTR token on normalized CTR, and the Relevance token on normalized Relevance under single-objective inference. When the system merges candidates from all three, it achieves the highest average normalized business score.
That does not mean the paper proved a universal best setup for all ad stacks. It means the tested system benefited from explicit controllability: separate business preferences could be called on and recombined later instead of being forced into one compromise policy.
- Online relative lifts versus production: CTR +0.71%, revenue per search +1.58%, revenue +1.32%.
- Operational trade-off: P99 latency +2.5%.
- Controllability result: objective tokens changed generation behavior in the expected direction.
What this does not tell you about your own platform
The authors are affiliated with Meituan, Beijing, China, and the dataset comes from that company’s industrial advertising platform. That means the evidence is operationally interesting but still single-company evidence. Traffic mix, auctions, advertiser density, and latency budgets may differ sharply in retail media, travel, lead generation, or B2B search.
There are also comparison limits. The online A/B test is against the current production model only, not a live head-to-head against every offline baseline. So you should not read the production gains as proof that UniPolicy beat each simpler baseline online.
Finally, the offline business tables are normalized relative to SFT and the single-objective GRPO baselines, so they cannot be read as raw business levels. The method is also tested on only three objectives, and training computation rises roughly linearly with the number of policies.
- Material disclosure: all authors list Meituan affiliation.
- External validity is limited to one in-house platform and dataset.
- Offline business scores are normalized, not absolute.
- Evidence covers only eCPM, CTR, and Relevance.
- Training cost scales roughly with the number of policies.
The source and evidence
UniPolicy: Unified Objective-Specific Policies for Generative Search Advertising
Study limitations and disclosures
- Material disclosure: the authors state "Affiliation: Meituan, Beijing, China" on the title page, so the evidence comes from a single company’s in-house system.
- The online A/B test comparator is only described as "the current production model," not the offline baselines, limiting head-to-head causal comparison across all methods in production; locator: Section 5.5.
- Offline business-value results are normalized rather than raw, which constrains direct estimation of absolute business impact; locator: Section 5.2.2.
- The method study covers only three objectives—eCPM, CTR, and Relevance—so evidence for broader objective sets is not provided; locator: Section 3, Eq. (4).
Evidence behind this briefing
[c1] In the paper’s real-system online experiment, UniPolicy produced small but positive lifts in click-through, RPS, and revenue, with a modest latency increase versus the production model.
Section 5.5, Table 4 · Read the study
[c2] Offline, UniPolicy achieved the best NDCG@10 and stronger balanced multi-objective business performance than the compared baselines, which matters if a marketer wants both relevance and monetization signals improved together.
Section 5.3 · Read the study
[c3] The paper reports that objective tokens can steer generation toward different business preferences, and combining them gave the highest average normalized business score.
Section 5.4.2, Table 3 · Read the study
[c4] UniPolicy is designed to optimize three advertiser-relevant objectives in one model—commercial value, click propensity, and relevance—rather than collapsing them into one reward.
Section 3, Eq. (4) · Read the study
[c5] The training method explicitly uses richer funnel feedback beyond clicks, which is relevant to marketers because it tries to learn from impressions and non-exposures, not just clicked ads.
Section 4.3.1 · Read the study
[c6] The paper’s business metrics in offline tables are normalized relative to SFT and single-objective GRPO, so marketers cannot read those values as raw CTR, revenue, or relevance levels.
Section 5.2.2 · Read the study
[c7] Training cost scales roughly with the number of policies, which is a practical limitation for adding more objectives or operating under tighter compute budgets.
Section 4.3.2 · Read the study
[c8] The study is based on one company’s industrial system and dataset, which limits direct generalization for marketers on other platforms or verticals.
Section 5.1.1 · Read the study