The practical brief
The finding
When assessing AI for legal research, ask whether it retrieves the law applicable to the event date—not just a relevant citation. A date-aware research agent achieved stronger statute-text reproduction and a higher mixed-task score than tested alternatives on a Chinese-law benchmark. [c1] [c3] [c4] [c5] [c6]
Why it matters
An answer can cite a real statute yet use the wrong amendment version, creating a risk for businesses researching historical disputes or obligations.
| Measure and scope | LegalSearch-R1 | DeepSeek-V3.2 |
|---|---|---|
| Seven-task in-domain average | 55.90 mixed-metric score; 896 Chinese-law test examples | 50.56 mixed-metric score; 896 Chinese-law test examples [c1] [c4] |
| Statute recitation | 96.73 character-level ROUGE-L; 128 statute-recitation test examples | 63.30 character-level ROUGE-L; 128 statute-recitation test examples [c1] [c4] |
| Six unseen examination tasks | 63.67% average accuracy; 768 Chinese-law examination questions | 66.41% average accuracy; 768 Chinese-law examination questions [c1] [c4] |
What to try
BLURSOR’s practical interpretation
Our interpretation: evaluate version selection separately from citation presence before relying on a legal research system.
Study boundary
This was an academic Chinese-language benchmark, not a test of commercial legal advice, brand visibility or customer trust.
The purchasing question is whether dates change retrieval
For a business choosing an AI legal research tool, citations alone are an incomplete acceptance test. The relevant question is whether the system can distinguish the law effective when an event occurred from later amendments.
The study tests a system built around that distinction. Its local statute collection attaches effective-date windows to individual provisions, allowing retrieval to filter for the period specified in a question. That is a more specific capability than browsing for a generally relevant legal page.
What the researchers built and tested
LegalSearch-R1 combines web research with retrieval-augmented generation, or RAG: fetching source material for a model to use in its answer. Web search handles broader legal research; a curated local corpus supplies precise statutory text across historical versions of 16 major Chinese statutes.
The researchers trained the agent using reinforcement learning—rewarding desired answers—on 3,584 examples across seven tasks. They evaluated it on 896 in-domain examples and 768 unseen professional-examination questions across six additional tasks.
The agent uses a Qwen2.5-7B backbone, but it is not a standalone small-model comparison: its browsing and statute tools also use Qwen3-30B-A3B. Evaluation used one deterministic response per prompt.
The strongest result was precise statute reproduction
The agent scored 55.90 on the seven-task in-domain average, compared with 50.56 for DeepSeek-V3.2 and 39.19 for DeepPlanner. This is not an overall accuracy percentage: the average combines accuracy, judged consultation quality and statute-text similarity.
Its largest advantage was statute recitation, measured with character-level ROUGE-L, a text-overlap metric. It scored 96.73 against DeepSeek-V3.2’s 63.30. However, the recitation examples deliberately target significantly revised provisions from three laws, rather than an unrestricted sample of statutes.
On the unseen examination tasks, the agent’s 63.67% average accuracy exceeded tested deep-research and legal-specialist systems, but remained below DeepSeek-V3.2’s 66.41%. The results support a narrower conclusion than universal legal superiority.
A realistic interpretation for buyers and publishers
Our interpretation is to test historical-version handling explicitly rather than treating a citation as proof of correctness. The study supports examining a tool’s source architecture, but does not prove that any single publishing change improves AI visibility.
For businesses publishing legal explainers, retaining effective dates and revision history is a reasonable source-quality practice. The study’s webpage reader was instructed to preserve those details; their effect on commercial answer-engine citations was not measured.
- Ask vendors which statutes and historical versions their retrieval corpus actually covers.
- Include dated questions in evaluations and have qualified reviewers check the applicable version.
- Keep benchmark performance separate from suitability for a live, high-stakes workflow.
The source and evidence
Study limitations and disclosures
- Testing covered Chinese law and Chinese-language tasks only. Multilingual and common-law performance remain unassessed, and the local corpus contains statutes rather than judicial precedents.
- Baseline coverage was incomplete: one system lacked public weights, while others failed the required output format. Evaluation used one deterministic response per prompt, not repeated stochastic trials.
- The authors release the framework and benchmark solely for academic research and acknowledge high-stakes misapplication risks. They disclose public academic datasets, API-license compliance and law-student annotation paid at US$15 per hour.
Evidence behind this briefing
[c1] The benchmark uses 3,584 training samples, 896 in-domain test samples and 768 out-of-domain test samples. Its headline in-domain average mixes statute-text ROUGE-L, LLM-judged consultation scores and accuracy; it is not a uniform accuracy measure. Training rewards also differ from evaluation metrics.
Section 6.2, Datasets and Metrics · Read the study
[c2] The nominally 7B agent also uses Qwen3-30B-A3B inside its browsing and statute-retrieval tools. Evaluation uses deterministic greedy decoding and one rollout per prompt, rather than repeated stochastic evaluations.
Appendix A, Implementation Details · Read the study
[c3] The retrieval design separates broad web research from precise statute retrieval. Its local corpus covers 16 major Chinese statutes with historical amendment versions and article validity windows; an auxiliary query analyzer extracts dates before temporal filtering and hybrid retrieval.
Section 5.1, Temporal-Enhanced Legal Statute RAG · Read the study
[c4] On the Chinese benchmark, the agent reports a 55.90 mixed-metric in-domain average versus DeepPlanner's 39.19 and DeepSeek-V3.2's 50.56. Its largest advantage is statute recitation: 96.73 ROUGE-L versus 63.30 for DeepSeek-V3.2. On 768 unseen examination questions across six domains, its 63.67 accuracy average exceeds the tested deep-research and legal-specialist baselines but remains below Qwen-3-30B-A3B's 63.80 and DeepSeek-V3.2's 66.41. These are benchmark scores, not evidence about brand visibility or commercial answer engines.
Section 6.3, Main Results; Table 2; sample denominators in Section 6.2 · Read the study
[c5] Generalization is limited by exclusively Chinese-language, Chinese-law testing, unassessed multilingual performance and a statute-only RAG corpus. Baseline coverage excludes LRAS because weights were unavailable and excludes Fuzi-Mingcha and Hanfei because their outputs failed the required answer-format protocol.
Limitations · Read the study
[c6] The authors disclose public academic datasets, law-student statute annotations paid at 15 USD per hour, no personally identifiable information in those annotations, API-license compliance and release solely for academic research because of high-stakes legal misapplication risks.
Ethics Statement; annotation compensation also disclosed in Appendix G · Read the study
[c7] The training/evaluation prompt requires search-grounded answers and routes specific statutory questions to RAG, while routing legal theory, case analysis and judicial practice to web search. These are configured instructions, not evidence that the system always complied.
Figure 7, System Prompt for LegalSearch-R1, Tool Routing Rules · Read the study
[c8] The webpage-reading agent is instructed to preserve original Chinese legal provisions verbatim, extract judicial material as completely as possible, and retain effective dates, versions and revision histories. This is an extraction design rather than measured extraction accuracy.
Figure 8, Reading Agent Extraction Prompt, Special Requirements · Read the study
[c9] RAG query analysis explicitly extracts temporal information, expands partial dates into date ranges, and returns an empty list when the query contains no time period.
Figure 9, RAG Query Analysis Prompt, Temporal Information · Read the study
[c10] The configured local RAG corpus covers temporal versions of 16 Chinese statutes; this is a bounded statutory corpus, not a claim of comprehensive coverage of all Chinese law.
Figure 10, Tool Schema Configuration, rag_retrieve · Read the study
[c11] LAR annotation selects significantly revised provisions from three named laws, requiring either a greater-than-20% character-count difference or a clear change in legal meaning or applicability. Consequently, these examples target revised provisions rather than an unrestricted sample of statutory text.
Figure 13, Annotation Guidelines for Legal Article Recitation, Objective and Selection Criteria · Read the study
[c12] The inheritance-dispute case study illustrates combining web search, statute retrieval and temporal reasoning to select the pre-Civil Code legal regime. It is an illustrative trace, not an aggregate accuracy measurement or causal demonstration of tool effectiveness.
Figure 14, caption · Read the study