The practical brief

The finding

LawCompass reported higher retrieval scores than either original-query or rewritten-query retrieval on 20 legal queries; the evaluation measured useful evidence passages, not answer accuracy. [c1] [c3] [c4] [c5] [c7] [c8]

Why it matters

Businesses assessing AI research tools need to distinguish finding relevant sources, making claims checkable, and reaching correct conclusions.

What the evidence supports—and what it leaves open
Evidence typeSupported conclusionNot established
Passage-relevance evaluationCombined retrieval scored higher than either query path alone.Generated answers were more accurate. [c4] [c5]
Clickable source catalogueUsers can inspect cited metadata and original excerpts.Every cited claim is correct or adequately supported. [c3]
Task-based user ratingsParticipants reported favorable perceptions of trust and research value.Verified accuracy or an advantage over another tool. [c6]
Qualitative interpretation of the study's evaluations; not additional measurements. Study source

What to try

BLURSOR’s practical interpretation

Our interpretation: evaluate source coverage and claim-to-source support separately before relying on an AI-generated research report.

Study boundary

This domain-specific demonstration does not establish marketing visibility gains, reliable legal advice, or deployment economics.

Buy verification capability, not a citation-shaped guarantee

If you are choosing an AI tool for business research, do not treat linked citations as proof that its conclusions are correct. This study supports a narrower proposition: a specialized legal assistant can retrieve useful authorities and make its sources easier to inspect.

LawCompass is a demonstration system for Chinese legal research, not a test of brand inclusion in general-purpose AI answers. Its relevance to business owners is the distinction between evidence discovery and evidence verification—two capabilities that can look identical in a polished report.

Evidence [c1] [c3] [c4]

The system searches a defined legal collection

The prototype uses DeepSeek-v4 as its primary reasoning engine, with 220,000 statutory provisions from Lawyee and 100,000 Chinese court cases from China Judgements Online. Tavily web search supplements that local collection.

It uses retrieval-augmented generation: finding external material before generating an answer. Search combines the user's original query with a rewritten version, using both keyword matching and vector search, which matches text by semantic similarity. Keyword results favor recency; retrieved statutes and cases are then ordered by legal authority.

Its Deep Research mode divides a question among software agents with planning, evidence-gathering and synthesis roles. A source catalogue supplies citation tags, and clickable links expose metadata and original excerpts. These are mechanisms for checking output, not measured guarantees of correctness.

Evidence [c1] [c2] [c3]

Retrieval improved; correct answers were not measured

Two legally trained annotators independently assessed the top five passages for each of 20 queries: 100 passages per retrieval method. They judged whether passages supplied useful legal authority, labeling them relevant, partially relevant or irrelevant.

LawCompass reported Precision@5 of 79.0, versus 76.0 for original-query retrieval and 75.0 for rewritten-query retrieval. Precision@5 measures relevance among the first five results. Its nDCG@5—a measure that rewards relevant results appearing higher—was 90.9, versus 85.5 and 85.9. The table does not specify percentage units or report uncertainty or significance tests.

After three tasks, users gave overall mean ratings of 4.2/5 for evidence trust and 4.4/5 for Deep Research value. These are subjective assessments. Participant counts, variability and a comparison system are absent, and the description ambiguously mixes scale ratings with rating open-ended feedback.

Evidence [c4] [c5] [c6]

A realistic interpretation for purchasing and publishing

Our interpretation is to assess AI research tools in separate layers rather than accepting a single trust score. The study shows why source selection matters: this system prioritizes legal authority, not semantic similarity alone. That does not establish how a commercial website can win inclusion.

For marketers, the defensible lesson is about auditability, not a proven optimization tactic. Inspectable source text and metadata make checking possible; this paper does not show that adding them increases citations, recommendations or sales.

  • Ask which databases the tool searches and how recently they were updated.
  • Inspect whether a cited excerpt actually supports the claim beside it.
  • Evaluate completed answers separately from retrieved passages and user satisfaction.

Evidence [c2] [c3] [c4] [c6] [c7]

Coverage and operating costs remain open questions

The authors acknowledge that new laws, recent cases and local regulations may be missing when databases update slowly. Multi-stage research also adds latency and computational cost, but neither is quantified. Whether those trade-offs are acceptable depends on the intended workflow.

The sample is small and domain-specific. It cannot establish reliable legal advice, general commercial discoverability or the economic value of deployment. The appendix's dismissal-compensation report is abridged, not independently accuracy-scored, and explicitly says it is not formal legal advice.

Evidence [c1] [c4] [c7] [c8]

The source and evidence

LawCompass: Navigating from Legal QA to Multi-Agent Deep Research with Grounded Evidence

Xiaoxia Cheng, Linnan Wang, Jiahao Ma, Zhichuan Ye, Xuemei Zhou, Chuanyu Tong, Bo Jiang, Qing Zhu · 2026-10-01

Study limitations and disclosures

  • The retrieval evaluation covers 20 legal queries and the top five passages per query. It measures passage relevance, not generated-answer accuracy; uncertainty and significance tests are not reported.
  • User ratings are subjective. Participant counts, rating variability and a comparator are not reported, and the account of scale ratings versus rated open-ended feedback is ambiguous.
  • Database completeness and update frequency limit coverage. Multi-agent research adds latency and computational cost, without reported measurements.
  • The study concerns a Chinese legal-research demonstration, not marketing interventions or brand visibility. Its abridged example report is not an accuracy evaluation and disclaims formal legal advice.
  • Seven authors list Anhui University affiliations; Chuanyu Tong lists Tsinghua University. The authors evaluate their own system. The paper does not state funding or commercial-interest disclosures; its use of named data and technology providers does not by itself establish sponsorship.

Evidence behind this briefing

[c1] The test system uses DeepSeek-v4 as its primary reasoning engine and a corpus of 220,000 statutory provisions and 100,000 Chinese court cases, supplemented by Tavily web search. This is a domain-specific implementation, not evidence about general commercial AI discovery.

§4.1 LLM and Data Services, Data Services; primary model disclosed in §4.1 LLM · Read the study

[c2] Discovery combines original and rewritten queries, keyword and vector search, recency ordering, and domain-authority ranking. Source selection therefore is not based on semantic similarity alone; the paper does not test how marketers can influence inclusion.

§4.3 Professional Retrieval, Dual-Path Retrieval · Read the study

[c3] Credibility is supported through a source catalogue, constrained citation tags, and clickable original excerpts. These are verification mechanisms, not measured guarantees that generated claims are accurate.

§4.5 Evidence-Grounded Generation, Mandatory Citation · Read the study

[c4] Retrieval was evaluated on 20 legal queries, with five retrieved passages per query (100 passages) independently judged by two legally trained annotators. Labels measured relevance of legal authority, not generated-answer accuracy.

§5.1 Retrieval Quality · Read the study

[c5] On the 20-query retrieval evaluation, LawCompass scored 79.0 P@5, 90.9 nDCG@5, and 1.28 average relevance, versus 76.0/85.5/1.10 for original-query retrieval and 75.0/85.9/1.11 for rewritten-query retrieval. The table does not specify percentage units or provide uncertainty or significance tests; this is a retrieval comparison, not a causal marketing result.

§5.1, Table 1 · Read the study

[c6] After three representative tasks, participants with and without legal backgrounds gave overall mean ratings of 4.2/5 for evidence trust, 4.4/5 for Deep Research value, and 4.3/5 for satisfaction. These are subjective ratings, not verified accuracy. Participant counts, variability, and a comparator are not reported; the prose also describes the averages as obtained by rating open-ended feedback.

§5.2 User Study, Table 2; task and feedback procedure in preceding prose · Read the study

[c7] The authors disclose that incomplete or slowly updated databases can miss new legal materials, and that multi-agent research adds latency and computational cost. No latency or cost measurements are supplied.

Limitations · Read the study

[c8] The illustrative dismissal-compensation report explicitly disclaims formal legal advice. Appendix A is an abridged demonstration with analyses and sources omitted for space, not an independently scored accuracy evaluation.

Appendix A, Table 3, Liability Disclaimer; abridgement disclosed in Table 3 caption · Read the study