The practical brief
The finding
On 209 deliberately difficult French tax-law questions, mean strict accuracy across eleven models was 2.7% with current-version retrieval and 98.3% with the authors’ end-to-end, date-aware retriever. [c1] [c2] [c3] [c4] [c5] [c6] [c7]
Why it matters
A real source link does not establish that the cited rule applied when the business event occurred.
| Answer condition | Mean strict accuracy |
|---|---|
| No retrieval | 3.0% across eleven models on 209 scored French tax-law questions [c2] [c3] |
| Current-version retrieval | 2.7% across eleven models on 209 scored French tax-law questions [c2] [c3] |
| End-to-end date-aware retrieval | 98.3% across eleven models on 209 scored French tax-law questions [c2] [c4] |
| Date-aware retrieval given the correct article | 99.1% across eleven models on 209 scored French tax-law questions [c2] [c4] |
What to try
BLURSOR’s practical interpretation
Our interpretation: when assessing date-sensitive AI answers or publishing changing guidance, check effective dates and preserve clearly identified historical versions rather than relying on a current page alone.
Study boundary
This selected benchmark is not an estimate of ordinary legal-AI failure rates or a test of marketing visibility. Most questions supplied the article identifier, and the headline retriever is proprietary.
Check the applicable date before accepting a citation
If you use an AI answer to review an earlier tax period, a clickable legal citation is not enough. The decision is whether the evidence applies to the period under review—not simply whether the source exists. This study isolates a risk the authors call temporal misgrounding: using today’s version of a rule when the question requires an earlier version.
For business owners and marketers, the distinction matters when assessing source-backed answers. A citation can identify the right legal article while supporting the wrong numerical claim for the requested date. The study measured that problem in French tax law; it did not measure customer trust, brand mentions or discoverability.
The experiment deliberately targeted changing, obscure rules
The authors built a corpus containing 32,436 article-versions across six French tax codes, preserving text and validity dates. They scored 209 author-written, expert-reviewed questions across 33 articles, with question dates from 1989–2025. Twelve additional released questions were excluded from scoring.
Eleven answer models were tested without retrieval, with current-version retrieval, and with date-aware retrieval. Retrieval-augmented generation, or RAG, means giving a model retrieved documents to help it answer. The current-version baseline received the correct article identifier; the end-to-end system had to find relevant articles and then select their date-applicable versions.
Strict accuracy required both the correct article and the exact date-applicable numerical value, checked by deterministic matching rather than another AI judge. A separate provenance measure checked whether retrieval included the applicable source version.
Selection strongly shaped the results: the questions targeted values models had not recovered during screening, and current text lacked the correct historical value for 208 of 209 scored questions. Low baseline accuracy was therefore largely an expected construction check, not a representative failure rate.
Date-aware retrieval changed performance on the same questions
Mean strict accuracy across eleven models was 3.0% without retrieval and 2.7% with current-version RAG. Their uncertainty intervals overlapped. Static retrieval included the date-applicable version in 0% of cases, illustrating why adding genuine references need not solve a date mismatch.
The end-to-end date-aware retriever reached 98.3% mean strict accuracy, with a cluster-bootstrap 95% confidence interval of 95.9–99.6%. It included the applicable version among its top five retrieved articles for 99% of questions. A comparison supplied with the correct article in advance reached 99.1% strict accuracy; the remaining gap was concentrated in finding the article.
The realistic interpretation is narrow but useful: on this selected task, retrieving the applicable version mattered far more than merely supplying current legal text. However, 175 of 209 questions already disclosed the article identifier, so finding the source was substantially assisted.
Treat version history as a credibility check, not a visibility tactic
Our interpretation is to evaluate date-sensitive evidence explicitly. For businesses publishing changing rules or guidance, preserving effective-date information is a reasonable application of the study’s mechanism—not a proven way to earn more AI citations.
- When reviewing an AI answer, compare the question’s date with the cited version’s validity period and verify the numerical claim in that version.
- When publishing revised guidance, distinguish when a rule applies from when the page was updated, and keep historical versions identifiable.
- When assessing a retrieval vendor, ask for date-sensitive tests and source-version provenance, not just demonstrations of clickable citations.
What this study cannot tell your business
The controlled evaluation covered current-law substitution, not future-rule leakage, amendment-relative questions, comparisons across versions or legal interpretation. It does not establish performance in other jurisdictions, ordinary customer queries or public AI search.
The strongest retriever cannot be independently rebuilt from the released artifacts. Its saved responses can be rescored, but that is different from reproducing its retrieval performance. The result supports further testing of version-aware systems, not removing professional review from tax decisions.
The source and evidence
Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law
Study limitations and disclosures
- The benchmark deliberately selected difficult, date-sensitive numerical questions. Current text lacked the historical answer for 208 of 209 scored questions, so baseline results should not be generalized to ordinary legal questions or marketing queries. Questions and answers were author-written and reviewed by a qualified French tax professional; twelve of 221 released questions were excluded as outside the answerable scope.
- The controlled study covered French tax law, 33 articles and question dates from 1989–2025. It tested substitution of current law for historical law, not future-law leakage, amendment-relative queries, cross-version comparisons or legal interpretation. It did not measure brand visibility, citation preference or marketing outcomes.
- Most questions supplied the article identifier: 175 of 209, or 83.7%. Retrieval performance was therefore conditioned on substantial identifier assistance; citing the article alone was weak evidence that retrieval succeeded.
- The end-to-end retriever’s encoder weights, multi-version index and inference code are proprietary and unreleased. Released responses allow independent score checking, but not reproduction of the headline retrieval system. An internal benchmark source pool is also proprietary. Encoder training overlapped four scored articles covering 28 questions; the authors report no historical gold values from those questions in training passages, with the checking script available on request.
- Rose Cymbler and Daniel Guez are affiliated with Talia; Laurent Fabre is affiliated with Databricks. These company affiliations and the proprietary retriever are material commercial context. The supplied paper does not state a funding source or provide a separate financial-conflict declaration; that absence should not be interpreted as proof of no funding or interests. Its Impact Statement calls for qualified professional verification and says outputs are not legal advice.
- Evaluation used one answer draw per question, with model-dependent sampling and reasoning settings. Provider constraints required Gemini and Qwen substitutions; 22 provider-side empty responses for Gemini and GLM across conditions counted as failures. Selection-time difficulty did not persist perfectly: at evaluation, at least one model supplied the correct value without retrieval for 37 of 209 questions.
- Retrieved extracts were capped at 8,000 characters per article, with applicable values checked for retention before scoring. The end-to-end condition supplied up to five articles, while the correct-article comparison supplied one. These controlled prompt conditions may differ from a live business workflow.
- The auxiliary court-citation dataset was not used in the controlled experiment. Its link-level precision audit does not establish exhaustive coverage: a separate manual audit found 53.4% article-level recall. Links generally select versions by decision date, which may differ from the date of the underlying facts.
Evidence behind this briefing
[c1] The corpus preserves article histories and explicit validity dates rather than treating each article identifier as a single current document. It contains 32,436 article-versions across six French tax codes; its 1938–2031 endpoints require the caveats in Appendix A.
§4.1, Pipeline; corpus totals in §4.1 Output and Appendix A, Table 2 · Read the study
[c2] Strict accuracy requires both the correct article and exact date-applicable numerical value, checked deterministically rather than by an LLM judge. Retrieval provenance is a separate metric: whether the applicable gold version appears in the retrieved set. These metrics should not be confused with citation frequency or persuasive credibility.
§5.2, Nugget-Based Scoring · Read the study
[c3] On the deliberately selected 209-question temporal-drift benchmark, mean strict accuracy across eleven models was 3.0% without retrieval and 2.7% with current-version RAG; their cluster-bootstrap 95% CIs were [1.4, 4.7] and [1.3, 4.8], respectively. Static RAG retrieved the applicable version in 0% of cases. This is not a general failure-rate estimate: selection excluded memorized values and screened for current-version divergence, making the low baseline scores largely construction checks.
§7.2, Controlled Experiment Results, opening results paragraph; selection criteria in §5.3 · Read the study
[c4] The end-to-end version-aware retriever achieved 98.3% mean strict accuracy (cluster-bootstrap 95% CI [95.9, 99.6]) and 99% top-5 applicable-version provenance on the same 209 questions across eleven models. Oracle-article retrieval achieved 99.1% strict accuracy (CI [98.4, 99.9]); the 0.8-percentage-point difference was concentrated in article recall. This controlled system comparison supports date-conditioned retrieval for this task, not a universal causal claim about AI discoverability.
§7.2, The operative result: end-to-end retrieval · Read the study
[c5] Most queries already disclosed the article identifier: 175 of 209 questions (83.7%). Article citation alone therefore provides weak evidence of successful retrieval, and the reported article-recall performance is conditioned on substantial identifier assistance. Value-only accuracy closely tracked strict accuracy.
§5.2, Article-nugget leakage · Read the study
[c6] The evidence is limited to French tax law and a small, expert-reviewed, deliberately difficult track: 221 released questions, 209 scored, spanning 33 articles and question dates from 1989–2025. Only current-law substitution was experimentally tested; future-law leakage, amendment-relative queries, comparisons across versions, and legal interpretation remain outside the controlled evaluation.
§9, Failure modes evaluated; scope in §9 and sample distribution in Appendix A, Table 3 · Read the study
[c7] The headline end-to-end retriever is proprietary and cannot be independently reproduced from the released artifacts, although its response scores can be rechecked. Encoder training overlapped four scored benchmark articles covering 28 questions; the authors report that none of those questions' historical gold values appeared in training passages, with the checking script available only on request.
Appendix D, Reproducibility, release exclusions and subsequent encoder-overlap disclosure · Read the study
[c8] The auxiliary jurisprudence links were not used in the controlled R3 experiment. Their 98–99% link-level precision from a 100-link audit does not establish exhaustive citation coverage: a separate 50-decision manual audit found 88.6% article-level precision and 53.4% article-level recall, despite 97% decision-level recall. Finding at least one grounding article is materially different from finding all cited articles.
Appendix B.2, Manual gold-standard recall; auxiliary status in Appendix B introduction · Read the study