The practical brief

The finding

The New York Times deployed a search agent that logged 4,500 questions from more than 100 journalists between launch and mid-June; at least 20 written reports contributed to published articles. [c3] [c4] [c7]

Why it matters

For businesses considering AI-assisted research or content production, the relevant lesson is a workflow built around inspectable evidence—not permission to publish generated answers unchecked.

Duplicate filtering: reliability versus coverage
Similarity thresholdPrecision: true duplicates / flagged pairsRecall: flagged duplicates / true duplicates
0.97 — strict0.64 fraction; evaluated nearest-neighbor pairs, population-weighted0.26 fraction; evaluated nearest-neighbor pairs, population-weighted [c6]
0.92 — deployed0.28 fraction; evaluated nearest-neighbor pairs, population-weighted0.86 fraction; evaluated nearest-neighbor pairs, population-weighted [c6]
0.80 — loose0.08 fraction; evaluated nearest-neighbor pairs, population-weighted1.00 fraction; evaluated nearest-neighbor pairs, population-weighted [c6]
Population-weighted Diff results for the evaluated nearest-neighbor page-pair population. These fractions measure duplicate detection, not answer accuracy. Study source

What to try

BLURSOR’s practical interpretation

Our interpretation: evaluate whether a research tool exposes original records and supports human verification before judging it by answer fluency or citation counts.

Study boundary

This was one newsroom investigation, not a controlled productivity study or a test of business visibility in public AI answers.

Buy a research aid, not an unchecked publishing pipeline

If you are choosing AI tools for research or marketing content, decide who will verify their evidence before anything reaches customers. This study documents a useful search workflow, but not a replacement for editorial judgment.

The Epstein Files Engine helped Times reporters investigate a large document release. Its outputs were intermediate research reports: journalists inspected cited originals, added reporting and decided what merited publication. Users were explicitly warned that reports were fallible and statistical claims should not be trusted.

Evidence [c3] [c4]

The system searched records and checked existing coverage

A large language model translated plain-English questions into SQL, the language used to query structured databases. It searched three collections: primary-source releases, the Times archive and external Epstein-related headlines. Extracted images received machine-generated tags and captions so visual material inside PDFs could also be searched.

The system assembled cited reports and compared findings with existing coverage. External headlines were a coverage check, not model-training data. The authors describe this architecture as emphasizing query planning rather than authoritative answer generation; they did not establish superiority over other retrieval systems.

Between launch and mid-June, logs recorded more than 100 journalists asking 4,500 questions. At least 20 reports contributed to articles. Those counts demonstrate use and contribution, not time saved or a question-to-publication conversion rate.

Evidence [c1] [c2] [c3]

Filtering repeated material introduced a measurable trade-off

Diff, the system’s duplicate detector, combined text similarity with visual fingerprints. Its retrospective evaluation sampled 400 nearest-neighbor page pairs from an inventory of 2,683,926 PDF pages.

At the deployed threshold, population-weighted precision was 0.28 and recall was 0.86. Precision measures how many flagged pairs were genuine duplicates; recall measures how many genuine duplicates in the evaluated nearest-neighbor population were flagged. These are not answer-accuracy scores.

A stricter threshold made duplicate flags more reliable but missed more duplicates. For a business evaluating document search, this illustrates why filtering settings deserve review: suppressing repetition can also hide distinct records.

Evidence [c5] [c6] [c8]

A realistic business interpretation: make evidence inspectable

The business relevance is narrower than an AI discoverability tactic. In this deployment, searchable source material and citations gave users something concrete to inspect, while prior coverage supplied context. Neither feature was isolated in a controlled test.

  • For internal research tools, our interpretation is to require access to the underlying records, not just a polished summary.
  • For content workflows, retain a named reviewer who checks supporting documents before publication.
  • For discoverability planning, treat this as a reason to consider evidence accessibility—not proof that any formatting change will win public AI citations.

Evidence [c1] [c4] [c7]

What the study cannot tell a business owner

The paper cannot quantify return on investment, establish productivity gains or predict brand recommendations. It reports one investigation at one newsroom, with activity counts that represent a floor rather than a complete usage distribution.

Diff was validated after deployment, using a manually tuned threshold. There was no comparison with text-only or visual-only versions, and its recall does not measure discovery of every duplicate anywhere in the collection. The results support a bounded deployment account, not a universal recipe.

Evidence [c7] [c8]

The source and evidence

Epstein Files Engine: Agentic Search for Investigative Journalism

Duy K. Nguyen, Teresa Mondr\'ia Terol, Dylan Freedman, Zach Seward · 2026-09-24

Study limitations and disclosures

  • This is a single-newsroom, single-investigation deployment account. It did not test marketing performance, brand discoverability or rankings in public AI answers. No controlled comparison isolated the benefit of the archive or external headlines.
  • Diff validation was retrospective, with a manually tuned threshold and no text-only or visual-only comparison. Labels came from one annotator; unsure pairs, including blank or fully redacted pages, represented 18% of the sampled population and were excluded. Recall concerns each page’s nearest composite neighbor, not every duplicate in the corpus.
  • At the deployed threshold, the 95% bootstrap intervals were 0.21–0.35 for precision and 0.69–1.00 for recall. Separately, assigning unsure pairs adversarially produced precision ranges of 0.21–0.43 and recall ranges of 0.59–0.83; these sensitivity ranges are not confidence intervals.
  • The authors describe their own institution’s deployment: Duy K. Nguyen, Dylan Freedman and Zach Seward were affiliated with The New York Times. Teresa Mondría Terol was affiliated with National Public Radio, with an author note stating that this work was done while at The New York Times. This institutional interest is material; the supplied paper contains no explicit funding statement or separate commercial-interest disclosure.

Evidence behind this briefing

[c1] The Engine queried three corpora: primary-source releases, the Times archive and external Epstein-related headlines. Image extraction and vision-language tags made embedded visual assets searchable. External headlines served as a coverage check, not LLM training data.

§3.1, The Engine’s underlying data · Read the study

[c2] The system used an LLM as a dynamic SQL query planner and report synthesizer, with cited records—not generated answers—as its source of truth. This is a design description, not comparative evidence that it outperforms RAG.

§3.2, The model as query planner · Read the study

[c3] From launch to mid-June, server logs recorded more than 100 journalists asking 4,500 questions. At least 20 written reports contributed to published articles; this is observed usage and contribution, not a causal productivity estimate or a conversion rate for all questions.

§4, Deployment and use · Read the study

[c4] Citation-backed output was an intermediate research aid: reporters inspected original documents and added original reporting before publication. Users were trained to regard output as fallible; the deployment also explicitly warned against trusting statistical claims.

§4.2, From query to published article; statistical-claim caution in §4 · Read the study

[c5] Diff was evaluated retrospectively on a stratified sample of 400 nearest-neighbor pairs from an inventory containing 2,683,926 PDF pages. Sampling deliberately emphasized similarity bands around the deployed 0.92 cutoff, and uncertainty was estimated with a stratified bootstrap.

§4.3, Validating Diff; Table 1 · Read the study

[c6] Diff’s population-weighted duplicate classification favored recall over precision. At the deployed threshold of 0.92, precision was 0.28, recall 0.86 and F1 0.42, with respective 95% bootstrap intervals [0.21,0.35], [0.69,1.00] and [0.34,0.50]. The strict 0.97 threshold yielded precision 0.64 and recall 0.26; the loose 0.80 threshold yielded precision 0.08 and recall 1.00. These are duplicate-detection metrics, not answer accuracy.

§4.3, Table 2 · Read the study

[c7] Usage counts represent a floor rather than a complete activity distribution. No controlled comparison withheld the archive or external headlines, so their contribution to turning retrieved material into leads rests on qualitative evidence rather than a causal test.

§5.1, Limitations, first paragraph · Read the study

[c8] Diff validation followed deployment, was not an ablation study and did not compare text-only or visual-only variants. Its threshold was manually tuned rather than trained against preregistered ground truth. Recall concerns each page’s nearest composite neighbor, not discovery of every duplicate anywhere in the corpus.

§5.1, Limitations, second and third paragraphs · Read the study