A Statutory Article Retrieval Dataset in French
Antoine LouisGerasimos Spanakis
Introduces the first native French statutory article retrieval benchmark containing expert-annotated legal questions paired with Belgian law statutes to evaluate dense and lexical information retrieval models on complex legal text.
Most citizens lack legal expertise and increasingly turn to standard internet search engines for guidance when confronted with legal issues. However, online searches frequently direct users to commercial services rather than practical answers, leaving those who cannot afford costly legal counsel underserved. Automated statutory article retrieval—matching layperson questions to relevant statutory law—can help bridge this access-to-justice gap, but progress has been stalled by a lack of native, large-scale, and expert-annotated datasets.
The article evaluates the feasibility of automated statutory retrieval by introducing and benchmarking the Belgian Statutory Article Retrieval Dataset (BSARD). The main objective is to establish how modern lexical and dense semantic retrieval models perform when matching ordinary natural language legal questions against a broad body of statutory provisions.
To construct BSARD, the authors gathered 22,633 Belgian statutory articles across 32 legal codes and paired them with 1,108 French-language questions collected between 2018 and 2021 by Droits Quotidiens, a legal advice organization. A team of six experienced jurists formulated and labeled these questions with exact statutory references. The dataset was split into training and test sets and used to benchmark baseline information retrieval methods, including traditional lexical algorithms (TF-IDF and BM25), off-the-shelf zero-shot embedding models, and supervised deep neural architectures.
The findings show that fine-tuned dense neural retrieval systems significantly outperform traditional and zero-shot approaches. The best-performing model—a supervised two-tower neural network using the French language model CamemBERT—achieved a recall of 74.8% within the top 100 retrieved articles, compared to 51.3% for optimized BM25 and only 4.2% for off-the-shelf CamemBERT. Traditional keyword matching through BM25 proved to be a respectable baseline, outperforming zero-shot neural models, which struggle because non-expert queries use very different vocabulary than formal statutes. Furthermore, out-of-the-box word-level embeddings such as word2vec (49.4% recall at 100) outperformed pretrained contextual transformers when no task-specific fine-tuning was applied.
These results demonstrate that semantic search models can successfully navigate the linguistic gap between citizen questions and formal statutory text, provided they undergo specialized supervised training. Deploying such systems can lower the cost and operational barrier of public legal assistance and improve the efficiency of legal research. However, because the top-performing recall of approximately 75% remains below the performance of human jurists, automated models cannot yet operate entirely without human supervision.
To advance the technology, research teams should focus on architectural improvements that account for the hierarchical structure of legal codes and expand neural models to handle lengthy statutory articles exceeding typical input limits. Decision-makers and developers aiming to support access to justice can leverage BSARD under its non-commercial open license to build public-interest assistive tools while ensuring models are properly monitored against potential misuse.
Key limitations include the dataset's focus on 32 core statutory codes, which excludes regional decrees, ordinances, and judicial case law needed for certain complex inquiries. In addition, the legal texts represent a snapshot of Belgian law as of May 2021, meaning the dataset serves as a benchmark for research and development rather than an up-to-date tool for live legal advice. While confidence in the benchmark evaluation is high, practical implementations require careful validation against current legislation and broader legal sources.
No sufficiently relevant recommendations were found.
- Paper: Legal Retrieval for Public Defenders, Dominik Stammbach et al. (2026). It carries the source’s evaluation of retrieval for real legal questions into public defenders’ practice, testing retrieval over their working materials with expert queries and relevance judgments.
