A Statutory Article Retrieval Dataset in French

Antoine LouisGerasimos Spanakis

article2022ACL57 citations

Introduces the first native French statutory article retrieval benchmark containing expert-annotated legal questions paired with Belgian law statutes to evaluate dense and lexical information retrieval models on complex legal text.

Listen

Most citizens lack legal expertise and increasingly turn to standard internet search engines for guidance when confronted with legal issues. However, online searches frequently direct users to commercial services rather than practical answers, leaving those who cannot afford costly legal counsel underserved. Automated statutory article retrieval—matching layperson questions to relevant statutory law—can help bridge this access-to-justice gap, but progress has been stalled by a lack of native, large-scale, and expert-annotated datasets.

The article evaluates the feasibility of automated statutory retrieval by introducing and benchmarking the Belgian Statutory Article Retrieval Dataset (BSARD). The main objective is to establish how modern lexical and dense semantic retrieval models perform when matching ordinary natural language legal questions against a broad body of statutory provisions.

To construct BSARD, the authors gathered 22,633 Belgian statutory articles across 32 legal codes and paired them with 1,108 French-language questions collected between 2018 and 2021 by Droits Quotidiens, a legal advice organization. A team of six experienced jurists formulated and labeled these questions with exact statutory references. The dataset was split into training and test sets and used to benchmark baseline information retrieval methods, including traditional lexical algorithms (TF-IDF and BM25), off-the-shelf zero-shot embedding models, and supervised deep neural architectures.

The findings show that fine-tuned dense neural retrieval systems significantly outperform traditional and zero-shot approaches. The best-performing model—a supervised two-tower neural network using the French language model CamemBERT—achieved a recall of 74.8% within the top 100 retrieved articles, compared to 51.3% for optimized BM25 and only 4.2% for off-the-shelf CamemBERT. Traditional keyword matching through BM25 proved to be a respectable baseline, outperforming zero-shot neural models, which struggle because non-expert queries use very different vocabulary than formal statutes. Furthermore, out-of-the-box word-level embeddings such as word2vec (49.4% recall at 100) outperformed pretrained contextual transformers when no task-specific fine-tuning was applied.

These results demonstrate that semantic search models can successfully navigate the linguistic gap between citizen questions and formal statutory text, provided they undergo specialized supervised training. Deploying such systems can lower the cost and operational barrier of public legal assistance and improve the efficiency of legal research. However, because the top-performing recall of approximately 75% remains below the performance of human jurists, automated models cannot yet operate entirely without human supervision.

To advance the technology, research teams should focus on architectural improvements that account for the hierarchical structure of legal codes and expand neural models to handle lengthy statutory articles exceeding typical input limits. Decision-makers and developers aiming to support access to justice can leverage BSARD under its non-commercial open license to build public-interest assistive tools while ensuring models are properly monitored against potential misuse.

Key limitations include the dataset's focus on 32 core statutory codes, which excludes regional decrees, ordinances, and judicial case law needed for certain complex inquiries. In addition, the legal texts represent a snapshot of Belgian law as of May 2021, meaning the dataset serves as a benchmark for research and development rather than an up-to-date tool for live legal advice. While confidence in the benchmark evaluation is high, practical implementations require careful validation against current legislation and broader legal sources.

arXiv: 2108.11792maastrichtlawtech/bsard

No sufficiently relevant recommendations were found.

  • Paper: Legal Retrieval for Public Defenders, Dominik Stammbach et al. (2026). It carries the source’s evaluation of retrieval for real legal questions into public defenders’ practice, testing retrieval over their working materials with expert queries and relevance judgments.
Cover for A Statutory Article Retrieval Dataset in French

Abstract

Statutory article retrieval is the task of automatically retrieving law articles relevant to a legal question. While recent advances in natural language processing have sparked considerable interest in many legal tasks, statutory article retrieval remains primarily untouched due to the scarcity of large-scale and high-quality annotated datasets. To address this bottleneck, we introduce the Belgian Statutory Article Retrieval Dataset (BSARD), which consists of 1,100+ French native legal questions labeled by experienced jurists with relevant articles from a corpus of 22,600+ Belgian law articles. Using BSARD, we benchmark several state-of-the-art retrieval approaches, including lexical and dense architectures, both in zero-shot and supervised setups. We find that fine-tuned dense retrieval models significantly outperform other systems. Our best performing baseline achieves 74.8% R@100, which is promising for the feasibility of the task and indicates there is still room for improvement. By the specificity of the domain and addressed task, BSARD presents a unique challenge problem for future research on legal information retrieval. Our dataset and source code are publicly available.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The Belgian Statutory Article Retrieval Dataset
  • 3.1 Dataset Collection
  • 3.2 Dataset Analysis
  • 4 Models
  • 4.1 Lexical Models
  • 4.2 Dense Models
  • 4.2.1 Zero-Shot Evaluation
  • 4.2.2 Training
  • 5 Experiments
  • 5.1 Experimental Setup
  • 5.2 Results
  • 6 Discussion
  • 6.1 Limitations
  • 6.2 Broader Impacts
  • 7 Conclusion
  • Acknowledgments
  • References
  • Appendix
  • A Legal Codes
  • B Evaluation Metrics
  • C Dataset Documentation
  • C.1 Dataset Nutrition Labels
  • C.2 Data Statement
  • C.3 Intended Uses
  • C.4 Hosting
  • C.5 Data Format
  • C.6 Reproducibility
  • C.7 Licensing
  • C.8 Maintenance

Knowls

  1. Knowl 1 — BSARD: a native French benchmark for statutory article retrieval

    definition

    The Belgian Statutory Article Retrieval Dataset (BSARD) is a benchmark for retrieving Belgian statutory articles relevant to a legal question expressed in French. It contains 1,108 anonymized, layperson-oriented questions based on recurring issues raised by Belgian citizens, paired with article identifiers judged relevant by experienced Belgian jurists, and a corpus of 22,633 articles from 32 Belgian legal codes. The questions and articles are in Belgian French as used in Wallonia and Brussels-Capital. Article records include the article text and code-related metadata; question records include topic tags and relevant article IDs. The dataset was released under the CC BY-NC-SA 4.0 license for non-commercial use.

  2. Knowl 2 — Construction of the questions, article corpus, and relevance labels

    model/method

    BSARD was assembled by collecting articles from 32 publicly available Belgian codes and extracting their hierarchical headings (such as book, chapter, and section names) alongside the article text. The article text was cleaned with regular expressions to remove amendment-related markup, and repealed articles were excluded, leaving 22,633 retrieval units. For questions, the authors partnered with Droits Quotidiens, which supplied more than 3,200 anonymized model questions, legal references, and categorization tags based on recurring citizen issues; jurists formulate these model questions in natural language rather than providing verbatim citizen emails. Roughly one-third of the collected questions were duplicates. Where tags supplied personal or location context, those descriptions were prepended to the question to distinguish duplicates and make the question more specific. The authors parsed the legal references and retained only questions whose references could be matched to articles in the selected codes; questions referring to non-article materials, uncollected legal acts, or codes outside the corpus were filtered out. The result is 1,108 questions, split into 886 training and 222 test questions.

  3. Knowl 3 — BSARD question and relevance distribution

    data/table

    The questions and their labels are unevenly distributed across topics and across the corpus. Questions are 5–44 words long (median 14); articles have a median length of 77 words, and 142 articles exceed 1,000 words, with the longest at 5,790 words. The topic breakdown is: family 30.6%, housing 27.4%, money 16.0%, justice 13.6%, foreigners 5.7%, social security 3.5%, and work 3.2%. In terms of labels, 75% of questions have fewer than five relevant articles, 18% have between 5 and 20, and 7% have more than 20; the maximum is 109 relevant articles for one question. Only 1,612 of the 22,633 corpus articles are labeled relevant to any question. Among those relevant articles, the median citation count across questions is 2, and fewer than 25% are cited more than five times. About 80% of the 1,612 relevant articles come from the Civil Code, Judicial Code, Code of Criminal Instruction, or Penal Code.

  4. Knowl 4 — Formalization of statutory article retrieval

    definition

    For a legal question qq and a corpus CC of statutory articles, a retrieval system returns a ranked subset F⊂CF \subset C containing articles it predicts to be relevant. In the benchmark, the system is evaluated by retrieving a fixed top-kk subset, where kk is much smaller than the corpus size, and comparing the ranking with the question’s expert-assigned relevant article IDs. The task therefore evaluates retrieval of potentially multiple relevant articles, rather than selection of a single answer.

  5. Knowl 5 — Dense bi-encoder baselines and zero-shot encoding

    model/method

    The dense retrievers encode a question and a statute article as pooled vectors and rank articles by vector similarity. The paper evaluates both a siamese architecture, in which the same encoder is used for questions and articles, and a two-tower architecture, in which separate encoders are used. For zero-shot evaluation, the siamese encoders use pretrained French word2vec (500-dimensional skip-gram), fastText (300-dimensional CBOW), or CamemBERT representations without retrieval-task fine-tuning. All zero-shot representations use mean pooling and cosine similarity. Because CamemBERT accepts at most 512 tokens per input, longer texts are split into 200-token chunks with 20-token overlap, and the chunk representations are pooled to represent the full text. For retrieval, article vectors are precomputed and each question is compared against the corpus articles.

  6. Knowl 6 — Fine-tuning dense retrievers with in-batch negatives

    algorithm

    The paper fine-tunes siamese and two-tower CamemBERT retrievers using a contrastive objective. A training instance is a question qiq_i paired with a relevant article ai+a_i^+; for a mini-batch, the positive articles paired with the other questions in that batch serve as negatives Ai−A_i^-. The model is trained to score the positive above the negatives using dot-product similarity sDs_D and temperature τ\tau:

    L(qi,ai+,Ai−)=−log⁡exp⁡(sD(qi,ai+)/τ)∑a∈Ai−∪{ai+}exp⁡(sD(qi,a)/τ).L(q_i,a_i^+,A_i^-)=-\log\frac{\exp(s_D(q_i,a_i^+)/\tau)}{\sum_{a\in A_i^-\cup\{a_i^+\}}\exp(s_D(q_i,a)/\tau)}.

    Here, LL is the per-instance loss, aa ranges over the positive and in-batch negative articles, and τ=0.05\tau=0.05. Training used batches of 22 question–article pairs for 100 epochs (about 20,500 steps), with AdamW, initial learning rate 2×10−52\times10^{-5}, β1=0.9\beta_1=0.9, β2=0.999\beta_2=0.999, weight decay 0.01, a 500-step warm-up, and linear learning-rate decay. For long texts, the model uses 200-token chunks with 20-token overlap. During training, article text is limited to its first 1,000 words; the article representation is the average of the chunk-level output [CLS] embeddings. This cap affects about 0.6% of corpus articles. The model was trained on a single 32-GB Tesla V100 GPU.

  7. Knowl 7 — Lexical retrieval baselines

    model/method

    The lexical baselines assign an article aa a score for question qq by summing term weights over the query terms. The paper evaluates TF-IDF and BM25:

    sL(q,a)=∑t∈qw(t,a).s_L(q,a)=\sum_{t\in q}w(t,a).

    For TF-IDF, w(t,a)=tf⁡(t,a)log⁡(∣C∣/df⁡(t))w(t,a)=\operatorname{tf}(t,a)\log(|C|/\operatorname{df}(t)), where tf⁡(t,a)\operatorname{tf}(t,a) is the number of occurrences of term tt in article aa, df⁡(t)\operatorname{df}(t) is the number of corpus articles containing tt, and ∣C∣|C| is the number of articles. The BM25 weight is

    w(t,a)=tf⁡(t,a)(k1+1)tf⁡(t,a)+k1(1−b+b∣a∣avglen⁡)log⁡∣C∣−df⁡(t)+0.5df⁡(t)+0.5,w(t,a)=\frac{\operatorname{tf}(t,a)(k_1+1)}{\operatorname{tf}(t,a)+k_1\left(1-b+b\frac{|a|}{\operatorname{avglen}}\right)}\log\frac{|C|-\operatorname{df}(t)+0.5}{\operatorname{df}(t)+0.5},

    where ∣a∣|a| is the length of article aa, avglen⁡\operatorname{avglen} is the average article length, and k1k_1 and bb are BM25 parameters. The values selected on the training set were k1=1.0k_1=1.0 and b=0.6b=0.6. Each lexical system scores the corpus and returns articles in descending score order.

  8. Knowl 8 — Evaluation protocol and metrics

    experimental setup

    Models were evaluated on the 222-question BSARD test split against the full 22,633-article corpus. The reported ranking metrics are macro-averaged recall at kk (R@kk), mean average precision at kk (MAP@kk), and mean reciprocal rank at kk (MRR@kk), with k∈{100,200,500}k\in\{100,200,500\} for recall and k=100k=100 for MAP and MRR. Recall measures the fraction of a question’s relevant articles retrieved within the cutoff; average precision summarizes precision at the ranks where relevant articles appear, normalized by the total number of relevant articles; reciprocal rank is the reciprocal of the rank of the first relevant article. Scores are averaged across questions. Precision@kk was not reported because questions have different numbers of relevant articles. The comparison also reports query latency in milliseconds.

  9. Knowl 9 — Retrieval performance on the BSARD test set

    empirical result

    Fine-tuned CamemBERT bi-encoders achieved the strongest overall retrieval results among the evaluated systems. The two-tower model had the best R@100 (74.78%) and MAP@100 (35.67%); the siamese model was slightly higher at R@200 (78.38%), R@500 (83.77%), and MRR@100 (43.52%). Both trained models substantially exceeded the lexical and zero-shot baselines. The table reports retrieval scores in percent, parameter counts, and query latency in milliseconds; “Train” indicates whether the model was fine-tuned on BSARD.

    Train Model / encoder Params Latency (ms) R@100 R@200 R@500 MAP@100 MRR@100
    No TF-IDF – 827 40.13 50.44 59.34 8.69 12.98
    No BM25 (official) – 1342 51.33 56.78 64.71 16.04 24.59
    No Siamese bi-encoder / word2vec – 4 49.41 61.76 71.57 12.90 21.49
    No Siamese bi-encoder / fastText – 3 32.93 41.33 49.26 6.29 11.78
    No Siamese bi-encoder / CamemBERT – 27 4.21 6.00 12.82 0.50 2.04
    Yes Siamese bi-encoder / CamemBERT 110M 28 71.63 78.38 83.77 35.44 43.52
    Yes Two-tower bi-encoder / CamemBERT 220M 26 74.78 78.04 83.39 35.67 42.46

    Among the zero-shot siamese encoders, word2vec was strongest, while un-fine-tuned CamemBERT performed poorly. This comparison indicates that pretrained embeddings alone do not guarantee good retrieval performance and that retrieval-task fine-tuning is important for the dense models.

  10. Knowl 10 — Scope and limitations of BSARD

    limitation

    BSARD covers articles from 32 selected Belgian codes, not all Belgian law. Decrees, directives, ordinances, and other uncollected legal acts are missing; questions referencing such materials were filtered out, so some retained questions may have only a subset of their original relevant provisions labeled. The remaining articles may still be appropriate but may not fully support the answer. In addition, some legal questions cannot be resolved from statutory text alone and may require case law or fact-specific reasoning, so suitability for statutory retrieval varies by question. The dataset reflects questions and law as collected in May 2021 and is intended as a research benchmark, not as an up-to-date source of legal advice.

Coverage note — The paper’s dataset hosting, CSV-format, and maintenance documentation is omitted because it is operational rather than a substantive retrieval contribution; the broader-impact discussion is not separated into a knowl because it does not add a distinct method or empirical finding.

References

  1. 1.Joshua Ainslie, Santiago Ontañón, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. 2020. ETC: encoding long and structured inputs in transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 268–284. Association for Computational Linguistics.
  2. 2.Nikolaos Aletras, Dimitrios Tsarapatsanis, Daniel Preotiuc-Pietro, and Vasileios Lampos. 2016. Predicting judicial decisions of the european court of human rights: a natural language processing perspective. PeerJ Computer Science, 2:e93.
  3. 3.Nigel J Balmer, Alexy Buck, Ash Patel, Catrina Denvir, and Pascoe Pleasence. 2010. Knowledge, capability and the experience of rights problems. London: PLEnet.
  4. 4.Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. CoRR, abs/2004.05150.
  5. 5.Emily M. Bender and Batya Friedman. 2018. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6:587–604.
  6. 6.Adam L. Berger, Rich Caruana, David Cohn, Dayne Freitag, and Vibhu O. Mittal. 2000. Bridging the lexical chasm: statistical approaches to answer-finding. In SIGIR 2000: Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 192–199. ACM.
  7. 7.Paheli Bhattacharya, Kaustubh Hiware, Subham Rajgaria, Nilay Pochhi, Kripabandhu Ghosh, and Saptarshi Ghosh. 2019. A comparative study of summarization algorithms applied to legal case judgments. In Advances in Information Retrieval - 41st European Conference on IR Research, volume 11437 of Lecture Notes in Computer Science, pages 413–428. Springer.
  8. 8.Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomás Mikolov. 2017. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5:135–146.
  9. 9.Lukasz Borchmann, Dawid Wisniewski, Andrzej Gretkowski, Izabela Kosmala, Dawid Jurkiewicz, Lukasz Szalkiewicz, Gabriela Palka, Karol Kaczmarek, Agnieszka Kaliska, and Filip Gralinski. 2020. Contract discovery: Dataset and a few-shot semantic retrieval challenge with competitive baselines. In Findings of the Association for Computational Linguistics: EMNLP 2020, volume EMNLP 2020 of Findings of ACL, pages 4254–4268. Association for Computational Linguistics.
  10. 10.Ilias Chalkidis, Ion Androutsopoulos, and Nikolaos Aletras. 2019a. Neural legal judgment prediction in english. In Proceedings of the 57th Conference of the Association for Computational Linguistics, pages 4317–4323. Association for Computational Linguistics.
  11. 11.Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis, and Ion Androutsopoulos. 2019b. Large-scale multi-label text classification on EU legislation. In Proceedings of the 57th Conference of the Association for Computational Linguistics, pages 6314–6322. Association for Computational Linguistics.
  12. 12.Huajie Chen, Deng Cai, Wei Dai, Zehui Dai, and Yadong Ding. 2019. Charge-based prison term prediction with deep gating network. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, pages 6361–6366. Association for Computational Linguistics.
  13. 13.Ting Chen, Yizhou Sun, Yue Shi, and Liangjie Hong. 2017. On sampling strategies for neural network-based collaborative filtering. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 767–776.
  14. 14.Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M. Voorhees. 2020. Overview of the TREC 2019 deep learning track. CoRR, abs/2003.07820.
  15. 15.Catrina Denvir. 2016. Online and in the know? Public legal education, young people and the internet. Computers & Education, 92-93:204–220.
  16. 16.Xingyi Duan, Baoxin Wang, Ziyue Wang, Wentao Ma, Yiming Cui, Dayong Wu, Shijin Wang, Ting Liu, Tianxiang Huo, Zhen Hu, Heng Wang, and Zhiyuan Liu. 2019. CJRC: A reliable human-annotated benchmark dataset for chinese judicial reading comprehension. In 18th China National Conference on Chinese Computational Linguistics, volume 11856 of Lecture Notes in Computer Science, pages 439–451. Springer.
  17. 17.Leigh S Estabrook, G Evans Witt, and Harrison Rainie. 2007. Information searches that solve problems: How people use the Internet, libraries, and government agencies when they need help. Pew Internet & American Life Project.
  18. 18.Jean-Philippe Fauconnier. 2015. French word embeddings.
  19. 19.Biralatei Fawei, Adam Z. Wyner, and Jeff Z. Pan. 2016. Passing a USA national bar exam: a first corpus for experimentation. In Proceedings of the Tenth International Conference on Language Resources and Evaluation, pages 3373–3378. European Language Resources Association (ELRA).
  20. 20.Daniel Gillick, Alessandro Presta, and Gaurav Singh Tomar. 2018. End-to-end retrieval in continuous space. CoRR, abs/1811.08008.
  21. 21.Edouard Grave, Piotr Bojanowski, Prakhar Gupta, Armand Joulin, and Tomás Mikolov. 2018. Learning word vectors for 157 languages. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation, LREC 2018. European Language Resources Association (ELRA).
  22. 22.Ben Hachey and Claire Grover. 2006. Extractive summarisation of legal texts. Artificial Intelligence and Law, 14(4):305–345.
  23. 23.Margaret Hagan and Yue Li. 2020. Legal help search audit: Are search engines effective brokers of legal information? Available at SSRN 3623333.
  24. 24.Hamza Harkous, Kassem Fawaz, Rémi Lebret, Florian Schaub, Kang G. Shin, and Karl Aberer. 2018. Polisis: Automated analysis and presentation of privacy policies using deep learning. In 27th USENIX Security Symposium, pages 531–548. USENIX Association.
  25. 25.Matthew L. Henderson, Rami Al-Rfou, Brian Strope, Yun-Hsuan Sung, László Lukács, Ruiqi Guo, Sanjiv Kumar, Balint Miklos, and Ray Kurzweil. 2017. Efficient natural language response suggestion for smart reply. CoRR, abs/1705.00652.
  26. 26.Dan Hendrycks, Collin Burns, Anya Chen, and Spencer Ball. 2021. CUAD: An expert-annotated nlp dataset for legal contract review. In Advances in Neural Information Processing Systems 31.
  27. 27.Sarah Holland, Ahmed Hosny, Sarah Newman, Joshua Joseph, and Kasia Chmielinski. 2018. The dataset nutrition label: A framework to drive higher data quality standards. arXiv preprint arXiv:1805.03677.
  28. 28.Nils Holzenberger, Andrew Blair-Stanek, and Benjamin Van Durme. 2020. A dataset for statutory reasoning in tax law entailment and question answering. In Proceedings of the Natural Legal Language Processing Workshop 2020 co-located with the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD 2020), volume 2645 of CEUR Workshop Proceedings, pages 31–38. CEUR-WS.org.
  29. 29.Zikun Hu, Xiang Li, Cunchao Tu, Zhiyuan Liu, and Maosong Sun. 2018. Few-shot charge prediction with discriminative legal attributes. In Proceedings of the 27th International Conference on Computational Linguistics, pages 487–498. Association for Computational Linguistics.
  30. 30.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, pages 6769–6781. Association for Computational Linguistics.
  31. 31.Mi-Young Kim, Juliano Rabelo, and Randy Goebel. 2019. Statute law information retrieval and entailment. In Proceedings of the 6th Competition on Legal Information Retrieval and Entailment Workshop in association with the Seventeenth International Conference on Artificial Intelligence and Law, pages 283–289. ACM.
  32. 32.Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain question answering. In Proceedings of the 57th Conference of the Association for Computational Linguistics, pages 6086–6096. Association for Computational Linguistics.
  33. 33.Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021. Datasets: A community library for natural language processing. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 175–184. Association for Computational Linguistics.
  34. 34.Marco Lippi, Przemyslaw Palka, Giuseppe Contissa, Francesca Lagioia, Hans-Wolfgang Micklitz, Giovanni Sartor, and Paolo Torroni. 2019. CLAUDETTE: an automated detector of potentially unfair clauses in online terms of service. Artificial Intelligence and Law, 27(2):117–139.
  35. 35.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  36. 36.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In Proceedings of the 7th International Conference on Learning Representations.
  37. 37.Bingfeng Luo, Yansong Feng, Jianbo Xu, Xiang Zhang, and Dongyan Zhao. 2017. Learning to predict charges for criminal cases with legal basis. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2727–2736. Association for Computational Linguistics.
  38. 38.Christopher D. Manning, Prabhakar Raghavan, and Hinrich Schütze. 2008. Introduction to information retrieval. Cambridge University Press.
  39. 39.Laura Manor and Junyi Jessy Li. 2019. Plain English summarization of contracts. In Proceedings of the Natural Legal Language Processing Workshop 2019, pages 1–11. Association for Computational Linguistics.
  40. 40.Louis Martin, Benjamin Müller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, and Benoît Sagot. 2020. Camembert: a tasty french language model. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, pages 7203–7219. Association for Computational Linguistics.
  41. 41.Tomás Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013a. Efficient estimation of word representations in vector space. In 1st International Conference on Learning Representations, ICLR 2013.
  42. 42.Tomás Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013b. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems 26, pages 3111–3119.
  43. 43.Ha-Thanh Nguyen, Hai-Yen Thi Vuong, Phuong Minh Nguyen, Tran Binh Dang, Quan Minh Bui, Vu Trong Sinh, Chau Minh Nguyen, Vu D. Tran, Ken Satoh, and Minh Le Nguyen. 2020. JNLP team: Deep learning for legal processing in COLIEE 2020. CoRR, abs/2011.08071.
  44. 44.Alejandro Ponce, Sarah Chamness Long, Elizabeth Andersen, Camilo Gutierrez Patino, Matthew Harman, Jorge A Morales, Ted Piccone, Natalia Rodriguez Cajamarca, Adriana Stephan, Kirssy Gonzalez, Jennifer VanRiper, Alicia Evangelides, Rachel Martin, Priya Khosla, Lindsey Bock, Erin Campbell, Emily Gray, Amy Gryskiewicz, Ayyub Ibrahim, Leslie Solis, Gabriel Hearn-Desautels, and Francesca Tinucci. 2019. Global Insights on Access to Justice 2019: Findings from the World Justice Project General Population Poll in 101 Countries. World Justice Project.
  45. 45.Juliano Rabelo, Mi-Young Kim, Randy Goebel, Masaharu Yoshioka, Yoshinobu Kano, and Ken Satoh. 2020. COLIEE 2020: Methods for legal document retrieval and entailment. In New Frontiers in Artificial Intelligence - JSAI-isAI 2020 Workshops, JURISIN, LENLS 2020 Workshops, volume 12758 of Lecture Notes in Computer Science, pages 196–210. Springer.
  46. 46.Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, pages 3980–3990. Association for Computational Linguistics.
  47. 47.Stephen E. Robertson, Steve Walker, Susan Jones, Micheline Hancock-Beaulieu, and Mike Gatford. 1994. Okapi at TREC-3. In Proceedings of The Third Text REtrieval Conference, TREC 1994, volume 500-225 of NIST Special Publication, pages 109–126. National Institute of Standards and Technology (NIST).
  48. 48.Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  49. 49.Chaojun Xiao, Haoxi Zhong, Zhipeng Guo, Cunchao Tu, Zhiyuan Liu, Maosong Sun, Yansong Feng, Xianpei Han, Zhen Hu, Heng Wang, and Jianfeng Xu. 2018. CAIL2018: A large-scale legal dataset for judgment prediction. CoRR, abs/1807.02478.
  50. 50.Chaojun Xiao, Haoxi Zhong, Zhipeng Guo, Cunchao Tu, Zhiyuan Liu, Maosong Sun, Tianyang Zhang, Xianpei Han, Zhen Hu, Heng Wang, and Jianfeng Xu. 2019. CAIL2019-SCM: A dataset of similar case matching in legal domain. CoRR, abs/1911.08962.
  51. 51.Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021. Approximate nearest neighbor negative contrastive learning for dense text retrieval. In 9th International Conference on Learning Representations, ICLR 2021. OpenReview.net.
  52. 52.Yinfei Yang, Daniel Cer, Amin Ahmad, Mandy Guo, Jax Law, Noah Constant, Gustavo Hernández Ábrego, Steve Yuan, Chris Tar, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2020. Multilingual universal sentence encoder for semantic retrieval. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, ACL 2020, pages 87–94. Association for Computational Linguistics.
  53. 53.Haoxi Zhong, Zhipeng Guo, Cunchao Tu, Chaojun Xiao, Zhiyuan Liu, and Maosong Sun. 2018. Legal judgment prediction via topological learning. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3540–3549. Association for Computational Linguistics.
  54. 54.Haoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang, Zhiyuan Liu, and Maosong Sun. 2020. JEC-QA: A legal-domain question answering dataset. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, volume 34(05), pages 9701–9708. AAAI Press.

Citation

MLA
Louis, A., and G. Spanakis. “A Statutory Article Retrieval Dataset in French”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 6789–803, https://doi.org/10.18653/v1/2022.acl-long.468.
APA
Louis, A., & Spanakis, G. (2022). A Statutory Article Retrieval Dataset in French. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6789–6803. https://doi.org/10.18653/v1/2022.acl-long.468
Chicago
Louis, A., and G. Spanakis. 2022. “A Statutory Article Retrieval Dataset in French”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6789–6803. https://doi.org/10.18653/v1/2022.acl-long.468.
Harvard
Louis, A. and Spanakis, G. (2022) “A Statutory Article Retrieval Dataset in French”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6789–6803. Available at: https://doi.org/10.18653/v1/2022.acl-long.468.
Vancouver
1. Louis A, Spanakis G (2022) A Statutory Article Retrieval Dataset in French. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6789–6803

BibTeX

@inproceedings{louis-spanakis-2022-statutory,
    title = "A Statutory Article Retrieval Dataset in {F}rench",
    author = "Louis, Antoine  and
      Spanakis, Gerasimos",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.468/",
    doi = "10.18653/v1/2022.acl-long.468",
    pages = "6789--6803"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/