Duplicate Record Detection: A Survey
Ahmed K. ElmagarmidPanagiotis G. IpeirotisVassilios S. Verykios
Synthesizes decades of research on duplicate record detection by systematically analyzing field-level similarity metrics, multi-attribute matching algorithms, scalability methods, and practical data cleaning tools.
Modern operational environments and business systems rely heavily on high-quality database information, yet data quality is frequently degraded by human error, inconsistent conventions, and the lack of universal identifiers across disparate systems. When consolidating records, organizations face substantial costs and operational risks if multiple, non-identical entries referring to the same real-world entity are left unresolved. This article presents an extensive survey of the duplicate record detection literature, evaluating how organizations can systematically identify, match, and merge approximate duplicate records across structured database environments.
The article synthesizes decades of foundational research across database management, artificial intelligence, and statistics to compare field-matching metrics, multi-field record matching algorithms, and computational scaling techniques. By reviewing the evolution from classical probabilistic record linkage to advanced machine learning and relational database indexing, the article maps how different methodological paradigms balance matching accuracy against execution speed.
The findings establish that duplicate record detection requires a structured pipeline starting with essential data preparation—including parsing, transformation, and standardization—which resolves surface-level structural variations before record comparison begins. For field-level comparisons, character-based metrics excel at catching typographical mistakes, token-based approaches best handle word rearrangements, and phonetic encodings mitigate pronunciation-based errors, with hybrid token-frequency metrics demonstrating superior overall performance. For multi-field matching, probabilistic and supervised machine learning models achieve the highest matching accuracy, though they depend heavily on labeled training data or manual clerical reviews for ambiguous edge cases. In contrast, ad hoc database methods and distance-based heuristics run significantly faster and scale to millions of records, but sacrifice matching precision. Finally, computational bottlenecks are effectively mitigated through efficiency techniques such as blocking, sorted neighborhood sliding windows, and canopy clustering, which prevent exhaustive pairwise record comparisons.
These insights demonstrate that no single algorithm or distance metric provides an optimal solution across all data environments. High-accuracy statistical and probabilistic models remain computationally prohibitive for massive datasets, while simple rule-based and distance metrics risk accumulating errors if deployed without careful domain adaptation. The persistence of duplicate records degrades analytical reporting, increases operating costs, and can cause systemic failures across customer relationship management, healthcare record linkage, and enterprise compliance.
Organizations addressing data deduplication should deploy a multi-stage approach: first standardize and clean incoming data, apply computationally light filtering methods like canopies or blocking to narrow candidate pairs, and then deploy domain-adapted matching algorithms for detailed evaluation. Active learning tools should be leveraged to minimize the expensive manual effort required to label training data. Future strategic efforts must prioritize the development of standardized, large-scale benchmark datasets to rigorously compare matching models and the creation of adaptive, continuous monitoring pipelines capable of handling noisy data extracted from web and text sources.
While the article provides high confidence in its comparative architectural findings, it notes clear limitations: existing evaluations largely depend on small, proprietary, or domain-specific datasets, and current duplicate detection techniques for non-textual numeric fields remain primitive. Decision-makers should validate candidate deduplication pipelines on representative organizational data before full-scale deployment.
- Paper: Winnowing: local algorithms for document fingerprinting, S. Schleimer et al. (2003). Read this earlier fingerprinting method first to understand the local-hash techniques and detection guarantees that inform approaches to finding duplicate text.
- Paper: Deduplicating Training Data Makes Language Models Better, Katherine Lee et al. (2022). This study carries deduplication into large-scale language-model training, testing scalable exact and approximate matching methods and measuring their effects on model quality and memorization.
- Paper: The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data Only, Guilherme Penedo et al. (2023). This work extends deduplication to a web-scale training-data pipeline, combining exact and fuzzy matching with other cleaning stages and testing the resulting corpus through language-model training.
- Paper: DataComp-LM: In search of the next generation of training sets for language models, Jeffrey Li et al. (2024). This benchmark continues the study of scalable dataset curation by isolating deduplication alongside extraction, filtering, and data mixing in controlled language-model experiments.
- Paper: Deduplicating Training Data Mitigates Privacy Risks in Language Models, Nikhil Kandpal et al. (2022). This study applies sequence-level deduplication to a consequential downstream problem, measuring how removing repeated training data changes privacy leakage from language models.
