LightEA: A Scalable, Robust, and Interpretable Entity Alignment Framework via Three-view Label Propagation
Xin MaoWenting WangYuanbin WuMan Lan
Proposes a non-neural entity alignment framework using three-view label propagation and sparse Sinkhorn iteration to achieve state-of-the-art alignment accuracy with a fraction of the computational time and full model interpretability.
Knowledge graphs organize real-world information into structured networks of entities and relationships, powering critical tools such as search engines and dialogue systems. Because these graphs are typically developed independently across different organizations and sources, integrating them through entity alignment—identifying matching entities across distinct graphs—is essential for expanding knowledge coverage. However, prevailing alignment techniques rely heavily on complex graph neural networks that require heavy iterative model training. Consequently, existing methods suffer from prohibitive computation times when applied to large real-world datasets and operate largely as uninterpretable black boxes.
The article introduces and evaluates LightEA, a non-neural framework designed to achieve highly scalable, robust, and interpretable entity alignment. LightEA adapts the classical label propagation technique to heterogeneous graphs using three primary components: generating compact random orthogonal label vectors to represent known alignments, propagating entity and relation labels across three structural graph perspectives, and resolving one-to-one entity matches using a sparse assignment algorithm.
To establish its effectiveness, the framework was evaluated across four public benchmark suites representing diverse conditions, including cross-lingual pairs, sparse structures, medium-sized graphs with 100,000 entities, and a million-scale dataset comprising over one million entities and nearly ten million relationships. Performance was measured on alignment accuracy and execution runtime against state-of-the-art neural and non-neural baselines.
The evaluation revealed several key findings. First, LightEA delivered substantial execution speedups, requiring under one-tenth the runtime of leading methods—completing alignment on standard benchmarks in 7 to 35 seconds and processing the million-entity dataset in under four minutes on a single graphic processing unit, where many existing systems fail to run entirely. Second, despite eliminating trainable parameters, the framework matched or surpassed the accuracy of leading neural approaches across benchmark datasets. Third, incorporating literal features, such as translated entity names, boosted top-one alignment accuracy above 95% on cross-lingual benchmarks without requiring pre-aligned training pairs. Finally, because the label propagation mechanism is entirely linear, alignment outcomes and errors can be directly audited and traced back to specific graph neighborhoods.
These findings indicate that complex neural architectures are not strictly necessary to achieve high-accuracy graph integration. By replacing intensive neural network training with linear label propagation, organizations can drastically cut computational hardware costs, accelerate integration timelines, and audit alignment decisions for higher confidence and accountability. When deploying the framework, practitioners should tailor the configuration to available resources: use the basic configuration when speed is paramount, apply iterative alignment when pre-labeled data is scarce, or incorporate textual names when textual attributes are present. Future research should prioritize refactoring the framework into high-performance computing languages, testing linear speedups across multi-processor environments, and refining ways to preserve model interpretability during large-scale execution without excessive memory overhead.
- Paper: LightGCN: Simplifying and Powering Graph Convolution Network for Recommendation, Xiangnan He et al. (2020). LightGCN establishes the paradigm that stripping away non-linearities and neural parameters in favor of pure linear neighborhood aggregation dramatically boosts scalability and efficiency in graph representation learning, directly inspiring LightEA's non-neural label propagation approach.
- Paper: Modeling Relational Data with Graph Convolutional Networks, Michael Schlichtkrull et al. (2018). This foundational paper introduces relational message-passing over heterogeneous knowledge graphs, providing the primary neural baseline architecture and relational representation concepts that LightEA seeks to replace with efficient linear propagation.
- Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). This work establishes first-order semi-supervised spectral graph convolutions, providing the mathematical framework for localized graph propagation that underlies both GCN baselines and simplified label propagation techniques.
- Paper: Revisiting Semi-Supervised Learning with Graph Embeddings, Zhilin Yang et al. (2016). This paper explores the foundational mechanics of semi-supervised learning and label propagation on graph structures, which LightEA extends to multi-view entity alignment without iterative parameter training.
- Paper: Translating Embeddings for Modeling Multi-relational Data, Antoine Bordes et al. (2013). TransE introduces the fundamental translation-based knowledge graph representation paradigm that benchmarked entity alignment and relational structure learning.
- Paper: A Survey on Knowledge Graphs: Representation, Acquisition, and Applications, Shaoxiong Ji et al. (2020). This survey provides a structured overview of knowledge graph representation learning, completion, and cross-graph integration challenges addressed by LightEA.
- Paper: Simple and Efficient Heterogeneous Graph Neural Network, Xiaocheng Yang et al. (2023). This paper extends the pursuit of parameter-free and efficient heterogeneous graph modeling by removing neighbor attention bottlenecks in complex multi-relational graphs.
- Paper: Linkless Link Prediction via Relational Distillation, Zhichun Guo et al. (2023). This work continues the trend toward fast, scalable graph inference by distilling complex relational graph representations into ultra-fast non-relational architectures.
- Paper: PRODIGY: Enabling In-context Learning Over Graphs, Qian Huang et al. (2023). This work advances beyond fixed graph training by investigating zero-parameter-update in-context learning over heterogeneous knowledge graphs.
