Compressed-VFL: Communication-Efficient Learning with Vertically Partitioned Data
Timothy J. CastigliaAnirban DasShiqiang WangStacy Patterson
Establishes theoretical convergence guarantees and practical techniques for vertical federated learning with compressed intermediate embeddings, cutting communication costs by over 90% without sacrificing model accuracy.
Organizations such as hospitals, banks, and insurers frequently need to train collaborative machine learning models on the same set of individuals without sharing sensitive, locally held features. This setup, known as vertical federated learning, enables collaborative modeling across partitioned feature sets while preserving data privacy and complying with strict regulations. However, exchanging the necessary intermediate vector representations, known as embeddings, across distributed networks creates massive communication bottlenecks that can require terabytes of bandwidth and significantly slow down training.
The article aims to evaluate whether introducing message compression and multiple local training updates to vertically partitioned federated learning can substantially reduce communication overhead without degrading model convergence or predictive accuracy.
To evaluate this framework, named Compressed Vertical Federated Learning, the authors established theoretical convergence proofs for complex, non-linear server models and non-convex objectives. They derived specific parameter bounds for common compression methods, including uniform scalar quantization, lattice vector quantization, and top-k sparsification. They then conducted extensive empirical experiments across diverse benchmark datasets, including healthcare records (MIMIC-III), 3D computer-aided design views (ModelNet10), and image classification sets (CIFAR-10 and ImageNet-100), simulating different party counts and network latencies.
The study established several key findings. First, theoretical analysis proved that the compressed vertical framework maintains standard convergence rates when compression error is properly bounded over training. Second, empirical tests demonstrated that compressing embeddings down to as few as 2 to 4 bits per component reduced overall communication costs by over 90% compared to uncompressed baselines while attaining nearly identical predictive accuracy and F1-scores. Third, across the evaluated compression schemes, 2-dimensional lattice vector quantization consistently delivered the strongest performance and reconstruction stability. Finally, executing multiple local iterations per communication round significantly cut the elapsed training time needed to hit target accuracy metrics, delivering the greatest speedups in high-latency network environments.
These findings indicate that organizations can train high-capacity vertical federated models across geographically dispersed entities at a fraction of standard bandwidth costs and time delays. By compressing intermediate representations rather than transmitting full-precision data, teams can overcome infrastructure limitations without compromising model quality or regulatory compliance.
Decision-makers should consider adopting compressed vertical federated learning when deploying distributed models across feature-partitioned organizations, prioritizing lattice vector quantization as the primary compression technique. Engineering teams should tune the number of local update iterations based on prevailing network latency, opting for higher local iteration counts when operating over high-latency connections. Additional pilot validation is advised before applying the approach in environments with unconstrained or unbounded embedding distributions, and future evaluations should explore adaptive compression mechanisms.
- Paper: Federated Machine Learning, Qiang Yang et al. (2019). Its distinction between horizontal and vertical federated learning provides the setup needed to understand why this paper compresses exchanged representations rather than conventional model updates.
- Paper: Federated Learning: Strategies for Improving Communication Efficiency, Jakub Konečný et al. (2016). Its account of structured and sketched update compression establishes the communication-reduction approaches that motivate compressing messages in federated training.
- Paper: Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training, Yujun Lin et al. (2018). Its analysis of aggressive compression and convergence safeguards gives useful grounding for the source’s bounds on compression error.
- Paper: Communication-Efficient Learning of Deep Networks from Decentralized Data, H. B. McMahan et al. (2016). Its FederatedAveraging framework explains how multiple local updates reduce communication rounds, a strategy the source adapts to vertically partitioned learning.
No sufficiently relevant recommendations were found.
