UniPROT: Uniform Prototype Selection via Partial Optimal Transport with Submodular Guarantees
Prateek ChandaPrayas AgrawalKarthik S. GurumoorthyGanesh RamakrishnanBamdev MishraPratik Jawanpuria
Introduces UniPROT, a subset selection framework that reformulates optimal transport into a submodular objective with a (1 - 1/e) greedy approximation guarantee, effectively preserving minority-class representation in imbalanced classification and language model training.
Modern artificial intelligence workflows increasingly rely on subset selection to identify small, representative groups of data points—known as prototypes—to summarize massive datasets and accelerate model training. Existing subset selection methods implicitly assign uneven importance scores to chosen examples. Under real-world conditions where data is imbalanced, this behavior disproportionately favors majority categories, leaving minority classes and rare sub-domains severely underrepresented and poorly modeled.
The article introduces and evaluates UniPROT, a framework designed to select an equally weighted, highly representative prototypical subset from source data to match a target data distribution. Its primary objective is to demonstrate that enforcing uniform prototype importance improves minority representation and training efficiency across imbalanced computer vision and natural language processing tasks.
To overcome the computational intractability of enforcing equal prototype weights, the authors reformulate the underlying optimal transport problem using a partial optimal transport approach with submodular properties. This reformulation mathematically proves that a fast, greedy selection algorithm can achieve a guaranteed theoretical approximation ratio while maintaining computational costs comparable to standard clustering techniques. The authors validate UniPROT across multiple settings, including nearest-neighbor image classification on skewed benchmarks such as MNIST and CIFAR-10, parameter-efficient fine-tuning of open-source language models like Phi-2, Phi-3, and Zephyr-3B on the highly imbalanced MathInstruct dataset, and pretraining LLaMA architectures on web text.
The experimental findings show clear advantages across several dimensions. First, UniPROT consistently boosts minority-class classification accuracy in long-tailed vision benchmarks without compromising majority-class performance. Second, in language model instruction fine-tuning at a 50% batch reduction budget, UniPROT outperforms full-batch fine-tuning and state-of-the-art data selection competitors across both in-domain and out-of-domain mathematical reasoning evaluations. Third, as data pruning becomes more aggressive—reducing mini-batch budgets to 25% or 12.5%—UniPROT maintains stable validation loss, whereas competing selection baselines degrade significantly. Finally, in language model pretraining, UniPROT achieves lower validation perplexity than standard full-batch training while processing fewer tokens.
These results demonstrate that enforcing equal exemplar weighting is an effective strategy for preventing data selection algorithms from neglecting critical minority groups. In practical deployments, this enables organizations to substantially cut compute expenses, memory overhead, and training timelines without sacrificing model quality or safety on rare edge cases. The framework also improves model interpretability by ensuring all selected exemplars contribute equally.
Organizations training models on heterogeneous, long-tailed data should consider adopting uniform prototype selection to construct high-quality training batches and coresets. When deploying the framework, engineering teams should configure the internal regularization parameter carefully, as empirical results confirm that tighter regularization yields superior transport fidelity and downstream model accuracy.
While the mathematical guarantees and empirical results are robust across vision and language benchmarks, the framework requires computing gradient or feature representations to establish pairwise similarities, which introduces a modest initial computational cost. Confidence in the reported performance gains is high for fine-tuning and classification tasks, though broader application to multi-modal architectures and distributed trillion-token training regimes remains an area for further validation.
- Paper: Optimal Transport for Domain Adaptation, Nicolas Courty et al. (2014). Establishes the foundational framework of aligning source and target distributions via optimal transport, providing key mathematical context for transport-based data alignment.
- Paper: Active Learning for Convolutional Neural Networks: A Core-Set Approach, Ozan Sener et al. (2018). Formulates representative data subset selection as a geometric core-set problem, providing essential background for discrete optimization and greedy subset guarantees.
- Paper: Prototypical Networks for Few-shot Learning, Jake Snell et al. (2017). Introduces the foundational concept of representing target data distributions and categories using prototypical exemplars in an embedding space.
No sufficiently relevant recommendations were found.
