Minimizing the Accumulated Trajectory Error to Improve Dataset Distillation
Jiawei DuYidi JiangVincent Y. F. TanJoey Tianyi ZhouHaizhou Li
Proposes Flat Trajectory Distillation to regularize training toward flat parameter trajectories, mitigating the accumulated trajectory errors between distillation and evaluation to boost synthetic dataset performance across benchmarks like CIFAR-100 and ImageNet.
Training modern deep neural networks requires massive real-world datasets, which creates significant computing, storage, and financial burdens during model training and architecture exploration. Dataset distillation addresses this issue by condensing a large dataset into a compact set of synthetic examples capable of training models to similar performance levels. Leading distillation methods match optimization trajectories between real and synthetic data during synthesis. However, these methods suffer from accumulated trajectory error, which arises because models start from slightly perturbed model parameters during evaluation compared to distillation, and these discrepancies compound over subsequent training iterations.
The main objective of the article is to demonstrate that regularizing training trajectories toward flatter loss landscapes during real data pre-training mitigates accumulated trajectory error and significantly improves the quality of distilled datasets. The authors propose a method called Flat Trajectory Distillation (FTD) to systematically minimize these compounded errors.
The researchers evaluated their approach across standard image classification benchmarks of varying resolutions, including CIFAR-10, CIFAR-100, Tiny ImageNet, and four selected subsets of ImageNet at higher resolutions. They tested different synthetic dataset sizes, specifically 1, 10, and 50 images per class, and assessed generalization by evaluating the distilled data across multiple neural network architectures such as standard convolutional networks, ResNet, VGG, and AlexNet. Additionally, they applied the synthetic data as a proxy in neural architecture search across 720 candidate network configurations.
The analysis produced several key findings. First, FTD consistently outperformed state-of-the-art distillation baselines across all datasets; for instance, on CIFAR-100, FTD improved test accuracy by 3.5 percentage points with 10 images per class and reached 50.7% accuracy with 50 images per class compared to 47.7% for the baseline trajectory matching method. Second, FTD delivered substantial gains on higher-resolution datasets, improving accuracy by over 4 percentage points on ImageNet subsets such as ImageNette. Third, the synthetic datasets generated by FTD generalized effectively across diverse network architectures rather than overfitting to the architecture used during distillation. Fourth, when used as a proxy for neural architecture search, FTD achieved a Spearman rank correlation of 0.87 for the top five network designs relative to full dataset training, substantially exceeding the 0.41 correlation achieved by the baseline method while reducing the search time from roughly 6,804 minutes to 360 minutes.
These findings indicate that actively stabilizing optimization paths resolves a fundamental limitation in gradient-matching distillation without requiring larger synthetic datasets. For organizations deploying deep learning, this translates into lower computational costs and faster development timelines for tasks like network architecture search and continual learning. Simple direct robustness strategies, such as adding random noise during distillation, degrade performance at low image budgets because they overload the synthetic dataset's capacity. In contrast, regularizing the reference trajectory achieves robustness at no added informational cost.
Organizations evaluating dataset distillation for acceleration pipelines should adopt flat-trajectory regularization principles when preparing reference models for distillation. When using distilled datasets as evaluation proxies, practitioners should prioritize trajectory-regularized methods to ensure reliable correlation with real data performance. Future work should investigate optimization techniques that allow teacher trajectories to run for more epochs without exacerbating trajectory accumulation errors, extending these methods to full-scale, complex datasets.
The findings are supported by consistent empirical improvements across diverse image benchmarks and model types, providing high confidence in the method's effectiveness for computer vision tasks. However, the evaluation was conducted on selected subsets of high-resolution datasets and convolutional architectures; stakeholders should validate performance before deploying these distilled datasets on distinct domains or non-convolutional architectures.
- Paper: Sharpness-Aware Training for Free, Jiawei Du et al. (2022). It provides foundational principles on flatness regularization and mitigating sharp loss landscapes during optimization, which directly underpins the trajectory flatness objective used in Flat Trajectory Distillation.
- Paper: Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks, Zhiwei Deng et al. (2022). It establishes essential concepts and baselines in dataset distillation by compressing datasets into compact representations for efficient neural network re-training.
- Paper: Gradient Norm Aware Minimization Seeks First-Order Flatness and Improves Generalization, Xingxuan Zhang et al. (2023). It extends flatness-aware optimization by formalizing first-order flatness via gradient norm awareness, advancing beyond the loss-landscape curvature techniques applied in trajectory-based distillation.
