Minimizing the Accumulated Trajectory Error to Improve Dataset Distillation

Jiawei DuYidi JiangVincent Y. F. TanJoey Tianyi ZhouHaizhou Li

article2023CVPR171 citations

Proposes Flat Trajectory Distillation to regularize training toward flat parameter trajectories, mitigating the accumulated trajectory errors between distillation and evaluation to boost synthetic dataset performance across benchmarks like CIFAR-100 and ImageNet.

Listen

Training modern deep neural networks requires massive real-world datasets, which creates significant computing, storage, and financial burdens during model training and architecture exploration. Dataset distillation addresses this issue by condensing a large dataset into a compact set of synthetic examples capable of training models to similar performance levels. Leading distillation methods match optimization trajectories between real and synthetic data during synthesis. However, these methods suffer from accumulated trajectory error, which arises because models start from slightly perturbed model parameters during evaluation compared to distillation, and these discrepancies compound over subsequent training iterations.

The main objective of the article is to demonstrate that regularizing training trajectories toward flatter loss landscapes during real data pre-training mitigates accumulated trajectory error and significantly improves the quality of distilled datasets. The authors propose a method called Flat Trajectory Distillation (FTD) to systematically minimize these compounded errors.

The researchers evaluated their approach across standard image classification benchmarks of varying resolutions, including CIFAR-10, CIFAR-100, Tiny ImageNet, and four selected subsets of ImageNet at higher resolutions. They tested different synthetic dataset sizes, specifically 1, 10, and 50 images per class, and assessed generalization by evaluating the distilled data across multiple neural network architectures such as standard convolutional networks, ResNet, VGG, and AlexNet. Additionally, they applied the synthetic data as a proxy in neural architecture search across 720 candidate network configurations.

The analysis produced several key findings. First, FTD consistently outperformed state-of-the-art distillation baselines across all datasets; for instance, on CIFAR-100, FTD improved test accuracy by 3.5 percentage points with 10 images per class and reached 50.7% accuracy with 50 images per class compared to 47.7% for the baseline trajectory matching method. Second, FTD delivered substantial gains on higher-resolution datasets, improving accuracy by over 4 percentage points on ImageNet subsets such as ImageNette. Third, the synthetic datasets generated by FTD generalized effectively across diverse network architectures rather than overfitting to the architecture used during distillation. Fourth, when used as a proxy for neural architecture search, FTD achieved a Spearman rank correlation of 0.87 for the top five network designs relative to full dataset training, substantially exceeding the 0.41 correlation achieved by the baseline method while reducing the search time from roughly 6,804 minutes to 360 minutes.

These findings indicate that actively stabilizing optimization paths resolves a fundamental limitation in gradient-matching distillation without requiring larger synthetic datasets. For organizations deploying deep learning, this translates into lower computational costs and faster development timelines for tasks like network architecture search and continual learning. Simple direct robustness strategies, such as adding random noise during distillation, degrade performance at low image budgets because they overload the synthetic dataset's capacity. In contrast, regularizing the reference trajectory achieves robustness at no added informational cost.

Organizations evaluating dataset distillation for acceleration pipelines should adopt flat-trajectory regularization principles when preparing reference models for distillation. When using distilled datasets as evaluation proxies, practitioners should prioritize trajectory-regularized methods to ensure reliable correlation with real data performance. Future work should investigate optimization techniques that allow teacher trajectories to run for more epochs without exacerbating trajectory accumulation errors, extending these methods to full-scale, complex datasets.

The findings are supported by consistent empirical improvements across diverse image benchmarks and model types, providing high confidence in the method's effectiveness for computer vision tasks. However, the evaluation was conducted on selected subsets of high-resolution datasets and convolutional architectures; stakeholders should validate performance before deploying these distilled datasets on distinct domains or non-convolutional architectures.

Cover for Minimizing the Accumulated Trajectory Error to Improve Dataset Distillation

Abstract

Model-based deep learning has achieved astounding successes due in part to the availability of large-scale real-world data. However, processing such massive amounts of data comes at a considerable cost in terms of computations, storage, training and the search for good neural architectures. Dataset distillation has thus recently come to the fore. This paradigm involves distilling information from large real-world datasets into tiny and compact synthetic datasets such that processing the latter ideally yields similar performances as the former. State-of-the-art methods primarily rely on learning the synthetic dataset by matching the gradients obtained during training between the real and synthetic data. However, these gradient-matching methods suffer from the so-called accumulated trajectory error caused by the discrepancy between the distillation and subsequent evaluation. To mitigate the adverse impact of this accumulated trajectory error, we propose a novel approach that encourages the optimization algorithm to seek a flat trajectory. We show that the weights trained on synthetic data are robust against the accumulated errors perturbations with the regularization towards the flat trajectory. Our method, called Flat Trajectory Distillation (FTD), is shown to boost the performance of gradient-matching methods by up to 4.7% on a subset of images of the ImageNet dataset with higher resolution images. We also validate the effectiveness and generalizability of our method with datasets of different resolutions and demonstrate its applicability to neural architecture search. Code is available at https://github.com/AngusDujw/FTD-distillation.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries and Related Work
  • 3. Methodology
  • 3.1. Matching Training Trajectories (MTT)
  • 3.2. Accumulated Trajectory Error
  • 3.3. Flat Trajectory helps reduce the accumulated trajectory error
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Results
  • 4.3. Ablation and Parameter Studies
  • 4.4. Neural Architecture Search (NAS)
  • 5. Conclusion and Future Work
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Accumulated Trajectory Error in Dataset Distillation

    equation

    In trajectory-matching dataset distillation, a synthetic dataset S={(si,yi)}i=1∣S∣\mathcal{S} = \{(s_i, y_i)\}_{i=1}^{|\mathcal{S}|} is optimized to match segments of parameter trajectories obtained by training a neural network fθf_\theta on a real dataset T={(xi,yi)}i=1∣T∣\mathcal{T} = \{(x_i, y_i)\}_{i=1}^{|\mathcal{T}|}. During the distillation phase, the tt-th segment of the student trajectory starts from an exact teacher checkpoint θt,0∗\theta^*_{t,0}. However, during evaluation, student updates proceed sequentially such that the initial weight of segment t+1t+1 is the output of segment tt, θ^t+1,0=θ^t,n\hat{\theta}_{t+1,0} = \hat{\theta}_{t,n}.

    The accumulated trajectory error ϵt+1\epsilon_{t+1} between the student parameter θ^t+1,n\hat{\theta}_{t+1,n} and the teacher parameter θt+1,m∗\theta^*_{t+1,m} after segment t+1t+1 satisfies the recurrence:

    ϵt+1=ϵt+I(θt,m∗,ϵt)+δt+1\epsilon_{t+1} = \epsilon_t + \mathcal{I}(\theta^*_{t,m}, \epsilon_t) + \delta_{t+1}

    where:

    • ϵt=θ^t,n−θt,m∗\epsilon_t = \hat{\theta}_{t,n} - \theta^*_{t,m} is the accumulated weight error at segment tt,
    • δt+1=A[∇θLS(fθt,m∗),n]−A[∇θLT(fθt,m∗),m]\delta_{t+1} = \mathcal{A}[\nabla_\theta L_\mathcal{S}(f_{\theta^*_{t,m}}), n] - \mathcal{A}[\nabla_\theta L_\mathcal{T}(f_{\theta^*_{t,m}}), m] is the matching error of the gradient optimizer A\mathcal{A} over nn synthetic steps versus mm real steps,
    • I(θt,m∗,ϵt)=A[∇θLS(fθt,m∗+ϵt),n]−A[∇θLS(fθt,m∗),n]\mathcal{I}(\theta^*_{t,m}, \epsilon_t) = \mathcal{A}[\nabla_\theta L_\mathcal{S}(f_{\theta^*_{t,m}+\epsilon_t}), n] - \mathcal{A}[\nabla_\theta L_\mathcal{S}(f_{\theta^*_{t,m}}), n] is the initialization error incurred by the weight perturbation ϵt\epsilon_t entering the student update.
  2. Knowl 2 — Flat Trajectory Distillation (FTD) Buffer-Phase Optimization

    model/method

    Flat Trajectory Distillation (FTD) mitigates the accumulated trajectory error ϵt\epsilon_t in trajectory matching by regularizing the teacher trajectory generation (the buffer phase) to find trajectories that lie in flat loss regions robust to parameter perturbations, rather than modifying the student distillation updates directly.

    In the buffer phase, the teacher network parameters θ∗\theta^* are optimized via a minimax objective based on surrogate gap / sharpness-aware minimization:

    θ∗=arg⁡min⁡θ{LT(fθ)+αS(θ)}\theta^* = \arg\min_\theta \big\{ L_\mathcal{T}(f_\theta) + \alpha S(\theta) \big\}

    where the loss sharpness S(θ)S(\theta) is defined over a perturbation ball Ψ={ϵ:∥ϵ∥2≤ρ}\Psi = \{\epsilon : \|\epsilon\|_2 \le \rho\} with perturbation radius ρ>0\rho > 0:

    S(θ)≜max⁡ϵ∈Ψ[LT(fθ+ϵ)−LT(fθ)]S(\theta) \triangleq \max_{\epsilon \in \Psi} \big[ L_\mathcal{T}(f_{\theta+\epsilon}) - L_\mathcal{T}(f_\theta) \big]

    Here, α\alpha is a balancing coefficient (default α=1.0\alpha = 1.0) and ρ\rho controls the permissible perturbation norm (default ρ=0.01\rho = 0.01). Generating teacher trajectories under this objective ensures that the trajectory states θt∗\theta^*_t exhibit low loss curvature, reducing the initialization error I(θt∗,ϵt)\mathcal{I}(\theta^*_t, \epsilon_t) encountered when the student deviates during evaluation.

  3. Knowl 3 — Relation Between Trajectory Initialization Error and Loss Hessian Curvature

    theoretical result

    Under a first-order optimizer A\mathcal{A} with learning rate η\eta (such that ∂A∂∇θLS=η\frac{\partial \mathcal{A}}{\partial \nabla_\theta L_\mathcal{S}} = \eta) and small parameter perturbation norm ∥ϵt∥2\|\epsilon_t\|_2, the first-order Taylor approximation of the trajectory initialization error I(θt∗,ϵt)=A[∇θLS(fθt∗+ϵt),n]−A[∇θLS(fθt∗),n]\mathcal{I}(\theta^*_t, \epsilon_t) = \mathcal{A}[\nabla_\theta L_\mathcal{S}(f_{\theta^*_t + \epsilon_t}), n] - \mathcal{A}[\nabla_\theta L_\mathcal{S}(f_{\theta^*_t}), n] yields:

    arg⁡min⁡θt∥I(θt∗,ϵt)∥22≈arg⁡min⁡θt∥∇θ2LS(fθt∗)∥22\arg\min_{\theta_t} \|\mathcal{I}(\theta^*_t, \epsilon_t)\|_2^2 \approx \arg\min_{\theta_t} \big\| \nabla^2_\theta L_\mathcal{S}(f_{\theta^*_t}) \big\|_2^2

    Minimizing the ℓ2\ell_2 norm of the initialization error is approximately equivalent to minimizing the spectral norm (largest eigenvalue) of the loss Hessian ∇θ2LS(fθt∗)\nabla^2_\theta L_\mathcal{S}(f_{\theta^*_t}). Because computing the exact Hessian spectral norm across trajectory checkpoints is computationally intractable, minimizing the sharpness of the loss landscape S(θ)=max⁡∥ϵ∥2≤ρ[LT(fθ+ϵ)−LT(fθ)]S(\theta) = \max_{\|\epsilon\|_2 \le \rho} [L_\mathcal{T}(f_{\theta+\epsilon}) - L_\mathcal{T}(f_\theta)] acts as an effective surrogate during teacher trajectory generation.

  4. Knowl 4 — Benchmark Classification Performance on CIFAR and Tiny ImageNet

    data/table

    The performance of 3-layer ConvNets (on CIFAR-10/100) and 4-layer ConvNets (on Tiny ImageNet) trained on distilled synthetic datasets generated by FTD and baseline methods across varying numbers of images per class (ipc∈{1,10,50}\text{ipc} \in \{1, 10, 50\}) is summarized below. Test accuracies (mean ±\pm standard deviation over 5 random initializations for 1000 evaluation steps) demonstrate that FTD consistently outperforms prior state-of-the-art distillation approaches:

    CIFAR-10 CIFAR-100 Tiny ImageNet
    ipc 1 10 50 1 10 50 1 10
    Real dataset 84.8±0.184.8 \pm 0.1 56.2±0.356.2 \pm 0.3 37.6±0.437.6 \pm 0.4
    DC 28.3±0.528.3 \pm 0.5 44.9±0.544.9 \pm 0.5 53.9±0.553.9 \pm 0.5 12.8±0.312.8 \pm 0.3 25.2±0.325.2 \pm 0.3 - - -
    DM 26.0±0.826.0 \pm 0.8 48.9±0.648.9 \pm 0.6 63.0±0.463.0 \pm 0.4 11.4±0.311.4 \pm 0.3 29.7±0.329.7 \pm 0.3 43.6±0.443.6 \pm 0.4 3.9±0.23.9 \pm 0.2 12.9±0.412.9 \pm 0.4
    DSA 28.8±0.728.8 \pm 0.7 52.1±0.552.1 \pm 0.5 60.6±0.560.6 \pm 0.5 13.9±0.313.9 \pm 0.3 32.3±0.332.3 \pm 0.3 42.8±0.442.8 \pm 0.4 - -
    CAFE 30.3±1.130.3 \pm 1.1 46.3±0.646.3 \pm 0.6 55.5±0.655.5 \pm 0.6 12.9±0.312.9 \pm 0.3 27.8±0.327.8 \pm 0.3 37.9±0.337.9 \pm 0.3 - -
    CAFE+DSA 31.6±0.831.6 \pm 0.8 50.9±0.550.9 \pm 0.5 62.3±0.462.3 \pm 0.4 14.0±0.314.0 \pm 0.3 31.5±0.231.5 \pm 0.2 42.9±0.242.9 \pm 0.2 - -
    PP 46.4±0.646.4 \pm 0.6 65.5±0.365.5 \pm 0.3 71.9±0.271.9 \pm 0.2 24.6±0.124.6 \pm 0.1 43.1±0.343.1 \pm 0.3 48.4±0.348.4 \pm 0.3 - -
    MTT 46.2±0.846.2 \pm 0.8 65.4±0.765.4 \pm 0.7 71.6±0.271.6 \pm 0.2 24.3±0.324.3 \pm 0.3 39.7±0.439.7 \pm 0.4 47.7±0.247.7 \pm 0.2 8.8±0.38.8 \pm 0.3 23.2±0.223.2 \pm 0.2
    MTT+Robust 45.8±0.745.8 \pm 0.7 63.2±0.763.2 \pm 0.7 72.7±0.272.7 \pm 0.2 24.1±0.324.1 \pm 0.3 39.4±0.439.4 \pm 0.4 47.9±0.247.9 \pm 0.2 - -
    FTD 46.8±0.3\mathbf{46.8 \pm 0.3} 66.6±0.3\mathbf{66.6 \pm 0.3} 73.8±0.2\mathbf{73.8 \pm 0.2} 25.2±0.2\mathbf{25.2 \pm 0.2} 43.4±0.3\mathbf{43.4 \pm 0.3} 50.7±0.3\mathbf{50.7 \pm 0.3} 10.4±0.3\mathbf{10.4 \pm 0.3} 24.5±0.2\mathbf{24.5 \pm 0.2}
  5. Knowl 5 — Dataset Distillation Performance on High-Resolution ImageNet Subsets

    data/table

    FTD and Matching Training Trajectories (MTT) were evaluated on four 10-class subsets of ImageNet-1k with image resolutions of 128×128128 \times 128 (ImageNette, ImageWoof, ImageFruit, ImageMeow) using a 5-layer ConvNet (ConvNet-5) at ipc∈{1,10}\text{ipc} \in \{1, 10\}. FTD achieves improvements of over 4% on ImageNette and consistently outperforms MTT across all subsets:

    ImageNette ImageWoof ImageFruit ImageMeow
    ipc 1 10 1 10 1 10 1 10
    Real dataset 87.4±1.087.4 \pm 1.0 67.0±1.367.0 \pm 1.3 63.9±2.063.9 \pm 2.0 66.7±1.166.7 \pm 1.1
    MTT 47.7±0.947.7 \pm 0.9 63.0±1.363.0 \pm 1.3 28.6±0.828.6 \pm 0.8 35.8±1.835.8 \pm 1.8 26.6±0.826.6 \pm 0.8 40.3±1.340.3 \pm 1.3 30.7±1.630.7 \pm 1.6 40.4±2.240.4 \pm 2.2
    FTD 52.2±1.0\mathbf{52.2 \pm 1.0} 67.7±0.7\mathbf{67.7 \pm 0.7} 30.1±1.0\mathbf{30.1 \pm 1.0} 38.8±1.4\mathbf{38.8 \pm 1.4} 29.1±0.9\mathbf{29.1 \pm 0.9} 44.9±1.5\mathbf{44.9 \pm 1.5} 33.8±1.5\mathbf{33.8 \pm 1.5} 43.3±0.6\mathbf{43.3 \pm 0.6}
  6. Knowl 6 — Cross-Architecture Generalization of FTD Distilled Datasets

    data/table

    To evaluate cross-architecture transferability, synthetic datasets distilled from CIFAR-10 with ipc=50\text{ipc} = 50 using a 3-layer ConvNet were used to train networks of distinct architectures from scratch: ConvNet (3-layer), ResNet18, VGG11, and AlexNet. The table below presents test accuracy (mean ±\pm standard deviation) comparing FTD to baseline distillation methods:

    Method ConvNet ResNet18 VGG11 AlexNet
    DC 53.9±0.553.9 \pm 0.5 20.8±1.020.8 \pm 1.0 38.8±1.138.8 \pm 1.1 28.7±0.728.7 \pm 0.7
    CAFE 55.5±0.455.5 \pm 0.4 25.3±0.925.3 \pm 0.9 40.5±0.840.5 \pm 0.8 34.0±0.634.0 \pm 0.6
    MTT 71.6±0.271.6 \pm 0.2 61.9±0.761.9 \pm 0.7 55.4±0.855.4 \pm 0.8 48.2±1.048.2 \pm 1.0
    FTD 73.8±0.2\mathbf{73.8 \pm 0.2} 65.7±0.3\mathbf{65.7 \pm 0.3} 58.4±1.6\mathbf{58.4 \pm 1.6} 53.8±0.9\mathbf{53.8 \pm 0.9}

    FTD preserves higher transfer accuracy across all target architectures, demonstrating that regularizing the teacher trajectory captures generalizable visual representations rather than architecture-specific artifacts.

  7. Knowl 7 — Neural Architecture Search (NAS) Benchmark using Distilled Datasets

    data/table

    To evaluate proxy dataset quality for Neural Architecture Search (NAS), 720 candidate ConvNets varying across depth, width, activation functions, normalization layers, and pooling operations were trained for 200 epochs on: (1) the full real CIFAR-10 training set (50,000 images), (2) the MTT synthetic set (500 images, ipc=50\text{ipc} = 50), and (3) the FTD synthetic set (500 images, ipc=50\text{ipc} = 50).

    The table below reports the Spearman rank correlation between architecture performance rankings obtained on each synthetic proxy dataset versus the real dataset for the top 5, 10, and 20 architectures, along with total search runtime:

    Proxy Top 5 Top 10 Top 20 Time (min) Image Count
    Real Dataset 1.001.00 1.001.00 1.001.00 6,8046,804 50,00050,000
    MTT 0.410.41 0.360.36 −0.04-0.04 360360 500500
    FTD 0.87\mathbf{0.87} 0.68\mathbf{0.68} 0.54\mathbf{0.54} 360360 500500

    FTD achieves a 0.87 correlation among top-5 architecture candidates while providing a nearly 19×19\times speedup over searching on the full dataset.

  8. Knowl 8 — Ineffectiveness of Flat Minima Optimization during Evaluation

    empirical result

    Applying Sharpness-Aware Minimization (SAM) during the student evaluation phase on synthetic data (termed "MTT + Flat Minimum") across a range of perturbation radii ρ∈{0.005,0.01,0.03,0.05,0.1}\rho \in \{0.005, 0.01, 0.03, 0.05, 0.1\} does not noticeably improve test accuracy over standard MTT evaluation (e.g., remaining near 39.7%39.7\% accuracy on CIFAR-100 with ipc=10\text{ipc} = 10, compared to 43.2%43.2\% achieved by FTD).

    This demonstrates that the performance gain of FTD originates specifically from suppressing the accumulated trajectory error along the optimization path via a flat teacher trajectory in the buffer phase, rather than merely converging to a flat minimum during the evaluation phase.

  9. Knowl 9 — Failure Mode of Direct Robust Learning during Distillation

    empirical result

    A naive approach to mitigating initialization discrepancy is robust learning during distillation, where Gaussian noise ϵ∼N(0,σ2I)\epsilon \sim \mathcal{N}(\mathbf{0}, \sigma^2 \mathbf{I}) is added directly to initial weights during student trajectory matching:

    S=arg⁡min⁡SEθ0∼Pθ0,ϵ∼N(0,σ2I)∥A[LS(fθ0+ϵ),n]−A[LT(fθ0),m]∥22\mathcal{S} = \arg\min_{\mathcal{S}} \mathbb{E}_{\theta_0 \sim P_{\theta_0}, \epsilon \sim \mathcal{N}(\mathbf{0}, \sigma^2 \mathbf{I})} \big\| \mathcal{A}[L_\mathcal{S}(f_{\theta_0 + \epsilon}), n] - \mathcal{A}[L_\mathcal{T}(f_{\theta_0}), m] \big\|_2^2

    Empirically, while this robust learning objective slightly improves MTT at ipc=50\text{ipc} = 50 (from 71.6%71.6\% to 72.7%72.7\% on CIFAR-10, and 47.7%47.7\% to 47.9%47.9\% on CIFAR-100), it degrades performance in low-data regimes where ipc∈{1,10}\text{ipc} \in \{1, 10\} (e.g., CIFAR-10 ipc=10\text{ipc} = 10 drops from 65.4%65.4\% to 63.2%63.2\%).

    This degradation occurs because mapping a perturbed parameter distribution Pθ0+ϵP_{\theta_0+\epsilon} into the optimal parameter space requires more capacity and information per class than small synthetic budgets (ipc∈{1,10}\text{ipc} \in \{1, 10\}) can encode.

  10. Knowl 10 — Ablation on Exponential Moving Average in FTD

    empirical result

    In the synthetic image optimization phase of FTD, an Exponential Moving Average (EMA) with decay factor β=0.999\beta = 0.999 is applied to accelerate convergence. Ablation experiments show that while EMA provides minor gains, FTD without EMA still outperforms MTT:

    • On CIFAR-100 with ipc=10\text{ipc} = 10: MTT achieves 39.7±0.4%39.7 \pm 0.4\%, FTD without EMA achieves 43.4±0.3%43.4 \pm 0.3\%, and FTD with EMA achieves 43.2±0.3%43.2 \pm 0.3\%.
    • On CIFAR-100 with ipc=50\text{ipc} = 50: MTT achieves 47.7±0.2%47.7 \pm 0.2\%, FTD without EMA achieves 49.8±0.3%49.8 \pm 0.3\%, and FTD with EMA achieves 50.7±0.3%50.7 \pm 0.3\%.
    • On Tiny ImageNet with ipc=1\text{ipc} = 1: MTT achieves 8.8±0.3%8.8 \pm 0.3\%, FTD without EMA achieves 9.8±0.2%9.8 \pm 0.2\%, and FTD with EMA achieves 10.0±0.2%10.0 \pm 0.2\%.
    • On Tiny ImageNet with ipc=10\text{ipc} = 10: MTT achieves 23.2±0.2%23.2 \pm 0.2\%, FTD without EMA achieves 24.1±0.3%24.1 \pm 0.3\%, and FTD with EMA achieves 24.5±0.2%24.5 \pm 0.2\%.

    These results confirm that the primary driver of performance improvement in FTD is the flat teacher trajectory regularization in the buffer phase rather than EMA.

Coverage note — None was omitted; all contributed methodology, theoretical derivations, experimental tables, ablations, NAS experiments, and limitation discussions were captured into self-contained knowls.

References

  1. 1.George Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A Efros, and Jun-Yan Zhu. Dataset distillation by matching training trajectories. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4750–4759, 2022.
  2. 2.Jang Hyun Cho and Bharath Hariharan. On the efficacy of knowledge distillation. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4794–4802, 2019.
  3. 3.Justin Cui, Ruochen Wang, Si Si, and Cho-Jui Hsieh. Dcbench: Dataset condensation benchmark. arXiv preprint arXiv:2207.09639, 2022.
  4. 4.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  6. 6.Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio. Sharp minima can generalize for deep nets. In International Conference on Machine Learning, pages 1019–1028. PMLR, 2017.
  7. 7.Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio. Sharp minima can generalize for deep nets. In International Conference on Machine Learning, pages 1019–1028. PMLR, 2017.
  8. 8.Tian Dong, Bo Zhao, and Lingjuan Lyu. Privacy for free: How does dataset condensation help privacy? arXiv preprint arXiv:2206.00240, 2022.
  9. 9.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  10. 10.Jiawei Du, Hanshu Yan, Jiashi Feng, Joey Tianyi Zhou, Liangli Zhen, Rick Siow Mong Goh, and Vincent YF Tan. Efficient sharpness-aware minimization for improved training of neural networks. arXiv preprint arXiv:2110.03141, 2021.
  11. 11.Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. Sharpness-aware minimization for efficiently improving generalization. In International Conference on Learning Representations, 2020.
  12. 12.Spyros Gidaris and Nikos Komodakis. Dynamic few-shot visual learning without forgetting. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4367–4375, 2018.
  13. 13.Jack Goetz and Ambuj Tewari. Federated learning via synthetic data. arXiv preprint arXiv:2008.04489, 2020.
  14. 14.Jianping Gou, Baosheng Yu, Stephen J Maybank, and Dacheng Tao. Knowledge distillation: A survey. International Journal of Computer Vision, 129(6):1789–1819, 2021.
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  16. 16.Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2(7), 2015.
  17. 17.Sepp Hochreiter and J¨urgen Schmidhuber. Simplifying neural nets by discovering flat minima. In Proceedings of the 8th International Conference on Neural Information Processing Systems, pages 529–536, 1995.
  18. 18.Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio. Fantastic generalization measures and where to find them. In International Conference on Learning Representations, 2019.
  19. 19.Haifeng Jin, Qingquan Song, and Xia Hu. Auto-keras: An efficient neural architecture search system. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining, pages 1946–1956, 2019.
  20. 20.Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. On largebatch training for deep learning: Generalization gap and sharp minima. arXiv preprint arXiv:1609.04836, 2016.
  21. 21.Nitish Shirish Keskar, Jorge Nocedal, Ping Tak Peter Tang, Dheevatsa Mudigere, and Mikhail Smelyanskiy. On largebatch training for deep learning: Generalization gap and sharp minima. In International Conference on Learning Representations, 2017.
  22. 22.Yoon Kim and Alexander M Rush. Sequence-level knowledge distillation. arXiv preprint arXiv:1606.07947, 2016.
  23. 23.Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton. CIFAR10 and CIFAR-100 datasets. URl: https://www. cs. toronto. edu/kriz/cifar. html, 6(1):1, 2009.
  24. 24.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012.
  25. 25.Ya Le and Xuan Yang. Tiny imagenet visual recognition challenge. CS 231N, 7(7):3, 2015.
  26. 26.Shiye Lei and Dacheng Tao. A comprehensive survey to dataset distillation. arXiv preprint arXiv:2301.05603, 2023.
  27. 27.Guang Li, Ren Togo, Takahiro Ogawa, and Miki Haseyama. Soft-label anonymous gastric x-ray image distillation. In 2020 IEEE International Conference on Image Processing (ICIP), pages 305–309. IEEE, 2020.
  28. 28.Guang Li, Ren Togo, Takahiro Ogawa, and Miki Haseyama. Dataset distillation using parameter pruning. arXiv preprint arXiv:2209.14609, 2022.
  29. 29.Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein. Visualizing the loss landscape of neural nets. Proceedings of the 32nd International Conference on Neural Information Processing Systems, 31, 2018.
  30. 30.Tengyuan Liang, Tomaso Poggio, Alexander Rakhlin, and James Stokes. Fisher–Rao metric, geometry, and complexity of neural networks. In The 22nd International Conference on Artificial Intelligence and Statistics, pages 888–896. PMLR, 2019.
  31. 31.Chen Liu, Mathieu Salzmann, Tao Lin, Ryota Tomioka, and Sabine S¨ustrunk. On the loss landscape of adversarial training: Identifying challenges and how to overcome them. Proceedings of the 34th International Conference on Neural Information Processing Systems, 33:21476–21487, 2020.
  32. 32.Dougal Maclaurin, David Duvenaud, and Ryan Adams. Gradient-based hyperparameter optimization through reversible learning. In International conference on machine learning, pages 2113–2122. PMLR, 2015.
  33. 33.David A McAllester. Pac-bayesian model averaging. In Proceedings of the twelfth annual conference on Computational learning theory, pages 164–170, 1999.
  34. 34.Timothy Nguyen, Zhourong Chen, and Jaehoon Lee. Dataset meta-learning from kernel ridge-regression. arXiv preprint arXiv:2011.00050, 2020.
  35. 35.Timothy Nguyen, Roman Novak, Lechao Xiao, and Jaehoon Lee. Dataset distillation with infinitely wide convolutional networks. Advances in Neural Information Processing Systems, 34:5186–5198, 2021.
  36. 36.Timothy Nguyen, Roman Novak, Lechao Xiao, and Jaehoon Lee. Dataset distillation with infinitely wide convolutional networks. Advances in Neural Information Processing Systems, 34:5186–5198, 2021.
  37. 37.Hieu Pham, Melody Guan, Barret Zoph, Quoc Le, and Jeff Dean. Efficient neural architecture search via parameters sharing. In International conference on machine learning, pages 4095–4104. PMLR, 2018.
  38. 38.Pengzhen Ren, Yun Xiao, Xiaojun Chang, Po-Yao Huang, Zhihui Li, Xiaojiang Chen, and Xin Wang. A comprehensive survey of neural architecture search: Challenges and solutions. ACM Computing Surveys (CSUR), 54(4):1–34, 2021.
  39. 39.Andrea Rosasco, Antonio Carta, Andrea Cossu, Vincenzo Lomonaco, and Davide Bacciu. Distilled replay: Overcoming forgetting through synthetic samples. arXiv preprint arXiv:2103.15851, 2021.
  40. 40.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  41. 41.Felipe Petroski Such, Aditya Rawal, Joel Lehman, Kenneth Stanley, and Jeffrey Clune. Generative teaching networks: Accelerating neural architecture search by learning to generate synthetic training data. In International Conference on Machine Learning, pages 9206–9216. PMLR, 2020.
  42. 42.Ilia Sucholutsky and Matthias Schonlau. Soft-label dataset distillation and text dataset distillation. In 2021 International Joint Conference on Neural Networks (IJCNN), pages 1–8. IEEE, 2021.
  43. 43.Paul Vicol, Jonathan P Lorraine, Fabian Pedregosa, David Duvenaud, and Roger B Grosse. On implicit bias in overparameterized bilevel optimization. In International Conference on Machine Learning, pages 22234–22259. PMLR, 2022.
  44. 44.Kai Wang, Bo Zhao, Xiangyu Peng, Zheng Zhu, Shuo Yang, Shuo Wang, Guan Huang, Hakan Bilen, Xinchao Wang, and Yang You. Cafe: Learning to condense dataset by aligning features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12196–12205, 2022.
  45. 45.Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba, and Alexei A Efros. Dataset distillation. arXiv preprint arXiv:1811.10959, 2018.
  46. 46.Ross Wightman. Pytorch image models. https : / / github . com / rwightman / pytorch - image - models, 2019.
  47. 47.Ruonan Yu, Songhua Liu, and Xinchao Wang. Dataset distillation: A comprehensive review. arXiv preprint arXiv:2301.07014, 2023.
  48. 48.Bo Zhao and Hakan Bilen. Dataset condensation with differentiable siamese augmentation. In International Conference on Machine Learning, pages 12674–12685. PMLR, 2021.
  49. 49.Bo Zhao and Hakan Bilen. Dataset condensation with distribution matching. arXiv preprint arXiv:2110.04181, 2021.
  50. 50.Bo Zhao, Konda Reddy Mopuri, and Hakan Bilen. Dataset condensation with gradient matching. ICLR, 1(2):3, 2021.
  51. 51.Juntang Zhuang, Boqing Gong, Liangzhe Yuan, Yin Cui, Hartwig Adam, Nicha Dvornek, Sekhar Tatikonda, James Duncan, and Ting Liu. Surrogate gap minimization improves sharpness-aware training. arXiv preprint arXiv:2203.08065, 2022.

Citation

MLA
Du, J., et al. “Minimizing the Accumulated Trajectory Error to Improve Dataset Distillation”. arXiv, 2022, http://arxiv.org/abs/2211.11004v3.
APA
Du, J., Jiang, Y., Tan, V. Y. F., Zhou, J. T., & Li, H. (2022). Minimizing the Accumulated Trajectory Error to Improve Dataset Distillation. arXiv. http://arxiv.org/abs/2211.11004v3
Chicago
Du, J., Y. Jiang, V. Y. F. Tan, J. T. Zhou, and H. Li. 2022. “Minimizing the Accumulated Trajectory Error to Improve Dataset Distillation”. arXiv. http://arxiv.org/abs/2211.11004v3.
Harvard
Du, J. et al. (2022) “Minimizing the Accumulated Trajectory Error to Improve Dataset Distillation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.11004v3.
Vancouver
1. Du J, Jiang Y, Tan VYF, Zhou JT, Li H (2022) Minimizing the Accumulated Trajectory Error to Improve Dataset Distillation. arXiv

BibTeX

@article{du2022minimizing,
  title = {Minimizing the Accumulated Trajectory Error to Improve Dataset Distillation},
  author = {Du, Jiawei and Jiang, Yidi and Tan, Vincent Y. F. and Zhou, Joey Tianyi and Li, Haizhou},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.11004v3},
  eprint = {2211.11004}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE