DAVINZ: Data Valuation using Deep Neural Networks at Initialization

Zhaoxuan WuYao ShuBryan Kian Hsiang Low

article2022ICML73 citations

Proposes a training-free data valuation framework that leverages neural tangent kernel theory and domain-aware generalization bounds to accurately quantify dataset contributions to deep neural networks at initialization without expensive model training.

Listen

Assessing the exact value of data contributors' submissions is essential for fair compensation in collaborative machine learning and commercial data marketplaces. However, conventional data valuation methods, such as the Shapley value, rely heavily on evaluating how well a deep neural network performs after complete model training. For large, complex modern architectures, repeatedly training models across various data subsets is computationally prohibitive and creates severe operational bottlenecks.

The article develops and evaluates a training-free data valuation method called Data Valuation at Initialization (DAVINZ). Its primary objective is to accurately and reliably estimate the value of data subsets on deep neural networks at initial parameter setup, bypassing the need for long-term model optimization.

The authors derived a theoretical generalization bound using neural tangent kernel theory that explicitly accounts for discrepancies between training data and target validation objectives. This bound serves as a scoring function that combines in-domain complexity and out-of-domain distributional divergence, which is then integrated into standard game-theoretic valuation frameworks. The approach was validated through extensive experiments on classification and regression benchmarks (including MNIST, CIFAR-10, Tiny ImageNet, and physical simulation datasets) evaluated against baseline deep architectures such as VGG and ResNet.

Key findings show that DAVINZ provides estimated valuation scores with a strong Pearson correlation of up to 0.954 relative to ground-truth validation accuracy. When assessing contributions, it maintains high correlation with ground-truth values while reducing computational runtime by over 30-fold compared to conventional training-based validation. It also vastly outperforms existing training-free alternatives like influence functions and robust volume metrics, which often degrade on deep non-convex architectures or ignore target domain preferences. Furthermore, the analysis demonstrates that DAVINZ rigorously preserves four crucial operational qualities: sensitivity to the consumer's target dataset preferences, proper reward scaling for data volume, numerical stability against input noise, and robustness across differing neural network models and random parameter initializations.

These results establish that data valuation for deep learning can be deployed without prohibitive compute budgets or extensive hyperparameter tuning. Bypassing iterative training mitigates operational risks associated with model training instability, drastically compresses valuation timelines from days to minutes, and enables practical multi-party data marketplaces and efficient dataset curation.

Organizations operating collaborative AI pipelines or commercial data exchanges should consider training-free valuation frameworks to establish fair compensation structures and prune redundant training data. Practical implementations should leverage diagonal block approximations for matrix calculations to maintain memory efficiency on commercial hardware.

Readers should note that the underlying theoretical bounds assume wide neural network formulations, and empirical scaling relies on block-diagonal approximations. Nonetheless, the consistent alignment across diverse image and regression benchmarks provides strong confidence in adopting the method for practical enterprise data workflows.

Wu et al (2022).pdf

No sufficiently relevant recommendations were found.

  • Paper: Data Shapley in One Training Run, Jiachen T. Wang et al. (2025). It carries data valuation forward from estimating contribution at initialization to assigning Shapley values during a single training run, extending the practical effort to make contribution measurement computationally feasible.
Cover for DAVINZ: Data Valuation using Deep Neural Networks at Initialization

Abstract

Recent years have witnessed a surge of interest in developing trustworthy methods to evaluate the value of data in many real-world applications (e.g., collaborative machine learning, data marketplaces). Existing data valuation methods typically valuate data using the generalization performance of converged machine learning models after their long-term model training, hence making data valuation on large complex deep neural networks (DNNs) unaffordable. To this end, we theoretically derive a domain-aware generalization bound to estimate the generalization performance of DNNs without model training. We then exploit this theoretically derived generalization bound to develop a novel training-free data valuation method named data valuation at initialization (DAVINZ) on DNNs, which consistently achieves remarkable effectiveness and efficiency in practice. Moreover, our training-free DAVINZ, surprisingly, can even theoretically and empirically enjoy the desirable properties that training-based data valuation methods usually attain, thus making it more trustworthy in practice.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Backgrounds and Notations
  • 3.1. Data Valuation
  • 3.2. Neural Tangent Kernel
  • 4. Data Valuation at Initialization (DAVINZ)
  • 4.1. Domain-Aware Generalization Bound for DNNs
  • 4.2. Generalization Bound as Scoring Function
  • 4.3. Training-free Data Valuation Algorithm
  • 5. Properties of DAVINZ
  • 5.1. Awareness of Data Quantity
  • 5.2. Stability to Noise
  • 5.3. Robustness to Model
  • 6. Experiments
  • 6.1. Valid Scoring Function in Practice
  • 6.2. Effective and Efficient DAVINZ
  • 6.3. Awareness of Data Preference
  • 6.4. Awareness of Data Quantity
  • 6.5. Stability to Noise
  • 6.6. Robustness to Model
  • 6.7. Application: Large-scale Shapley Value
  • 6.8. Application: Data Summarization
  • 7. Conclusion & Discussion
  • Acknowledgements
  • References
  • A. Proofs
  • A.1. Proof of Theorem 1
  • A.2. Proof of Proposition 1
  • A.3. Proof of Proposition 2
  • A.4. Proof of Proposition 3
  • B. Efficient Approximations of Θ0
  • B.1. Diagonal Block Approximation
  • B.2. Diagonal Block Approximation with Permutations
  • B.3. L-block-banded Block Matrix Approximation
  • B.4. Invertibility of Θ0
  • C. Additional Literature
  • D. Experimental Setup
  • D.1. Dataset, Model and Ground Truth
  • D.2. Validation Performance
  • D.3. DAVINZ Setups
  • D.4. Influence Functions
  • D.5. Robust Volume
  • E. More Results
  • E.1. Robustness to Different Initializations
  • E.2. |f(x, θ0)| at Initialization
  • E.3. DAVINZ on Large-scale Datasets

Knowls

  1. Knowl 1 — Domain-aware generalization bound for initialized neural networks

    theoretical result

    Let S={(xi,yi)}i=1mSS=\{(x_i,y_i)\}_{i=1}^{m_S} be independently sampled from a source domain DSD_S, and let T={(xi′,yi′)}i=1mTT=\{(x'_i,y'_i)\}_{i=1}^{m_T} be independently sampled from a target domain DTD_T. For a predictor ff, write LD(f)L_D(f) for its expected loss on domain DD, and let f∗f^* minimize LDS(f)+LDT(f)L_{D_S}(f)+L_{D_T}(f). Assume the network output and labels lie in [0,1][0,1], inputs satisfy ∥x∥2≤1\|x\|_2\leq 1, and there is a function hh in a function class H\mathcal H such that h(x)≤1h(x)\leq 1 and ∣f(x)−f∗(x)∣≤h(x)|f(x)-f^*(x)|\leq h(x). Define the population domain discrepancy by

    dH(DT,DS)=sup⁡h∈H∣Ex∼DT[h(x)]−Ex∼DS[h(x)]∣,d_{\mathcal H}(D_T,D_S)=\sup_{h\in\mathcal H}\left|\mathbb E_{x\sim D_T}[h(x)]-\mathbb E_{x\sim D_S}[h(x)]\right|,

    and its empirical estimate by replacing each expectation with its sample average on TT or SS, denoted dH(T,S)d_{\mathcal H}(T,S). Use squared loss ℓ(f,y)=(f−y)2/2\ell(f,y)=(f-y)^2/2. Let Θ0\Theta_0 be the neural tangent kernel (NTK) matrix on SS at initialization θ0\theta_0, and let Θ∞\Theta_\infty be its infinite-width limit. Suppose λmin⁡(Θ0)>0\lambda_{\min}(\Theta_0)>0 and ∥∇θf(x,θ0)∥2≤ρ\|\nabla_\theta f(x,\theta_0)\|_2\leq\rho on SS. For sufficiently large width n>Nn>N, gradient descent with learning rate

    η<min⁡{2n(λmin⁡(Θ∞)+λmax⁡(Θ∞)),  mSλmax⁡(Θ0)}\eta<\min\left\{\frac{2}{n(\lambda_{\min}(\Theta_\infty)+\lambda_{\max}(\Theta_\infty))},\;\frac{m_S}{\lambda_{\max}(\Theta_0)}\right\}

    produces, at any time t>0t>0, a predictor ftf_t satisfying with probability at least 1−2δ1-2\delta:

    LDT(ft)≤LS(ft)+2ρy^TΘ0−1y^mS+dH(T,S)+ε.L_{D_T}(f_t)\leq L_S(f_t)+2\rho\sqrt{\frac{\widehat y^{\mathsf T}\Theta_0^{-1}\widehat y}{m_S}}+d_{\mathcal H}(T,S)+\varepsilon.

    Here LSL_S is empirical source loss, y^=y−f(S,θ0)\widehat y=y-f(S,\theta_0) is the vector of initialization residuals, and λmin⁡\lambda_{\min} and λmax⁡\lambda_{\max} denote extreme eigenvalues. The remainder is ε=2c/n+4log⁡(4/δ)/(2mS)+log⁡(4/δ)/(2mT)+LDT(f∗)+LDS(f∗)\varepsilon=2c/\sqrt n+4\sqrt{\log(4/\delta)/(2m_S)}+\sqrt{\log(4/\delta)/(2m_T)}+L_{D_T}(f^*)+L_{D_S}(f^*), for a constant c>0c>0. The bound separates an NTK-based source-domain complexity term from a source–target discrepancy term, so it accounts for domain shift rather than requiring the training and validation data to share a distribution.

  2. Knowl 2 — Training-free score from the domain-aware bound

    model/method

    DAVINZ estimates the usefulness of a dataset SS for a target validation set TT by using a simplified version of the domain-aware generalization bound as its score. For a DNN initialized at θ0\theta_0, let Θ0\Theta_0 be its NTK matrix on SS, let yy be the label vector, and let y^=y−f(S,θ0)\widehat y=y-f(S,\theta_0). With mS=∣S∣m_S=|S| and empirical discrepancy dH(T,S)d_{\mathcal H}(T,S), the score is

    ν(S)=−κy^TΘ0−1y^mS−dH(T,S).\nu(S)=-\kappa\sqrt{\frac{\widehat y^{\mathsf T}\Theta_0^{-1}\widehat y}{m_S}}-d_{\mathcal H}(T,S).

    A higher score favors a smaller NTK complexity term and closer alignment with the validation domain. The score omits the empirical training loss and the bound remainder: the paper argues that converged DNN training loss is usually near zero and that the remainder is constant for relative comparisons among datasets. In practice, H\mathcal H is chosen from a reproducing-kernel Hilbert space and the discrepancy is estimated using multiple-kernel maximum mean discrepancy (MK-MMD), providing a computationally tractable domain comparison.

    The trade-off parameter κ\kappa balances the two score components. For contributor datasets SiS_i, it is set to equalize their aggregate scales:

    κ=∑i=1KdH(T,Si)∑i=1Ky^SiTΘ0,Si−1y^Si/mSi.\kappa=\frac{\sum_{i=1}^{K}d_{\mathcal H}(T,S_i)}{\sum_{i=1}^{K}\sqrt{\widehat y_{S_i}^{\mathsf T}\Theta_{0,S_i}^{-1}\widehat y_{S_i}/m_{S_i}}}.

    Here Θ0,Si\Theta_{0,S_i} and y^Si\widehat y_{S_i} are computed on SiS_i at initialization. The paper also allows averaging this calibration across random initializations.

  3. Knowl 3 — Coalition valuation with DAVINZ

    algorithm

    DAVINZ combines its initialization-based score with standard marginal-contribution valuations. Given KK contributors with datasets S1,…,SKS_1,\ldots,S_K, let A={1,…,K}A=\{1,\ldots,K\} and let SC=⋃j∈CSjS_C=\bigcup_{j\in C}S_j be the pooled data for coalition C⊆AC\subseteq A. For contributor ii and coalition C⊆A∖{i}C\subseteq A\setminus\{i\}, the marginal contribution is Δi,C=ν(SC∪{i})−ν(SC)\Delta_{i,C}=\nu(S_{C\cup\{i\}})-\nu(S_C), where each score uses the initialized DNN, target validation set, and score definition given in this method. The value assigned to contributor ii is ϕi=∑C⊆A∖{i}αCΔi,C\phi_i=\sum_{C\subseteq A\setminus\{i\}}\alpha_C\Delta_{i,C}.

    The procedure avoids training or fine-tuning a DNN for any coalition; it evaluates coalition scores from initialization and aggregates their marginal differences. Choosing αC=∣C∣!(K−∣C∣−1)!/K!\alpha_C=|C|!(K-|C|-1)!/K! gives the Shapley value; choosing αC=2−(K−1)\alpha_C=2^{-(K-1)} gives the Banzhaf value; and choosing αC=1\alpha_C=1 only for C=A∖{i}C=A\setminus\{i\} and zero otherwise gives leave-one-out valuation.

    Input: Contributor datasets S_1,...,S_K; validation set T; DNN and initialization θ_0; discrepancy kernel; weights α_C
    Output: Contributor values φ_1,...,φ_K
    For each coalition C needed by the selected valuation:
        Compute ν(S_C) using the initialization-based score
    For each contributor i in {1,...,K}:
        Set φ_i = 0
        For each coalition C contained in A \ {i}:
            Set Δ_i,C = ν(S_(C ∪ {i})) - ν(S_C)
            Set φ_i = φ_i + α_C Δ_i,C
    Return φ_1,...,φ_K
  4. Knowl 4 — Benchmark effectiveness and computational cost

    data/table

    The paper compares DAVINZ with validation performance (VP), influence functions (IF), and robust volume (RV), using leave-one-out values from converged DNNs as ground truth. Correlations are mean ± standard error over five evaluations; costs are minutes for evaluating 11 LOO scores over 10 datasets. VP uses 300 training epochs, while DAVINZ evaluates at initialization. The results show that DAVINZ is consistently more effective than the other training-free baselines and is much faster than VP, while its correlation with the converged-model ground truth is often comparable to VP.

    Could not parse LaTeX table
  5. Knowl 5 — Diagonal-block NTK approximation for practical evaluation

    model/method

    Exact NTK evaluation can be impractical because it requires pairwise inner products of per-example parameter gradients and may require storing an m×pm\times p gradient matrix for mm examples and pp model parameters. DAVINZ therefore uses a diagonal-block approximation: partition the dataset, in its chosen sample order, into batches of size ll; compute the exact l×ll\times l NTK submatrix within each batch; and set every between-batch NTK entry to zero. The resulting matrix is block diagonal, so its inverse is obtained by separately inverting each block. This limits gradient storage to a batch and enables batched, parallel per-sample gradient computation. The paper's large-model setup uses 100 diagonal blocks.

    On the Ising regression valuation experiment, the approximation retained high correlation with the ground truth across block counts. Cost is minutes for 11 LOO scores over 10 datasets. Increasing the number of blocks did not materially improve correlation, and the authors recommend using a large batch size that fits GPU memory.

    Could not parse LaTeX table
  6. Knowl 6 — Initialization score tracks converged validation accuracy

    empirical result

    To test whether the score can substitute for trained-model validation performance, the authors evaluated 200 datasets, each containing up to 10,000 randomly bootstrapped MNIST images, with a DNN consisting of two convolutional layers followed by a fully connected layer. The score from the domain-aware bound had a Pearson correlation of 0.954 with validation accuracy from converged DNNs and showed a nearly linear relationship. The score slightly underestimated true validation accuracy in the plotted comparison, consistent with its derivation from an upper bound on generalization error. This result supports using the score in marginal-contribution valuation without repeatedly training the predictive model.

  7. Knowl 7 — Validation-domain discrepancy gives awareness of data preference

    empirical result

    The preference experiment used 10 training datasets of 10,000 images each, formed from different mixtures of MNIST and MNIST-M; the validation set contained only MNIST-M images. Thus, increasing the MNIST-M share brought each training dataset closer to the consumer's validation domain. DAVINZ scores followed a trend similar to converged-model ground truth and trained validation performance, with Pearson correlation 0.960 against ground truth. In contrast, a score using only the in-domain NTK term did not follow the ground-truth trend, and validation-free robust volume also gave an inconsistent trend. The comparison demonstrates that the target-domain discrepancy term is important when data contributors' distributions differ from the consumer's preferred validation domain.

  8. Knowl 8 — Theoretical awareness of data quantity

    theoretical result

    Let SS contain mSm_S samples from a zero-mean input distribution in dimension dd. The paper assumes the distribution satisfies the stated data-scaling conditions—E∥x∥2=Θ(d)\mathbb E\|x\|_2=\Theta(\sqrt d), E∥x∥22=Θ(d)\mathbb E\|x\|_2^2=\Theta(d), and E∥x−Ex∥22=Ω(d)\mathbb E\|x-\mathbb E x\|_2^2=\Omega(d)—and Lipschitz concentration: for every Lipschitz function gg, deviations from its expectation have tails bounded by 2exp⁡(−ct2/∥g∥Lip2)2\exp(-ct^2/\|g\|_{\mathrm{Lip}}^2) for an absolute constant c>0c>0. It also assumes d=Θ(mα)d=\Theta(m^\alpha) for some α>0\alpha>0. Under these conditions, there is a constant β>0\beta>0 such that, with high probability, the DAVINZ score obeys

    ν(S)≥−κβmS−α/2−dH(T,S).\nu(S)\geq-\kappa\beta m_S^{-\alpha/2}-d_{\mathcal H}(T,S).

    The result implies that increasing the sample count raises this lower bound when the empirical discrepancy changes only slightly. The paper argues that discrepancy typically changes little when adding samples from an already well-sampled source domain, connecting the result to the desired preference for larger useful datasets.

  9. Knowl 9 — Stability bound for small input perturbations

    theoretical result

    Consider a dataset SS and a perturbed version SrS_r with the same labels and inputs satisfying ∥xi−xi,r∥2≤r\|x_i-x_{i,r}\|_2\leq r, where r∈[0,1]r\in[0,1]. Let Θ0\Theta_0 and Θ0,r\Theta_{0,r} be their initialization NTK matrices, and let y^\widehat y and y^r\widehat y_r be their residual vectors. Suppose ∣f(x,θ0)∣≤τ|f(x,\theta_0)|\leq\tau on both datasets, both NTK minimum eigenvalues exceed λ>0\lambda>0, and both quadratic forms y^TΘ0−1y^\widehat y^{\mathsf T}\Theta_0^{-1}\widehat y and y^rTΘ0,r−1y^r\widehat y_r^{\mathsf T}\Theta_{0,r}^{-1}\widehat y_r are at least γ>0\gamma>0. Define Δ(S,Sr)=∣dH(T,S)−dH(T,Sr)∣\Delta(S,S_r)=|d_{\mathcal H}(T,S)-d_{\mathcal H}(T,S_r)|. With high probability, the score difference satisfies

    ∣ν(Sr)−ν(S)∣≤κ2γ(O(τ)+βmS3/2rλ2)+Δ(S,Sr),|\nu(S_r)-\nu(S)|\leq\frac{\kappa}{2\sqrt{\gamma}}\left(O(\tau)+\frac{\beta m_S^{3/2}r}{\lambda^2}\right)+\Delta(S,S_r),

    where β\beta is a constant and O(τ)O(\tau) uses the paper's asymptotic notation. Thus, small perturbations yield close scores when the domain-discrepancy change is also small and the NTK matrices are adequately conditioned.

  10. Knowl 10 — Robustness bound across DNN models

    theoretical result

    For the same dataset SS of size mSm_S, let ff and f′f' be two DNNs with initialization NTK matrices Θ0,f\Theta_{0,f} and Θ0,f′\Theta_{0,f'} and residual vectors y^f\widehat y_f and y^f′\widehat y_{f'}. Assume both initialized outputs have magnitude at most τ\tau, both NTK minimum eigenvalues exceed λ>0\lambda>0, and both residual quadratic forms are at least γ>0\gamma>0. If ∥Θ0,f−Θ0,f′∥2≤ϵ\|\Theta_{0,f}-\Theta_{0,f'}\|_2\leq\epsilon, then, with high probability,

    ∣ν(S;f′)−ν(S;f)∣≤κ2γ(O(τ)+mS ϵλ2).|\nu(S;f')-\nu(S;f)|\leq\frac{\kappa}{2\sqrt{\gamma}}\left(O(\tau)+\frac{\sqrt{m_S}\,\epsilon}{\lambda^2}\right).

    The domain-discrepancy component cancels because it does not depend on the model. Consequently, similar NTKs give similar scores under these conditions; the bound also indicates that a sufficiently large minimum eigenvalue can limit score variation even when the NTK difference is substantial.

  11. Knowl 11 — Scalable Shapley valuation and data summarization

    empirical result

    The paper demonstrates two applications enabled by training-free coalition scores. For large-scale Shapley valuation, it used ResNet18 to value 100 CIFAR-10 contributors with dataset sizes increasing from 5 to 500 samples. Truncated Monte Carlo Shapley estimation used 1,000 contributor orderings (approximately the cost of 20,000 trained-model evaluations under VP); DAVINZ completed the calculation in a few hours, with individual contribution-score evaluations taking seconds. The resulting Shapley values increased overall with contributor dataset size.

    For data summarization, the authors modified the valuation procedure to evaluate only each dataset's marginal contribution relative to the empty coalition. This reduces score evaluations to linear growth in the number of datasets. In a 10-dataset MNIST experiment, removing the seven lowest-valued datasets retained validation accuracy comparable to that from the full training set; adding only the three highest-valued datasets achieved relatively high validation accuracy. These results support removing low-valued datasets first under compute constraints and selecting high-valued datasets first under a purchase budget.

Coverage note — The detailed empirical checks of data-quantity trends, noise sensitivity, cross-model valuations, and random-initialization sensitivity were not included separately; the formal property results and the principal score, preference, and efficiency experiments capture the paper's main claims without reproducing all auxiliary corroboration.

References

  1. 1.Agarwal, A., Dahleh, M., and Sarkar, T. A marketplace for data: An algorithmic solution. In Proc. ACM EC, pp. 701–726, 2019.
  2. 2.Agussurja, L., Xu, X., and Low, B. K. H. On the convergence of the Shapley value in parametric Bayesian learning games. In Proc. ICML, 2022.
  3. 3.Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R., and Wang, R. On exact computation with an infinitely wide neural net. In Proc. NeurIPS, pp. 8139–8148, 2019a.
  4. 4.Arora, S., Du, S. S., Hu, W., Li, Z., and Wang, R. Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks. In Proc. ICML, pp. 322–332, 2019b.
  5. 5.Asif, A. and Moura, J. Block matrices with l-block-banded inverse: inversion algorithms. IEEE Transactions on Signal Processing, 53(2):630–642, 2005.
  6. 6.Banzhaf, J. F. I. Weighted voting doesn’t work: A mathematical analysis. Rutgers Law Review, 19:317, 1964.
  7. 7.Basu, S., Pope, P., and Feizi, S. Influence functions in deep learning are fragile. In Proc. ICLR, 2021.
  8. 8.Ben-David, S., Blitzer, J., Crammer, K., Kulesza, A., Pereira, F., and Vaughan, J. W. A theory of learning from different domains. Machine Learning, 79(1-2):151–175, 2010.
  9. 9.Bishop, C. M. Training with noise is equivalent to Tikhonov regularization. Neural Computation, 7(1):108–116, 1995.
  10. 10.Cao, Y. and Gu, Q. Generalization bounds of stochastic gradient descent for wide and deep neural networks. In Proc. NeurIPS, pp. 10836–10846, 2019.
  11. 11.Cook, R. D. Detection of influential observation in linear regression. Technometrics, 19(1):15–18, 1977.
  12. 12.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. ImageNet: A large-scale hierarchical image database. In Proc. CVPR, pp. 248–255, 2009.
  13. 13.Dubey, P. and Shapley, L. S. Mathematical properties of the Banzhaf power index. Mathematics of Operations Research, 4(2):99–131, 1979.
  14. 14.Ganin, Y., Ustinova, E., Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., March, M., and Lempitsky, V. Domain-adversarial training of neural networks. JMLR, 17(1):2096–2030, 2016.
  15. 15.Ghorbani, A. and Zou, J. Data Shapley: Equitable valuation of data for machine learning. In Proc. ICML, pp. 2242–2251, 2019.
  16. 16.Ghorbani, A., Kim, M., and Zou, J. A distributional framework for data valuation. In Proc. ICML, pp. 3535–3544, 2020.
  17. 17.Ghorbani, A., Zou, J., and Esteva, A. Data Shapley valuation for efficient batch active learning. arXiv:2104.08312, 2021.
  18. 18.Goodfellow, I. J., Shlens, J., and Szegedy, C. Explaining and harnessing adversarial examples. In Proc. ICLR, 2015.
  19. 19.Gretton, A., Borgwardt, K. M., Rasch, M. J., Scholkopf, B., and Smola, A. A kernel two-sample test. JMLR, 13(25):723–773, 2012a.
  20. 20.Gretton, A., Sriperumbudur, B. K., Sejdinovic, D., Strathmann, H., Balakrishnan, S., Pontil, M., and Fukumizu, K. Optimal kernel choice for large-scale two-sample tests. In Proc. NeurIPS, pp. 1205–1213, 2012b.
  21. 21.Han, D., Wooldridge, M., Rogers, A., Tople, S., Ohrimenko, O., and Tschiatschek, S. Replication-robust payoff-allocation for machine learning data markets. arXiv:2006.14583, 2021.
  22. 22.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proc. CVPR, pp. 770–778, 2016.
  23. 23.IMDA. Guide to data valuation for data sharing. Technical report, The Infocomm Media Development Authority Singapore, 2019.
  24. 24.Jacot, A., Gabriel, F., and Hongler, C. Neural Tangent Kernel: Convergence and generalization in neural networks. In Proc. NeurIPS, pp. 8580–8589, 2018.
  25. 25.Jia, R., Dao, D., Wang, B., Hubis, F. A., Gurel, N. M., Li, B., Zhang, C., Spanos, C., and Song, D. Efficient task-specific data valuation for nearest neighbor algorithms. Proc. VLDB Endowment, 12(11):1610–1623, 2019a.
  26. 26.Jia, R., Dao, D., Wang, B., Hubis, F. A., Hynes, N., Gurel, N. M., Li, B., Zhang, C., Song, D., and Spanos, C. J. Towards efficient data valuation based on the Shapley value. In Proc. AISTATS, pp. 1167–1176, 2019b.
  27. 27.Koh, P. W. and Liang, P. Understanding black-box predictions via influence functions. In Proc. ICML, pp. 1885–1894, 2017.
  28. 28.Koh, P. W., Ang, K.-S., Teo, H., and Liang, P. S. On the accuracy of influence functions for measuring group effects. In Proc. NeurIPS, pp. 5254–5264, 2019.
  29. 29.Krizhevsky, A., Hinton, G., et al. Learning multiple layers of features from tiny images. Technical report, 2009.
  30. 30.Laurent, B. and Massart, P. Adaptive estimation of a quadratic functional by model selection. Annals of Statistics, 28:1302–1338, 2000.
  31. 31.Lecun, Y., Bottou, L., Bengio, Y., and Haffner, P. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  32. 32.Lee, J., Xiao, L., Schoenholz, S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J. Wide neural networks of any depth evolve as linear models under gradient descent. In Proc. NeurIPS, pp. 8572–8583, 2019.
  33. 33.Lo, B. and DeMets, D. L. Incentives for clinical trialists to share data. New England Journal of Medicine, 375(12):1112–1115, 2016.
  34. 34.Long, M., Cao, Y., Wang, J., and Jordan, M. I. Learning transferable features with deep adaptation networks. In Proc. ICML, pp. 97–105, 2015.
  35. 35.Mangalam, K. and Prabhu, V. U. Do deep neural networks learn shallow learnable examples first? In Proc. ICML Workshop on Identifying and Understanding Deep Learning Phenomena, 2019.
  36. 36.Mills, K. and Tamblyn, I. Big graphene dataset. https://nrc-digital-repository.canada.ca/eng/view/object/?id=9f09901d-0736-4204-a35d-0c88ffb8da3b, 2019.
  37. 37.Nguyen, Q., Mondelli, M., and Montufar, G. F. Tight bounds on the smallest eigenvalue of the neural tangent kernel for deep ReLU networks. In Proc. ICML, pp. 8119–8129, 2021.
  38. 38.Sejdinovic, D., Sriperumbudur, B., Gretton, A., and Fukumizu, K. Equivalence of distance-based and RKHS-based statistics in hypothesis testing. Annals of Statistics, 41(5):2263–2291, 2013.
  39. 39.Shapley, L. S. A value for n-person games. In Contributions to the Theory of Games (AM-28), Volume II, chapter 17, pp. 307–318. Princeton University Press, 1953.
  40. 40.Shu, Y., Dai, Z., Wu, Z., and Low, B. K. H. Unifying and boosting gradient-based training-free neural architecture search. arXiv:2201.09785, 2022.
  41. 41.Sim, R. H. L., Zhang, Y., Chan, M. C., and Low, B. K. H. Collaborative machine learning with incentive-aware model rewards. In Proc. ICML, pp. 8927–8936, 2020.
  42. 42.Sim, R. H. L., Xu, X., and Low, B. K. H. Data valuation in machine learning: “ingredients”, strategies, and open challenges. In Proc. IJCAI, 2022.
  43. 43.Simonyan, K. and Zisserman, A. Very deep convolutional networks for large-scale image recognition. In Proc. ICLR, 2015.
  44. 44.Tay, S. S., Xu, X., Foo, C. S., and Low, B. K. H. Incentivizing collaboration in machine learning via synthetic data rewards. In Proc. AAAI, 2022.
  45. 45.Wang, T., Yang, Y., and Jia, R. Learnability of learning performance and its application to data valuation. arXiv:2107.06336, 2021a.
  46. 46.Wang, T., Zeng, Y., Jin, M., and Jia, R. A unified framework for task-driven data quality management. arXiv:2106.05484, 2021b.
  47. 47.Xu, X., Lyu, L., Ma, X., Miao, C., Foo, C. S., and Low, B. K. H. Gradient driven rewards to guarantee fairness in collaborative machine learning. In Proc. NeurIPS, pp. 16104–16117, 2021a.
  48. 48.Xu, X., Wu, Z., Foo, C. S., and Low, B. K. H. Validation free and replication robust volume-based data valuation. In Proc. NeurIPS, pp. 10837–10848, 2021b.
  49. 49.Yang, G. and Littwin, E. Tensor programs IIb: Architectural universality of neural tangent kernel training dynamics. In Proc. ICML, pp. 11762–11772, 2021.
  50. 50.Yoon, J., Arik, S. O., and Pfister, T. Data valuation using reinforcement learning. In Proc. ICML, pp. 10842–10851, 2020.
  51. 51.Zaheer, M., Kottur, S., Ravanbakhsh, S., Poczos, B., Salakhutdinov, R. R., and Smola, A. J. Deep sets. In Proc. NeurIPS, pp. 3394–3404, 2017.

Citation

MLA
Wu, Z., et al. “DAVINZ: Data Valuation Using Deep Neural Networks at Initialization”. International Conference on Machine Learning, vol. 162, 2022, pp. 24150–76, https://proceedings.mlr.press/v162/wu22j.html.
APA
Wu, Z., Shu, Y., & Low, B. K. H. (2022). DAVINZ: Data Valuation using Deep Neural Networks at Initialization. International Conference on Machine Learning, 162, 24150–24176. https://proceedings.mlr.press/v162/wu22j.html
Chicago
Wu, Z., Y. Shu, and B. K. H. Low. 2022. “DAVINZ: Data Valuation Using Deep Neural Networks at Initialization”. International Conference on Machine Learning 162: 24150–76. https://proceedings.mlr.press/v162/wu22j.html.
Harvard
Wu, Z., Shu, Y. and Low, B.K.H. (2022) “DAVINZ: Data Valuation using Deep Neural Networks at Initialization”, International Conference on Machine Learning. PMLR, pp. 24150–24176. Available at: https://proceedings.mlr.press/v162/wu22j.html.
Vancouver
1. Wu Z, Shu Y, Low BKH (2022) DAVINZ: Data Valuation using Deep Neural Networks at Initialization. In: International Conference on Machine Learning. PMLR, pp 24150–24176

BibTeX

@InProceedings{pmlr-v162-wu22j,
  title = 	 {{DAVINZ}: Data Valuation using Deep Neural Networks at Initialization},
  author =       {Wu, Zhaoxuan and Shu, Yao and Low, Bryan Kian Hsiang},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {24150--24176},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/wu22j/wu22j.pdf},
  url = 	 {https://proceedings.mlr.press/v162/wu22j.html},
  abstract = 	 {Recent years have witnessed a surge of interest in developing trustworthy methods to evaluate the value of data in many real-world applications (e.g., collaborative machine learning, data marketplaces). Existing data valuation methods typically valuate data using the generalization performance of converged machine learning models after their long-term model training, hence making data valuation on large complex deep neural networks (DNNs) unaffordable. To this end, we theoretically derive a domain-aware generalization bound to estimate the generalization performance of DNNs without model training. We then exploit this theoretically derived generalization bound to develop a novel training-free data valuation method named data valuation at initialization (DAVINZ) on DNNs, which consistently achieves remarkable effectiveness and efficiency in practice. Moreover, our training-free DAVINZ, surprisingly, can even theoretically and empirically enjoy the desirable properties that training-based data valuation methods usually attain, thus making it more trustworthy in practice.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/