Deep Unlearning via Randomized Conditionally Independent Hessians

Ronak MehtaSourav PalVikas SinghSathya N. Ravi

article2022CVPR119 citations

Proposes a scalable approximate machine unlearning method that avoids full Hessian inversion by using a conditional independence coefficient to identify and update only the most relevant parameter subsets.

Listen

Privacy regulations such as GDPR and CCPA, alongside enforcement actions by bodies like the Federal Trade Commission, increasingly mandate the "right to be forgotten." This requires organizations to remove specific user data from trained artificial intelligence models. While completely retraining models from scratch without the deleted data satisfies these rules, doing so for large-scale production architectures is computationally prohibitive, slow, and expensive. Conventional approximate "unlearning" techniques attempt to reverse parameter updates using second-order loss curvature, known as the Hessian matrix. However, these methods fail to scale because inverting a Hessian matrix across millions of parameters requires intractable computation.

The article demonstrates an efficient framework for approximate machine unlearning in deep neural networks by selectively updating only the small subset of parameters most tied to the data targeted for deletion. The main objective is to eliminate the need for full-network matrix inversions, thereby enabling rapid data scrubbing in previously infeasible large-scale vision and natural language processing models without sacrificing overall predictive performance.

To achieve this, the authors introduce a randomized conditional independence metric called L-CODEC, integrated into a feature selection procedure called L-FOCI. By adding slight perturbations to an input sample and tracking internal network activations, the method identifies the "Markov Blanket"—the minimal subset of parameters sufficient to explain the model's output on that specific sample. The unlearning step then computes and applies an approximate Hessian update solely to this parameter subset using block-coordinate updates combined with differential privacy noise. The approach was evaluated across standard benchmarks (MNIST and CIFAR-10), large image models (ResNet-50 on Market-1501 for person re-identification and VGGFace with over 24 million parameters), and natural language transformers (DistilBERT on legal provisions).

The evaluation yielded several key findings. First, parameter selection via L-FOCI significantly reduces computational complexity from cubic in total model size to cubic in the much smaller selected subset size, completing an unlearning step on a 24-million parameter ResNet-50 in approximately three minutes. Second, the method successfully eliminates the target sample's influence—causing prediction accuracy or F1 score on the scrubbed classes to drop sharply—while maintaining baseline validation performance on remaining data. Third, unlearning capacity is constrained by target privacy budgets: under loose privacy constraints, models supported more than 100 sample deletions without degradation, whereas strict privacy constraints limited deletions to approximately 6 to 21 samples before accuracy was impacted. Finally, theoretical analysis confirmed that unlearning updates converge exponentially fast relative to target precision gaps, with residual gradient error decaying rapidly at a rate inversely proportional to the square of the sample size.

These findings indicate that organizations can comply with strict privacy mandates and regulatory model-deletion orders without incurring the recurring financial and computational costs of full retraining. By isolating updates to functionally relevant sub-networks, practitioners can operationalize data removal in complex production models while balancing privacy guarantees against model utility.

Senior leaders should consider integrating selective parameter unlearning pipelines into compliance and data governance workflows, particularly for high-risk applications like biometric identification. When deploying these techniques, organizations must determine operational trade-offs: strict privacy budgets require periodic full retraining after a small number of deletions, while moderate privacy requirements permit sustained, continuous unlearning updates. Organizations should conduct pilot tests on internal models to evaluate how parameter selection thresholds affect their specific performance metrics.

Confidence in these findings is supported by theoretical proofs and consistent empirical performance across diverse computer vision and language tasks. Nonetheless, decision-makers should note certain limitations: the method provides approximate rather than exact mathematical deletion, relies on regularization assumptions to navigate non-convex deep learning landscapes, and requires careful tuning of noise and sample-perturbation parameters for optimal performance.

arXiv: 2204.07655
Cover for Deep Unlearning via Randomized Conditionally Independent Hessians

Abstract

Recent legislation has led to interest in machine unlearning, i.e., removing specific training samples from a predictive model as if they never existed in the training dataset. Unlearning may also be required due to corrupted/adversarial data or simply a user’s updated privacy requirement. For models which require no training (k-NN), simply deleting the closest original sample can be effective. But this idea is inapplicable to models which learn richer representations. Recent ideas leveraging optimization-based updates scale poorly with the model dimension d, due to inverting the Hessian of the loss function. We use a variant of a new conditional independence coefficient, L-CODEC, to identify a subset of the model parameters with the most semantic overlap on an individual sample level. Our approach completely avoids the need to invert a (possibly) huge matrix. By utilizing a Markov blanket selection, we premise that L-CODEC is also suitable for deep unlearning, as well as other applications in vision. Compared to alternatives, L-CODEC makes approximate unlearning possible in settings that would otherwise be infeasible, including vision models used for face recognition, person re-identification and NLP models that may require unlearning samples identified for exclusion. Code is available at https://github.com/vsingh-group/LCODEC-deep-unlearning

Table of Contents

  • 1. Introduction
  • 2. Problem Setup for Unlearning
  • 3. Related Work
  • 4. Randomized Markovian Block Coordinate Unlearning
  • 4.1. Efficient Subset Selection that is also Sufficient for Predictive Purposes
  • 5. Deep Unlearning via L-FOCI Hessians
  • 5.1. Theoretical Analysis
  • 6. L-FOCI in Generic ML Settings
  • 7. L-FOCI for Machine Unlearning
  • 7.1. Compare to Full Hessian Computation
  • 7.2. Removal in NLP models
  • 7.3. Removal from Pretrained Models
  • 7.4. Removal from Person re-identification model
  • 8. Conclusion
  • References

Knowls

  1. Knowl 1 — L-CODEC Randomized Conditional Dependence Measure

    model/method

    The Coefficient of Conditional Dependence (CODEC) tests conditional independence between variables AA and BB given CC by evaluating nearest-neighbor rankings. In settings with discrete or highly repeated values, tie-breaking in metric search structures such as kdkd-trees requires expanding all equivalent nodes, leading to prohibitive computational overhead.

    L-CODEC resolves this by injecting continuous Gaussian perturbations to break ties without altering relative rankings across distinct values:

    TL:=T(B~,C~∣A~)T_L := T\left(\tilde{B}, \tilde{C} \mid \tilde{A}\right)

    where A~=A+N(0,σ2)\tilde{A} = A + \mathcal{N}(0, \sigma^2), B~=B+N(0,σ2)\tilde{B} = B + \mathcal{N}(0, \sigma^2), and C~=C+N(0,σ2)\tilde{C} = C + \mathcal{N}(0, \sigma^2). The noise standard deviation σ\sigma is chosen to be strictly smaller than the minimum nonzero pairwise distance across the data points. In expectation, TLT_L preserves the underlying dependence measure while guaranteeing unique nearest neighbors. This perturbation aligns with the randomization criterion for conditional independence on Borel spaces, where A⊥B∣CA \perp B \mid C if and only if A=a.s.h(B,U)A \stackrel{\text{a.s.}}{=} h(B, U) for a measurable function hh and an independent uniform variable U∼Uniform(0,1)U \sim \text{Uniform}(0, 1).

  2. Knowl 2 — Parameter Slice Selection for Deep Networks via Conditional Independence

    model/method

    For a deep neural network hypothesis w∈Ww \in \mathcal{W} parameterized by Θ={1,…,d}\Theta = \{1, \dots, d\} trained on dataset SS, approximate machine unlearning assumes that for any target sample z′∈Sz' \in S, there exists a sufficient parameter subset P∗⊂ΘP^* \subset \Theta satisfying:

    f(z′)⊥wΘ∖P∗economics∗∣wP∗∗f(z') \perp w_{\Theta \setminus P^* economics}^* \mid w_{P^*}^*

    Because parameters and samples are deterministic post-training, empirical distributions are generated by perturbing the target sample z′=(x,y)z' = (x, y) with random Gaussian noise ξj∼N(0,σ2)\xi^j \sim \mathcal{N}(0, \sigma^2) across j=1,…,mj = 1, \dots, m perturbations, yielding perturbed inputs xj=x+ξjx^j = x + \xi^j. For each perturbation, the model computes layer-wise linear activation slices and the resulting task loss (l(xj),a1j,…,a∣L∣j)(l(x^j), a_1^j, \dots, a_{|\mathcal{L}|}^j), where an activation coordinate al,ka_{l, k} corresponds to a weight slice wl[:,k]w_l[:, k] in layer l∈Ll \in \mathcal{L}.

    The conditional independence condition is evaluated over activations:

    f(z′)⊥aΘ∖P∗∣aP∗f(z') \perp a_{\Theta \setminus P}^* \mid a_P^*

    Applying Feature Ordering by Conditional Independence with L-CODEC (L-FOCI) over the collection of perturbation tuples identifies the Markov Blanket P⊆ΘP \subseteq \Theta of parameters that directly account for the model's loss on sample z′z'.

  3. Knowl 3 — Machine Unlearning via Conditional Dependence Block Selection

    algorithm

    This procedure unlearns a specific sample z′z' from a converged model w^\hat{w} without full dataset retraining or full-network Hessian inversion. It applies forward perturbations to isolate the Markov Blanket subset of parameters P∗⊆ΘP^* \subseteq \Theta via L-FOCI, estimates the Hessian on P∗P^*, and performs a blockwise Newton step with differential privacy noise.

    Input: Trained model w^\hat{w}, dataset size nn, training gradients ∇F(w^)\nabla F(\hat{w}), sample to unlearn z′∈Sz' \in S, number of perturbations mm, perturbation variance σ2\sigma^2
    Output: Unlearned model parameters w′w'
    for j=1j = 1 to mm do
        Sample perturbation ξj∼N(0,σ2)\xi^j \sim \mathcal{N}(0, \sigma^2)
        Set perturbed sample z′,j=z′+ξjz'^{,j} = z' + \xi^j
        Compute activations and loss: lj,aj=f(z′,j)l^j, a^j = f(z'^{,j})
    end for
    Compute sufficient parameter subset P∗=L-FOCI({lj,aj}j=1m)P^* = \text{L-FOCI}(\{l^j, a^j\}_{j=1}^m)
    Compute Hessian ∇P∗2f(w^,z′)\nabla^2_{P^*} f(\hat{w}, z') via finite differences
    Compute subset Hessian adjustment:
        HP∗′=1n−1(n∇P∗2F(w^)−∇P∗2f(w^,z′))H'_{P^*} = \frac{1}{n-1}\left(n \nabla^2_{P^*} F(\hat{w}) - \nabla^2_{P^*} f(\hat{w}, z')\right)
    Update selected parameters:
        wP∗′=w^P∗+1n−1(HP∗′)−1∇f(w^,z′)P∗w'_{P^*} = \hat{w}_{P^*} + \frac{1}{n-1}(H'_{P^*})^{-1} \nabla f(\hat{w}, z')_{P^*}
    Set unselected parameters:
        wΘ∖P∗′=w^Θ∖P∗w'_{\Theta \setminus P^*} = \hat{w}_{\Theta \setminus P^*}
    return w′w'
  4. Knowl 4 — Residual Gradient Norm Gap for Subset Unlearning

    theoretical result

    Let wFoci−w^-_{\text{Foci}} be the parameter vector obtained by executing L-FOCI parameter subset unlearning on sample z′z', and let wFull−w^-_{\text{Full}} be the parameter vector obtained by applying the full Newton-based unlearning update over all dd parameters. Evaluated over the remaining dataset D′=S∖{z′}D' = S \setminus \{z'\}, the gap in residual gradient ℓ2\ell_2-norms satisfies:

    ∣∥∇F(wFoci−,D′)∥2−∥∇F(wFull−,D′)∥2∣=O(1n2)\left| \|\nabla F(w^-_{\text{Foci}}, D')\|_2 - \|\nabla F(w^-_{\text{Full}}, D')\|_2 \right| = \mathcal{O}\left(\frac{1}{n^2}\right)

    where nn is the total number of training samples before deletion. Because parameter modifications outside the selected subset PP propagate to other layers scaled by 1/n1/n, a Taylor expansion about the updated layer activations confirms that localized subset unlearning does not disrupt performance on the retained dataset.

  5. Knowl 5 — Forgetting Guarantee and Convergence Rate of Slice-Based Unlearning

    theoretical result

    Assume that the layer-wise sampling probabilities computed by L-CODEC are strictly non-zero across all model layers. For specified target forgetting parameters ϵ>0\epsilon > 0 and δ>0\delta > 0, the randomized slice-based unlearning procedure satisfies (ϵ′,δ′)(\epsilon', \delta')-forgetting:

    P(U(A(S),z′)∈W)≤eϵ′P(A(S∖{z′})∈W)+δ′\mathbb{P}(\mathcal{U}(\mathcal{A}(S), z') \in \mathcal{W}) \le e^{\epsilon'} \mathbb{P}(\mathcal{A}(S \setminus \{z'\}) \in \mathcal{W}) + \delta'

    where ϵ′>ϵ\epsilon' > \epsilon and δ′>δ\delta' > \delta define the desired precision tolerances. Furthermore, iteratively applying the randomized unlearning update converges exponentially fast in expectation with respect to the precision gaps gϵ=ϵ′−ϵ>0g_\epsilon = \epsilon' - \epsilon > 0 and gδ=δ′−δ>0g_\delta = \delta' - \delta > 0, requiring at most:

    O(log⁡1gϵlog⁡1gδ)\mathcal{O}\left(\log \frac{1}{g_\epsilon} \log \frac{1}{g_\delta}\right)

    iterations to output a valid (ϵ′,δ′)(\epsilon', \delta')-unlearned hypothesis.

  6. Knowl 6 — Computational Complexity of L-FOCI Deep Unlearning

    model/method

    Standard one-shot approximate unlearning schemes based on Newton updates invert the full empirical Hessian matrix of size d×dd \times d, incurring a time complexity of O(d3)\mathcal{O}(d^3), where dd is the total parameter count. For deep networks where dd is in the millions or hundreds of millions, this operation is computationally intractable.

    L-FOCI unlearning reduces the dimension of the parameter update by identifying a sufficient subnetwork block P⊆ΘP \subseteq \Theta of size p=∣P∣≪dp = |P| \ll d. Computing mm forward perturbation passes, running L-FOCI feature ordering, and inverting the subset Hessian matrix requires:

    O(md+dmlog⁡m+p3)\mathcal{O}(md + dm \log m + p^3)

    time operations. Because p≪dp \ll d and m≪dm \ll d, this formulation makes second-order unlearning updates feasible on deep neural network architectures.

  7. Knowl 7 — Markov Blanket Identification Benchmark on 3D-Bullseye Data

    data/table

    Markov Blanket identification was evaluated on the 3D-Bullseye synthetic distribution benchmark, comparing standard Conditional Mutual Information (CIT), L-CODEC combined with CIT, and L-CODEC combined with L-FOCI across raw data and learned feature maps.

    Raw Data Feature Maps
    Method TPR FPR Time (s) TPR FPR Time (s)
    CIT 0.75 0.50 5124.22 0.875 0.00 516.19
    L-CODEC + CIT 1.00 0.50 402.10 0.750 0.00 117.29
    L-CODEC + L-FOCI N/A N/A N/A 0.833 0.50 0.464

    Replacing CMI with L-CODEC reduces runtime by over an order of magnitude on raw data (from 5124.22 s5124.22\text{ s} to 402.10 s402.10\text{ s}) while achieving higher True Positive Rate (1.001.00 vs 0.750.75). Applying L-FOCI directly on feature maps completes subset identification in 0.464 s0.464\text{ s}—over 1100×1100\times faster than baseline CIT—with competitive feature recovery.

  8. Knowl 8 — Transformer Provision Unlearning on LEDGAR Benchmark

    empirical result

    Unlearning was evaluated on a fine-tuned DistilBERT transformer classifier trained on the LEDGAR multilabel contract dataset (110,156110{,}156 provisions across the top 13 provision classes). Provisions from specific classes were scrubbed under varying differential privacy budgets ϵ\epsilon.

    The maximum number of supported sequential removals before substantial model degradation depends directly on ϵ\epsilon:

    • For ϵ=0.1\epsilon = 0.1: >100> 100 supported removals for both Governing Laws and Terminations.
    • For ϵ=0.01\epsilon = 0.01: >100> 100 supported removals for both Governing Laws and Terminations.
    • For ϵ=0.001\epsilon = 0.001: 1818 supported removals for Governing Laws and 2121 for Terminations.
    • For ϵ=0.0005\epsilon = 0.0005: 66 supported removals for Governing Laws and 77 for Terminations.

    Across removals, the Micro F1 score on the scrubbed class drops rapidly toward zero, while overall Micro F1 score across remaining classes exhibits only a gradual degradation.

  9. Knowl 9 — Pretrained Vision Model Scrubbing in Face Recognition and Person Re-Identification

    empirical result

    L-FOCI unlearning was applied to large pretrained vision networks for identity removal:

    1. Face Recognition (VGGFace): A model classifying 26222622 celebrity identities from ≈1\approx 1 million images contains large fully connected layers (25088×409625088 \times 4096). A memory-efficient variant of L-FOCI selecting the single weight slice with highest conditional dependence on the output loss was evaluated at ϵ=10−5\epsilon = 10^{-5}. Scrubbing consecutive images of an individual causes targeted class accuracy to collapse rapidly while residual class accuracy remains stable.
    2. Person Re-Identification (Market-1501): A ResNet-50 backbone (approx24M\\approx 24\text{M} parameters) executes a complete unlearning step in approximately 33 minutes. At ϵ=0.1\epsilon = 0.1, all target samples of a designated identity can be scrubbed without degrading overall mean Average Precision (mAP). At ϵ=0.0005\epsilon = 0.0005, up to 1010 sequential removals per identity are supported. Network activation maps show that post-unlearning representations for the scrubbed identity become diffuse and non-informative, while feature activations for non-scrubbed individuals remain invariant.
  10. Knowl 10 — Spurious Feature Regularization via L-FOCI Markov Blanket Selection

    model/method

    To prevent deep neural networks from learning spurious correlations with outside confounding factors S\mathcal{S}, adding regularization penalties ∑S∈SRS(θ)\sum_{S \in \mathcal{S}} R_S(\theta) over every known attribute degrades primary task optimization when ∣S∣|\mathcal{S}| is large.

    Instead, L-FOCI identifies the minimal Markov Blanket subset MB(Y)⊂S\text{MB}(Y) \subset \mathcal{S} such that the target label YY is conditionally independent of all other factors given MB(Y)\text{MB}(Y):

    Y⊥S∖MB(Y)∣MB(Y)Y \perp \mathcal{S} \setminus \text{MB}(Y) \mid \text{MB}(Y)

    Regularization (such as Gradient Reversal Layers) is then applied strictly to attributes within MB(Y)\text{MB}(Y). On the CelebA dataset predicting the No Beard attribute, regularizing only over the L-FOCI-identified attribute subset maintained higher overall classification accuracy and balanced predictive performance across correlated subsets compared to regularizing over all features or over randomly chosen feature subsets.

Coverage note — None was omitted. All principal contributions—including the L-CODEC randomized statistic, parameter slice formulation, the unlearning algorithm, theoretical bounds, complexity analysis, Markov Blanket benchmarks, and empirical evaluations on transformers, face recognition, and person re-ID—are captured.

References

  1. 1.Mona Azadkia and Sourav Chatterjee. A simple measure of conditional dependence. arXiv preprint arXiv:1910.12327, 2019. 3, 4
  2. 2.Samyadeep Basu, Phil Pope, and Soheil Feizi. Influence functions in deep learning are fragile. In International Conference on Learning Representations, 2021. 6
  3. 3.David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba. Network dissection: Quantifying interpretability of deep visual representations. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6541–6549, 2017. 2
  4. 4.Lucas Bourtoule, Varun Chandrasekaran, Christopher A Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot. Machine unlearning. In 2021 IEEE Symposium on Security and Privacy (SP), pages 141–159. IEEE, 2021. 3
  5. 5.Yinzhi Cao and Junfeng Yang. Towards making systems forget with machine unlearning. In 2015 IEEE Symposium on Security and Privacy, pages 463–480. IEEE, 2015. 3
  6. 6.Nicholas Carlini, Samuel Deng, Sanjam Garg, Somesh Jha, Saeed Mahloujifar, Mohammad Mahmoody, Shuang Song, Abhradeep Thakurta, and Florian Tramer. An attack on instahide: Is private learning possible with instance encoding? arXiv preprint arXiv:2011.05315, 2020. 1
  7. 7.Nicholas Carlini, Chang Liu, Ulfar Erlingsson, Jernej Kos, and Dawn Song. The secret sharer: Evaluating and testing unintended memorization in neural networks. In 28th {USENIX} Security Symposium ({USENIX} Security 19), pages 267–284, 2019. 1
  8. 8.Sourav Chatterjee. A new coefficient of correlation. Journal of the American Statistical Association, 0(0):1–21, 2020. 4
  9. 9.Jelena Diakonikolas and Lorenzo Orecchia. The approximate duality gap technique: A unified theory of first-order methods. SIAM Journal on Optimization, 29(1):660–689, 2019. 5
  10. 10.Mahyar Fazlyab, Alexander Robey, Hamed Hassani, Manfred Morari, and George J Pappas. Efficient and accurate estimation of lipschitz constants for deep neural networks. In Advances in Neural Information Processing Systems (NeurIPS), 2019. 6
  11. 11.Ruth Fong and Andrea Vedaldi. Net2vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8730–8738, 2018. 2
  12. 12.FTC. California company settles ftc allegations it deceived consumers about use of facial recognition in photo storage app, Jan 2021. 1
  13. 13.A Ginart, M Guan, G Valiant, and J Zou. Making ai forget you: Data deletion in machine learning. Advances in neural information processing systems, 2019. 2, 3
  14. 14.Aditya Golatkar, Alessandro Achille, Avinash Ravichandran, Marzia Polito, and Stefano Soatto. Mixed-privacy forgetting in deep networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 792–801, June 2021. 2, 3
  15. 15.Aditya Golatkar, Alessandro Achille, and Stefano Soatto. Eternal sunshine of the spotless net: Selective forgetting in deep networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9304–9312, 2020. 3
  16. 16.Aditya Golatkar, Alessandro Achille, and Stefano Soatto. Forgetting outside the box: Scrubbing deep networks of information accessible from input-output observations. In European Conference on Computer Vision, pages 383–398. Springer, 2020. 2, 3
  17. 17.Eduard Gorbunov, Filip Hanzely, and Peter Richtárik. A unified theory of sgd: Variance reduction, sampling, quantization and coordinate descent. In International Conference on Artificial Intelligence and Statistics, pages 680–690. PMLR, 2020. 5
  18. 18.Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik. Sgd: General analysis and improved rates. In International Conference on Machine Learning, pages 5200–5209. PMLR, 2019. 5
  19. 19.Chuan Guo, Tom Goldstein, Awni Hannun, and Laurens Van Der Maaten. Certified data removal from machine learning models. In International Conference on Machine Learning, pages 3832–3842. PMLR, 2020. 3
  20. 20.Jules. Harvey, Adam. LaPlace. Exposing.ai, 2021. 1
  21. 21.Gary B Huang, Marwan Mattar, Tamara Berg, and Eric Learned-Miller. Labeled faces in the wild: A database forstudying face recognition in unconstrained environments. In Workshop on faces in’Real-Life’Images: detection, alignment, and recognition, 2008. 7
  22. 22.Zachary Izzo, Mary Anne Smart, Kamalika Chaudhuri, and James Zou. Approximate data deletion from machine learning models. In International Conference on Artificial Intelligence and Statistics, pages 2008–2016. PMLR, 2021. 3
  23. 23.Masayuki Karasuyama and Ichiro Takeuchi. Multiple incremental decremental learning of support vector machines. Advances in neural information processing systems, 22:907–915, 2009. 3
  24. 24.Kate Kaye. The ftc’s new enforcement weapon spells death for algorithms, Mar 2022. 1
  25. 25.Seth Neel, Aaron Roth, and Saeed Sharifi-Malvajerdi. Descent-to-delete: Gradient-based methods for machine unlearning. In Algorithmic Learning Theory, pages 931–962. PMLR, 2021. 3
  26. 26.Peter Orbanz. Probability theory, Spring 2016. 4
  27. 27.Omkar M. Parkhi, Andrea Vedaldi, and Andrew Zisserman. Deep face recognition. In British Machine Vision Conference, 2015. 7
  28. 28.Enrique Romero, Ignacio Barrio, and Lluı́s Belanche. Incremental and decremental learning for linear support vector machines. In International Conference on Artificial Neural Networks, pages 209–218. Springer, 2007. 3
  29. 29.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019. 7
  30. 30.Ayush Sekhari, Jayadev Acharya, Gautam Kamath, and Ananda Theertha Suresh. Remember what you want to forget: Algorithms for machine unlearning, 2021. 2, 5, 6
  31. 31.Umut Simsekli, Lingjiong Zhu, Yee Whye Teh, and Mert Gurbuzbalaban. Fractional underdamped langevin dynamics: Retargeting sgd with momentum under heavy-tailed gradient noise. In International Conference on Machine Learning, pages 8970–8980. PMLR, 2020. 5
  32. 32.Yiyou Sun, Sathya N. Ravi, and Vikas Singh. Adaptive activation thresholding: Dynamic routing type behavior for interpretability in convolutional neural networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2019. 2
  33. 33.Yann Traonmilin and Jean-Franc¸ois Aujol. The basins of attraction of the global minimizers of the non-convex sparse spike estimation problem. Inverse Problems, 36(4):045003, feb 2020. 6
  34. 34.Don Tuggener, Pius von Däniken, Thomas Peetz, and Mark Cieliebak. Ledgar: a large-scale multi-label corpus for text classification of legal provisions in contracts. In 12th Language Resources and Evaluation Conference (LREC) 2020, pages 1228–1234. European Language Resources Association, 2020. 7
  35. 35.Stephen Wright, Jorge Nocedal, et al. Numerical optimization. Springer Science, 35(67-68):7, 1999. 6
  36. 36.Alan Yang, AmirEmad Ghassami, Maxim Raginsky, Negar Kiyavash, and Elyse Rosenbaum. Model-augmented conditional mutual information estimation for feature selection. In Conference on Uncertainty in Artificial Intelligence, pages 1139–1148. PMLR, 2020. 4, 6
  37. 37.Zhanpeng Zeng, Yunyang Xiong, Sathya Ravi, Shailesh Acharya, Glenn M Fung, and Vikas Singh. You only sample (almost) once: Linear cost self-attention via bernoulli sampling. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 12321–12332. PMLR, 18–24 Jul 2021. 6
  38. 38.Liang Zheng, Liyue Shen, Lu Tian, Shengjin Wang, Jingdong Wang, and Qi Tian. Scalable person re-identification: A benchmark. In Proceedings of the IEEE international conference on computer vision, pages 1116–1124, 2015. 8

Citation

MLA
Mehta, R., et al. “Deep Unlearning via Randomized Conditionally Independent Hessians”. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, Pp. 10422-10431, 2022, http://arxiv.org/abs/2204.07655v2.
APA
Mehta, R., Pal, S., Singh, V., & Ravi, S. N. (2022). Deep Unlearning via Randomized Conditionally Independent Hessians. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, Pp. 10422-10431. http://arxiv.org/abs/2204.07655v2
Chicago
Mehta, R., S. Pal, V. Singh, and S. N. Ravi. 2022. “Deep Unlearning via Randomized Conditionally Independent Hessians”. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, Pp. 10422-10431. http://arxiv.org/abs/2204.07655v2.
Harvard
Mehta, R. et al. (2022) “Deep Unlearning via Randomized Conditionally Independent Hessians”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 10422-10431 [Preprint]. Available at: http://arxiv.org/abs/2204.07655v2.
Vancouver
1. Mehta R, Pal S, Singh V, Ravi SN (2022) Deep Unlearning via Randomized Conditionally Independent Hessians. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 10422-10431

BibTeX

@article{mehta2022deep,
  title = {Deep Unlearning via Randomized Conditionally Independent Hessians},
  author = {Mehta, Ronak and Pal, Sourav and Singh, Vikas and Ravi, Sathya N.},
  year = {2022},
  journal = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 10422-10431},
  url = {http://arxiv.org/abs/2204.07655v2},
  eprint = {2204.07655}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE