Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack

Tiansheng HuangSihao HuFatih IlhanSelim F. TekinLing Liu

article2024NeurIPS78 citations

Proposes a bi-state optimization method incorporating a proximal constraint to protect large language models against harmful fine-tuning attacks without degrading downstream task performance.

Listen

Commercial fine-tuning services allow users to customize large language models with proprietary data. However, this creates a major security and liability risk: if user data inadvertently or maliciously includes harmful content, it can override previous safety alignment and cause the model to generate unsafe responses. Existing defenses either fail when fine-tuning involves extensive steps or require expensive, full-scale retraining for every individual request.

The article evaluates a computationally lightweight defense implemented directly during the fine-tuning stage. It demonstrates a novel method designed to retain safety guardrails against harmful fine-tuning data without undermining task performance on downstream user applications.

The authors conducted empirical experiments and convergence analyses to investigate multi-task optimization between safe alignment data and user datasets. They evaluated the framework across three foundation models (Llama2-7B, Opt-2.7B, and Mistral-7B) on four benchmark tasks (SST2, AGNEWS, GSM8K, and AlpacaEval) under varying ratios of harmful data. To address optimization instability caused by alternating between safety and user data, the proposed method—Lazy Safety Alignment (Lisa)—adds a proximal constraint to limit model drift between training phases.

The primary findings demonstrate that simple alternating optimization between alignment and user datasets degrades if computational steps allocated to safety are limited, which causes the model parameters to drift excessively. By introducing a proximal penalty to control this drift, Lisa reduces the average harmful output rate by 7.07% compared to alignment-stage defenses (Vaccine-SFT) and by 3.68% compared to data-mixing baselines (Vlguard), while keeping task accuracy intact within a 0.59% variance. Across distinct architectures, Lisa reduced harmful response rates by 11.2% to 11.9% on reasoning tasks and remained robust even when harmful training data ratios approached 100%. Furthermore, system measurements confirmed that Lisa introduces modest overhead, requiring only about 8.3% more execution time and 3.14 GB of additional GPU memory compared to standard fine-tuning.

These results show that service providers can mitigate liability and safety risks without sacrificing customization quality or incurring heavy computational penalties. Unlike purely preventive alignment strategies, Lisa operates during the customization phase and functions effectively even when user data filtering exhibits false negatives.

Organizations providing fine-tuning services should consider adopting proximal-constrained alternating optimization in their training workflows. For stronger protection, combining input data filtering with Lisa is recommended to eliminate residual risks. Future work should expand the method beyond supervised fine-tuning to evaluate compatibility with reinforcement learning from human feedback and test deployments on interactive agent applications.

While the theoretical convergence and experimental results are consistent across multiple benchmarks, evaluations were limited to supervised fine-tuning setups on medium-scale models. Stakeholders should conduct pilot assessments on production workloads and higher-parameter models to determine exact resource overheads and performance trade-offs.

Cover for Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack

Abstract

Recent studies show that Large Language Models (LLMs) with safety alignment can be jail-broken by fine-tuning on a dataset mixed with harmful data. For the first time in the literature, we show that the jail-break effect can be mitigated by separating two states in the fine-tuning stage to respectively optimize over the alignment and user datasets. Unfortunately, our subsequent study shows that this simple Bi-State Optimization (BSO) solution experiences convergence instability when steps invested in its alignment state is too small, leading to downgraded alignment performance. By statistical analysis, we show that the excess drift towards the switching iterates of the two states could be a probable reason for the instability. To remedy this issue, we propose Lazy(i) safety alignment (Lisa), which introduces a proximal term to constraint the drift of each state. Theoretically, the benefit of the proximal term is supported by the convergence analysis, wherein we show that a sufficient large proximal factor is necessary to guarantee Lisa’s convergence. Empirically, our results on four downstream fine-tuning tasks show that Lisa with a proximal term can significantly increase alignment performance while maintaining the LLM’s accuracy on the user tasks. Code is available at https://github.com/git-disl/Lisa.

Disclaimer: This document contains content that some may find disturbing or offensive, including content that is hateful or violent in nature.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Preliminaries
  • 4 Methodology
  • 4.1 Bi-State Optimization
  • 4.2 Lazy Safety Alignment
  • 5 Experiments
  • 5.1 Setup
  • 5.2 Main Results
  • 5.3 Statistical/System Evaluation
  • 5.4 Hyper-parameters Analysis and Ablation Study
  • 5.5 Alternative Design
  • 5.6 Visualization
  • 6 Conclusion
  • 7 Acknowledgment
  • References
  • A Missing Information for Experiments
  • A.1 Detailed Setup
  • A.2 Baselines and its Description
  • A.3 More Results
  • A.4 More Visualizations
  • B Missing contents in theoretical analysis
  • B.1 Preliminaries
  • B.2 Assumptions
  • B.3 Facts
  • B.4 Theorems
  • B.5 Missing Proof of Theorem 2
  • B.5.1 Key Lemmas
  • B.5.2 Formal Proof
  • B.6 Missing Proof of Theorem 3
  • B.6.1 Key Lemmas
  • B.6.2 Formal Proof
  • C Broader Impact
  • D Limitations
  • NeurIPS Paper Checklist

Knowls

  1. Knowl 1 — Lazy Safety Alignment (Lisa) Algorithm

    algorithm

    Lazy Safety Alignment (Lisa) is a fine-tuning defense algorithm designed to preserve safety alignment in Large Language Models (LLMs) when fine-tuning on user data that may be mixed with harmful examples. It formulates fine-tuning as an alternating bi-state optimization problem with proximal regularization to prevent the parameter iterates from drifting excessively between states.

    Let f(w)f(w) denote the standard cross-entropy loss over a curated safety alignment dataset, and let h(w)h(w) denote the cross-entropy loss over the user's downstream fine-tuning dataset, parameterized by model weights w∈Rdw \in \mathbb{R}^d. Lisa splits the training into TT alternating cycles. In each cycle t∈{0,…,T−1}t \in \{0, \dots, T-1\}, it performs K1K_1 gradient steps on the alignment loss regularized by a proximal term centered at the previous cycle's checkpoint wtw_t, producing an intermediate checkpoint w~t+1\tilde{w}_{t+1}. It then performs K2K_2 gradient steps on the fine-tuning loss regularized by a proximal term centered at w~t+1\tilde{w}_{t+1}, producing wt+1w_{t+1}.

    Input: Alignment step count K1K_1, Fine-tuning step count K2K_2, Total cycles TT, Proximal penalty ρ>0\rho > 0, Initial model weights w0w_0
    Output: Fine-tuned model weights wTw_T
    for t=0t = 0 to T−1T-1 do
        // State 1: Alignment state with proximal regularization
        Initialize local weight wt,0=wtw_{t,0} = w_t
        for k=0k = 0 to K1−1K_1 - 1 do
            Sample batch (xt,k,yt,k)(x_{t,k}, y_{t,k}) from alignment dataset
            gt,k=∇f(wt,k;xt,k,yt,k)+ρ(wt,k−wt)g_{t,k} = \nabla f(w_{t,k}; x_{t,k}, y_{t,k}) + \rho (w_{t,k} - w_t)
            wt,k+1=Optimizer_Step(wt,k,gt,k)w_{t,k+1} = \text{Optimizer\_Step}(w_{t,k}, g_{t,k})
        end for
        w~t+1=wt,K1\tilde{w}_{t+1} = w_{t,K_1}
        // State 2: User fine-tuning state with proximal regularization
        Initialize local weight wt,0′=w~t+1w'_{t,0} = \tilde{w}_{t+1}
        for k=0k = 0 to K2−1K_2 - 1 do
            Sample batch (xt,k′,yt,k′)(x'_{t,k}, y'_{t,k}) from user fine-tuning dataset
            gt,k′=∇h(wt,k′;xt,k′,yt,k′)+ρ(wt,k′−w~t+1)g'_{t,k} = \nabla h(w'_{t,k}; x'_{t,k}, y'_{t,k}) + \rho (w'_{t,k} - \tilde{w}_{t+1})
            wt,k+1′=Optimizer_Step(wt,k′,gt,k′)w'_{t,k+1} = \text{Optimizer\_Step}(w'_{t,k}, g'_{t,k})
        end for
        wt+1=wt,K2′w_{t+1} = w'_{t,K_2}
    end for
    return wTw_T

    In standard implementations, the model weights ww represent Low-Rank Adaptation (LoRA) adaptor parameters (rank 8), the optimizer is AdamW with learning rate 1×10−51\times 10^{-5}, default step allocation is K1=100K_1 = 100 and K2=900K_2 = 900, and the default proximal penalty is ρ=1\rho = 1.

  2. Knowl 2 — Convergence Rates of Lisa under Kurdyka-Łojasiewicz Geometry

    theoretical result

    Let f(w)f(w) and h(w)h(w) be proper, closed, and LL-smooth loss functions on Rd\mathbb{R}^d. Let the potential function D(w~t,wt)=f(w~t)+h(wt)+ρ2∥w~t−wt∥2\mathcal{D}(\tilde{w}_t, w_t) = f(\tilde{w}_t) + h(w_t) + \frac{\rho}{2}\|\tilde{w}_t - w_t\|^2 satisfy the Kurdyka-Łojasiewicz (KL) property with desingularizing function φ(v)=cv1−θ\varphi(v) = c v^{1-\theta} for θ∈[0,1)\theta \in [0, 1) and constant c>0c > 0. If the proximal intensity parameter satisfies ρ>L\rho > L, and a subsequence of the sequence of iterates (w~t,wt)(\tilde{w}_t, w_t) generated by Lisa converges to a cluster point (w~∗,w∗)(\tilde{w}^*, w^*), then the residual rt=D(w~t,wt)−D(w~∗,w∗)r_t = \mathcal{D}(\tilde{w}_t, w_t) - \mathcal{D}(\tilde{w}^*, w^*) and the stationarity metric ∥∇f(w~T)+∇h(wT)∥\|\nabla f(\tilde{w}_T) + \nabla h(w_T)\| satisfy the following convergence rates depending on θ\theta:

    1. Case θ=0\theta = 0 (Finite termination): For all iterations T>t0T > t_0 (where t0t_0 is a finite integer threshold): ∥∇f(w~T)+∇h(wT)∥=0\|\nabla f(\tilde{w}_T) + \nabla h(w_T)\| = 0

    2. Case θ∈(0,1/2]\theta \in (0, 1/2] (Linear convergence): For all iterations T>t0′T > t'_0: ∥∇f(w~T)+∇h(wT)∥≤2ρρ−L(1−ρ−Lρ2c2(1−θ)2)T−t0′rt0′\|\nabla f(\tilde{w}_T) + \nabla h(w_T)\| \le \frac{\sqrt{2\rho}}{\sqrt{\rho - L}} \sqrt{\left(1 - \frac{\rho - L}{\rho^2 c^2 (1 - \theta)^2}\right)^{T - t'_0} r_{t'_0}}

    3. Case θ∈(1/2,1)\theta \in (1/2, 1) (Sub-linear convergence): For all iterations T>0T > 0: ∥∇f(w~T)+∇h(wT)∥≤2ρρ−LT(2θ−1)ρ−Lρ2c2(1−θ)22−4θ\|\nabla f(\tilde{w}_T) + \nabla h(w_T)\| \le \frac{\sqrt{2\rho}}{\sqrt{\rho - L}} \sqrt[2-4\theta]{T (2\theta - 1) \frac{\rho - L}{\rho^2 c^2 (1 - \theta)^2}}

    Setting ρ>L\rho > L is strictly necessary; as ρ→L\rho \to L, the right-hand bound approaches infinity and the gradient norm of the last iterate becomes unbounded.

  3. Knowl 3 — Subsequence Convergence of Lisa to a Stationary Point

    theoretical result

    Assume that the alignment loss f:Rd→Rf: \mathbb{R}^d \to \mathbb{R} and the user task loss h:Rd→Rh: \mathbb{R}^d \to \mathbb{R} are proper, closed, and LL-smooth (L>0L > 0), such that ∥∇f(w)−∇f(w′)∥≤L∥w−w′∥\|\nabla f(w) - \nabla f(w')\| \le L\|w - w'\| and ∥∇h(w)−∇h(w′)∥≤L∥w−w′∥\|\nabla h(w) - \nabla h(w')\| \le L\|w - w'\| for all w,w′∈Rdw, w' \in \mathbb{R}^d. Let the proximal penalty satisfy ρ≥L\rho \ge L.

    If there exists a subsequence (w~tj,wtj)(\tilde{w}_{t_j}, w_{t_j}) of the sequence (w~t,wt)(\tilde{w}_t, w_t) generated by Lisa that converges to a cluster point (w~∗,w∗)(\tilde{w}^*, w^*), then: lim⁡j→∞(w~tj+1,wtj+1)=lim⁡j→∞(w~tj,wtj)=(w~∗,w∗)\lim_{j \to \infty} (\tilde{w}_{t_{j+1}}, w_{t_{j+1}}) = \lim_{j \to \infty} (\tilde{w}_{t_j}, w_{t_j}) = (\tilde{w}^*, w^*) Furthermore, this cluster point is a stationary point of the joint optimization problem min⁡wf(w)+h(w)\min_w f(w) + h(w), satisfying: ∇f(w~∗)+∇h(w∗)=0\nabla f(\tilde{w}^*) + \nabla h(w^*) = 0

  4. Knowl 4 — Potential Function for Lisa Optimization

    definition

    For the alternating minimization scheme of Lisa with proximal parameter ρ>0\rho > 0, alignment loss f(w)f(w), and user fine-tuning loss h(w)h(w), the potential function D:Rd×Rd→R\mathcal{D}: \mathbb{R}^d \times \mathbb{R}^d \to \mathbb{R} at cycle tt is defined as: D(w~t,wt)=f(w~t)+h(wt)+ρ2∥w~t−wt∥2\mathcal{D}(\tilde{w}_t, w_t) = f(\tilde{w}_t) + h(w_t) + \frac{\rho}{2}\|\tilde{w}_t - w_t\|^2 where w~t\tilde{w}_t is the intermediate model checkpoint resulting from the alignment step optimization and wtw_t is the model checkpoint resulting from the user fine-tuning step optimization.

    Along the sequence of iterates generated by Lisa, the potential function achieves sufficient and non-increasing descent whenever ρ≥L\rho \ge L (LL being the Lipschitz smoothness constant of ff and hh): D(w~t+1,wt+1)−D(w~t,wt)≤−ρ−L2(∥wt+1−wt∥2+∥w~t+1−w~t∥2)\mathcal{D}(\tilde{w}_{t+1}, w_{t+1}) - \mathcal{D}(\tilde{w}_t, w_t) \le -\frac{\rho - L}{2} \left( \|w_{t+1} - w_t\|^2 + \|\tilde{w}_{t+1} - \tilde{w}_t\|^2 \right)

  5. Knowl 5 — Excess Drift Phenomenon in Asymmetric Bi-State Optimization

    definition

    In Bi-State Optimization (BSO) for language model fine-tuning, training alternates between K1K_1 steps on a safety alignment dataset and K2K_2 steps on a user fine-tuning dataset. Excess drift is defined as the Euclidean distance between the model weight vector obtained at the end of one state and that obtained at the end of the other state within a cycle, ∥wFT−wAlign∥\|w_{\text{FT}} - w_{\text{Align}}\|.

    When computation is allocated symmetrically (K1=K2=500K_1 = K_2 = 500), BSO converges stably. However, under asymmetric computation designed to reduce training cost where fewer steps are allocated to alignment (K1≪K2K_1 \ll K_2, e.g., K1=100,K2=900K_1 = 100, K_2 = 900), unconstrained gradient updates during the user fine-tuning state create large parameter drift away from the alignment state switching point. This excess drift causes the joint gradient norm ∥∇f(wt)+∇h(wt)∥\|\nabla f(w_t) + \nabla h(w_t)\| and alignment loss to diverge after an initial drop, degrading safety alignment performance and increasing the harmful response rate of the resulting model by up to 17.6%17.6\%.

  6. Knowl 6 — Bi-State Optimization (BSO) Algorithm

    algorithm

    Bi-State Optimization (BSO) is a multi-task fine-tuning defense algorithm that mitigates safety degradation by alternating between optimization on a safety alignment dataset and optimization on a user-provided downstream dataset.

    Input: Alignment step count K1K_1, Fine-tuning step count K2K_2, Number of cycles TT, Initial model weights w0,0w_{0,0}
    Output: Final fine-tuned weights wT,0w_{T,0}
    for t=0t = 0 to T−1T-1 do
        for k=0k = 0 to K1+K2−1K_1 + K_2 - 1 do
            if k<K1k < K_1 then
                Sample batch (xt,k,yt,k)(x_{t,k}, y_{t,k}) from safety alignment dataset
                gt,k=∇f(wt,k;xt,k,yt,k)g_{t,k} = \nabla f(w_{t,k}; x_{t,k}, y_{t,k})
            else
                Sample batch (xt,k,yt,k)(x_{t,k}, y_{t,k}) from user fine-tuning dataset
                gt,k=∇h(wt,k;xt,k,yt,k)g_{t,k} = \nabla h(w_{t,k}; x_{t,k}, y_{t,k})
            end if
            wt,k+1=Optimizer_Step(wt,k,gt,k)w_{t,k+1} = \text{Optimizer\_Step}(w_{t,k}, g_{t,k})
        end for
        wt+1,0=wt,K1+K2w_{t+1,0} = w_{t, K_1 + K_2}
    end for
    return wT,0w_{T,0}

    BSO directly solves min⁡wf(w)+h(w)\min_w f(w) + h(w) without proximal terms. It performs well when K1=K2K_1 = K_2, but suffers from excess drift and convergence instability when K1≪K2K_1 \ll K_2.

  7. Knowl 7 — Empirical Performance of Lisa Across Harmful Fine-Tuning Ratios

    data/table

    Evaluation of fine-tuning defense methods on a Llama2-7B base model fine-tuned on the SST-2 sentiment classification task (n=5000n = 5000 training samples) mixed with varying proportions pp of harmful samples from the BeaverTails dataset. Baseline methods evaluated include NonAligned-SFT (NA-SFT), Supervised Fine-Tuning (SFT), Elastic Weight Consolidation (EWC), Vaccine-SFT, VLGuard, Bi-State Optimization (BSO, with K1=100,K2=900K_1=100, K_2=900), and Lisa (with K1=100,K2=900,ρ=1K_1=100, K_2=900, \rho=1).

    Methods Harmful Score (%) ↓\downarrow Finetune Accuracy (%) ↑\uparrow
    (n=5000n=5000) clean p=0.05p=0.05 p=0.1p=0.1 p=0.2p=0.2 p=0.3p=0.3 Average clean p=0.05p=0.05 p=0.1p=0.1 p=0.2p=0.2 p=0.3p=0.3 Average
    NA-SFT 24.10 54.90 55.00 55.70 53.50 48.64 94.84 95.07 95.41 95.64 95.07 95.21
    SFT 34.60 49.10 51.60 52.50 53.60 48.28 95.30 95.30 94.95 95.76 95.30 95.32
    EWC 38.30 41.70 41.80 46.40 46.50 42.94 45.18 13.88 11.58 8.72 11.81 18.23
    Vaccine-SFT 26.60 48.50 52.70 53.50 53.00 46.86 95.30 93.92 94.27 94.50 94.38 94.47
    Vlguard 33.80 42.00 43.90 47.40 49.30 43.28 95.64 94.72 95.18 95.64 95.53 95.34
    BSO 34.30 46.00 49.00 50.70 50.60 46.12 95.53 94.72 95.18 95.30 95.18 95.18
    Lisa 34.90 36.60 40.20 42.60 43.60 39.58 95.07 95.18 94.84 94.61 94.04 94.75

    Lisa reduces the average Harmful Score by 7.07%7.07\% compared to Vaccine-SFT, by 3.68%3.68\% compared to VLGuard, and by 6.54%6.54\% compared to vanilla BSO, while maintaining downstream fine-tuning accuracy within 0.59%0.59\% of VLGuard and improving by 0.28%0.28\% over Vaccine-SFT. EWC achieves low harmful scores only at the cost of severe downstream utility degradation (average accuracy of 18.23%18.23\%).

  8. Knowl 8 — Ablation Study of Lisa Components

    data/table

    An ablation study on Llama2-7B fine-tuned on SST-2 (n=5000n=5000) isolates the effects of the two core mechanisms of Lisa: alternating optimization on alignment/user data (BSO) and proximal regularization.

    Methods Harmful Score (%) ↓\downarrow Finetune Accuracy (%) ↑\uparrow
    (n=5000n=5000) clean p=0.05p=0.05 p=0.1p=0.1 p=0.2p=0.2 p=0.3p=0.3 Average clean p=0.05p=0.05 p=0.1p=0.1 p=0.2p=0.2 p=0.3p=0.3 Average
    SFT 34.60 49.10 51.60 52.50 53.60 48.28 95.30 95.30 94.95 95.76 95.30 95.32
    Lisa (BSO only, ρ=0\rho=0) 34.30 46.00 49.00 50.70 50.60 46.12 95.53 94.72 95.18 95.30 95.18 95.18
    Lisa (proximal only) 31.40 31.60 32.30 32.90 34.10 32.46 88.19 88.88 87.27 85.78 84.98 87.02
    Lisa (BSO + proximal) 34.90 36.60 40.20 42.60 43.60 39.58 95.07 95.18 94.84 94.61 94.04 94.75

    When using BSO alone (ρ=0\rho = 0), excess drift causes the average Harmful Score to remain high (46.12%46.12\% vs. 39.58%39.58\% for full Lisa). When using proximal regularization alone without alignment data, the parameters remain tied to the initial model, yielding low Harmful Scores (32.46%32.46\%) but degrading downstream task generalization (87.02%87.02\% average accuracy vs. 94.75%94.75\% for Lisa). Both components are necessary to achieve safety without sacrificing task accuracy.

  9. Knowl 9 — Generalization of Lisa Across Language Models and Tasks

    data/table

    Empirical evaluation of Lisa and baselines across different LLM architectures on GSM8K (with poison ratio p=0.1p=0.1, n=5000n=5000) and across different fine-tuning tasks on Mistral-7B.

    GSM8K Task Opt-2.7B Llama2-7B Mistral-7B Average
    Methods HS ↓\downarrow FA ↑\uparrow HS ↓\downarrow FA ↑\uparrow HS ↓\downarrow FA ↑\uparrow HS ↓\downarrow FA ↑\uparrow
    NA-SFT 59.60 7.20 56.70 24.00 55.30 34.60 57.20 21.93
    SFT 53.80 6.60 52.30 22.10 46.90 24.30 51.00 17.67
    Vaccine-SFT 51.40 6.70 49.60 19.10 41.10 8.90 47.37 11.57
    Vlguard 49.20 7.10 48.90 21.50 46.10 22.80 48.07 17.13
    BSO 48.00 7.40 48.00 23.00 43.40 23.40 46.47 17.93
    Lisa 35.00 2.40 36.80 17.90 36.70 28.00 36.17 16.10
    Mistral-7B SST2 AGNEWS GSM8K AlpacaEval Average
    Methods HS ↓\downarrow FA ↑\uparrow HS ↓\downarrow FA ↑\uparrow HS ↓\downarrow FA ↑\uparrow HS ↓\downarrow FA ↑\uparrow HS ↓\downarrow FA ↑\uparrow
    NA-SFT 54.40 94.15 56.90 91.50 55.30 34.60 43.20 54.33 52.45 68.65
    SFT 51.20 93.92 52.40 91.90 46.90 24.30 38.50 45.63 47.25 63.94
    Vaccine-SFT 46.30 81.77 48.60 85.70 41.10 8.90 33.00 10.68 42.25 46.76
    Vlguard 43.90 95.18 43.90 90.00 46.10 22.80 37.50 43.27 42.85 62.81
    BSO 45.90 95.53 47.80 91.20 43.40 23.40 37.40 41.75 43.63 62.97
    Lisa 39.80 95.99 40.50 89.60 36.70 28.00 33.10 41.35 37.53 63.74

    Across models, Lisa reduces the average Harmful Score on GSM8K to 36.17%36.17\%, outperforming VLGuard (48.07%48.07\%) and Vaccine-SFT (47.37%47.37\%). Across tasks on Mistral-7B, Lisa achieves the lowest average Harmful Score (37.53%37.53\%) while retaining an average downstream accuracy of 63.74%63.74\%.

  10. Knowl 10 — Combining Lisa with Vaccine Alignment and Moderation-Based Data Filtering

    model/method

    Because Lisa modifies only the user fine-tuning stage, it can be combined orthogonally with alignment-stage defenses (such as Vaccine) and input-level data moderation filtering:

    1. Vaccine + Lisa (Vaccine-Lisa): Vaccine adds adversarial perturbations in the alignment stage to harden safety representations, while Lisa controls drift during fine-tuning. On Llama2-7B fine-tuned on SST-2 (n=5000n=5000), Vaccine-Lisa achieves an average Harmful Score of 32.86%32.86\% (across clean to p=0.3p=0.3), compared to 46.86%46.86\% for Vaccine-SFT and 39.58%39.58\% for standard SFT-Lisa, with fine-tuning accuracy remaining high (93.97%93.97\% vs. 94.47%94.47\% for Vaccine-SFT).

    2. Data Filtering + Lisa (Filter+Lisa): An input moderation model (e.g., BeaverTails moderation) is applied before fine-tuning to remove flagged harmful inputs. Due to the classifier's 7.71%7.71\% false negative rate, residual toxic samples leak into the training split. Deploying Lisa on the filtered dataset reduces the average Harmful Score to 34.18%34.18\% (compared to 39.08%39.08\% for Filter+SFT and 45.92%45.92\% for unmoderated SFT) across poison ratios up to p=1.0p=1.0.

  11. Knowl 11 — Computational and Memory Overhead of Lisa

    data/table

    System resource measurements for 1000 training steps on an NVIDIA H100 workstation using LoRA adaptors (rank 8) with batch size 5.

    Metric SFT VLGuard BSO Lisa
    Clock time (seconds) 115.53 122.34 119.22 125.11
    GPU Memory (GB) 47.97 48.48 50.85 51.11

    Lisa incurs an 8.3%8.3\% increase in clock time compared to standard SFT and a 4.9%4.9\% increase compared to BSO, caused by forward and backward computations on the proximal penalty term ρ2∥w−wref∥2\frac{\rho}{2}\|w - w_{\text{ref}}\|^2. GPU memory usage increases by 3.14 GB3.14\text{ GB} (6.5%6.5\%) over SFT due to caching the switching reference checkpoint. Crucially, these overheads scale with the number of trainable adapter parameters rather than with the size of the training dataset.

  12. Knowl 12 — Limitations of Lazy Safety Alignment

    limitation

    The methodology of Lisa exhibits three primary limitations:

    1. Fine-Tuning Stage Computation: Because Lisa operates during the user fine-tuning stage, the additional computational and memory overheads are incurred for every incoming user fine-tuning request, scaling with the volume of user requests rather than being amortized once as in alignment-stage defenses.
    2. Restriction to Supervised Fine-Tuning: The formulation and empirical verification are developed strictly on top of Supervised Fine-Tuning (SFT) workflows. The method was not adapted to or evaluated with Reinforcement Learning from Human Feedback (RLHF) pipelines.
    3. Evaluation Task Breadth: Downstream evaluation was confined to standard academic classification, reasoning, and instruction benchmarks (SST-2, AGNEWS, GSM8K, AlpacaEval), without testing on complex multi-agent architectures or open-ended conversational agents.

Coverage note — None was omitted. All principal contributions—including the BSO baseline, the excess drift phenomenon, the Lisa algorithm, the complete theoretical convergence results and potential function definition, extensive empirical evaluations across models, tasks, sample sizes, and poisoning ratios, ablation studies, combined defense pipelines (Vaccine-Lisa and Filter+Lisa), system overhead benchmarks, and limitations—have been completely covered.

References

  1. 1.Acar, D. A. E., Zhao, Y., Navarro, R. M., Mattina, M., Whatmough, P. N., and Saligrama, V. Federated learning based on dynamic regularization. arXiv preprint arXiv:2111.04263, 2021.
  2. 2.Anonymous. Identifying and tuning safety neurons in large language models. In Submitted to The Thirteenth International Conference on Learning Representations, 2024a. URL https://openreview.net/forum?id=yR47RmND1m. under review.
  3. 3.Anonymous. Safety alignment shouldn't be complicated. In Submitted to The Thirteenth International Conference on Learning Representations, 2024b. URL https://openreview.net/forum?id=9H91juqfgb. under review.
  4. 4.Anonymous. SaloRA: Safety-alignment preserved low-rank adaptation. In Submitted to The Thirteenth International Conference on Learning Representations, 2024c. URL https://openreview.net/forum?id=GOoVzE9nSj. under review.
  5. 5.Anonymous. Your task may vary: A systematic understanding of alignment and safety degradation when fine-tuning LLMs. In Submitted to The Thirteenth International Conference on Learning Representations, 2024d. URL https://openreview.net/forum?id=vQ0zFYJaMo. under review.
  6. 6.Attouch, H., Bolte, J., Redont, P., and Soubeyran, A. Proximal alternating minimization and projection methods for nonconvex problems: An approach based on the kurdyka-łojasiewicz inequality. Mathematics of operations research, 35(2):438–457, 2010.
  7. 7.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022.
  8. 8.Bianchi, F., Suzgun, M., Attanasio, G., Röttger, P., Jurafsky, D., Hashimoto, T., and Zou, J. Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions. arXiv preprint arXiv:2309.07875, 2023.
  9. 9.Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., Koh, P. W. W., Ippolito, D., Tramer, F., and Schmidt, L. Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems, 36, 2024.
  10. 10.Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E. Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419, 2023.
  11. 11.Chen, C., Huang, B., Li, Z., Chen, Z., Lai, S., Xu, X., Gu, J.-C., Gu, J., Yao, H., Xiao, C., et al. Can editing llms inject harm? arXiv preprint arXiv:2407.20224, 2024.
  12. 12.Chen, Z., Zhou, Y., Xu, T., and Liang, Y. Proximal gradient descent-ascent: Variable convergence under k {\L} geometry. arXiv preprint arXiv:2102.04653, 2021.
  13. 13.Choi, H. K., Du, X., and Li, Y. Safety-aware fine-tuning of large language models. arXiv preprint arXiv:2410.10014, 2024.
  14. 14.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  15. 15.Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y. Safe rlhf: Safe reinforcement learning from human feedback. arXiv preprint arXiv:2310.12773, 2023.
  16. 16.Dong, H., Xiong, W., Goyal, D., Pan, R., Diao, S., Zhang, J., Shum, K., and Zhang, T. Raft: Reward ranked finetuning for generative foundation model alignment. arXiv preprint arXiv:2304.06767, 2023.
  17. 17.Du, Y., Zhao, S., Cao, J., Ma, M., Zhao, D., Fan, F., Liu, T., and Qin, B. Towards secure tuning: Mitigating security risks arising from benign instruction fine-tuning. arXiv preprint arXiv:2410.04524, 2024.
  18. 18.Eiras, F., Petrov, A., Torr, P. H., Kumar, M. P., and Bibi, A. Mimicking user data: On mitigating fine-tuning risks in closed large language models. arXiv preprint arXiv:2406.10288, 2024.
  19. 19.Fernando, H., Shen, H., Ram, P., Zhou, Y., Samulowitz, H., Baracaldo, N., and Chen, T. Mitigating forgetting in llm supervised fine-tuning and preference learning. arXiv preprint arXiv:2410.15483, 2024.
  20. 20.Griffith, S., Subramanian, K., Scholz, J., Isbell, C. L., and Thomaz, A. L. Policy shaping: Integrating human feedback with reinforcement learning. Advances in neural information processing systems, 26, 2013.
  21. 21.Halawi, D., Wei, A., Wallace, E., Wang, T. T., Haghtalab, N., and Steinhardt, J. Covert malicious finetuning: Challenges in safeguarding llm adaptation. arXiv preprint arXiv:2406.20053, 2024.
  22. 22.He, L., Xia, M., and Henderson, P. What's in your" safe" data?: Identifying benign data that breaks safety. arXiv preprint arXiv:2404.01099, 2024.
  23. 23.Hsu, C.-Y., Tsai, Y.-L., Lin, C.-H., Chen, P.-Y., Yu, C.-M., and Huang, C.-Y. Safe lora: the silver lining of reducing safety risks when fine-tuning large language models. arXiv preprint arXiv:2405.16833, 2024.
  24. 24.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  25. 25.Hu, S., Huang, T., ˙Ilhan, F., Tekin, S. F., and Liu, L. Large language model-powered smart contract vulnerability detection: New perspectives. In 2023 5th IEEE International Conference on Trust, Privacy and Security in Intelligent Systems and Applications (TPS-ISA), pp. 297–306. IEEE, 2023a.
  26. 26.Hu, S., Zhang, Z., Luo, B., Lu, S., He, B., and Liu, L. Bert4eth: A pre-trained transformer for ethereum fraud detection. In Proceedings of the ACM Web Conference 2023, pp. 2189–2197, 2023b.
  27. 27.Hu, S., Huang, T., Chow, K.-H., Wei, W., Wu, Y., and Liu, L. Zipzap: Efficient training of language models for large-scale fraud detection on blockchain. In Proceedings of the ACM on Web Conference 2024, pp. 2807–2816, 2024a.
  28. 28.Hu, S., Huang, T., Ilhan, F., Tekin, S., Liu, G., Kompella, R., and Liu, L. A survey on large language model-based game agents. arXiv preprint arXiv:2404.02039, 2024b.
  29. 29.Huang, T., Shen, L., Sun, Y., Lin, W., and Tao, D. Fusion of global and local knowledge for personalized federated learning. arXiv preprint arXiv:2302.11051, 2023.
  30. 30.Huang, T., Bhattacharya, G., Joshi, P., Kimball, J., and Liu, L. Antidote: Post-fine-tuning safety alignment for large language models against harmful fine-tuning. arXiv preprint arXiv:2408.09600, 2024a.
  31. 31.Huang, T., Hu, S., Chow, K.-H., Ilhan, F., Tekin, S., and Liu, L. Lockdown: backdoor defense for federated learning with isolated subspace training. Advances in Neural Information Processing Systems, 36, 2024b.
  32. 32.Huang, T., Hu, S., Ilhan, F., Tekin, S. F., and Liu, L. Booster: Tackling harmful fine-tuning for large language models via attenuating harmful perturbation. arXiv preprint arXiv:2409.01586, 2024c.
  33. 33.Huang, T., Hu, S., Ilhan, F., Tekin, S. F., and Liu, L. Harmful fine-tuning attacks and defenses for large language models: A survey. arXiv preprint arXiv:2409.18169, 2024d.
  34. 34.Huang, T., Hu, S., and Liu, L. Vaccine: Perturbation-aware alignment for large language model. arXiv preprint arXiv:2402.01109, 2024e.
  35. 35.Idelbayev, Y. and Carreira-Perpinán, M. A. Low-rank compression of neural nets: Learning the rank of each layer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8049–8059, 2020.
  36. 36.Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Sun, R., Wang, Y., and Yang, Y. Beavertails: Towards improved safety alignment of llm via a human-preference dataset. arXiv preprint arXiv:2307.04657, 2023.
  37. 37.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  38. 38.Karimi, H., Nutini, J., and Schmidt, M. Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition. In Joint European conference on machine learning and knowledge discovery in databases, pp. 795–811. Springer, 2016.
  39. 39.Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526, 2017.
  40. 40.Leong, C. T., Cheng, Y., Xu, K., Wang, J., Wang, H., and Li, W. No two devils alike: Unveiling distinct mechanisms of fine-tuning attacks. arXiv preprint arXiv:2405.16229, 2024.
  41. 41.Lermen, S., Rogers-Smith, C., and Ladish, J. Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b. arXiv preprint arXiv:2310.20624, 2023.
  42. 42.Li, G. and Pong, T. K. Douglas–rachford splitting for nonconvex optimization with application to nonconvex feasibility problems. Mathematical programming, 159(1):371–401, 2016.
  43. 43.Li, S., Ngai, E. C.-H., Ye, F., and Voigt, T. Peft-as-an-attack! jailbreaking language models during federated parameter-efficient fine-tuning. arXiv preprint arXiv:2411.19335, 2024.
  44. 44.Li, T., Sahu, A. K., Zaheer, M., Sanjabi, M., Talwalkar, A., and Smith, V. Federated optimization in heterogeneous networks. arXiv preprint arXiv:1812.06127, 2018.
  45. 45.Li, X., Huang, K., Yang, W., Wang, S., and Zhang, Z. On the convergence of fedavg on non-iid data. arXiv preprint arXiv:1907.02189, 2019.
  46. 46.Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B. Alpacaeval: An automatic evaluator of instruction-following models. https://github.com/tatsu-lab/alpaca_eval, 2023.
  47. 47.Liu, G., Lin, W., Huang, T., Mo, R., Mu, Q., and Shen, L. Targeted vaccine: Safety alignment for large language models against harmful fine-tuning via layer-wise perturbation. arXiv preprint arXiv:2410.09760, 2024a.
  48. 48.Liu, H., Sferrazza, C., and Abbeel, P. Chain of hindsight aligns language models with feedback. arXiv preprint arXiv:2302.02676, 3, 2023a.
  49. 49.Liu, R., Yang, R., Jia, C., Zhang, G., Zhou, D., Dai, A. M., Yang, D., and Vosoughi, S. Training socially aligned language models in simulated human society. arXiv preprint arXiv:2305.16960, 2023b.
  50. 50.Liu, X., Liang, J., Ye, M., and Xi, Z. Robustifying safety-aligned large language models through clean data curation. arXiv preprint arXiv:2405.19358, 2024b.
  51. 51.Lyu, K., Zhao, H., Gu, X., Yu, D., Goyal, A., and Arora, S. Keeping llms aligned after fine-tuning: The crucial role of prompt templates. arXiv preprint arXiv:2402.18540, 2024.
  52. 52.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  53. 53.Ozdayi, M. S., Kantarcioglu, M., and Gel, Y. R. Defending against backdoors in federated learning with robust learning rate. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pp. 9268–9276, 2021.
  54. 54.Peng, S., Chen, P.-Y., Hull, M., and Chau, D. H. Navigating the safety landscape: Measuring risks in finetuning large language models. arXiv preprint arXiv:2405.17374, 2024.
  55. 55.Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P. Fine-tuning aligned language models compromises safety, even when users do not intend to! arXiv preprint arXiv:2310.03693, 2023.
  56. 56.Qi, X., Huang, K., Panda, A., Henderson, P., Wang, M., and Mittal, P. Visual adversarial examples jailbreak aligned large language models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pp. 21527–21536, 2024a.
  57. 57.Qi, X., Panda, A., Lyu, K., Ma, X., Roy, S., Beirami, A., Mittal, P., and Henderson, P. Safety alignment should be made more than just a few tokens deep. arXiv preprint arXiv:2406.05946, 2024b.
  58. 58.Qi, X., Wei, B., Carlini, N., Huang, Y., Xie, T., He, L., Jagielski, M., Nasr, M., Mittal, P., and Henderson, P. On evaluating the durability of safeguards for open-weight llms. arXiv preprint arXiv:2412.07097, 2024c.
  59. 59.Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290, 2023.
  60. 60.Rajeswaran, A., Finn, C., Kakade, S. M., and Levine, S. Meta-learning with implicit gradients. Advances in neural information processing systems, 32, 2019.
  61. 61.Rosati, D., Edkins, G., Raj, H., Atanasov, D., Majumdar, S., Rajendran, J., Rudzicz, F., and Sajjad, H. Defending against reverse preference attacks is difficult. arXiv preprint arXiv:2409.12914, 2024a.
  62. 62.Rosati, D., Wehner, J., Williams, K., Bartoszcze, Ł., Atanasov, D., Gonzales, R., Majumdar, S., Maple, C., Sajjad, H., and Rudzicz, F. Representation noising effectively prevents harmful fine-tuning on llms. arXiv preprint arXiv:2405.14577, 2024b.
  63. 63.Rosati, D., Wehner, J., Williams, K., Bartoszcze, Ł., Batzner, J., Sajjad, H., and Rudzicz, F. Immunization against harmful fine-tuning attacks. arXiv preprint arXiv:2402.16382, 2024c.
  64. 64.Shen, H., Chen, P.-Y., Das, P., and Chen, T. Seal: Safety-enhanced aligned llm fine-tuning via bilevel data selection. arXiv preprint arXiv:2410.07471, 2024.
  65. 65.Shen, L., Sun, P., Wang, Y., Liu, W., and Zhang, T. An algorithmic framework of variable metric over-relaxed hybrid proximal extra-gradient method. In International Conference on Machine Learning, pp. 4634–4643. PMLR, 2018.
  66. 66.Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pp. 1631–1642, 2013.
  67. 67.Song, F., Yu, B., Li, M., Yu, H., Huang, F., Li, Y., and Wang, H. Preference ranking optimization for human alignment. arXiv preprint arXiv:2306.17492, 2023.
  68. 68.Sun, Y., Shen, L., Huang, T., Ding, L., and Tao, D. Fedspeed: Larger local interval, less communication round, and higher generalization accuracy. arXiv preprint arXiv:2302.10429, 2023.
  69. 69.Tamirisa, R., Bharathi, B., Phan, L., Zhou, A., Gatti, A., Suresh, T., Lin, M., Wang, J., Wang, R., Arel, R., et al. Tamper-resistant safeguards for open-weight llms. arXiv preprint arXiv:2408.00761, 2024.
  70. 70.Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. Alpaca: A strong, replicable instruction-following model. Stanford Center for Research on Foundation Models. https://crfm. stanford. edu/2023/03/13/alpaca. html, 3(6):7, 2023.
  71. 71.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  72. 72.Wang, J., Li, J., Li, Y., Qi, X., Chen, M., Hu, J., Li, Y., Li, B., and Xiao, C. Mitigating fine-tuning jailbreak attack with backdoor enhanced alignment. arXiv preprint arXiv:2402.14968, 2024.
  73. 73.Wei, A., Haghtalab, N., and Steinhardt, J. Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2024a.
  74. 74.Wei, B., Huang, K., Huang, Y., Xie, T., Qi, X., Xia, M., Mittal, P., Wang, M., and Henderson, P. Assessing the brittleness of safety alignment via pruning and low-rank modifications. arXiv preprint arXiv:2402.05162, 2024b.
  75. 75.Wu, B., Shen, L., Zhang, T., and Ghanem, B. Map inference via ℓ2-sphere linear program reformulation. International Journal of Computer Vision, 128(7):1913–1936, 2020.
  76. 76.Wu, Q., Bansal, G., Zhang, J., Wu, Y., Zhang, S., Zhu, E., Li, B., Jiang, L., Zhang, X., and Wang, C. Autogen: Enabling next-gen llm applications via multi-agent conversation framework. arXiv preprint arXiv:2308.08155, 2023a.
  77. 77.Wu, T., Zhu, B., Zhang, R., Wen, Z., Ramchandran, K., and Jiao, J. Pairwise proximal policy optimization: Harnessing relative feedback for llm alignment. arXiv preprint arXiv:2310.00212, 2023b.
  78. 78.Xu, J., Wang, S., Wang, L., and Yao, A. C.-C. Fedcm: Federated learning with client-level momentum. arXiv preprint arXiv:2106.10874, 2021.
  79. 79.Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D. Shadow alignment: The ease of subverting safely-aligned language models. arXiv preprint arXiv:2310.02949, 2023.
  80. 80.Ye, R., Chai, J., Liu, X., Yang, Y., Wang, Y., and Chen, S. Emerging safety attack and defense in federated instruction tuning of large language models. arXiv preprint arXiv:2406.10630, 2024.
  81. 81.Ye, S., Feng, X., Zhang, T., Ma, X., Lin, S., Li, Z., Xu, K., Wen, W., Liu, S., Tang, J., et al. Progressive dnn compression: A key to achieve ultra-high weight pruning and quantization rates using admm. arXiv preprint arXiv:1903.09769, 2019.
  82. 82.Ye, S., Jo, Y., Kim, D., Kim, S., Hwang, H., and Seo, M. Selfee: Iterative self-revising llm empowered by self-feedback generation. Blog post, May, 3, 2023.
  83. 83.Yi, J., Ye, R., Chen, Q., Zhu, B., Chen, S., Lian, D., Sun, G., Xie, X., and Wu, F. On the vulnerability of safety alignment in open-access llms. In Findings of the Association for Computational Linguistics ACL 2024, pp. 9236–9260, 2024a.
  84. 84.Yi, X., Zheng, S., Wang, L., Wang, X., and He, L. A safety realignment framework via subspace-oriented model fusion for large language models. arXiv preprint arXiv:2405.09055, 2024b.
  85. 85.Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F. Rrhf: Rank responses to align language models with human feedback without tears. arXiv preprint arXiv:2304.05302, 2023.
  86. 86.Zhan, Q., Fang, R., Bindu, R., Gupta, A., Hashimoto, T., and Kang, D. Removing rlhf protections in gpt-4 via fine-tuning. arXiv preprint arXiv:2311.05553, 2023.
  87. 87.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  88. 88.Zhang, X., Zhao, J., and LeCun, Y. Character-level convolutional networks for text classification. Advances in neural information processing systems, 28, 2015.
  89. 89.Zhao, J., Deng, Z., Madras, D., Zou, J., and Ren, M. Learning and forgetting unsafe examples in large language models. arXiv preprint arXiv:2312.12736, 2023.
  90. 90.Zhou, P., Yuan, X., Xu, H., Yan, S., and Feng, J. Efficient meta learning via minibatch proximal update. Advances in Neural Information Processing Systems, 32, 2019.
  91. 91.Zhu, M., Yang, L., Wei, Y., Zhang, N., and Zhang, Y. Locking down the finetuned llms safety. arXiv preprint arXiv:2410.10343, 2024.
  92. 92.Zong, Y., Bohdal, O., Yu, T., Yang, Y., and Hospedales, T. Safety fine-tuning at (almost) no cost: A baseline for vision large language models. arXiv preprint arXiv:2402.02207, 2024.
  93. 93.Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., et al. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405, 2023.

Citation

MLA
Huang, T., et al. “Lisa: Lazy Safety Alignment for Large Language Models Against Harmful Fine-tuning Attack”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 104521–55, https://proceedings.neurips.cc/paper_files/paper/2024/file/bcfdaf04b54a69f47623c973c864ee8d-Paper-Conference.pdf.
APA
Huang, T., Hu, S., Ilhan, F., Tekin, S. F., & Liu, L. (2024). Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack. Advances in Neural Information Processing Systems, 37, 104521–104555. https://proceedings.neurips.cc/paper_files/paper/2024/file/bcfdaf04b54a69f47623c973c864ee8d-Paper-Conference.pdf
Chicago
Huang, T., S. Hu, F. Ilhan, S. F. Tekin, and L. Liu. 2024. “Lisa: Lazy Safety Alignment for Large Language Models Against Harmful Fine-tuning Attack”. Advances in Neural Information Processing Systems 37: 104521–55. https://proceedings.neurips.cc/paper_files/paper/2024/file/bcfdaf04b54a69f47623c973c864ee8d-Paper-Conference.pdf.
Harvard
Huang, T. et al. (2024) “Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 104521–104555. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/bcfdaf04b54a69f47623c973c864ee8d-Paper-Conference.pdf.
Vancouver
1. Huang T, Hu S, Ilhan F, Tekin SF, Liu L (2024) Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 104521–104555

BibTeX

@inproceedings{huang2024lisa,
  title = {Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack},
  author = {Huang, Tiansheng and Hu, Sihao and Ilhan, Fatih and Tekin, Selim F. and Liu, Ling},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {104521-104555},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/bcfdaf04b54a69f47623c973c864ee8d-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors