Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack
Tiansheng HuangSihao HuFatih IlhanSelim F. TekinLing Liu
Proposes a bi-state optimization method incorporating a proximal constraint to protect large language models against harmful fine-tuning attacks without degrading downstream task performance.
Commercial fine-tuning services allow users to customize large language models with proprietary data. However, this creates a major security and liability risk: if user data inadvertently or maliciously includes harmful content, it can override previous safety alignment and cause the model to generate unsafe responses. Existing defenses either fail when fine-tuning involves extensive steps or require expensive, full-scale retraining for every individual request.
The article evaluates a computationally lightweight defense implemented directly during the fine-tuning stage. It demonstrates a novel method designed to retain safety guardrails against harmful fine-tuning data without undermining task performance on downstream user applications.
The authors conducted empirical experiments and convergence analyses to investigate multi-task optimization between safe alignment data and user datasets. They evaluated the framework across three foundation models (Llama2-7B, Opt-2.7B, and Mistral-7B) on four benchmark tasks (SST2, AGNEWS, GSM8K, and AlpacaEval) under varying ratios of harmful data. To address optimization instability caused by alternating between safety and user data, the proposed method—Lazy Safety Alignment (Lisa)—adds a proximal constraint to limit model drift between training phases.
The primary findings demonstrate that simple alternating optimization between alignment and user datasets degrades if computational steps allocated to safety are limited, which causes the model parameters to drift excessively. By introducing a proximal penalty to control this drift, Lisa reduces the average harmful output rate by 7.07% compared to alignment-stage defenses (Vaccine-SFT) and by 3.68% compared to data-mixing baselines (Vlguard), while keeping task accuracy intact within a 0.59% variance. Across distinct architectures, Lisa reduced harmful response rates by 11.2% to 11.9% on reasoning tasks and remained robust even when harmful training data ratios approached 100%. Furthermore, system measurements confirmed that Lisa introduces modest overhead, requiring only about 8.3% more execution time and 3.14 GB of additional GPU memory compared to standard fine-tuning.
These results show that service providers can mitigate liability and safety risks without sacrificing customization quality or incurring heavy computational penalties. Unlike purely preventive alignment strategies, Lisa operates during the customization phase and functions effectively even when user data filtering exhibits false negatives.
Organizations providing fine-tuning services should consider adopting proximal-constrained alternating optimization in their training workflows. For stronger protection, combining input data filtering with Lisa is recommended to eliminate residual risks. Future work should expand the method beyond supervised fine-tuning to evaluate compatibility with reinforcement learning from human feedback and test deployments on interactive agent applications.
While the theoretical convergence and experimental results are consistent across multiple benchmarks, evaluations were limited to supervised fine-tuning setups on medium-scale models. Stakeholders should conduct pilot assessments on production workloads and higher-parameter models to determine exact resource overheads and performance trade-offs.
- Paper: Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!, Xiangyu Qi et al. (2023). This foundational paper demonstrates how even minor fine-tuning on downstream datasets compromises pre-trained safety alignment, defining the core vulnerability that Lisa is engineered to resolve.
- Paper: Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations, Hakan Inan et al. (2023). This work introduces baseline guardrail and moderation mechanisms for language models, providing essential context for the input-output safety defenses compared against Lisa.
- Paper: Llama 2: Open Foundation and Fine-Tuned Chat Models, Hugo Touvron et al. (2023). This technical report presents the Llama 2 architecture and its supervised safety tuning procedures, which serve as the primary foundation and evaluation testbed in Lisa.
- Paper: On the Exploitability of Instruction Tuning, Manli Shu et al. (2023). This study demonstrates how instruction fine-tuning datasets can be exploited via data poisoning to subvert model behavior, establishing the threat model addressed by proximal-constrained defenses.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). This paper analyzes the structural failure modes of safety alignment under competing objectives, illuminating the theoretical reasons why fine-tuning causes alignment degradation.
- Paper: Safe RLHF: Safe Reinforcement Learning from Human Feedback, Josef Dai et al. (2024). This work formulates safety alignment as a constrained optimization problem balancing helpfulness and harmlessness during RLHF, extending beyond Lisa's supervised fine-tuning scope.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). This framework standardizes adversarial red-teaming and robust refusal benchmarks across various harm categories, providing an extensive testbed to validate defenses like Lisa.
- Paper: Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks, Samyak Jain et al. (2024). This research mechanistically analyzes how fine-tuning alters representation layers rather than core capabilities, offering theoretical depth to explain parameter drift during customization.
- Paper: BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents, Yifei Wang et al. (2024). This paper explores fine-tuning vulnerabilities in autonomous interactive agents, providing a concrete continuation into the complex agent deployments highlighted for future exploration in Lisa.
- Paper: Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression, Junyuan Hong et al. (2024). This study scrutinizes how post-training parameter modifications and compression techniques impact aligned trustworthiness across diverse benchmarks.
