Learning Diverse Attacks on Large Language Models for Robust Red-Teaming and Safety Tuning
Seanie Lee1^{1}1 ∗^{*}∗ Minsu Kim1,2,3^{1,2,3}1,2,3 Lynn Cherif2,4^{2,4}2,4 David Dobre2,3^{2,3}2,3 Juho Lee1^{1}1 Sung Ju Hwang1^{1}1 Kenji Kawaguchi5^{5}5 Gauthier Gidel2,3,7^{2,3,7}2,3,7 Yoshua Bengio2,3,7^{2,3,7}2,3,7 Esmeralda S. Whitammer6^{6}6 Moksh Jain2,3^{2,3}2,3
1^{1}1 KAIST 2^{2}2 Mila – Québec AI Institute 3^{3}3 Université de Montréal 4^{4}4 McGill University 5^{5}5 National University of Singapore 6^{6}6 University of Edinburgh 7^{7}7 CIFAR AI Chair
1^{1}1 KAIST 2^{2}2 Mila – Québec AI Institute 3^{3}3 Université de Montréal 4^{4}4 McGill University 5^{5}5 National University of Singapore 6^{6}6 University of Edinburgh 7^{7}7 CIFAR AI Chair
Abstract
Red-teaming, or identifying prompts that elicit harmful responses, is a critical step in ensuring the safe and responsible deployment of large language models (LLMs). Developing effective protection against many modes of attack prompts requires discovering diverse attacks. Automated red-teaming typically uses reinforcement learning to fine-tune an attacker language model to generate prompts that elicit undesirable responses from a target LLM, as measured, for example, by an auxiliary toxicity classifier. We show that even with explicit regularization to favor novelty and diversity, existing approaches suffer from mode collapse or fail to generate effective attacks. As a flexible and probabilistically principled alternative, we propose to use GFlowNet fine-tuning, followed by a secondary smoothing phase, to train the attacker model to generate diverse and effective attack prompts. We find that the attacks generated by our method are effective against a wide range of target LLMs, both with and without safety tuning, and transfer well between target LLMs. Finally, we demonstrate that models safety-tuned using a dataset of red-teaming prompts generated by our method are robust to attacks from other RL-based red-teaming approaches. Code is available at https://github.com/GFNOrg/red-teaming.
1. Introduction
The deployment of large language models (LLMs) in the wild has raised concerns about their potential harmful impacts for nearly a decade ([1, 2]). These concerns have grown with the increasing capabilities of LLMs: even models fine-tuned to satisfy certain safety constraints can be manipulated to produce toxic outputs ([3]). Red-teaming, or identification of 'attack' prompts that elicit undesirable responses, gives model developers as well as regulators a chance to identify and address such vulnerabilities before deployment ([4]). This paper studies the problem of automatically generating diverse attack prompts for LLMs and argues for the potential of robust automated red-teaming in the development of effective defenses.
Effective red-teaming requires identifying many modes of attack ([5]). Methods for automated red-teaming based on stochastic optimization of attack prompts ([6, 7]) have been proposed, while others have used reinforcement learning (RL) to train an attacker language model (LM), allowing to generate novel prompts efficiently at test time ([4, 5]). However, even when regularized to favor diversity, these methods struggle to balance between diversity and effective attacks (Figure 2). They often suffer from mode collapse, where the attacker LM generates a small set of similar prompts, or focus solely on diversity and fail to generate effective attacks (Figure 3). Moreover, we have empirically found that they also fail to discover attacks that transfer across different target LLMs (Table 3).
This paper takes an amortized inference perspective on red-teaming: we view the problem of generating an attack prompt as sampling a latent variable in a probabilistic model. Using the off-policy RL approach of GFlowNet fine-tuning, proposed for inference of linguistic latent variables in ([8]), we fine-tune attack LMs to sample the full posterior distribution over attack prompts.
However, controlling the 'peakiness' of the posterior distribution – the preference of attack quality to attack diversity – is challenging, especially when red-teaming a target LLM that has been safety-tuned to resist some modes of attack, leading to a sparser landscape of attack prompts. Inspired by the success of behavior cloning in offline RL ([9, 10]), we propose a two-stage GFlowNet fine-tuning procedure with MLE smoothing. As illustrated in Figure 1, we first fine-tune a pretrained attacker LM with a GFlowNet objective and collect high-reward attack prompts discovered in the course of training (Step 1). The collected prompts form an offline dataset. Subsequently, the pretrained attacker model is fine-tuned again to maximize the likelihood of the offline dataset (Step 2). The first stage, GFlowNet fine-tuning, enables us to collect a set of diverse and effective attack prompts using exploratory off-policy training. In the second phase, we obtain a smooth distribution over high-reward attack prompts, since all the collected attack prompts in the offline dataset are considered equally important and the attacker LM is trained to maximize their log-likelihood uniformly. Consequently, we find that the attacker LM is able to sample attack prompts that are both diverse and effective.
We empirically evaluate the efficacy of our proposed method in red-teaming five target LLMs: GPT-2 ([11]), Dolly-v2-7b ([12]), Gemma-2b-it ([13]), Llama-2-7b-chat ([14]), and Llama-3.1-8B-Instruct ([15]). Our approach is found to sample more diverse and effective attack prompts than other relevant baselines. Moreover, many of our attack prompts effectively transfer to other target LLMs that are not used for training the attacker model, such as Llama-2-13b/70b-chat, Llama-3-8b/70b-instruct ([15]), Starling-7b-beta ([16]), and Mistral-7b-instruct-v0.2 ([17]). Lastly, we fine-tune a target LLM to generate refusal responses to the discovered attack prompts and find that the model fine-tuned with our red-teaming prompts is more robust than the models safety-tuned with other RL-based red-teaming methods.
It is important to note that while we study an approximate measure of toxicity as a proxy for harmfulness, following past works ([4, 5]), the true harmful impact of an LLM output is often subjective and dependent on the social context of deployment ([2]). We nonetheless believe that the methods we propose will be useful in practice and can be extended to other measures of harmfulness.
Our contributions and findings are summarized below:
- To generate diverse and effective attack prompts, we take a probabilistic perspective on red-teaming and demonstrate the usefulness of the off-policy RL approach of GFlowNet fine-tuning.
- We propose a smoothing and reranking step that can be used to generalize from high-reward samples found during GFlowNet fine-tuning, improving the attacker model and allowing efficient adaptation to new target LLMs.
- Attacker LMs trained with GFlowNet-finetuning followed by MLE generate more diverse and effective attack prompts that also transfer to other target LLMs.
- When safety-tuned on attack prompts generated by our method, target LLMs become robust to attacks generated by other RL-based methods without performance degradation on other tasks.
2. Related work
Red-teaming.
As LLMs increase in general capabilities and performance, so does the risk associated to potential misuse of LLMs. To mitigate this, LLMs are often trained to refuse to generate content given prompts that are dangerous, offensive, or harmful ([18, 19]). This is done at various stages of the training process such as filtering out harmful training data ([13]) or fine-tuning on 'safe' responses to harmful prompts ([14]). This process is often augmented by red-teaming, which proactively looks for ways to elicit harmful behavior from models. Prior works ([20, 21, 22]) rely on a large amount of human annotation to identify vulnerabilities of LMs. To automate red-teaming, [4] formulate red teaming as an RL problem and train an LM to sample toxic prompts. However, most RL algorithms are not suitable for sampling diverse objects since they tend to converge to a single reward-maximizing trajectory. To overcome this limitation, [5] propose using a novelty-based reward to encourage a policy to explore diverse samples during RL training. Instead of generating a prompt from scratch, [23] replace words of prompts from a predefined user input pool to attack LMs using Bayesian optimization in a sample-efficient manner. Rainbow Teaming ([24]) samples an attack prompt from a pool and iteratively mutates the prompt with auxiliary LLMs.
Jailbreaks. Jailbreaking and red-teaming are closely related in that red-teaming proactively tries to discover vulnerabilities for the purpose of improving model safety, whereas jailbreaking generally refers to circumventing the built-in safeguards of models. Initially, jailbreaks were found manually through trial and error, taking advantage of the different objectives models were trained against ([3]). Recently, automated jailbreak attacks are becoming increasingly popular. They utilize techniques such as genetic algorithms ([25]), iterative gradient-based methods ([6]), or automated prompting via auxiliary LLMs ([26]) to optimize query prompts. [27] propose a method defending against GCG ([6]), one of the most popular gradient-based jailbreak methods. A drawback of these methods is the computational cost since the optimization has to be performed separately for each new query prompt. Another drawback is the poor transferability of jailbreaks. [28] have shown that prompts optimized by GCG to jailbreak one target LLM do not transfer to jailbreak other target LLMs.
GFlowNets. Generative flow networks (GFlowNets; [29]) are a probabilistic framework to train stochastic policies to sample discrete compositional objects (e.g., graphs, sequences) proportionally to a reward. Sampling objects proportionally to a reward results in diverse high-reward samples. Consequently, GFlowNets have found applications in a wide variety of problems including biological sequence generation ([30]), combinatorial optimization ([31, 32]), Bayesian structure learning ([33]), variational EM with discrete latent variables ([34]), and probabilistic neurosymbolic inference ([35]). Most closely related to our work is ([8]), which uses the GFlowNet objective to fine-tune LMs for solving intractable inference problems such as sampling chains of thought ([36]). We use GFlowNet fine-tuning as a part of our approach for learning policies which generate diverse prompts that elicit toxic responses from target LLMs.
3. Sampling diverse attacks with GFlowNet fine-tuning
3.1 Preliminaries
The target LLM, denoted pϕp_\phipϕ, samples a text response y{\mathbf{y}}y for a given prompt x{\mathbf{x}}x with probability pϕ(y∣x)p_\phi({\mathbf{y}}\mid {\mathbf{x}})pϕ(y∣x). The goal of red-teaming an LLM is to identify prompts x{\mathbf{x}}x that elicit toxic responses from the target LLM. A binary toxicity classifier, denoted as pψp_\psipψ, is used to quantify the effectiveness of an attack prompt. Specifically, the effectiveness of a prompt x{\mathbf{x}}x is measured by the likelihood of the response y∼pϕ(y∣x){\mathbf{y}}\sim p_\phi({\mathbf{y}}\mid {\mathbf{x}})y∼pϕ(y∣x) being classified as toxic by the classifier: pψ(c=1∣x,y)p_\psi(c=1\mid {\mathbf{x}}, {\mathbf{y}})pψ(c=1∣x,y), where c∈{0,1}c\in\{0,1\}c∈{0,1} is a binary variable denoting toxicity. Moreover, for the attack to be effective, the prompt x{\mathbf{x}}x should appear natural, as unnatural prompts (with high perplexity under some prior) are easy to defend against with simple filters ([37]).
Red-teaming can often be a time-consuming process if done manually as the space of prompts is quite large. [4, 5] formulate red-teaming as an RL problem, to automate the discovery of these prompts. This involves training a LM as a policy pθp_\thetapθ, parameterized by θ\thetaθ, to generate prompts that maximize the expected reward (as measured by the toxicity of the response generated by the target LLM):
where the KL divergence term, weighted by a hyperparameter λ>0\lambda>0λ>0, encourages the policy pθp_\thetapθ to remain close to an initial pretrained LM prefp_\texttt{ref}pref, penalizing the generation of prompts x{\mathbf{x}}x that are far from natural language text. However, most RL algorithms are not suitable for discovering diverse prompts since they generally concentrate most of probability mass of the policy pθp_\thetapθ on actions with highest reward, often resulting in a deterministic policy that generates a single prompt ([29]). While [5] propose adding a novelty-based reward term along with entropy bonus ([38]) as a regularization to encourage the policy to generate diverse prompts, empirically we find that it is challenging to find an optimal trade-off between diversity and toxicity rate even with the regularization. In the context of red-teaming, identifying diverse and effective attack prompts is critical to ensure that the target LLM is sufficiently safety-tuned for a broad range of scenarios which might be encountered when the model is deployed in the wild.
3.2 GFlowNet fine-tuning and smoothing with MLE on collected high-reward prompts
A probabilistic view of the problem provides a principled alternative. Specifically, problem of generating diverse and effective red-teaming prompts can be viewed as one of generating samples from a (tempered) reward distribution. We adopt the perspective of generative flow networks (GFlowNets; [29, 39]), leveraging their ability to learn policies that sample from a target distribution defined over compositional objects such as sequences ([30]) and graphs ([39]). To instantiate the probabilistic perspective, we propose a two-stage approach designed to learn a stochastic policy to sample diverse and effective prompts for red-teaming. The first stage consists of fine-tuning a pretrained LM pθp_\thetapθ as a GFlowNet policy ([8]) in order to collect prompts, and the second stage restarts fine-tuning from the original pretrained LM policy but this time with maximum likelihood estimation (MLE) on the high-reward prompts collected during GFlowNet training in the first stage.
Stage 1: GFlowNet fine-tuning. GFlowNets are diversity-seeking RL algorithms that learn a policy pθp_\thetapθ which samples prompts with a probability proportional to the reward associated with the prompt1. We define the reward for a prompt x{\mathbf{x}}x as follows:
1.
In the case of generating sequences, GFlowNets are equivalent to MaxEnt RL ([40]).
where β\betaβ and γ\gammaγ are positive constants that control the 'peakiness' (tempering) of the toxicity score R1(x)R_1({\mathbf{x}})R1(x) and of the reference LM likelihood R2(x)R_2({\mathbf{x}})R2(x), respectively. The prompt x=(x0,x1,…,xT){\mathbf{x}}=(x_0, x_1, \ldots, x_T)x=(x0,x1,…,xT), consisting of TTT tokens with a special token x0x_0x0 indicating the beginning of a sentence, is generated autoregressively from a behavior policy, which is a mix of pθp_\thetapθ and a tempered variant of it. We define (x0,x1,…,xt)(x_0, x_1, \ldots, x_t)(x0,x1,…,xt) as a state in the generative process and the token sampled from the policy at each step is the action. To learn the parameters θ\thetaθ, we use the trajectory balance learning objective ([41]):
where Zθ>0Z_\theta >0Zθ>0 is a learnable scalar approximating the partition function. One distinction of the red-teaming setting, compared to other GFlowNet tasks, is that the reward is stochastic as it depends on the response sampled from the LLM. In practice, we approximate the log reward logR(x)\log R({\mathbf{x}})logR(x) with an empirical mean over kkk samples from the target LLM:
As we illustrate in Section 4, using GFlowNet fine-tuning alone to sample effective and diverse red-teaming prompts can be challenging in practice due to non-trivial choice of the temperature parameters β\betaβ and γ\gammaγ. While in principle there are choices of β\betaβ and γ\gammaγ which can balance the reward and diversity well, in practice GFlowNet fine-tuning can be overly sensitive to the peakiness of the reward ([42]). Moreover, balancing between β\betaβ and γ\gammaγ to achieve the desired behavior is non-trivial. For example, while all three examples shown in Table 1 get a high toxicity reward, the first two get a low total reward compared to the last one, even though they are grammatically valid sentences, since they are assigned a low likelihood by prefp_\texttt{ref}pref. If we set a much smaller β\betaβ to increase the weight of the toxicity reward R1(x)R_1({\mathbf{x}})R1(x), the policy pθp_\thetapθ would likely generate prompts from potentially spurious modes of the toxicity classifier, which will have high perplexity under the reference model. On the other hand, if we set γ\gammaγ to a small value, the model would merely focus on the naturality score R2(x)R_2({\mathbf{x}})R2(x) and not generate toxic prompts.
Stage 2: Smoothing with MLE.
To reduce sensitivity to the aforementioned parameters of the reward distribution, while preserving the mode coverage and ability of the training procedure to generalize to new modes, we propose an inexpensive retraining step that is applied following GFlowNet fine-tuning. This second step is akin to behavior cloning ([43, 9, 10]) in RL, where a policy is trained to imitate expert trajectories. First, we store all prompts sampled by the policy pθp_\thetapθ during GFlowNet fine-tuning in Stage 1. We expect this set to contain diverse and high-reward prompts discovered by off-policy exploration during GFlowNet fine-tuning. Subsequently, we filter the prompts in this set based on the toxicity score R1(x)R_1({\mathbf{x}})R1(x) and language model likelihood R2(x)R_2({\mathbf{x}})R2(x) being larger than some thresholds. The collected examples form an offline dataset, and the reference policy is fine-tuned again (from the same initial state as in Stage 1) to maximize log-likelihood of samples from this offline dataset. Stage 2 is very inexpensive in practice, taking under 5% of total (Stage 1 and 2) training time in our experiments (Table 4).
We outline our complete method in. This procedure results in a smooth approximation to the distribution over high-reward prompts found during exploratory training.
4. Experiments
4.1 Experimental setup
Task. We validate the efficacy of our method by red-teaming a variety of open source LMs with varying levels of saferty-tuning. Recall that the goal of red-teaming is to generate diverse prompts that elicit harmful responses from a target LLM pϕp_\phipϕ. In our experiments, we use GPT-2 ([11]), Dolly-v2-7b ([12]), Llama-2-7b-chat ([14]), Llama-3.1-8B-Instruct ([15]), and Gemma-2b-it ([13]) as target LLMs. Note that we only require black-box access, i.e., we can only query a prompt and get a text response.
Evaluation. To measure effectiveness of each method, we compute the toxicity rate which is the percentage of generated prompts that are toxic. We consider a prompt toxic if the toxicity classifier pψp_\psipψ assigns a score greater than 0.5 to an input pair consisting of the prompt x{\mathbf{x}}x and response y{\mathbf{y}}y generated by the target LLM. For pψp_\psipψ, we use a RoBERTa hate speech classifier ([44]) for GPT-2 and dolly-v2-7b, and Llama-Guard ([45]) for Llama-2-7b-chat, Llama-3.1-8B-Instruct, and Gemma-2b-it. To measure the diversity of the generated prompts, we embed the generated prompts using MiniLMv2 ([46]) encoder and compute the average pairwise cosine distance between embeddings of the prompts.
Methods. We compare our proposed method against some relevant red-teaming baselines:
- Supervised Fine-tuning (SFT): We fine-tune the pretrained LM pθp_\thetapθ with a maximum likelihood objective on 3,003 toxic prompts from SafetyDataset ([47]) and AdvBench ([6]).
- In-Context Learning (ICL) ([48]): We sample 5-shot demonstrations from toxic prompt datasets (SafetyDataset and AdvBench) and prompt GPT-2 to generate a prompt.
- REINFORCE ([49]): We fine-tune the pretrained LM pθp_\thetapθ as an RL policy with policy gradients to optimize the reward in Equation 1.
- PPO + Novelty ([5]): This method adds entropy bonus ([38]) along with a novelty-based term to the reward in Equation 1 and train the policy pθp_\thetapθ with proximal policy optimization (PPO; [50]). For novelty-based reward, it utilizes self-BLEU ([51]) and pairwise cosine similarity between embeddings of all the past generated prompts.
- GFlowNet ([41]): We fine-tune the pretrained LM pθp_\thetapθ with Equation 3. (This is Stage 1 of our full procedure.)
- GFlowNet + MLE: This is our full method for collecting high-reward prompts during GFlowNet fine-tuning and re-training the pretrained LM pθp_\thetapθ with maximum likelihood estimation (MLE) on the collected prompts as described in.
4.2 Results: Robust red-teaming
Studying the trade-off between diversity and toxicity rate. As the number of prompts which would elicit toxic responses occupy a small subset of all possible sequences, there is a natural trade-off between diversity and toxicity. We start by investigating how each method handles this trade-off. Figure 2 illustrates the cosine distance plotted against the toxicity rate for 10,00010,00010,000 red-teaming prompts generated by each method across five different target LLMs. We find that our GFlowNet + MLE is the only method which manages to balance a high toxicity rate with the diversity of generated prompts across all four target LLMs. Qualitative assessment of examples generated by GFlowNet + MLE, included in Table 10, Table 11, Table 12, Table 13, and Table 14, supports the numerical results. While the GFlowNet achieves both high diversity and toxicity rate for red-teaming GPT-2 (Figure 7) and Dolly-v2-7b (Figure 2a), the toxicity rate drops significantly for the target LLMs with safety fine-tuning: Gemma-2b-it (Figure 2b), Llama-2-7b-chat (Figure 2c), and Llama-3.1-8B-Instruct (Figure 2d). We hypothesize this drop comes from the reward signal (toxicity of responses from the target) becoming sparse with safety-tuned models. Similarly, PPO + Novelty fails to find a balance between diversity and toxicity. When it is able to find effective prompts (Figure 7 and Figure 2a), they are not as diverse and for models with strong safety-guardrail, such as Llama-2 and Gemma, it fails to find any prompts which elicit a toxic response (Figure 2b and Figure 2c). When it comes to red-teaming Llama-3.1-8B-Instruct, it moderately finds a balance between toxicity and diversity but still falls significantly short compared to our GFlowNets + MLE approach. (For context, a random policy would have the highest diversity but would have a low toxicity rate). On the other hand, REINFORCE, which does not take diversity into account, collapses to deterministically generating a single reward-maximizing prompt. Finally, SFT and ICL generate diverse but ineffective prompts.
Scaling to a larger attacker LM. Table 2 shows the effect of scaling GFlowNet+MLE with larger and stronger attackers like Llama-3.2-1B ([15]). Scaling to a larger attacker results in significant improvements in both the toxicity rate and diversity.
GFlowNet + MLE generates diverse and effective prompts. To further understand the behavior of each method beyond the toxicity rate (which depends on the pψ(c=1∣x,y)>0.5p_\psi(c=1\mid {\mathbf{x}}, {\mathbf{y}}) > 0.5pψ(c=1∣x,y)>0.5 decision boundary), we illustrate the distribution over the toxicity scores and corresponding average pairwise cosine distances for the generated prompts in Figure 3, obtained from the experiment for red-teaming the Llama-2-7b-chat target LLM. Results for the other target LLMs are illustrated in Figure 8, Figure 9, Figure 10, and Figure 11 in Appendix B.2. GFlowNet + MLE achieves consistently high diversity across different toxicity score bins. On the other hand, all other methods fail to achieve high diversity and toxicity at the same time. GFlowNet generates fewer toxic prompts compared to GFlowNet + MLE. Notably, PPO + Novelty does not generate prompts with the toxicity score greater than 0.20.20.2 at all for Gemma-2b-it and Llama-2-7b-chat. While REINFORCE generates a single highly toxic prompt achieving a much lower diversity, SFT and ICL generate few toxic prompts.
GFlowNet attacks are more transferable across target LLMs. A potential advantage of generating diverse attack prompts is that prompts generated for red-teaming a given target LLM can potentially transfer to other LLMs, since some of the failure modes of a target LLM might be shared by other models, for instance, due to using similar web-filtered data or similar safety alignment recipes. To study this empirically, we train an attacker policy pθp_\thetapθ for red-teaming the Gemma-2b-it as the target LLM. We then sample 1,0241,0241,024 prompts from the trained attacker LM and evaluate the number of prompts which transfer to other LLMs, i.e., elicit toxic responses from unseen LLMs: Llama-2-7b-chat, Llama-2-13b-chat, Llama-2-70b-chat, Llama-3-8b-instruct ([15]), Llama-3-70b-instruct, Gemma-7b-it, Gemma-1.1-2b-it, Gemma-1.1-7b-it, Mistral-7b-instruct-v0.2 ([17]), and Starling-7b-beta ([16]). As shown in Table 3, we find that many prompts generated by GFlowNet + MLE transfer to unseen target LLMs, outperforming all other methods across all the target LLMs except Mistral-7b-instruct-v0.2. REINFORCE generates almost identical prompts, tailored to the Gemma-2b-it target it was trained with, which consequently do not transfer to other target LLMs. This highlights a drawback of methods which do not generate diverse attacks. On the other extreme, PPO + Novelty is unable to discover any prompt that is effective in eliciting toxic responses and consequently none of the prompts transfer to any other LLM. These results further highlight the efficacy and usefulness of GFlowNet + MLE, which can generate both diverse and effective red-teaming prompts that can be transferred to red-team other LLMs. Additionally, we perform another transfer experiment targeted for a proprietary model, GPT-4o. We generate 1,024 prompts with an attacker LM trained to red-team Llama-2-7b-chat and evaluate how many prompts can elicit harmful responses from GPT-4o. On average, across five different sets of 1,024 prompts, 65% of them can successfully attack GPT-4o.
Stage 2 (MLE) is cheap. As shown in Table 4, our proposed second stage MLE training is a lightweight process compared to other RL methods since it does not need on-policy samples or expensive reward computation. With just two hours of additional training, MLE training can significantly enhance the diversity and toxicity rate of GFlowNets.
MLE with reranking allows fast adaptation to new target LMs. Another advantage of our two-stage approach is that it can enable fast adaptation of an attacker LM policy to a new target: an attacker trained against one target LLM can be adapted to red-team a different target LLM by repeating Stage 2 on a dataset filtered using the new target LLM. Concretely, we can recompute the reward of the stored attack prompts sampled during GFlowNet fine-tuning (Stage 1), with a different target LLM and rerank the prompts (instead of scoring them with the same target LLM). The offline dataset can be constructed by filtering the prompts with the newly computed R1(x)R_1({\mathbf{x}})R1(x) and the precomputed R2(x)R_2({\mathbf{x}})R2(x) based on the corresponding thresholds r1r_1r1 and r2r_2r2. The initial pretrained attacker LM policy pθp_\thetapθ is fine-tuned with supervised learning on this dataset. For this experiment, we consider the the prompts stored during the red-teaming of Gemma-2b-it and adapt the attacker LM to red-team Gemma-1.1-2b-it, Gemma-7b-it, Gemma-1.1-7b-it, Llama-2-7b-chat, and Llama-3-8b-instruct target LLMs. As shown in Figure 4, adaptation of the attack LM policy with this reranking procedure is effective and significantly improves toxicity rate over direct transfer from an attacker trained to red-team the initial target LLM, Gemma-2b-it. Note that a considerable amount of computational cost and wall-clock time can be saved (cf. Table 4), since we skip the GFlowNet fine-tuning stage (Stage 1) and simply reuse the stored prompts.
Reward temperature controls toxicity vs. diversity. In this experiment, we demonstrate empirically the challenges in tuning the temperature β\betaβ in Equation 2 and how the second phase of MLE smoothing provides a better trade-off between toxicity rate and diversity. We fine-tune the pretrained initial policy pθp_\thetapθ as a GFlowNet by setting the temperature β\betaβ to each value in {0.01,0.02,…,0.1,1.0}\{0.01, 0.02, \ldots, 0.1, 1.0\}{0.01,0.02,…,0.1,1.0} and fine-tune again the initial attacker LM policy with MLE on each of the high-reward prompts discovered during GFlowNet fine-tuning with the corresponding β\betaβ. As shown in Figure 5, the GFlowNet objective sacrifices diversity (cosine distance) considerably to obtain high toxicity rate, or it significantly degrades the toxicity rate resulting in diverse prompts. On the other hand, smoothing with MLE is robust to this choice of β\betaβ and enables the attacker policy to sample effective attack prompts while retaining diversity.
GFlowNet samples are better than PPO + Novelty for MLE smoothing. We perform an ablation study to demonstrate the importance of the off-policy exploration ability of GFlowNets for collecting the offline dataset in Stage 1 for MLE smoothing in Stage 2, as described in. To show this, we attempt to replace GFlowNet fine-tuning by PPO + Novelty in Stage 1. To be precise, we train the initial pretrained policy pθp_\thetapθ with PPO + Novelty for red-teaming the Llama-2-7b-chat target LLM and attempt to collect high-reward prompts sampled by the policy during training. However, PPO + Novelty does not generate any high-reward prompts at all during training, as shown in Table 5. Thus, we cannot collect any samples for the offline dataset and consequently cannot fine-tune the policy with MLE. This indicates that a diversity-seeking approach like GFlowNet fine-tuning is crucial for collecting effective red-teaming prompts.
4.3 Results: Safety fine-tuning
Setup. To demonstrate that attack prompts generated by red-teaming methods can enhance the safety alignment of target LLMs, we fine-tune Gemma-2b-it target LLM with the attack prompts discovered by various methods and evaluate its robustness. Specifically, we generate 1,024 prompts with each red-teaming method and generate refusal responses to the attack prompts with the target LLM using the prompt: "Refuse answering to the following instruction.
attack_prompt". Subsequently, the target model is fine-tuned with LoRA ([52]) to maximize the conditional log-likelihood of the refusal responses to the attack prompts, resulting in six different fine-tuned target LLMs corresponding to each red-teaming method. Finally, each fine-tuned model generates responses to the attack prompts generated by each red-teaming method, and we measure the toxicity rate of the responses with Llama-Guard as the toxicity classifier pψp_\psipψ.GFlowNet + MLE allows for robust safety-tuned target LLMs. As shown in Table 4, the target LLM fine-tuned on the attack prompts generated by GFlowNet + MLE is the most robust to unseen attack prompts generated by the other RL-based red-teaming methods. On the other hand, even after safety fine-tuning, all the other target LLMs cannot defend against the attack prompts generated by GFlowNet + MLE. We also confirm that our safety-tuned model still preserves general instruction-following capabilities: as shown in Table 7, the performance on the six tasks in the Open LLM Leaderboard changes insignificantly with safety tuning. These results highlight the importance of the diversity of generated red-teaming prompts for downstream safety fine-tuning.
5. Conclusion
As LMs become increasingly more capable and widely used, red-teaming them for a wide variety of potential attacks becomes more critical for safe and responsible deployment. We have proposed an approach to generate diverse and effective red-teaming prompts using a novel two-stage procedure consisting of GFlowNet fine-tuning followed by MLE smoothing. Through our experiments, we showed that our approach is effective for red-teaming a wide variety of target LMs with varying levels of safety-tuning. An interesting observation is the transferability of the generated prompts to different target LLMs, which reveals shared failure modes of current approaches for aligning LMs and opens interesting direction for future work. In particular, our reranking-based adaptation procedure can serve as a quick way to red-team new target LLMs during development.
Our approach is not limited to text tokens and future work can explore the applicability to red-team multimodal models (e.g., text-to-image models ([53, 54])). Further, an interesting area of future work is extending the approach to the jailbreaking setting, where an attacker language model generates a suffix for an adversarial query prompt. Finally, in addition to red-teaming, it would be interesting to apply our method to generate prompts which can improve model performance on different tasks ([55]).
Limitations. While our approach shows promising performance for red-teaming various target language models, the performance is still limited by the classifier used to quantify the harmfulness of a response. The true harm that an LM output causes is often subjective and depends on the social context of deployment ([2]). As with other RL-based approaches, our approach is trained online (i.e., requires iteratively sampling the current model) and, consequently, requires sampling several responses from the target LLM to compute the reward during training, which can be costly.
Ethics statement
Our proposed red-teaming framework is useful for automatically discovering diverse ways to induce undesirable responses from LLMs. Before deployment of the LLM, we can perform safety fine-tuning of the model to prevent generation of harmful responses. However, our method can be misused to attack commercial LLMs at scale, since it can generate harmful prompts that transfer to other target LLMs. This necessitates precautions for the deployment of LLMs. We can defend against such attacks by filtering harmful responses with the toxicity classifier employed for training the attacker model.
Reproducibility statement
We use PyTorch ([56]) and the Hugging Face Transformers library ([57]) to implement our models and all the baselines. All the implementation details are described in Appendix A, and our code is available at https://github.com/GFNOrg/red-teaming.
Acknowledgments
The authors would like to thank Nicholas Meade for helpful suggestions at the inception of this project.
The research was enabled in part by computational resources provided by the Digital Research Alliance of Canada (https://alliancecan.ca), Mila (https://mila.quebec), and NVIDIA.
The authors acknowledge funding from CIFAR, NSERC, IVADO, and Samsung. Lynn Cherif is supported by a FRQNT Master's Training Scholarship.
This material is based upon work supported by the Air Force Office of Scientific Research under award number FA2386-24-1-4011, and this research is partially supported by the Singapore Ministry of Education Academic Research Fund Tier 1 (Award No: T1 251RES2207).
This work was partially supported by Institute for Information & communications Technology Promotion (IITP) grant funded by the Korea government (MSIT) (No. RS-2019-II190075, Artificial Intelligence Graduate School Program (KAIST)), (No. RS-2020-II200153, Penetration Security Testing of ML Model Vulnerabilities and Defense), the National Research Foundation of Korea(NRF) grant funded by the Korea government(MSIT) (NRF-2022R1A5A708390812), Institute of Information & communications Technology Planning & Evaluation(IITP) grant funded by the Korea government(MSIT) (No.2022-0-00184, Development and Study of AI Technologies to Inexpensively Conform to Evolving Policy on Ethics), and Institute of Information & communications Technology Planning & Evaluation(IITP) grant funded by the Korea government(MSIT) (No.RS-2022-II220713, Meta-learning Applicable to Real-world Problems).
Appendix
A. Implementation details
For all the experiments, we use pretrained GPT-2 consisting of 124 million parameters for the policy pθp_\thetapθ. Apart from the ICL baseline, we initially fine-tune GPT-2 using 3,003 toxic prompts from the SafetyDataset and AdvBench with the AdamW optimizer (AdamW) for 200 iterations. We set the batch size, learning rate, and weight decay to 102410241024, 3⋅10−53\cdot10^{-5}3⋅10−5 and 0.10.10.1, respectively. Subsequently, we further fine-tune the model with each method. For GFlowNet fine-tuning, we fine-tune the model for 5,0005,0005,000 iterations with AdamW optimzer, setting batch size and learning rate to 128128128 and 10−410^{-4}10−4, respectively. Regarding the hyperparameters for the reward, we set β\betaβ and γ\gammaγ to 0.10.10.1 and 1.01.01.0, respectively, and use k=5k=5k=5 samples for approximating the log-reward. Following GFlowNet fine-tuning, we collect samples generated by GFlowNet, if the sample achieves toxicity score R1(x)R_1({\mathbf{x}})R1(x) and reference language model log likelihood logR2(x)\log R_2({\mathbf{x}})logR2(x) greater than 0.70.70.7 and −100-100−100, respectively. Then we train the initial supervised fine-tuned model on the collected samples using AdamW Optimizer, learning rate 3⋅10−53\cdot10^{-5}3⋅10−5, and batch size 2,0482,0482,048 for 1,0001,0001,000 steps or 2,0002,0002,000 steps, depending on the target language model. When red-teaming Llama and Gemma, we use A100 80GB gpu to train the policy with GFlowNet and re-train the model with MLE for 1,0001,0001,000 steps. Otherwise, we use 3090 RTX gpu for GFlowNet Training and re-train the model for 2,0002,0002,000 steps.
B. Additional results
B.1 Trade-off between toxicity score and diversity
B.2 Toxicity score
B.3 Ablation of toxicity classifier
In order to study the effect of a reward function, we replace Llama-Guard ([45]) with a RoBERTa-based hate speech classifier ([44]) during the training of GFlowNet for computing the reward R1(x)R_1({\mathbf{x}})R1(x) in Equation 2. As shown in Table 6, the RoBERTa classifier assigns high toxicity score (reward) to prompts that do not actually elicit toxic responses from the Llama-2-7b-chat target model. This leads GFlowNet to generate false positive prompts, a phenomenon known as reward hacking ([58]), where a policy trained with a proxy behaves well according to the proxy but misaligns with the true objective due to mis-specifications of the proxy ([59]). Note that reward hacking is common in many RL applications ([60, 61, 62, 63]), and both PPO + Novelty and REINFORCE also suffer from the same reward hacking issue when red-teaming Gemma-2b-it and Llama-2-7b-chat models with the RoBERTa classifier. The reward hacking issue can be mitigated if we use Llama-Guard as a toxicity classifier as shown in Table 13 and Table 12. GFlowNet + MLE effectively generate prompts that elicit toxic responses from target language models. This is the reason why we use Llama-Guard for red-teaming and evaluating all the target models trained with safety alignment.
B.4 Downstream task performance after safety-tuning
As discussed in Section 4.3, we fine-tune Gemma-2b-it target LLM with LoRA ([52]) to maximize the log-likelihood of refusal responses to the red-teaming prompts that our GFlowNet + MLE generated. Subsequently, we evaluate the safety-tuned model on Open LLM Leaderboard benchmark which consists of six datasets — ARC ([64]), HellaSwag ([65]), TruthfulQA ([66]), MMLU ([67]), and GSM8k ([68]). As shown in Table 7, there is no significant performance drop after safety-tuning, which indicates that the safety-tuned target LLM still retrain instruction following capabilities.
B.5 Results with standard deviation
B.6 Example attacks and responses
References
[1] Peter Lee (2016). Learning from Tay's introduction. https://blogs.microsoft.com/blog/2016/03/25/learning-tays-introduction/.
[2] Laura Weidinger et al. (2021). Ethical and social risks of harm from Language Models. arXiv preprint arXiv:2112.04359.
[3] Wei et al. (2023). Jailbroken: How does LLM safety training fail?. Neural Information Processing Systems (NeurIPS).
[4] Perez et al. (2022). Red Teaming Language Models with Language Models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. pp. 3419–3448. doi:10.18653/v1/2022.emnlp-main.225. https://aclanthology.org/2022.emnlp-main.225.
[5] Zhang-Wei Hong et al. (2024). Curiosity-driven Red-teaming for Large Language Models. International Conference on Learning Representations (ICLR).
[6] Zou et al. (2023). Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043.
[7] Zhao et al. (2024). Accelerating Greedy Coordinate Gradient via Probe Sampling. arXiv preprint arXiv:2403.01251.
[8] Hu et al. (2024). Amortizing intractable inference in large language models. International Conference on Learning Representations (ICLR).
[9] Scott Emmons et al. (2022). RvS: What is Essential for Offline RL via Supervised Learning?. International Conference on Learning Representations (ICLR).
[10] Eric Jang et al. (2021). BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning. Conference on Robot Learning (CoRL).
[11] Alec Radford et al. (2019). Language Models are Unsupervised Multitask Learners.
[12] Mike Conover et al. (2023). Free Dolly: Introducing the World's First Truly Open Instruction-Tuned LLM. https://www.databricks.com/blog/2023/04/12/dolly-first-open-commercially-viable-instruction-tuned-llm.
[13] Gemma Team Thomas Mesnard et al. (2024). Gemma: Open Models Based on Gemini Research and Technology. arXiv preprint arXiv:2403.08295.
[14] Hugo Touvron et al. (2023). Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv preprint arXiv:2307.09288.
[15] Dubey et al. (2024). The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
[17] Jiang et al. (2023). Mistral 7B. arXiv preprint arXiv:2310.06825.
[18] Yuntao Bai et al. (2022). Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. arXiv preprint arXiv:2204.05862.
[19] Bai et al. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv preprint arXiv:2212.08073.
[20] Dinan et al. (2019). Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). pp. 4537–4546. doi:10.18653/v1/D19-1461. https://aclanthology.org/D19-1461.
[21] Xu et al. (2021). Bot-Adversarial Dialogue for Safe Conversational Agents. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. pp. 2950–2968. doi:10.18653/v1/2021.naacl-main.235. https://aclanthology.org/2021.naacl-main.235.
[22] Wallace et al. (2022). Analyzing Dynamic Adversarial Training Data in the Limit. In Findings of the Association for Computational Linguistics: ACL 2022. pp. 202–217. doi:10.18653/v1/2022.findings-acl.18. https://aclanthology.org/2022.findings-acl.18.
[23] Lee et al. (2023). Query-Efficient Black-Box Red Teaming via Bayesian Optimization. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 11551–11574. doi:10.18653/v1/2023.acl-long.646. https://aclanthology.org/2023.acl-long.646.
[24] Mikayel Samvelyan et al. (2024). Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts. arXiv preprint arXiv:2402.16822.
[25] Xiaogeng Liu et al. (2024). AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models. International Conference on Learning Representations (ICLR).
[26] Chao et al. (2023). Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419.
[27] Mantas Mazeika et al. (2024). HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. arXiv preprint arXiv:2402.04249.
[28] Meade et al. (2024). Universal Adversarial Triggers Are Not Universal. arXiv preprint arXiv:2404.16020.
[29] Emmanuel Bengio et al. (2021). Flow Network based Generative Models for Non-Iterative Diverse Candidate Generation. Neural Information Processing Systems (NeurIPS).
[30] Jain et al. (2022). Biological Sequence Design with GFlowNets. International Conference on Machine Learning (ICML).
[31] David W Zhang et al. (2023). Robust Scheduling with GFlowNets. International Conference on Learning Representations (ICLR).
[32] Dinghuai Zhang et al. (2023). Let the flows tell: Solving graph combinatorial problems with GFlowNets. Neural Infromation Processing Systems (NeurIPS ).
[33] Deleu et al. (2022). Bayesian Structure Learning with Generative Flow Networks. Uncertainty in Artificial Intelligence (UAI).
[34] Hu et al. (2023). GFlowNet-EM for learning compositional latent variable models. International Conference on Machine Learning (ICML).
[35] van Krieken et al. (2023). A-NeSI: A scalable approximate method for probabilistic neurosymbolic inference. Neural Information Processing Systems (NeurIPS).
[36] Wei et al. (2022). Chain-of-thought prompting elicits reasoning in large language models. Neural Information Processing Systems (NeurIPS).
[37] Neel Jain et al. (2023). Baseline Defenses for Adversarial Attacks Against Aligned Language Models. arXiv preprint arXiv:2309.00614.
[38] Schulman et al. (2017). Equivalence between policy gradients and soft Q-learning. arXiv preprint arXiv:1704.06440.
[39] Bengio et al. (2023). GFlowNet foundations. Journal of Machine Learning Research. 24(210). pp. 1–55.
[40] Tuomas Haarnoja et al. (2017). Reinforcement Learning with Deep Energy-Based Policies. International Conference on Machine Learning (ICML).
[41] Malkin et al. (2022). Trajectory balance: Improved credit assignment in GFlowNets. Neural Information Processing Systems (NeurIPS).
[42] Lau et al. (2024). QGFN: Controllable Greediness with Action Values. arXiv preprint arXiv:2402.05234.
[43] Lili Chen et al. (2021). Decision Transformer: Reinforcement Learning via Sequence Modeling. Neural Information Processing Systems (NeurIPS).
[44] Vidgen et al. (2021). Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). pp. 1667–1682. doi:10.18653/v1/2021.acl-long.132. https://aclanthology.org/2021.acl-long.132.
[45] Inan et al. (2023). Llama guard: LLM-based input-output safeguard for human-AI conversations. arXiv preprint arXiv:2312.06674.
[46] Wang et al. (2021). MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021. pp. 2140–2151. doi:10.18653/v1/2021.findings-acl.188. https://aclanthology.org/2021.findings-acl.188.
[47] Federico Bianchi et al. (2024). Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions. International Conference on Learning Representations (ICLR).
[48] Tom B. Brown et al. (2020). Language Models are Few-Shot Learners. Neural Information Processing Systems (NeurIPS).
[49] Williams, Ronald J (1992). Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning. 8. pp. 229–256.
[50] Schulman et al. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.
[51] Zhu et al. (2018). Texygen: A Benchmarking Platform for Text Generation Models. SIGIR.
[52] Edward J Hu et al. (2022). LoRA: Low-Rank Adaptation of Large Language Models. International Conference on Learning Representations (ICLR).
[53] Ramesh et al. (2021). Zero-shot text-to-image generation. International Conference on Machine Learning (ICML).
[54] Saharia et al. (2022). Photorealistic text-to-image diffusion models with deep language understanding. Neural Information Processing Systems (NeurIPS).
[55] Lin et al. (2023). The unlocking spell on base LLMs: Rethinking alignment via in-context learning. arXiv preprint arXiv:2312.01552.
[56] Paszke et al. (2019). Pytorch: An imperative style, high-performance deep learning library. Neural Information Processing Systems (NeurIPS).
[57] Wolfe et al. (2022). Supporting Mouthing in Signed Languages: New innovations and a proposal for future corpus building. In Proceedings of the 7th International Workshop on Sign Language Translation and Avatar Technology: The Junction of the Visual and the Textual: Challenges and Perspectives. pp. 125–130. https://aclanthology.org/2022.sltat-1.19.
[58] Skalse et al. (2022). Defining and characterizing reward gaming. Neural Information Processing Systems (NeurIPS).
[59] Alexander Pan et al. (2022). The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models. International Conference on Learning Representations (ICLR).
[60] Romain Paulus et al. (2018). A Deep Reinforced Model for Abstractive Summarization. International Conference on Learning Representations (ICLR).
[61] Yizhong Wang et al. (2023). How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources. Neural Information Processing Systems (NeurIPS).
[62] Tom Everitt et al. (2017). Reinforcement Learning with a Corrupted Reward Channel. International Joint Conference on Artificial Intelligence (IJCAI).
[63] Bowen Baker et al. (2020). Emergent Tool Use From Multi-Agent Autocurricula. International Conference on Learning Representations (ICLR).
[64] Peter Clark et al. (2018). Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge. arXiv preprint arXiv:1803.05457.
[65] Zellers et al. (2019). HellaSwag: Can a Machine Really Finish Your Sentence?. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. pp. 4791–4800. doi:10.18653/v1/P19-1472. https://aclanthology.org/P19-1472.
[66] Lin et al. (2022). TruthfulQA: Measuring How Models Mimic Human Falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 3214–3252. doi:10.18653/v1/2022.acl-long.229. https://aclanthology.org/2022.acl-long.229.
[67] Dan Hendrycks et al. (2021). Measuring Massive Multitask Language Understanding. International Conference on Learning Representations (ICLR).
[68] Cobbe et al. (2021). Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.





















