Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation

Xianghe PangShuo TangRui YeYuxin XiongBolun ZhangYanfeng WangSiheng Chen

article2024ICML50 citations

Introduces MATRIX, a social scene simulator that aligns large language models to human values entirely through self-play role-playing and consequence evaluation, enabling a 13B model to surpass GPT-4 in value alignment without relying on external supervision.

Listen

As large language models become increasingly capable, aligning them with human values is essential to prevent severe harms like the generation of toxic content, cyberattack assistance, or dangerous misinformation. Existing alignment strategies depend heavily on external resources, such as labor-intensive human feedback or costly oversight from commercial models, while self-alignment methods typically rely on rigid, pre-defined rules that fail to handle complex real-world situations. The article demonstrates that language models can autonomously achieve value alignment through multi-agent social scene simulation, activating their latent understanding of social norms without requiring external supervision or sacrificing inference efficiency.

To achieve this, the researchers developed MATRIX, a social scene simulator that acts as a virtual rehearsal space. When presented with a potentially harmful user instruction, a single model role-plays multiple characters and objects involved in the scenario—similar to a solo theater performance—while an internal social modulator enforces realistic logical constraints and communication flow. By observing the negative multi-party consequences simulated in this rehearsal, the model generates tailored critiques to correct its initial response. To eliminate the computational runtime delay of running simulations during live use, the resulting simulated interactions and refined responses are compiled into a dataset used to fine-tune the model through standard supervised training.

Across extensive empirical evaluations spanning four major safety benchmarks, the fine-tuned model consistently outperformed over ten established training and inference baselines. Most notably, in a blind study of 875 human ratings on safety-critical queries, a fine-tuned 13-billion-parameter open-source model surpassed the alignment quality of GPT-4. Automated evaluations using GPT-4 as an impartial judge confirmed that the system achieves high win rates over existing commercial and research models across diverse domains, including hate speech, advice on criminal activity, and sensitive safety concerns, all while preserving general performance on standard knowledge and conversational benchmarks.

These findings indicate that organizations deploying language models can drastically reduce alignment expenses, operational latency, and third-party API dependencies by using simulated multi-stakeholder interactions. Rather than relying on abstract rulebooks that models struggle to apply, context-specific social feedback teaches the model to recognize concrete societal harms. Next steps for decision-makers and technical teams include integrating social scene simulation pipelines into model fine-tuning workflows and conducting pilot studies on domain-specific risks. Further research should evaluate how simulation frameworks handle subtle cultural biases and whether incorporating external digital tools into agent simulations can expand alignment capabilities to even broader operational domains.

Cover for Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation

Abstract

Aligning large language models (LLMs) with human values is imperative to mitigate potential adverse effects resulting from their misuse. Drawing from the sociological insight that acknowledging all parties’ concerns is a key factor in shaping human values, this paper proposes a novel direction to align LLMs by themselves: social scene simulation. To achieve this, we present MATRIX, a novel social scene simulator that emulates realistic scenes around a user’s input query, enabling the LLM to take social consequences into account before responding. MATRIX serves as a virtual rehearsal space, akin to a Monopolylogue, where the LLM performs diverse roles related to the query and practice by itself. To inject this alignment, we fine-tune the LLM with MATRIX-simulated data, ensuring adherence to human values without compromising inference speed. We theoretically show that the LLM with MATRIX outperforms Constitutional AI under mild assumptions. Finally, extensive experiments validate that our method outperforms over 10 baselines across 4 benchmarks. As evidenced by 875 user ratings, our tuned 13B-size LLM exceeds GPT-4 in aligning with human values. See our project page at https://shuotang123.github.io/MATRIX.

Table of Contents

  • 1. Introduction
  • 2. Proposed Self-Alignment System
  • 3. MATRIX: Social Scene Simulator
  • 3.1. Social Roles
  • 3.2. Social Modulator
  • 3.3. Social Scene Simulation
  • 3.4. Discussions
  • 4. Theoretical Analysis of MATRIX
  • 5. Experiments
  • 5.1. Experimental Setup
  • 5.2. Evaluation of the LLM with MATRIX
  • 5.3. Evaluation of MATRIX-Tuned LLM
  • 5.4. Ablation Study
  • 6. Related Work
  • 7. Conclusions
  • Acknowledgements
  • Impact Statement
  • References
  • A. Analysis and proofs
  • A.1. Definitions
  • A.2. Analysis for Assumption 4.4.
  • A.3. Lemmas
  • A.4. Proof to Lemma A.4
  • A.5. Proof to Lemma A.5
  • A.6. Proof to Lemma A.6
  • A.7. Analysis and proof to Theorem 4.5
  • B. Experiments
  • B.1. Fine-tuning Configurations
  • B.2. Baselines Implementation
  • B.3. Evaluation Details for GPT-4
  • B.4. Experiments on other models
  • B.5. Experiments on simulation with different models
  • B.6. Experiments on Time Consumption
  • B.7. Experiments on Mitigating Bias
  • C. Prompts and Examples
  • D. Human Evaluation Details
  • E. Qualitative Examples [Warning: Potentially Harmful Content!]

Knowls

  1. Knowl 1 — MATRIX Social Scene Simulation Framework

    model/method

    MATRIX is a social scene simulation framework that enables an unaligned large language model (LLM) to achieve value alignment without external human or AI feedback. Operating on the principle of a Monopolylogue—where a single actor embodies multiple roles—the unaligned LLM generates and executes all elements of an interactive social environment conditioned on a user instruction qq and its initial response rr.

    MATRIX comprises three main components:

    1. Social Roles: Instantiated by the LLM into living agents and non-living objects:

      • User Agent: Emulates the user by executing the action sequence deconstructed from the initial response rr.
      • Reactive Agents: Represent stakeholders and bystanders within the scenario, each assigned distinct personal traits, status, and self-interests.
      • Objects: Non-autonomous entities (e.g., banks, accounts, physical items) whose textual states are modified by agent actions.
    2. Social Modulator: A central coordinator driven by the same LLM, equipped with a textual memory system and two core mechanisms:

      • Action Feasibility Determination: Checks whether an agent's proposed action conforms to physical constraints, real-world common sense, and past memory. Infeasible actions are rejected.
      • Message Distribution: Contextually filters and routes information, deciding which agents observe specific actions or object state changes based on perceptual feasibility.
      • Critique Generation and Memory Logging: For every feasible action, the modulator evaluates potential social harm, logging the action description and its corresponding critique into textual memory.
    3. Simulation Lifecycle & Consequence Summarization: The simulation proceeds through sequential turns of agent action, feasibility verification, critique logging, and message routing. It concludes either upon natural narrative convergence (agents produce no further actions) or upon premature termination if a user action irreversibly violates real-world logic. After termination, the modulator synthesizes the memory log into a comprehensive textual summary of social consequences, which is used by the LLM to generate an instruction-specific critique and revise its response into a socially aligned output.

  2. Knowl 2 — Two-Stage Self-Alignment Pipeline via Supervised Fine-Tuning

    model/method

    The MATRIX self-alignment system transforms an unaligned base language model into a socially aligned model in two phases:

    1. Self-Generation of Consequence-Aware Responses: Given a set of user instructions, the unaligned base LLM generates an initial response rr. MATRIX simulates the multi-agent social interactions resulting from rr, logging critiques and compiling a social consequence summary. Conditioning on this consequence summary, the base LLM performs self-critique and revises its initial response into a harmless and helpful response oo.

    2. Supervised Fine-Tuning (SFT): To eliminate the computational latency of running multi-agent simulations at inference time, an SFT dataset is constructed containing:

      • Helpful instruction-response pairs,
      • Harmful instruction pairs paired with their MATRIX-revised consequence-aware responses, and
      • Dialogue interaction traces generated during the MATRIX simulations.

    The base LLM is fine-tuned on this dataset using parameter-efficient fine-tuning (e.g., QLoRA for 3 epochs with cosine learning rate schedule, learning rate 2×10−52 \times 10^{-5}, sequence length 1024, and DeepSpeed ZeRO-2). The resulting fine-tuned LLM directly outputs aligned responses during standard inference without invoking the simulation environment or suffering runtime overhead.

  3. Knowl 3 — Probabilistic Formulation of Critique-Based Alignment

    definition

    Let SS denote the set of token sequences, M:S→SM: S \to S be a language model, and O+⊆SO^+ \subseteq S be the target set of harmless and helpful sequences. A critique-based alignment method TM(⋅)T_M(\cdot) maps an input instruction q∈Sq \in S to an output o∈So \in S via two steps:

    1. Initial Step: The model generates an initial response r=M(q)r = M(q) drawn from distribution P(r∣q)P(r \mid q).
    2. Revise Step: Conditioned on a critique cc, the model revises rr to produce o=M(q,r,c)o = M(q, r, c) drawn from distribution P(o∣q,r,c)P(o \mid q, r, c).

    A method T1T_1 is defined to be better aligned than T2T_2 with respect to O+O^+ (T1⪰O+T2T_1 \succeq_{O^+} T_2) if for all q∈Sq \in S:

    P(T1(q)∈O+∣q)≥P(T2(q)∈O+∣q)P(T_1(q) \in O^+ \mid q) \ge P(T_2(q) \in O^+ \mid q)

    For the standard Critique-Revise baseline TMCRT_M^{\text{CR}} (as in Constitutional AI), a fixed, human-predefined rule critique cCRc^{\text{CR}} is applied:

    P(TMCR(q)∈O+∣q)=∑rP(r∣q)P(o∈O+∣q,r,cCR)P(T_M^{\text{CR}}(q) \in O^+ \mid q) = \sum_{r} P(r \mid q) P(o \in O^+ \mid q, r, c^{\text{CR}})

    For MATRIX TMMT_M^M, conditioned on an NN-step simulation generating interactions (i1,…,iN)(i_1, \dots, i_N) and critiques (c1M,…,cNM)(c_1^M, \dots, c_N^M):

    P(TMM(q)∈O+∣q)=∑rP(r∣q)∑ik,k∈[N](∏j=1NP(ij∣ij−1,…,i1,r))∑ck,k∈[N]P(o∈O+∣c1M,…,cNM,q,r)∏k=1NP(ckM∣ik,q)P(T_M^M(q) \in O^+ \mid q) = \sum_{r} P(r \mid q) \sum_{i_k, k \in [N]} \left( \prod_{j=1}^N P(i_j \mid i_{j-1}, \dots, i_1, r) \right) \sum_{c_k, k \in [N]} P(o \in O^+ \mid c_1^M, \dots, c_N^M, q, r) \prod_{k=1}^N P(c_k^M \mid i_k, q)

    For any η∈R+\eta \in \mathbb{R}^+, the η\eta-bounded critique set is:

    C(η,q,r)={c∣P(o∈O+∣q,r,c)≤η}C_{(\eta, q, r)} = \{ c \mid P(o \in O^+ \mid q, r, c) \le \eta \}

    The maximum effectiveness ξ\xi of a critique cc across all queries is defined as:

    ξ=inf⁡ηsup⁡q∈S{η∣c∈C(η,q,M(q))}\xi = \inf_{\eta} \sup_{q \in S} \{ \eta \mid c \in C_{(\eta, q, M(q))} \}

  4. Knowl 4 — Theoretical Superiority of MATRIX over Rule-Based Critique-Revise

    theoretical result

    Let MM be a language model. Let TMCRT_M^{\text{CR}} denote the critique-revise method using fixed human-predefined critiques cCRc^{\text{CR}} with maximum effectiveness ξCR\xi_{\text{CR}}, and let TMMT_M^M denote the LLM aligned via MATRIX generating collective critiques c[1:n]={ci}i=1nc_{[1:n]} = \{c_i\}_{i=1}^n.

    Assumptions:

    1. Collective advantage: For any single critique cic_i (i∈[n]i \in [n]) generated during simulation, revising with the collective critique c[1:n]c_{[1:n]} is more effective: P(M(q,r,ci)∈O+∣q,r,ci)≤P(M(q,r,c[1:n])∈O+∣q,r,c[1:n])P(M(q, r, c_i) \in O^+ \mid q, r, c_i) \le P(M(q, r, c_{[1:n]}) \in O^+ \mid q, r, c_{[1:n]}).
    2. Stable critique generating: There exists λ∈R+\lambda \in \mathbb{R}^+ such that for all n∈N∗n \in \mathbb{N}^*:

    1nlog⁡(P(o∈O+∣q,r,c[1:n])P(o[1:n]∈O+∣q,r,c[1:n]))≤λ\frac{1}{n} \log \left( \frac{P(o \in O^+ \mid q, r, c_{[1:n]})}{P(o_{[1:n]} \in O^+ \mid q, r, c_{[1:n]})} \right) \le \lambda

    1. Alignment chance: For all q∈Sq \in S, there exists ϵ>0\epsilon > 0 such that P(o∈O+∣q)>ϵP(o \in O^+ \mid q) > \epsilon.

    Theorem: If the maximum effectiveness ξCR\xi_{\text{CR}} of the human-predefined critique satisfies:

    ξCR<1−1−e−λ\sqrt{\xi_{\text{CR}}} < 1 - \sqrt{1 - e^{-\lambda}}

    then TMM⪰O+TMCRT_M^M \succeq_{O^+} T_M^{\text{CR}}, meaning that for all instructions q∈Sq \in S:

    P(TMM(q)∈O+∣q)≥P(TMCR(q)∈O+∣q)P(T_M^M(q) \in O^+ \mid q) \ge P(T_M^{\text{CR}}(q) \in O^+ \mid q)

    This confirms that because abstract constitutional rules have low instruction-specific effectiveness ξCR\xi_{\text{CR}}, the dynamic multi-agent simulation critiques in MATRIX provide superior alignment guarantees.

  5. Knowl 5 — Inference-Time Alignment Performance of LLM with MATRIX

    data/table

    Pairwise win, tie, and lose rates of Wizard-Vicuna-30B-Uncensored enhanced with MATRIX at inference time against 7 baseline methods across four safety benchmarks: HH-RLHF, PKU-SafeRLHF, AdvBench, and HarmfulQA. Comparisons were evaluated by GPT-4 as an automated judge on 100 randomly sampled questions per dataset, alternating response presentation order to eliminate position bias.

    Evaluation Dataset HH Safe-RLHF AdvBench HarmfulQA
    LLM with MATRIX vs. Win Tie Lose Win Tie Lose Win Tie Lose Win Tie Lose
    Vanilla 81% 12% 7% 91% 6% 3% 83% 12% 5% 82% 11% 7%
    Self-Align 96% 2% 2% 89% 3% 8% 73% 11% 16% 93% 2% 5%
    Context Distillation 82% 8% 10% 91% 6% 3% 67% 16% 17% 82% 9% 9%
    Critique-Revise 94% 5% 1% 89% 6% 5% 80% 8% 12% 81% 12% 7%
    RAIN 70% 0% 30% 100% 0% 0% 70% 20% 10% 90% 0% 10%
    LLM Debate 77% 11% 12% 88% 7% 5% 71% 10% 19% 77% 14% 9%
    Best-of-N Sampling 78% 8% 14% 84% 6% 10% 75% 17% 8% 77% 10% 13%
    ChatGPT (GPT-3.5-Turbo) 57% 10% 33% 71% 8% 21% 91% 3% 6% 54% 6% 40%

    The LLM with MATRIX consistently outperforms all self-alignment baselines, multi-agent debate, Best-of-N sampling, and GPT-3.5-Turbo across all datasets.

  6. Knowl 6 — Safety Alignment and General Capability Retention of MATRIX-Tuned LLMs

    data/table

    Pairwise safety comparisons evaluated by GPT-4 for the MATRIX-tuned 30B LLM (Wizard-Vicuna-30B-Uncensored fine-tuned via SFT on MATRIX-generated data) against 8 baselines across four safety datasets, alongside general capability scores on Vicuna-Bench and MT-Bench (rated on a 1–10 scale).

    Evaluation Dataset HH Safe-RLHF AdvBench HarmfulQA
    MATRIX-Tuned LLM vs. Win Tie Lose Win Tie Lose Win Tie Lose Win Tie Lose
    Vanilla 84% 10% 6% 80% 9% 11% 84% 14% 2% 82% 9% 9%
    Self-Align 89% 5% 6% 93% 4% 3% 71% 17% 12% 96% 1% 3%
    Context Distillation 84% 5% 11% 82% 9% 9% 76% 15% 9% 81% 8% 11%
    Critique-Revise 89% 6% 5% 81% 8% 11% 76% 18% 6% 90% 4% 6%
    Stable Alignment 80% 16% 4% 80% 13% 7% 78% 9% 13% 73% 14% 13%
    Mistake Analysis 87% 5% 8% 80% 2% 8% 77% 18% 5% 83% 9% 8%
    RLcd 66% 13% 21% 62% 20% 18% 44% 15% 41% 50% 17% 33%
    RLAIF 84% 6% 10% 80% 5% 15% 71% 19% 10% 72% 18% 10%
    ChatGPT (GPT-3.5-Turbo) 65% 9% 26% 64% 6% 30% 74% 5% 21% 58% 9% 33%
    Evaluation Benchmark Vicuna-Bench MT-Bench
    Vanilla 8.37 6.99
    Self-Align 5.79 4.11
    Context Distillation 8.12 6.80
    Critique-Revise (Harmful) 6.31 5.46
    Critique-Revise (Harmful Helpful) 8.14 6.92
    Stable Alignment 8.40 6.78
    Mistake Analysis 8.38 6.87
    RLcd 8.47 6.95
    RLAIF 7.41 6.60
    MATRIX-Tuned LLM 8.49 6.99

    The MATRIX-tuned model achieves substantial alignment gains over all baselines without incurring alignment tax, achieving the highest Vicuna-Bench score (8.49) and matching the vanilla base model on MT-Bench (6.99).

  7. Knowl 7 — Human Evaluation of MATRIX-Tuned LLMs vs. GPT-4

    empirical result

    In a blinded human study on 100 harmful prompts sampled across 14 safety categories from PKU-SafeRLHF, 35 non-author volunteers evaluated pairwise model responses (each prompt received at least 8 independent ratings, yielding 875 total ratings per model comparison and 1750 ratings across 13B and 30B models).

    Both the MATRIX-tuned 13B LLM (Wizard-Vicuna-13B-Uncensored fine-tuned with MATRIX data) and the MATRIX-tuned 30B LLM achieved higher aggregate win rates than GPT-4 across the 14 harmful prompt categories, marking the first demonstration of open 13B-size and 30B-size models aligned via self-simulation outperforming GPT-4 on human value alignment.

  8. Knowl 8 — Training Data Composition and Iterative Self-Improvement

    data/table

    Ablation study evaluating the contribution of Harmful data (Ha., 6K samples), Helpful data (He., 6K samples), and MATRIX Simulation dialogue data (Si.) when fine-tuning Wizard-Vicuna-30B. Models were evaluated on MT-Bench for general capability, on pairwise win rate against the base LLM on the HH dataset after SFT, and on pairwise win rate when the fine-tuned model was further paired with inference-time MATRIX simulation.

    Harmful (Ha.) Helpful (He.) Simulation (Si.) MT-Bench Tuned vs. Base Further MATRIX vs. Base
    - - - 6.99 - 92.0%
    ✓ - - 6.50 89.3% 21.4%
    ✓ ✓ - 6.94 85.2% 63.3%
    ✓ ✓ ✓ 6.99 93.3% 95.6%

    Training solely on harmful data degrades general capabilities (MT-Bench drops from 6.99 to 6.50) and diminishes subsequent simulation capability (win rate drops to 21.4%). Combining harmful, helpful, and simulation data preserves general ability (6.99), yields the highest direct alignment win rate (93.3%), and enables continuous self-improvement through additional MATRIX simulation turns (95.6% win rate).

  9. Knowl 9 — Impact of Agent Quantity, Interaction Depth, and Model Scale

    empirical result

    Ablation studies on the simulation parameters of MATRIX and base model scale reveal the following trends:

    1. Number of Simulated Agents: Increasing the number of social roles in MATRIX from 1 to 2, 4, and 7 increases the harmless win rate against the unaligned base model from ~55% (1 agent) to >90% (7 agents). Involving multiple diverse perspectives enriches the model's understanding of multi-stakeholder consequences.
    2. Number of Interactions: Increasing the interaction rounds in the simulation from 1 to 4, 8, 12, and 16 monotonically increases the harmless win rate against the base model from ~54% (1 interaction) to >90% (16 interactions), demonstrating that deeper simulation rollouts uncover delayed and compounding social harms.
    3. Model Scale Scaling: The win rate of both LLM with MATRIX and MATRIX-tuned LLM against GPT-3.5-Turbo scales positively with parameter size across 7B, 13B, and 30B Wizard-Vicuna base models (e.g., MATRIX-tuned LLM achieves ~48% win rate at 7B, ~63% at 13B, and ~75% at 30B).
  10. Knowl 10 — Inference Latency and Quality Comparison

    data/table

    Inference runtime and generation quality measured across 10 samples from the HH dataset on a single NVIDIA RTX 3090 GPU comparing the original unaligned base LLM (Wizard-Vicuna-30B), LLM with inference-time MATRIX, the SFT MATRIX-Tuned LLM, and RAIN (an inference-time alignment method relying on token-level search and backtracking).

    Metric RAIN Original LLM LLM with MATRIX MATRIX-Tuned LLM
    Inference Time (s) 2826.35 3.46 341.63 6.05
    Win / Tie / Lose vs. RAIN - 30 / 10 / 60 80 / 0 / 20 60 / 40 / 0

    While inference-time MATRIX takes 341.63s per sample (substantially faster than RAIN's 2826.35s while achieving an 80% win rate against it), SFT fine-tuning on MATRIX simulation data transfers this alignment into a model that executes in 6.05s per sample (nearly matching the vanilla base model's 3.46s) while outperforming RAIN (60% win, 40% tie, 0% lose).

  11. Knowl 11 — Social Stereotype and Bias Mitigation on BBQ

    data/table

    Evaluation of social bias reduction on the Bias Benchmark for QA (BBQ) for the vanilla Wizard-Vicuna-7B model versus its MATRIX-Tuned version. Accuracy measures the percentage of correct, non-stereotypical answers chosen in ambiguous contexts.

    Model Gender Accuracy Age Accuracy Race Accuracy
    Vanilla Wizard-Vicuna-7B 0.42 0.34 0.44
    MATRIX-Tuned Wizard-Vicuna-7B 0.46 0.48 0.50

    Fine-tuning on MATRIX-generated consequence simulations reduced stereotypical choices across all three demographic categories, increasing gender accuracy by 0.04, age accuracy by 0.14, and race accuracy by 0.06.

  12. Knowl 12 — Scope and Tool-Use Limitations of MATRIX

    limitation

    The current MATRIX simulation framework is designed and evaluated strictly for value alignment tasks (such as harmlessness, helpfulness, and bias mitigation). The social roles simulated in MATRIX operate purely through natural language generation and do not incorporate external tool use, web browsing, or executable API interactions during the simulation turns. Extending Monopolylogue multi-role simulation to scenarios with tool integration remains an open direction for general LLM self-improvement.

Coverage note — None was omitted; all contributed models, theoretical results, primary benchmarks, human evaluation studies, ablation analyses, latency evaluations, debiasing results, and stated limitations are fully covered.

References

  1. 1.Social constructivism, 2024. URL https://en.wikipedia.org/wiki/Social_constructivism. Accessed February 2, 2024.
  2. 2.Symbolic interactionism, 2024. URL https://en.wikipedia.org/wiki/Symbolic_interactionism. Accessed February 2, 2024.
  3. 3.Anwar, U., Saparov, A., Rando, J., Paleka, D., Turpin, M., Hase, P., Lubana, E. S., Jenner, E., Casper, S., Sourbut, O., et al. Foundational challenges in assuring alignment and safety of large language models. arXiv preprint arXiv:2404.09932, 2024.
  4. 4.Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861, 2021.
  5. 5.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022a.
  6. 6.Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022b.
  7. 7.Barrett, C., Boyd, B., Bursztein, E., Carlini, N., Chen, B., Choi, J., Chowdhury, A. R., Christodorescu, M., Datta, A., Feizi, S., et al. Identifying and mitigating the security risks of generative ai. Foundations and Trends® in Privacy and Security, 6(1):1–52, 2023.
  8. 8.Bengio, Y., Hinton, G., Yao, A., Song, D., Abbeel, P., Darrell, T., Harari, Y. N., Zhang, Y.-Q., Xue, L., Shalev-Shwartz, S., et al. Managing extreme ai risks amid rapid progress. Science, pp. eadn0117, 2024.
  9. 9.Bhardwaj, R. and Poria, S. Red-teaming large language models using chain of utterances for safety-alignment. arXiv preprint arXiv:2308.09662, 2023.
  10. 10.Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., et al. Weak-to-strong generalization: Eliciting strong capabilities with weak supervision. arXiv preprint arXiv:2312.09390, 2023.
  11. 11.Chen, K., Wang, C., Yang, K., Han, J., HONG, L., Mi, F., Xu, H., Liu, Z., Huang, W., Li, Z., Yeung, D.-Y., and Shang, L. Gaining wisdom from setbacks: Aligning large language models via mistake analysis. In The Twelfth International Conference on Learning Representations, 2024a.
  12. 12.Chen, W., Su, Y., Zuo, J., Yang, C., Yuan, C., Chan, C.-M., Yu, H., Lu, Y., Hung, Y.-H., Qian, C., et al. Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors. In The Twelfth International Conference on Learning Representations, 2024b.
  13. 13.Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. https://vicuna. lmsys. org, 2023.
  14. 14.Cognitivecomputations. Wizard-vicuna-30b-uncensored. https://huggingface.co/cognitivecomputations/Wizard-Vicuna-30B-Uncensored, 2024.
  15. 15.Deshpande, A., Murahari, V., Rajpurohit, T., Kalyan, A., and Narasimhan, K. Toxicity in chatgpt: Analyzing persona-assigned language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 1236–1270, 2023.
  16. 16.Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. Qlora: Efficient finetuning of quantized llms. Advances in Neural Information Processing Systems, 36, 2023.
  17. 17.Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., and Mordatch, I. Improving factuality and reasoning in language models through multiagent debate. arXiv preprint arXiv:2305.14325, 2023.
  18. 18.Fan, Z., Hu, S., Yao, J., Niu, G., Zhang, Y., Sugiyama, M., and Wang, Y. Locally estimated global perturbations is better than local perturbations for federated sharpness-aware minimization. In international conference on machine learning. PMLR, 2024a.
  19. 19.Fan, Z., Yao, J., Han, B., Zhang, Y., Wang, Y., et al. Federated learning with bilateral curation for partially class-disjoint data. Advances in Neural Information Processing Systems, 36, 2024b.
  20. 20.Gao, C., Lan, X., Li, N., Yuan, Y., Ding, J., Zhou, Z., Xu, F., and Li, Y. Large language models empowered agent-based modeling and simulation: A survey and perspectives. arXiv preprint arXiv:2312.11970, 2023.
  21. 21.Gupta, S., Shrivastava, V., Deshpande, A., Kalyan, A., Clark, P., Sabharwal, A., and Khot, T. Bias runs deep: Implicit reasoning biases in persona-assigned llms. In The Twelfth International Conference on Learning Representations, 2024.
  22. 22.Hall, P. M. Symbolic interaction. In Blackwell Encyclopedia of Sociology. Blackwell, 2007. ISBN 9781405124331. doi: 10.1002/9781405165518.wbeoss310.
  23. 23.Hazell, J. Large language models can be used to effectively scale spear phishing campaigns. arXiv preprint arXiv:2305.06972, 2023.
  24. 24.Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Wang, J., Zhang, C., Wang, Z., Yau, S. K. S., Lin, Z., et al. Metagpt: Meta programming for multi-agent collaborative framework. In The Twelfth International Conference on Learning Representations, 2024.
  25. 25.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  26. 26.Hua, W., Fan, L., Li, L., Mei, K., Ji, J., Ge, Y., Hemphill, L., and Zhang, Y. War and peace (waragent): Large language model-based multi-agent simulation of world wars. arXiv preprint arXiv:2311.17227, 2023.
  27. 27.Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Chen, B., Sun, R., Wang, Y., and Yang, Y. Beavertails: Towards improved safety alignment of llm via a human-preference dataset. Advances in Neural Information Processing Systems, 36, 2023.
  28. 28.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  29. 29.Kang, D., Li, X., Stoica, I., Guestrin, C., Zaharia, M., and Hashimoto, T. Exploiting programmatic behavior of llms: Dual-use through standard security attacks. arXiv preprint arXiv:2302.05733, 2023.
  30. 30.Khanov, M., Burapacheep, J., and Li, Y. Alignment as reward-guided search. In The Twelfth International Conference on Learning Representations, 2024.
  31. 31.Lee, H., Phatale, S., Mansoor, H., Lu, K., Mesnard, T., Bishop, C., Carbune, V., and Rastogi, A. Rlaif: Scaling reinforcement learning from human feedback with ai feedback. arXiv preprint arXiv:2309.00267, 2023.
  32. 32.Li, G., Hammoud, H., Itani, H., Khizbullin, D., and Ghanem, B. Camel: Communicative agents for" mind" exploration of large language model society. Advances in Neural Information Processing Systems, 36, 2023.
  33. 33.Li, Y., Wei, F., Zhao, J., Zhang, C., and Zhang, H. Rain: Your language models can align themselves without fine-tuning. In The Twelfth International Conference on Learning Representations, 2024.
  34. 34.Lin, B. Y., Ravichander, A., Lu, X., Dziri, N., Sclar, M., Chandu, K., Bhagavatula, C., and Choi, Y. Urial: Aligning untuned llms with just the’write’amount of in-context learning. In The Twelfth International Conference on Learning Representations, 2024.
  35. 35.Liu, R., Yang, R., Jia, C., Zhang, G., Yang, D., and Vosoughi, S. Training socially aligned language models on simulated social interactions. In The Twelfth International Conference on Learning Representations, 2024.
  36. 36.Liu, Z., Yao, W., Zhang, J., Xue, L., Heinecke, S., Murthy, R., Feng, Y., Chen, Z., Niebles, J. C., Arpit, D., et al. Bolaa: Benchmarking and orchestrating llm-augmented autonomous agents. arXiv preprint arXiv:2308.05960, 2023.
  37. 37.McKinley, J. Critical argument and writer identity: Social constructivism as a theoretical framework for efl academic writing. Critical Inquiry in Language Studies, 12(3):184–207, 2015. doi: 10.1080/15427587.2015.1060558.
  38. 38.OpenAI. Our approach to ai safety. https://openai.com/blog/our-approach-to-ai-safety, 2023.
  39. 39.OpenAI. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  40. 40.OpenAssistant. Openassistant/reward-model-deberta-v3-large-v2. https://huggingface.co/OpenAssistant/reward-model-deberta-v3-large-v2, 2023.
  41. 41.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  42. 42.Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, pp. 1–22, 2023.
  43. 43.Parrish, A., Chen, A., Nangia, N., Padmakumar, V., Phang, J., Thompson, J., Htut, P. M., and Bowman, S. Bbq: A hand-built bias benchmark for question answering. In Findings of the Association for Computational Linguistics: ACL 2022, pp. 2086–2105, 2022.
  44. 44.Qian, C., Cong, X., Yang, C., Chen, W., Su, Y., Xu, J., Liu, Z., and Sun, M. Communicative agents for software development. arXiv preprint arXiv:2307.07924, 2023.
  45. 45.Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36, 2023.
  46. 46.Salewski, L., Alaniz, S., Rio-Torto, I., Schulz, E., and Akata, Z. In-context impersonation reveals large language models’ strengths and biases. arXiv preprint arXiv:2305.14930, 2023.
  47. 47.Sun, Z., Shen, Y., Zhou, Q., Zhang, H., Chen, Z., Cox, D. D., Yang, Y., and Gan, C. Principle-driven self-alignment of language models from scratch with minimal human supervision. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=p40XRfBX96.
  48. 48.Sun, Z., Shen, Y., Zhang, H., Zhou, Q., Chen, Z., Cox, D. D., Yang, Y., and Gan, C. Salmon: Self-alignment with principle-following reward models. In The Twelfth International Conference on Learning Representations, 2024.
  49. 49.Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. Stanford alpaca: an instruction-following llama model (2023). URL https://github. com/tatsu-lab/stanford_alpaca, 2023.
  50. 50.TheBloke. Wizard-vicuna-13b-uncensored-hf. https://huggingface.co/TheBloke/Wizard-Vicuna-13B-Uncensored-HF, 2024a.
  51. 51.TheBloke. Wizard-vicuna-7b-uncensored-hf. https://huggingface.co/TheBloke/Wizard-Vicuna-7B-Uncensored-HF, 2024b.
  52. 52.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  53. 53.Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., Lin, Q., and Jiang, D. Wizardlm: Empowering large pre-trained language models to follow complex instructions. In The Twelfth International Conference on Learning Representations, 2024.
  54. 54.Xu, Y., Wang, S., Li, P., Luo, F., Wang, X., Liu, W., and Liu, Y. Exploring large language models for communication games: An empirical study on werewolf. arXiv preprint arXiv:2309.04658, 2023.
  55. 55.Yang, K., Klein, D., Celikyilmaz, A., Peng, N., and Tian, Y. Rlcd: Reinforcement learning from contrastive distillation for lm alignment. In The Twelfth International Conference on Learning Representations, 2024.
  56. 56.Ye, R., Wang, W., Chai, J., Li, D., Li, Z., Xu, Y., Du, Y., Wang, Y., and Chen, S. Openfedllm: Training large language models on decentralized private data via federated learning. arXiv preprint arXiv:2402.06954, 2024.
  57. 57.Yi, J., Ye, R., Chen, Q., Zhu, B. B., Chen, S., Lian, D., Sun, G., Xie, X., and Wu, F. Open-source can be dangerous: On the vulnerability of value alignment in open-source LLMs, 2024. URL https://openreview.net/forum?id=NIouO0C0ex.
  58. 58.Yuan, H., Yuan, Z., Tan, C., Wang, W., Huang, S., and Huang, F. Rrhf: Rank responses to align language models with human feedback. Advances in Neural Information Processing Systems, 36, 2023.
  59. 59.Zhang, J., Xu, X., and Deng, S. Exploring collaboration mechanisms for llm agents: A social psychology view. arXiv preprint arXiv:2310.02124, 2023.
  60. 60.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36, 2023.
  61. 61.Zhou, X., Zhu, H., Mathur, L., Zhang, R., Yu, H., Qi, Z., Morency, L.-P., Bisk, Y., Fried, D., Neubig, G., et al. Sotopia: Interactive evaluation for social intelligence in language agents. In The Twelfth International Conference on Learning Representations, 2023.
  62. 62.Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M. Universal and transferable adversarial attacks on aligned language models, 2023. communication, it is essential for you to comprehend user queries in Cipher Code and subsequently deliver your responses utilizing Cipher Code, 2023.

Citation

MLA
Pang, X., et al. “Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation”. arXiv, 2024, http://arxiv.org/abs/2402.05699v3.
APA
Pang, X., Tang, S., Ye, R., Xiong, Y., Zhang, B., Wang, Y., & Chen, S. (2024). Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation. arXiv. http://arxiv.org/abs/2402.05699v3
Chicago
Pang, X., S. Tang, R. Ye, et al. 2024. “Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation”. arXiv. http://arxiv.org/abs/2402.05699v3.
Harvard
Pang, X. et al. (2024) “Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.05699v3.
Vancouver
1. Pang X, Tang S, Ye R, Xiong Y, Zhang B, Wang Y, Chen S (2024) Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation. arXiv

BibTeX

@article{pang2024self,
  title = {Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation},
  author = {Pang, Xianghe and Tang, Shuo and Ye, Rui and Xiong, Yuxin and Zhang, Bolun and Wang, Yanfeng and Chen, Siheng},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.05699v3},
  eprint = {2402.05699}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/