Guiding Large Language Models via Directional Stimulus Prompting

Zekun LiBaolin PengPengcheng HeMichel GalleyJianfeng GaoXifeng Yan

article2023NeurIPS146 citations

Proposes Directional Stimulus Prompting, a framework that trains a small tunable model to generate instance-specific prompt hints for black-box language models, significantly boosting task performance and reasoning accuracy with minimal labeled data.

Listen

Large language models provide powerful general-purpose natural language capabilities, but organizations frequently struggle to steer them toward precise, application-specific outputs. Direct model fine-tuning is often impossible because leading commercial models are closed black boxes, and modifying open models requires prohibitive computing infrastructure and extensive training data. Existing prompt engineering approaches rely primarily on static, broad task instructions that fail to guide the model on a case-by-case basis.

The article demonstrates a novel framework called Directional Stimulus Prompting to solve this steering challenge. The objective is to evaluate whether a small, tunable auxiliary model can generate dynamic, instance-specific hints that effectively guide frozen, black-box large language models toward desired behaviors across multiple complex language tasks.

To establish this framework, the authors train a compact, accessible policy model (such as a 220-million to 780-million parameter T5 model) using a two-stage approach. First, the small model undergoes supervised fine-tuning on a limited set of annotated examples to learn how to generate relevant hints from an input. Second, it is optimized using reinforcement learning, where rewards are tied directly to the performance and accuracy of the target language model's generated output. This architecture was evaluated across news summarization (CNN/Daily Mail), goal-oriented dialogue (MultiWOZ), and arithmetic chain-of-thought reasoning (MultiArith and AQuA) using prominent black-box models including ChatGPT, Codex, and InstructGPT.

The findings show that instance-specific hints significantly enhance model precision while requiring minimal training data. In goal-oriented dialogue, the framework boosted ChatGPT's overall combined performance score by a relative 41.4% using only 80 training dialogues, matching or exceeding fully supervised benchmark systems trained on thousands of dialogues. In news summarization, generating targeted keywords improved overlap and alignment metrics by 4% to 13% with as few as 1,000 to 4,000 training examples. In mathematical reasoning tasks, instance-specific reasoning triggers improved InstructGPT's zero-shot accuracy up to 84%, outperforming both expert-written prompts and existing automated prompt discovery techniques.

These results demonstrate that organizations can achieve highly reliable, task-aligned AI performance without expensive fine-tuning or massive data-labeling efforts. By shifting the optimization burden from the core language model to a lightweight, easily trainable controller, enterprise teams can significantly lower compute costs, reduce deployment timelines, and improve control over third-party API-based models. Furthermore, reinforcement learning proved essential, as it enabled the small policy model to discover subtle prompting strategies that supervised learning alone could not uncover.

Organizations seeking to deploy language models in specialized workflows should implement lightweight controller models to generate tailored, runtime prompt hints rather than relying exclusively on static, generic prompt templates. Teams should explore automated reinforcement learning optimization against downstream business metrics to continuously refine these hints. However, stakeholders should note that the framework was evaluated on standardized academic benchmarks, and initial supervised stages still depend on pseudo-labeled data or heuristic hint designs. While confidence in the core approach is high, production deployments should begin with targeted pilot testing on domain-specific workflows to determine optimal hint formats and validate performance gains.

Cover for Guiding Large Language Models via Directional Stimulus Prompting

Abstract

We introduce Directional Stimulus Prompting, a novel framework for guiding black-box large language models (LLMs) towards specific desired outputs. Instead of directly adjusting LLMs, our method employs a small tunable policy model (e.g., T5) to generate an auxiliary directional stimulus prompt for each input instance. These directional stimulus prompts act as nuanced, instance-specific hints and clues to guide LLMs in generating desired outcomes, such as including specific keywords in the generated summary. Our approach sidesteps the challenges of direct LLM tuning by optimizing the policy model to explore directional stimulus prompts that align LLMs with desired behaviors. The policy model can be optimized through 1) supervised fine-tuning using labeled data and 2) reinforcement learning from offline or online rewards based on the LLM’s output. We evaluate our method across various tasks, including summarization, dialogue response generation, and chain-of-thought reasoning. Our experiments indicate a consistent improvement in the performance of LLMs such as ChatGPT, Codex, and InstructGPT on these supervised tasks with minimal labeled data. Remarkably, by utilizing merely 80 dialogues from the MultiWOZ dataset, our approach boosts ChatGPT’s performance by a relative 41.4%, achieving or exceeding the performance of some fully supervised state-of-the-art models. Moreover, the instance-specific chain-of-thought prompt generated through our method enhances InstructGPT’s reasoning accuracy, outperforming both generalized human-crafted prompts and those generated through automatic prompt engineering. The code and data are publicly available.

Table of Contents

  • 1 Introduction
  • 2 Directional stimulus prompting
  • 2.1 Supervised fine-tuning
  • 2.2 Reinforcement learning
  • 3 Experiments
  • 3.1 Summarization
  • 3.2 Dialogue response generation
  • 3.3 Chain-of-Thought reasoning
  • 4 Related work
  • 5 Conclusions and future work
  • Acknowledgments and Disclosure of Funding
  • References
  • A Implementation Details
  • A.1 Summarization
  • A.2 Dialogue response generation
  • A.3 Chain of Thought reasoning
  • B Additional results
  • B.1 Summarization
  • B.2 Dialogue response generation
  • B.3 Chain-of-Thought reasoning
  • C Running examples
  • D Prompts

Knowls

  1. Knowl 1 — Directional Stimulus Prompting Framework

    model/method

    Directional Stimulus Prompting (DSP) is a framework designed to steer frozen, black-box large language models (LLMs) toward instance-specific target behaviors without updating LLM parameters.

    Given an input query x∈Xx \in \mathcal{X} drawn from distribution D\mathcal{D} and an output space Y\mathcal{Y}, standard prompting queries the frozen LLM directly as y∼pLLM(⋅∣x)y \sim p_{\text{LLM}}(\cdot \mid x). In contrast, DSP introduces a compact, tunable policy language model pPOL(z∣x)p_{\text{POL}}(z \mid x) (e.g., T5 or Flan-T5) that generates a discrete sequence of tokens zz, termed the directional stimulus (or hint), specific to instance xx.

    The generated stimulus zz is incorporated alongside xx into a structured prompt fed to the frozen black-box LLM:

    y∼pLLM(⋅∣x,z),z∼pPOL(⋅∣x)y \sim p_{\text{LLM}}(\cdot \mid x, z), \quad z \sim p_{\text{POL}}(\cdot \mid x)

    The stimulus zz provides fine-grained guidance—such as key concepts to mention in summarization, communicative dialogue acts in goal-oriented dialogue, or reasoning triggers in multi-step problem solving—enabling precise per-instance control over the generation process while keeping pLLMp_{\text{LLM}} completely frozen.

  2. Knowl 2 — Reinforcement Learning Objective for DSP Policy Models

    equation

    In Directional Stimulus Prompting (DSP), because the parameters of the black-box large language model pLLMp_{\text{LLM}} cannot be updated directly via gradients, the task alignment objective is transferred to optimizing the policy model pPOLp_{\text{POL}}.

    Given a downstream alignment or evaluation metric R(x,y)\mathcal{R}(x, y) on input xx and output yy, the LLM conditioned on stimulus zz is treated as an evaluation function:

    RLLM(x,z)=R(x,y),where y∼pLLM(⋅∣x,z)R_{\text{LLM}}(x, z) = \mathcal{R}(x, y), \quad \text{where } y \sim p_{\text{LLM}}(\cdot \mid x, z)

    The optimization objective for the policy model is:

    max⁡pPOLEx∼D, z∼pPOL(⋅∣x)[RLLM(x,z)]\max_{p_{\text{POL}}} \mathbb{E}_{x \sim \mathcal{D},\, z \sim p_{\text{POL}}(\cdot \mid x)} \left[ R_{\text{LLM}}(x, z) \right]

    To optimize this objective, the token-by-token generation of the stimulus z=(z1,z2,…,zT)z = (z_1, z_2, \dots, z_T) by the policy network π\pi is formulated as a Markov Decision Process ⟨S,A,r,P⟩\langle \mathcal{S}, \mathcal{A}, r, \mathcal{P} \rangle:

    • State space S\mathcal{S}: The input sequence concatenated with generated tokens, (x,z<t)(x, z_{<t}).
    • Action space A\mathcal{A}: Token selection from vocabulary V\mathcal{V}.
    • Reward rr: The evaluation score RLLM(x,z)R_{\text{LLM}}(x, z) penalized for distribution shift from the supervised reference model.
    • Transition P\mathcal{P}: Deterministic token concatenation until the end-of-sequence token is chosen.

    The policy model π\pi is optimized using Proximal Policy Optimization (PPO), specifically Natural Language Policy Optimization (NLPO), with top-pp token masking (p=0.9p = 0.9) to constrain the action space to high-probability tokens.

  3. Knowl 3 — Dynamically Adapted KL-Divergence Regularized Reward in DSP

    equation

    To prevent the policy model π\pi from deviating excessively from the initial supervised fine-tuned model pPOLp_{\text{POL}} during reinforcement learning, the per-instance reward function incorporates a Kullback-Leibler (KL) divergence penalty:

    r(x,z)=RLLM(x,z)−βlog⁡π(z∣x)pPOL(z∣x)r(x, z) = R_{\text{LLM}}(x, z) - \beta \log \frac{\pi(z \mid x)}{p_{\text{POL}}(z \mid x)}

    where RLLM(x,z)R_{\text{LLM}}(x, z) represents the task performance measure achieved by the frozen large language model when conditioned on input xx and stimulus zz, and β>0\beta > 0 is an adaptive regularization coefficient.

    The coefficient β\beta is dynamically adjusted across training iterations tt according to the error ete_t between the observed KL divergence KL(πt,pPOL)\text{KL}(\pi_t, p_{\text{POL}}) and a target divergence KLtarget\text{KL}_{\text{target}}:

    et=clip(KL(πt,pPOL)−KLtargetKLtarget,−0.2,0.2)e_t = \text{clip}\left(\frac{\text{KL}(\pi_t, p_{\text{POL}}) - \text{KL}_{\text{target}}}{\text{KL}_{\text{target}}}, -0.2, 0.2\right)

    βt+1=βt(1+Kβet)\beta_{t+1} = \beta_t (1 + K_\beta e_t)

    where clip(u,−0.2,0.2)\text{clip}(u, -0.2, 0.2) limits the proportional error to [−0.2,0.2][-0.2, 0.2], and KβK_\beta scales the update magnitude.

  4. Knowl 4 — Supervised Fine-Tuning of the DSP Policy Model

    model/method

    Before reinforcement learning optimization, the DSP policy model pPOLp_{\text{POL}} is initialized through supervised fine-tuning (SFT) on a small set of input-stimulus pairs D′={(x,z∗)}\mathcal{D}' = \{(x, z^*)\}. The pseudo-stimulus z∗z^* is constructed per instance depending on the downstream task:

    1. Text Summarization: Pseudo-stimuli are extracted using the TextRank algorithm on both the source article and reference summary, retaining only those keywords present in the reference summary, formatted as a semicolon-separated string: [Keyword1]; [Keyword2]; ...; [KeywordN].
    2. Task-Oriented Dialogue: Pseudo-stimuli are verbalized dialogue acts describing system communicative intent, represented as bracketed domain-intent-slot sequences (e.g., [hotel] [inform] choice type [request] area).
    3. Chain-of-Thought Reasoning: Candidate trigger prompts are collected from human-designed prompt pools by executing zero-shot reasoning over the training set and selecting prompt triggers that yield correct reasoning solutions for each instance.

    The policy model parameters are optimized by minimizing the negative log-likelihood loss:

    LSFT=−E(x,z∗)∼D′[log⁡pPOL(z∗∣x)]\mathcal{L}_{\text{SFT}} = -\mathbb{E}_{(x, z^*) \sim \mathcal{D}'} \left[ \log p_{\text{POL}}(z^* \mid x) \right]

  5. Knowl 5 — Task-Oriented Dialogue Response Generation Performance on MultiWOZ

    data/table

    In task-oriented dialogue response generation on MultiWOZ 2.0 and MultiWOZ 2.1, Directional Stimulus Prompting (DSP) uses a 780M parameter Flan-T5-Large policy model to generate dialogue act hints for frozen LLMs (ChatGPT and Codex). The policy model is trained on low-resource splits of 1% (80 dialogues) or 10% (800 dialogues).

    Evaluation measures include:

    • Inform: Percentage of dialogues where the system provides an entity matching user constraints.
    • Success (Succ.): Percentage of dialogues where all requested attribute information is successfully delivered.
    • BLEU: Corpus-level BLEU against reference responses.
    • Combined Score (Comb.): Defined as (Inform+Success)×0.5+BLEU(\text{Inform} + \text{Success}) \times 0.5 + \text{BLEU}.
    Method MultiWOZ 2.0 MultiWOZ 2.1
    Inform Succ. BLEU Comb. Inform Succ. BLEU Comb.
    Codex
    Standard Prompting 76.7 41.5 7.7 66.8 74.2 41.9 7.8 65.9
    DSP w/ SFT (1% / 80) 74.9 66.3 11.1 81.7 72.0 66.0 11.3 80.1
    DSP w/ SFT+RL (1% / 80) 91.0 76.0 9.8 93.3 89.7 78.6 9.4 93.4
    DSP w/ SFT (10% / 800) 79.4 71.9 11.3 87.0 72.0 67.0 13.1 82.6
    DSP w/ SFT+RL (10% / 800) 96.0 86.9 10.7 102.2 94.0 86.0 9.2 99.2
    ChatGPT
    Standard Prompting 71.8 44.1 10.5 68.4 72.8 44.2 10.4 68.9
    DSP w/ SFT (1% / 80) 76.6 66.5 11.2 82.8 76.0 64.3 11.3 81.4
    DSP w/ SFT+RL (1% / 80) 90.9 82.2 10.2 96.7 87.3 78.7 10.7 93.7
    DSP w/ SFT (10% / 800) 72.7 64.7 11.8 80.5 75.0 67.7 12.6 83.9
    DSP w/ SFT+RL (10% / 800) 95.3 82.3 10.9 99.6 95.0 84.0 10.7 100.2
    Fully Supervised Baselines (100% / 8438 dialogues)
    DAMD 76.3 60.4 16.6 85.0 - - - -
    MinTL 84.9 74.9 17.9 97.8 - - - -
    Soloist 85.5 72.9 16.5 95.7 - - - -
    SimpleTOD 84.4 70.1 15.0 92.3 85.0 70.5 15.2 93.0
    DoTS 86.6 74.1 15.1 95.5 86.7 74.2 15.9 96.3
    PPTOD 89.2 79.4 18.6 102.9 87.1 79.1 19.2 102.3
    UBAR 95.4 80.7 17.0 105.1 95.7 81.8 16.5 105.3
    GALAXY 94.4 85.3 20.5 110.4 95.3 86.2 20.0 110.8

    With 80 training dialogues (1%), DSP w/ SFT+RL boosts ChatGPT's combined score from 68.4 to 96.7 on MultiWOZ 2.0 (+41.4% relative improvement), exceeding the performance of fully supervised models DAMD, MinTL, and Soloist trained on the complete dataset (8,438 dialogues). With 800 dialogues (10%), DSP with ChatGPT achieves a combined score of 99.6–100.2, competitive with fully supervised state-of-the-art architectures.

  6. Knowl 6 — News Summarization Performance with Directional Stimulus Prompting

    empirical result

    On the CNN/Daily Mail news summarization benchmark, Directional Stimulus Prompting (DSP) utilizes a 780M parameter Flan-T5-Large policy model to provide keyword hints to ChatGPT (gpt-3.5-turbo). Evaluated on a 500-sample test split across training subsets of 1,000 (1K), 2,000 (2K), and 4,000 (4K) samples:

    • Standard Prompting: ROUGE-1 = 38.45, ROUGE-2 = 16.20, ROUGE-L = 25.52, BLEU = 8.06, METEOR = 31.59, BERTScore = 0.8807.
    • DSP w/ SFT:
      • 1K samples: ROUGE-1 39.30, ROUGE-2 16.82, ROUGE-L 26.09, BLEU 8.40, METEOR 32.05, BERTScore 0.8826.
      • 2K samples: ROUGE-1 39.32, ROUGE-2 16.88, ROUGE-L 26.20, BLEU 8.70, METEOR 32.22, BERTScore 0.8831.
      • 4K samples: ROUGE-1 39.47, ROUGE-2 17.08, ROUGE-L 26.25, BLEU 8.71, METEOR 32.24, BERTScore 0.8832.
    • DSP w/ SFT+RL:
      • 1K samples: ROUGE-1 39.99, ROUGE-2 17.45, ROUGE-L 26.69, BLEU 8.99, METEOR 32.59, BERTScore 0.8843.
      • 2K samples: ROUGE-1 40.06, ROUGE-2 17.50, ROUGE-L 26.70, BLEU 9.07, METEOR 32.70, BERTScore 0.8841.
      • 4K samples: ROUGE-1 40.19, ROUGE-2 17.70, ROUGE-L 26.84, BLEU 9.10, METEOR 32.72, BERTScore 0.8845.

    In pairwise GPT-4 evaluation over 500 test instances judging key-point alignment with reference summaries, DSP-guided summaries won in 51.0% (255 cases), standard prompting summaries won in 44.4% (222 cases), and 4.6% (23 cases) were ties.

  7. Knowl 7 — Zero-Shot Chain-of-Thought Reasoning with Instance-Specific Trigger Prompts

    data/table

    In zero-shot Chain-of-Thought (CoT) reasoning, Directional Stimulus Prompting (DSP) trains a 220M parameter T5-Base policy model to generate instance-specific trigger prompts tailored to each input problem, rather than relying on a single static trigger prompt across all instances. Evaluation on InstructGPT (text-davinci-002) across the MultiArith and AQuA arithmetic reasoning benchmarks yields the following accuracy results:

    Category Chain-of-Thought Trigger Prompt MultiArith (%) AQuA (%)
    Human-Designed Let’s think step by step. 79.6 31.9
    Human-Designed We should think about this step by step. 81.2 28.7
    Human-Designed First, 78.0 38.2
    Human-Designed Before we dive into the answer, 54.8 27.2
    Human-Designed Proof followed by the answer. 58.4 37.8
    Human-Designed Let’s think step by step in a realistic way. 59.6 33.9
    Human-Designed Let’s think step by step using common sense and knowledge. 80.0 34.3
    Human-Designed Let’s think like a detective step by step. 73.6 24.0
    Human-Designed Let’s think about this logically. 75.2 34.7
    Human-Designed Let’s think step by step. First, 78.8 32.3
    Human-Designed Let’s think 56.8 38.2
    Human-Designed Let’s solve this problem by splitting it into steps. 72.4 33.2
    Human-Designed The answer is after the proof. 42.8 34.3
    Human-Designed Let’s be realistic and think step by step. 69.6 29.9
    APE (Zhou et al., 2022) Let’s work this out in a step by step way to be sure we have the right answer. 81.6 34.3
    DSP w/ SFT Generated instance-specific prompt 75.2 35.8
    DSP w/ SFT+RL Generated instance-specific prompt 84.0 38.6

    The instance-specific prompts discovered by DSP w/ SFT+RL (such as "Let's think like a detective step by step. First," and "Let's think step by step using both the above information and the testing."), achieve 84.0% accuracy on MultiArith and 38.6% on AQuA, outperforming all 14 fixed human-crafted prompts and the prompt discovered by Automatic Prompt Engineer (APE).

  8. Knowl 8 — Linguistic Characteristics and Training Dynamics of Keyword Stimuli

    empirical result

    Analysis of keyword stimulus generation during DSP training on the CNN/Daily Mail summarization task reveals the following properties:

    1. Part-of-Speech (POS) Distribution: POS tagging with spaCy shows that proper nouns (PROPN, 37.64%) and nouns (NOUN, 34.47%) account for over 72% of all generated keywords. Adjectives (ADJ, 7.28%) and numerals (NUM, 7.45%) represent the majority of remaining tokens.
    2. Named Entity Recognition (NER) Distribution: NER tagging indicates that the most frequent entity classes in generated keywords are person names (PERSON, 29.67%), geopolitical entities (GPE, 14.79%), dates (DATE, 14.39%), organizations (ORG, 13.87%), and cardinal numerals (CARDINAL, 12.60%).
    3. Precision vs. Quantity Dynamics: Across reinforcement learning training on 4,000 samples, keyword precision (the proportion of generated keywords appearing in the reference summary) increases from 0.35 to 0.48, which directly correlates with the summary ROUGE-1 score rising from 39.3 to 40.0. However, when the total quantity of generated keywords is overly restricted, high precision alone does not lead to performance improvements.
  9. Knowl 9 — Robustness of Directional Stimulus Prompting Under Zero-Shot Evaluation

    empirical result

    When DSP policy models trained on 4,000 CNN/Daily Mail samples are evaluated under zero-shot prompting settings (i.e., presenting zero in-context demonstration examples to the black-box LLM during evaluation), performance gains over standard prompting are preserved:

    • Standard Zero-Shot Prompting: ROUGE-1 = 38.34, ROUGE-2 = 14.13, ROUGE-L = 24.43, BLEU = 4.86, METEOR = 29.85, BERTScore = 0.8798.
    • DSP (0-shot training, 0-shot evaluation): ROUGE-1 = 38.73, ROUGE-2 = 15.14, ROUGE-L = 25.26, BLEU = 5.22, METEOR = 31.21, BERTScore = 0.8820.
    • DSP (3-shot training, 0-shot evaluation): ROUGE-1 = 38.47, ROUGE-2 = 15.20, ROUGE-L = 24.82, BLEU = 5.17, METEOR = 31.29, BERTScore = 0.8815.

    These results indicate that DSP is robust to mismatches between the number of demonstration examples used during RL policy optimization and those used during final inference.

  10. Knowl 10 — Limitations of Directional Stimulus Prompting

    limitation

    Directional Stimulus Prompting (DSP) has several identified limitations:

    1. Dependence on Pseudo-Stimulus Heuristics: The supervised fine-tuning initialization relies on heuristically selected or annotated pseudo-stimuli (such as TextRank-extracted keywords or annotated dialogue acts), which may not represent the optimal directional stimulus for steering LLMs.
    2. Natural Language Interface Constraints: Restricting the stimulus to human-interpretable natural language tokens limits the communication bandwidth between the policy model and the LLM compared to potential learned machine-to-machine intermediate representations.
    3. Modality Limitation: The current DSP formulation is restricted strictly to textual prompt hints and does not accommodate continuous soft prompt embeddings or multimodal stimuli.

Coverage note — No substantial contributed material was omitted; all key methodology components (SFT, RL/NLPO, dynamic KL), empirical evaluations (summarization, dialogue, CoT reasoning), stimulus analyses, and stated limitations are covered.

References

  1. 1.Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., et al. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691, 2022.
  2. 2.An, S., Li, Y., Lin, Z., Liu, Q., Chen, B., Fu, Q., Chen, W., Zheng, N., and Lou, J.-G. Input-Tuning: Adapting unfamiliar inputs to frozen pretrained models. arXiv preprint arXiv:2203.03131, 2022.
  3. 3.Banerjee, S. and Lavie, A. METEOR: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pp. 65–72, 2005.
  4. 4.Bang, Y., Cahyawijaya, S., Lee, N., Dai, W., Su, D., Wilie, B., Lovenia, H., Ji, Z., Yu, T., Chung, W., et al. A multitask, multilingual, multimodal evaluation of ChatGPT on reasoning, hallucination, and interactivity. arXiv preprint arXiv:2302.04023, 2023.
  5. 5.Barrios, F., López, F., Argerich, L., and Wachenchauzer, R. Variations of the similarity function of textrank for automated summarization. arXiv preprint arXiv:1602.03606, 2016.
  6. 6.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  7. 7.Budzianowski, P., Wen, T.-H., Tseng, B.-H., Casanueva, I., Ultes, S., Ramadan, O., and Gašić, M. MultiWOZ – a large-scale multi-domain Wizard-of-Oz dataset for task-oriented dialogue modelling. arXiv preprint arXiv:1810.00278, 2018.
  8. 8.Card, D., Henderson, P., Khandelwal, U., Jia, R., Mahowald, K., and Jurafsky, D. With little power comes great responsibility. arXiv preprint arXiv:2010.06595, 2020.
  9. 9.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  10. 10.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. PaLM: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  11. 11.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022.
  12. 12.Dathathri, S., Madotto, A., Lan, J., Hung, J., Frank, E., Molino, P., Yosinski, J., and Liu, R. Plug and play language models: A simple approach to controlled text generation. arXiv preprint arXiv:1912.02164, 2019.
  13. 13.Deng, M., Wang, J., Hsieh, C.-P., Wang, Y., Guo, H., Shu, T., Song, M., Xing, E. P., and Hu, Z. RLPrompt: Optimizing discrete text prompts with reinforcement learning. arXiv preprint arXiv:2205.12548, 2022.
  14. 14.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  15. 15.Eric, M., Goel, R., Paul, S., Kumar, A., Sethi, A., Ku, P., Goyal, A. K., Agarwal, S., Gao, S., and Hakkani-Tur, D. MultiWOZ 2.1: A consolidated multi-domain dialogue dataset with state corrections and state tracking baselines. arXiv preprint arXiv:1907.01669, 2019.
  16. 16.Goyal, T., Li, J. J., and Durrett, G. News summarization and evaluation in the era of GPT-3. arXiv preprint arXiv:2209.12356, 2022.
  17. 17.Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., and Smith, N. A. Don’t stop pretraining: Adapt language models to domains and tasks. arXiv preprint arXiv:2004.10964, 2020.
  18. 18.Gutiérrez, B. J., McNeal, N., Washington, C., Chen, Y., Li, L., Sun, H., and Su, Y. Thinking about GPT-3 in-context learning for biomedical IE? Think again. arXiv preprint arXiv:2203.08410, 2022.
  19. 19.He, W., Dai, Y., Zheng, Y., Wu, Y., Cao, Z., Liu, D., Jiang, P., Yang, M., Huang, F., Si, L., et al. Galaxy: A generative pre-trained model for task-oriented dialog with semi-supervised learning and explicit policy injection. In Proceedings of the AAAI Conference on Artificial Intelligence, pp. 10749–10757, 2022.
  20. 20.Honnibal, M. and Montani, I. spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing. To appear, 2017.
  21. 21.Hosseini-Asl, E., McCann, B., Wu, C.-S., Yavuz, S., and Socher, R. A simple language model for task-oriented dialogue. Advances in Neural Information Processing Systems, 33: 20179–20191, 2020.
  22. 22.Hudeček, V. and Dušek, O. Are LLMs all you need for task-oriented dialogue? arXiv preprint arXiv:2304.06556, 2023.
  23. 23.Jeon, H. and Lee, G. G. Domain state tracking for a simplified dialogue system. arXiv preprint arXiv:2103.06648, 2021.
  24. 24.Keskar, N. S., McCann, B., Varshney, L. R., Xiong, C., and Socher, R. CTRL: A conditional transformer language model for controllable generation. arXiv preprint arXiv:1909.05858, 2019.
  25. 25.Khattab, O., Santhanam, K., Li, X. L., Hall, D., Liang, P., Potts, C., and Zaharia, M. Demonstrate-search-predict: Composing retrieval and language models for knowledge-intensive nlp. arXiv preprint arXiv:2212.14024, 2022.
  26. 26.Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213, 2022.
  27. 27.Krause, B., Gotmare, A. D., McCann, B., Keskar, N. S., Joty, S., Socher, R., and Rajani, N. F. GeDi: Generative discriminator guided sequence generation. arXiv preprint arXiv:2009.06367, 2020.
  28. 28.Kumar, G., Foster, G., Cherry, C., and Krikun, M. Reinforcement learning based curriculum optimization for neural machine translation. arXiv preprint arXiv:1903.00041, 2019.
  29. 29.Le, M. and Fokkens, A. Tackling error propagation through reinforcement learning: A case of greedy dependency parsing. arXiv preprint arXiv:1702.06794, 2017.
  30. 30.Lester, B., Al-Rfou, R., and Constant, N. The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691, 2021.
  31. 31.Li, J., Monroe, W., Ritter, A., Galley, M., Gao, J., and Jurafsky, D. Deep reinforcement learning for dialogue generation. arXiv preprint arXiv:1606.01541, 2016.
  32. 32.Li, X. L. and Liang, P. Prefix-tuning: Optimizing continuous prompts for generation. arXiv preprint arXiv:2101.00190, 2021.
  33. 33.Lin, C.-Y. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pp. 74–81, 2004.
  34. 34.Lin, Z., Madotto, A., Winata, G. I., and Fung, P. Mintl: Minimalist transfer learning for task-oriented dialogue systems. arXiv preprint arXiv:2009.12005, 2020.
  35. 35.Ling, W., Yogatama, D., Dyer, C., and Blunsom, P. Program induction by rationale generation: Learning to solve and explain algebraic word problems. arXiv preprint arXiv:1705.04146, 2017.
  36. 36.Liu, A., Sap, M., Lu, X., Swayamdipta, S., Bhagavatula, C., Smith, N. A., and Choi, Y. Dexperts: Decoding-time controlled text generation with experts and anti-experts. arXiv preprint arXiv:2105.03023, 2021.
  37. 37.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  38. 38.Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., and Zhu, C. Gpteval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634, 2023.
  39. 39.Lu, P., Qiu, L., Chang, K.-W., Wu, Y. N., Zhu, S.-C., Rajpurohit, T., Clark, P., and Kalyan, A. Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning. arXiv preprint arXiv:2209.14610, 2022.
  40. 40.Lu, X., Welleck, S., Jiang, L., Hessel, J., Qin, L., West, P., Ammanabrolu, P., and Choi, Y. Quark: Controllable text generation with reinforced unlearning. arXiv preprint arXiv:2205.13636, 2022.
  41. 41.Mihalcea, R. and Tarau, P. Textrank: Bringing order into text. In Proceedings of the 2004 conference on empirical methods in natural language processing, pp. 404–411, 2004.
  42. 42.Moradi, M., Blagec, K., Haberl, F., and Samwald, M. Gpt-3 models are poor few-shot learners in the biomedical domain. arXiv preprint arXiv:2109.02555, 2021.
  43. 43.Nallapati, R., Zhou, B., Gulcehre, C., Xiang, B., et al. Abstractive text summarization using sequence-to-sequence rnns and beyond. arXiv preprint arXiv:1602.06023, 2016.
  44. 44.Neu, G. and Szepesvári, C. Training parsers by inverse reinforcement learning. Machine Learning. 2009 Dec; 77: 303-37., 2009.
  45. 45.OpenAI. Gpt-4 technical report, 2023.
  46. 46.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  47. 47.Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp. 311–318, 2002.
  48. 48.Paulus, R., Xiong, C., and Socher, R. A deep reinforced model for abstractive summarization. arXiv preprint arXiv:1705.04304, 2017.
  49. 49.Peng, B., Li, C., Li, J., Shayandeh, S., Liden, L., and Gao, J. SOLOIST: Building task bots at scale with transfer learning and machine teaching. Transactions of the Association for Computational Linguistics, 9:807–824, 2021.
  50. 50.Petroni, F., Rocktäschel, T., Lewis, P., Bakhtin, A., Wu, Y., Miller, A. H., and Riedel, S. Language models as knowledge bases? arXiv preprint arXiv:1909.01066, 2019.
  51. 51.Post, M. A call for clarity in reporting bleu scores. arXiv preprint arXiv:1804.08771, 2018.
  52. 52.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  53. 53.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  54. 54.Ramamurthy, R., Ammanabrolu, P., Brantley, K., Hessel, J., Sifa, R., Bauckhage, C., Hajishirzi, H., and Choi, Y. Is reinforcement learning (not) for natural language processing?: Benchmarks, baselines, and building blocks for natural language policy optimization. arXiv preprint arXiv:2210.01241, 2022.
  55. 55.Reynolds, L. and McDonell, K. Prompt programming for large language models: Beyond the few-shot paradigm. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems, pp. 1–7, 2021.
  56. 56.Reynolds, L. and McDonell, K. Prompt programming for large language models: Beyond the few-shot paradigm. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems, pp. 1–7, 2021.
  57. 57.Roy, S. and Roth, D. Solving general arithmetic word problems. arXiv preprint arXiv:1608.01413, 2016.
  58. 58.Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilic, S., Hesslow, D., Castagné, R., Luccioni, A. S., Yvon, F., Gallé, M., et al. BLOOM: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022.
  59. 59.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  60. 60.Shi, W., Min, S., Yasunaga, M., Seo, M., James, R., Lewis, M., Zettlemoyer, L., and Yih, W.-t. Replug: Retrieval-augmented black-box language models. arXiv preprint arXiv:2301.12652, 2023.
  61. 61.Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. arXiv preprint arXiv:2010.15980, 2020.
  62. 62.Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021, 2020.
  63. 63.Su, Y., Shu, L., Mansimov, E., Gupta, A., Cai, D., Lai, Y.-A., and Zhang, Y. Multi-task pre-training for plug-and-play task-oriented dialogue system. arXiv preprint arXiv:2109.14739, 2021.
  64. 64.Sun, T., Shao, Y., Qian, H., Huang, X., and Qiu, X. Black-box tuning for language-model-as-a-service. In International Conference on Machine Learning, pp. 20841–20855. PMLR, 2022.
  65. 65.Suzgun, M., Melas-Kyriazi, L., and Jurafsky, D. Follow the wisdom of the crowd: Effective text generation via minimum bayes risk decoding. arXiv preprint arXiv:2211.07634, 2022.
  66. 66.Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al. LaMDA: Language models for dialog applications. arXiv preprint arXiv:2201.08239, 2022.
  67. 67.Vu, T., Lester, B., Constant, N., Al-Rfou, R., and Cer, D. Spot: Better frozen model adaptation through soft prompt transfer. arXiv preprint arXiv:2110.07904, 2021.
  68. 68.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022.
  69. 69.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E., Le, Q., and Zhou, D. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903, 2022.
  70. 70.Wu, J., Ouyang, L., Ziegler, D. M., Stiennon, N., Lowe, R., Leike, J., and Christiano, P. Recursively summarizing books with human feedback. arXiv preprint arXiv:2109.10862, 2021.
  71. 71.Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al. Google’s neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144, 2016.
  72. 72.Yang, Y., Li, Y., and Quan, X. UBAR: Towards fully end-to-end task-oriented dialog system with GPT-2. In Proceedings of the AAAI Conference on Artificial Intelligence, pp. 14230–14238, 2021.
  73. 73.Zhang, S., Diab, M., and Zettlemoyer, L. Democratizing access to large-scale language models with opt-175b. Meta AI, 2022.
  74. 74.Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675, 2019.
  75. 75.Zhang, T., Ladhak, F., Durmus, E., Liang, P., McKeown, K., and Hashimoto, T. B. Benchmarking large language models for news summarization. arXiv preprint arXiv:2301.13848, 2023.
  76. 76.Zhang, T., Wang, X., Zhou, D., Schuurmans, D., and Gonzalez, J. E. TEMPERA: Test-time prompt editing via reinforcement learning. In International Conference on Learning Representations, 2023.
  77. 77.Zhang, Y., Ou, Z., and Yu, Z. Task-oriented dialog systems that consider multiple appropriate responses under the same context. In Proceedings of the AAAI Conference on Artificial Intelligence, pp. 9604–9611, 2020.
  78. 78.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023.
  79. 79.Zhou, Y., Muresanu, A. I., Han, Z., Paster, K., Pitis, S., Chan, H., and Ba, J. Large language models are human-level prompt engineers. arXiv preprint arXiv:2211.01910, 2022.
  80. 80.Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593, 2019.

Citation

MLA
Li, Z., et al. “Guiding Large Language Models via Directional Stimulus Prompting”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 62630–56, https://proceedings.neurips.cc/paper_files/paper/2023/file/c5601d99ed028448f29d1dae2e4a926d-Paper-Conference.pdf.
APA
Li, Z., Peng, B., He, P., Galley, M., Gao, J., & Yan, X. (2023). Guiding Large Language Models via Directional Stimulus Prompting. Advances in Neural Information Processing Systems, 36, 62630–62656. https://proceedings.neurips.cc/paper_files/paper/2023/file/c5601d99ed028448f29d1dae2e4a926d-Paper-Conference.pdf
Chicago
Li, Z., B. Peng, P. He, M. Galley, J. Gao, and X. Yan. 2023. “Guiding Large Language Models via Directional Stimulus Prompting”. Advances in Neural Information Processing Systems 36: 62630–56. https://proceedings.neurips.cc/paper_files/paper/2023/file/c5601d99ed028448f29d1dae2e4a926d-Paper-Conference.pdf.
Harvard
Li, Z. et al. (2023) “Guiding Large Language Models via Directional Stimulus Prompting”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 62630–62656. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/c5601d99ed028448f29d1dae2e4a926d-Paper-Conference.pdf.
Vancouver
1. Li Z, Peng B, He P, Galley M, Gao J, Yan X (2023) Guiding Large Language Models via Directional Stimulus Prompting. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 62630–62656

BibTeX

@inproceedings{li2023guiding,
  title = {Guiding Large Language Models via Directional Stimulus Prompting},
  author = {Li, Zekun and Peng, Baolin and He, Pengcheng and Galley, Michel and Gao, Jianfeng and Yan, Xifeng},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {62630-62656},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/c5601d99ed028448f29d1dae2e4a926d-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission