InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction Tuning
Prakhar GuptaCathy JiaoYi-Ting YehShikib MehriMaxine EskénaziJeffrey P. Bigham
Introduces an instruction-tuning framework unifying 48 dialogue tasks across 59 datasets, establishing benchmarks that boost zero- and few-shot cross-task generalization while incorporating novel meta-tasks to ensure instruction adherence.
Building conversational artificial intelligence systems currently requires extensive, specialized datasets and costly fine-tuning for each distinct task. While instruction tuning—training models to follow natural language task descriptions—has shown strong zero-shot generalization across general language processing, its application to dialogue domains has remained unstandardized and fragmented. Dialogue systems must manage diverse capabilities ranging from intent detection and state tracking to open-domain conversation and safety moderation, making unified adaptation difficult.
The article introduces and evaluates InstructDial, an open-source instruction-tuning framework designed to standardize dialogue data and induce strong zero-shot and few-shot performance across unseen conversational tasks. InstructDial unifies 59 publicly available dialogue datasets into a repository of 48 distinct tasks using a consistent text-to-sequence format. To ensure models genuinely follow user instructions rather than relying on dataset artifacts, the framework introduces specialized meta-tasks that train models to map instructions directly to input-output pairs.
The evaluation evaluated two instruction-tuned architectures: DIAL-BART0 (406 million parameters) and DIAL-T0 (3 billion parameters), benchmarked against existing state-of-the-art baselines and GPT-3. The experiments yielded several major findings. First, instruction tuning on InstructDial boosted performance on unseen dialogue tasks, with DIAL-BART0 outperforming baseline models by approximately threefold on evaluation selection, relation classification, and initial phrase generation. Second, DIAL-T0 achieved an average Spearman correlation of 0.465 with human relevance judgments across 13 benchmark sets, outperforming specialized automatic evaluation metrics without task-specific training. Third, incorporating only 100 task instances in a few-shot setup yielded 12% to 50% relative performance gains across difficult tasks, and DIAL-BART0 achieved a 36.9-point F1 improvement over prior models in zero-shot slot filling on Restaurant8k. Finally, the analysis showed that smaller, efficiently tuned models can match or outperform larger models, as DIAL-BART0 outperformed the much larger PPTOD on few-shot intent prediction (84.30% accuracy) and remained competitive in dialogue state tracking.
These findings indicate that dialogue systems can achieve broad generalization without deploying massive, prohibitively expensive language models. Standardizing dialogue tasks through natural language instructions enables development teams to rapidly prototype, diagnose, and deploy new conversational functionalities at lower compute and operational costs. The results demonstrate that pre-training on general language instructions directly complements dialogue-specific instruction tuning, significantly lowering the barrier to multi-domain deployment.
Organizations developing conversational AI should adopt standardized instruction schemas and leverage targeted few-shot tuning (around 100 examples) when deploying systems to novel domains. Technical teams should also incorporate instruction-matching meta-tasks during training to mitigate task confusion and enforce instruction adherence. Future research must address negative task interference—where training on certain dialogue tasks degrades performance on others—and explore methods to make systems less sensitive to instruction phrasing.
Confidence in these findings is high regarding core task-oriented capabilities and automatic dialogue evaluation. However, stakeholders should exercise caution regarding zero-shot reliability on complex edge tasks like utterance infilling, sensitivity to instruction wording variations, and potential task interference when scaling task repositories.
- Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). Read FLAN first to understand the instruction-tuning approach InstructDial adapts from general NLP to dialogue tasks.
- Paper: Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System, Yixuan Su et al. (2022). PPTOD establishes the dialogue-specific, prompt-driven multi-task baseline that InstructDial compares against and seeks to improve.
- Paper: BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension, Mike Lewis et al. (2020). BART’s sequence-to-sequence pretraining provides essential architectural context for InstructDial’s DIAL-BART0 model.
No sufficiently relevant recommendations were found.
