Multi-Task Deep Neural Networks for Natural Language Understanding

Xiaodong LiuPengcheng HeWeizhu ChenJianfeng Gao

article2019ACL1,364 citationsBest Paper Award

Introduces MT-DNN, a framework that integrates multi-task learning with pre-trained BERT representations to achieve state-of-the-art results across ten natural language understanding benchmarks while requiring significantly fewer labeled examples for domain adaptation.

Listen

Developing accurate natural language understanding systems often requires large volumes of task-specific labeled data, which can be prohibitively expensive and time-consuming to collect. While pre-trained language models have improved text representation learning using unlabeled text, relying on individual task fine-tuning still exposes models to overfitting when labeled examples are scarce. Combining supervised multi-task learning across related domains with unsupervised language pre-training offers an effective strategy to build more robust, generalizable text representations.

The article demonstrates that combining pre-trained bidirectional transformer encoders with multi-task deep neural networks improves cross-task generalization and allows rapid adaptation to new domains. It evaluates this unified architecture against existing benchmarks across a diverse set of language understanding tasks.

The authors implemented the Multi-Task Deep Neural Network by initializing its shared text-encoding layers with a pre-trained language model, BERT, and adding task-specific output layers for single-sentence classification, pairwise text classification, semantic text similarity, and relevance ranking. The shared representations were jointly trained on multiple supervised datasets using mini-batch gradient descent and subsequently fine-tuned on individual target tasks. The framework was evaluated on the nine-task General Language Understanding Evaluation benchmark, as well as the Stanford Natural Language Inference and SciTail datasets, using varying proportions of training data to test domain adaptation efficiency.

The framework established new state-of-the-art performance across ten natural language understanding tasks, raising the overall General Language Understanding Evaluation benchmark score to 82.7%, an absolute improvement of 2.2% over large-scale BERT. Performance gains were most pronounced on tasks with limited labeled data, such as textual entailment and paraphrase detection, and the model even outperformed prior baselines without task-specific fine-tuning on most tasks. In domain adaptation experiments, the model demonstrated exceptional data efficiency: when trained on only 0.1% of available target data, it reached 82.1% accuracy on the Stanford dataset compared to 52.5% for BERT, and 81.9% compared to 51.2% on SciTail. Using full datasets, the architecture achieved benchmark scores of 91.6% on the Stanford dataset (a 1.5% improvement) and 95.0% on SciTail (a 6.7% improvement).

These findings indicate that multi-task learning provides a powerful regularizing effect, preventing models from overfitting to single tasks and generating text representations that transfer effectively across domains. For organizational leaders and technical teams, this approach significantly reduces the time, risk, and cost associated with acquiring large specialized datasets when deploying language models into new domains or applications.

Organizations should adopt joint multi-task training strategies to optimize text analysis pipelines, particularly when operating under tight data collection budgets. Engineering teams should also integrate specialized output modules, such as pairwise ranking loss for question-answering tasks, to maximize accuracy. Future work supported by the article includes evaluating model resilience against adversarial inputs, exploring methods to explicitly encode linguistic structure, and developing training techniques that leverage task relatedness.

The results carry high confidence across standard natural language benchmarks, though caution is warranted when deploying the model directly without fine-tuning on highly idiosyncratic datasets with small sample sizes, where cross-task learning may occasionally underfit without targeted adaptation.

arXiv: 1901.11504
Cover for Multi-Task Deep Neural Networks for Natural Language Understanding

Abstract

In this paper, we present a Multi-Task Deep Neural Network (MT-DNN) for learning representations across multiple natural language understanding (NLU) tasks. MT-DNN not only leverages large amounts of cross-task data, but also benefits from a regularization effect that leads to more general representations in order to adapt to new tasks and domains. MT-DNN extends the model proposed in Liu et al. (2015) by incorporating a pre-trained bidirectional transformer language model, known as BERT (Devlin et al., 2018). MT-DNN obtains new state-of-the-art results on ten NLU tasks, including SNLI, SciTail, and eight out of nine GLUE tasks, pushing the GLUE benchmark to 82.7% (2.2% absolute improvement). We also demonstrate using the SNLI and SciTail datasets that the representations learned by MT-DNN allow domain adaptation with substantially fewer in-domain labels than the pre-trained BERT representations. The code and pre-trained models are publicly available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Tasks
  • 3 The Proposed MT-DNN Model
  • 3.1 The Training Procedure
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Implementation details
  • 4.3 GLUE Main Results
  • 4.4 Domain Adaptation Results on SNLI and SciTail
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Multi-Task Deep Neural Network (MT-DNN) Architecture

    model/method

    The Multi-Task Deep Neural Network (MT-DNN) combines multi-task learning with a pre-trained bidirectional Transformer language model (such as BERT) to learn universal text representations across diverse natural language understanding (NLU) tasks.

    The model comprises shared bottom layers and task-specific top layers:

    1. Lexicon Encoder (l1l_1): Given an input sequence of tokens X=(x1,…,xm)X = (x_1, \ldots, x_m) (where x1x_1 is the special [CLS] token and sentence pairs (X1,X2)(X_1, X_2) are concatenated with a separating [SEP] token), each token is mapped to an input embedding constructed by summing its word embedding, segment embedding, and positional embedding.

    2. Transformer Encoder (l2l_2): A multi-layer bidirectional Transformer encoder maps the sequence of input representations into a sequence of contextual embedding vectors C∈Rd×mC \in \mathbb{R}^{d \times m}, where dd is the hidden representation dimension. Contextual representations in l2l_2 are shared across all tasks and jointly updated by multi-task objectives.

    3. Task-Specific Output Layers:

    • Single-Sentence Classification (e.g., SST-2, CoLA): The contextual embedding x∈Rdx \in \mathbb{R}^d corresponding to [CLS] represents the sentence. Class probability is computed via a linear projection with softmax: Pr⁡(c∣X)=softmax⁡(W⊤x)\operatorname{Pr}(c \mid X) = \operatorname{softmax}(W^\top x) where W∈Rd×CW \in \mathbb{R}^{d \times C} is a task-specific parameter matrix and CC is the number of classes.

    • Text Similarity Scoring (e.g., STS-B): For a sentence pair (X1,X2)(X_1, X_2), the contextual embedding xx of [CLS] is mapped to a real-valued similarity score: Sim⁡(X1,X2)=wSTS⊤x\operatorname{Sim}(X_1, X_2) = w_{\text{STS}}^\top x where wSTS∈Rdw_{\text{STS}} \in \mathbb{R}^d is a task-specific parameter vector.

    • Pairwise Text Classification (e.g., MNLI, RTE, MRPC, QQP): Implemented via a Stochastic Answer Network (SAN) module performing multi-step reasoning across contextual premise and hypothesis memories.

    • Relevance Ranking (e.g., QNLI): Formulated as ranking candidate answers AA for a query QQ by computing a relevance score Rel⁡(Q,A)=g(wQNLI⊤x)\operatorname{Rel}(Q, A) = g(w_{\text{QNLI}}^\top x) on the [CLS] embedding xx.

  2. Knowl 2 — Stochastic Answer Network (SAN) Output Module for Pairwise Text Classification

    model/method

    For pairwise text classification and Natural Language Inference (NLI), MT-DNN incorporates the answer module of the Stochastic Answer Network (SAN) to perform multi-step iterative reasoning instead of a direct single-step classification.

    Given a premise P=(p1,…,pm)P = (p_1, \ldots, p_m) and a hypothesis H=(h1,…,hn)H = (h_1, \ldots, h_n), the Transformer encoder yields premise working memory Mp∈Rd×mM^p \in \mathbb{R}^{d \times m} and hypothesis working memory Mh∈Rd×nM^h \in \mathbb{R}^{d \times n}. The SAN module carries out KK-step reasoning:

    1. Memory Initialization: The initial state s0∈Rds^0 \in \mathbb{R}^d is an attention-weighted summary of hypothesis memory MhM^h: s0=∑j=1nαjMjh,αj=exp⁡(w1⊤Mjh)∑i=1nexp⁡(w1⊤Mih)s^0 = \sum_{j=1}^n \alpha_j M_j^h, \quad \alpha_j = \frac{\exp(w_1^\top M_j^h)}{\sum_{i=1}^n \exp(w_1^\top M_i^h)} where w1∈Rdw_1 \in \mathbb{R}^d is a learnable parameter vector.

    2. Iterative Memory State Update: At reasoning step k∈{1,…,K−1}k \in \{1, \ldots, K-1\}, state sks^k is updated using a Gated Recurrent Unit (GRU): sk=GRU⁡(sk−1,xk)s^k = \operatorname{GRU}(s^{k-1}, x^k) where xk=∑j=1mβjMjpx^k = \sum_{j=1}^m \beta_j M_j^p is the attention-weighted context from premise memory MpM^p, with attention weights: β=softmax⁡(sk−1W2⊤Mp)\beta = \operatorname{softmax}(s^{k-1} W_2^\top M^p) where W2∈Rd×dW_2 \in \mathbb{R}^{d \times d} is a learnable parameter matrix.

    3. Step-Level Classification: At each step k∈{0,…,K−1}k \in \{0, \ldots, K-1\}, a classifier predicts relation distribution PrkP_r^k over relations R\mathcal{R}: Prk=softmax⁡(W3⊤[sk;  xk;  ∣sk−xk∣;  sk⊙xk])P_r^k = \operatorname{softmax}\left(W_3^\top \left[s^k;\; x^k;\; |s^k - x^k|;\; s^k \odot x^k\right]\right) where [⋅;⋅][\cdot; \cdot] indicates concatenation, ⊙\odot is element-wise multiplication, and W3W_3 is a learnable weight matrix.

    4. Aggregation: Final predictions average the distribution across all KK steps: Pr=1K∑k=0K−1PrkP_r = \frac{1}{K} \sum_{k=0}^{K-1} P_r^k During training, stochastic prediction dropout is applied before averaging; during decoding/inference, all KK step outputs are averaged.

  3. Knowl 3 — Pairwise Ranking Formulation and Loss for QNLI

    model/method

    In the GLUE benchmark, Question-answering NLI (QNLI, derived from SQuAD) is conventionally framed as a binary classification task to determine whether a candidate sentence contains the answer to a question. MT-DNN reformulates QNLI as a pairwise relevance ranking task.

    Given a query QQ and candidate answers A\mathcal{A} containing one positive answer A+A^+ (which contains the correct answer) and ∣A∣−1|\mathcal{A}| - 1 negative candidate answers, MT-DNN computes a continuous relevance score: Rel⁡(Q,A)=g(wQNLI⊤x)\operatorname{Rel}(Q, A) = g(w_{\text{QNLI}}^\top x) where x∈Rdx \in \mathbb{R}^d is the contextual embedding of the [CLS] token for the input pair (Q,A)(Q, A), wQNLI∈Rdw_{\text{QNLI}} \in \mathbb{R}^d is a learnable parameter vector, and g(⋅)g(\cdot) is an activation function.

    The model is trained by minimizing the negative log-likelihood of the positive candidate A+A^+ among all candidates for query QQ: Lrank(Θ)=−∑(Q,A+)log⁡Pr⁡(A+∣Q)\mathcal{L}_{\text{rank}}(\Theta) = - \sum_{(Q, A^+)} \log \operatorname{Pr}(A^+ \mid Q) Pr⁡(A+∣Q)=exp⁡(γRel⁡(Q,A+))∑A′∈Aexp⁡(γRel⁡(Q,A′))\operatorname{Pr}(A^+ \mid Q) = \frac{\exp(\gamma \operatorname{Rel}(Q, A^+))}{\sum_{A' \in \mathcal{A}} \exp(\gamma \operatorname{Rel}(Q, A'))} where γ\gamma is a scaling hyperparameter (set to γ=1\gamma = 1). During evaluation, candidate answers for a given query are ranked by Rel⁡(Q,A)\operatorname{Rel}(Q, A).

  4. Knowl 4 — Multi-Task Learning Training Procedure and Objectives

    algorithm

    The training procedure for MT-DNN consists of two stages: unsupervised pre-training of the shared Transformer layers followed by multi-task supervised training.

    Input: Datasets for TT tasks D1,…,DT\mathcal{D}_1, \ldots, \mathcal{D}_T, maximum epochs epochmax⁡\text{epoch}_{\max}, learning rate schedule
    Output: Jointly trained model parameter set Θ\Theta
    Initialize shared encoder parameters with pre-trained BERT (masked LM and next sentence prediction objectives)
    Initialize task-specific layer parameters randomly
    for t=1,2,…,Tt = 1, 2, \ldots, T do
        Partition dataset Dt\mathcal{D}_t into mini-batches
    end for
    for epoch=1,2,…,epochmax⁡\text{epoch} = 1, 2, \ldots, \text{epoch}_{\max} do
        Merge all mini-batches into a single pool: D=D1∪D2∪…∪DT\mathcal{D} = \mathcal{D}_1 \cup \mathcal{D}_2 \cup \ldots \cup \mathcal{D}_T
        Randomly shuffle the mini-batches in D\mathcal{D}
        for each mini-batch bt∈Db_t \in \mathcal{D} (where btb_t belongs to task tt) do
            if task tt is single-sentence or pairwise classification then
                Compute cross-entropy loss: L(Θ)=−∑c1(X,c)log⁡(Pr⁡(c∣X))\mathcal{L}(\Theta) = -\sum_c \mathbf{1}(X, c) \log(\operatorname{Pr}(c \mid X))
            else if task tt is text similarity regression then
                Compute mean squared error loss: L(Θ)=(y−Sim⁡(X1,X2))2\mathcal{L}(\Theta) = (y - \operatorname{Sim}(X_1, X_2))^2
            else if task tt is relevance ranking then
                Compute pairwise ranking loss: L(Θ)=−log⁡Pr⁡(A+∣Q)\mathcal{L}(\Theta) = -\log \operatorname{Pr}(A^+ \mid Q)
            end if
            Compute gradient: g=∇ΘL(Θ)g = \nabla_\Theta \mathcal{L}(\Theta)
            Update model parameters: Θ←Θ−ϵg\Theta \leftarrow \Theta - \epsilon g
        end for
    end for
    return Θ\Theta
  5. Knowl 5 — GLUE Benchmark Test Evaluation Results

    empirical result

    The Multi-Task Deep Neural Network (MT-DNN) was evaluated on the official General Language Understanding Evaluation (GLUE) benchmark test server using BERTLARGE\text{BERT}_{\text{LARGE}} to initialize its shared encoder. MT-DNN was tested both with and without task-specific fine-tuning.

    Model CoLA SST-2 MRPC STS-B QQP MNLI-m/mm QNLI RTE WNLI AX Score
    BiLSTM+ELMo+Attn 36.0 90.4 84.9/77.9 75.1/73.3 64.8/84.7 76.4/76.1 - 56.8 65.1 26.5 70.5
    Transformer (GPT) 45.4 91.3 82.3/75.7 82.0/80.0 70.3/88.5 82.1/81.4 - 56.0 53.4 29.8 72.8
    GPT on STILTs 47.2 93.1 87.7/83.7 85.3/84.8 70.1/88.1 80.8/80.6 - 69.1 65.1 29.4 76.9
    BERTLARGE\text{BERT}_{\text{LARGE}} 60.5 94.9 89.3/85.4 87.6/86.5 72.1/89.3 86.7/85.9 92.7 70.1 65.1 39.6 80.5
    MT-DNNno-fine-tune\text{MT-DNN}_{\text{no-fine-tune}} 58.9 94.6 90.1/86.4 89.5/88.8 72.7/89.6 86.5/85.8 93.1 79.1 65.1 39.4 81.7
    MT-DNN 62.5 95.6 91.1/88.2 89.5/88.8 72.7/89.6 86.7/86.0 93.1 81.4 65.1 40.3 82.7
    Human Performance 66.4 97.8 86.3/80.8 92.7/92.6 59.5/80.4 92.0/92.8 91.2 93.6 95.9 - 87.1

    Metrics reported: Matthews correlation for CoLA; Accuracy for SST-2, QNLI, RTE, and WNLI; Accuracy/F1 for MRPC and QQP; Pearson/Spearman correlation for STS-B; Accuracy on matched/mismatched sets for MNLI-m/mm; and overall average GLUE score.

    MT-DNN achieves an overall GLUE score of 82.7%, representing a 2.2% absolute improvement over BERTLARGE\text{BERT}_{\text{LARGE}} (80.5%) and outperforming all prior models on 8 out of 9 GLUE tasks (all except WNLI, where all models hit the majority baseline of 65.1). MT-DNN without fine-tuning (MT-DNNno-fine-tune\text{MT-DNN}_{\text{no-fine-tune}}) achieves an 81.7% score, also outperforming fine-tuned BERTLARGE\text{BERT}_{\text{LARGE}}. The largest improvements over BERT occur on tasks with small amounts of in-domain training data (e.g., RTE: +11.3% absolute; MRPC: +1.8%/+2.8% absolute).

  6. Knowl 6 — Single-Task DNN Ablation and GLUE Development Set Performance

    empirical result

    To evaluate the relative contributions of multi-task learning versus task-specific architectural improvements, MT-DNN is compared against Single-Task DNN (ST-DNN), which shares the exact same architecture as MT-DNN (including the SAN answer module and pairwise ranking loss) but initializes shared layers directly from BERTLARGE\text{BERT}_{\text{LARGE}} without multi-task refinement, fine-tuning solely on task-specific data.

    Model MNLI-m/mm QQP RTE QNLI (v1/v2) MRPC CoLA SST-2 STS-B
    BERTLARGE\text{BERT}_{\text{LARGE}} 86.3/86.2 91.1/88.0 71.1 90.5/92.4 89.5/85.8 61.8 93.5 89.6/89.3
    ST-DNN 86.6/86.3 91.3/88.4 72.0 96.1/– 89.7/86.4 – – –
    MT-DNN 87.1/86.7 91.9/89.2 83.4 97.4/92.9 91.0/87.5 63.5 94.3 90.7/90.6

    Key takeaways:

    1. On all four pairwise classification tasks (MNLI, QQP, RTE, MRPC), ST-DNN outperforms BERTLARGE\text{BERT}_{\text{LARGE}}, confirming the benefit of the SAN multi-step reasoning answer module over standard linear classification heads.
    2. On QNLI, ST-DNN (96.1% on v1) outperforms BERTLARGE\text{BERT}_{\text{LARGE}} (90.5%), demonstrating that formulating QA relevance as pairwise ranking is superior to binary classification.
    3. MT-DNN yields substantial additional gains across all tasks over ST-DNN (e.g., RTE increases from 72.0% to 83.4%; QNLI v1 increases from 96.1% to 97.4%), confirming the complementary value of multi-task representation learning.
  7. Knowl 7 — Few-Shot Domain Adaptation Performance on SNLI and SciTail

    empirical result

    Domain adaptation experiments evaluate the transferability and sample efficiency of shared representations by fine-tuning MT-DNN and BERT models on varying fractions of in-domain training data (0.1%, 1%, 10%, and 100%) on two external NLI benchmarks: SNLI and SciTail. Both models are based on the BERTBASE\text{BERT}_{\text{BASE}} architecture, and results are averaged over 5 random sampling runs.

    Model 0.1% 1% 10% 100%
    SNLI Dataset (Dev Accuracy %)
    # Training Data 549 5,493 54,936 549,367
    BERTBASE\text{BERT}_{\text{BASE}} 52.5 78.1 86.7 91.0
    MT-DNNBASE\text{MT-DNN}_{\text{BASE}} 82.1 85.2 88.4 91.5
    SciTail Dataset (Dev Accuracy %)
    # Training Data 23 235 2,359 23,596
    BERTBASE\text{BERT}_{\text{BASE}} 51.2 82.2 90.5 94.3
    MT-DNNBASE\text{MT-DNN}_{\text{BASE}} 81.9 88.3 91.1 95.7

    When training data is extremely limited (0.1% split), MT-DNNBASE\text{MT-DNN}_{\text{BASE}} outperforms BERTBASE\text{BERT}_{\text{BASE}} by 29.6% absolute on SNLI (82.1% vs 52.5% with 549 examples) and by 30.7% absolute on SciTail (81.9% vs 51.2% with only 23 examples). The performance of MT-DNNBASE\text{MT-DNN}_{\text{BASE}} with only 0.1% or 1.0% training data matches or exceeds कई prior full-dataset models, indicating that multi-task training produces more generalizable representations that adapt with far fewer target labels.

  8. Knowl 8 — Full-Dataset Benchmark Results on SNLI and SciTail

    empirical result

    When fine-tuned using the full training sets of SNLI and SciTail, MT-DNN achieves new state-of-the-art results across both datasets, outperforming prior models including GPT, specialized matching architectures, and BERT.

    Model Dev Accuracy (%) Test Accuracy (%)
    SNLI Dataset
    GPT (Radford et al., 2018) – 89.9
    Kim et al. (2018) – 90.1
    BERTBASE\text{BERT}_{\text{BASE}} 91.0 90.8
    MT-DNNBASE\text{MT-DNN}_{\text{BASE}} 91.5 91.1
    BERTLARGE\text{BERT}_{\text{LARGE}} 91.7 91.0
    MT-DNNLARGE\text{MT-DNN}_{\text{LARGE}} 92.2 91.6
    SciTail Dataset
    GPT (Radford et al., 2018) – 88.3
    BERTBASE\text{BERT}_{\text{BASE}} 94.3 92.0
    MT-DNNBASE\text{MT-DNN}_{\text{BASE}} 95.7 94.1
    BERTLARGE\text{BERT}_{\text{LARGE}} 95.7 94.4
    MT-DNNLARGE\text{MT-DNN}_{\text{LARGE}} 96.3 95.0

    On SNLI test accuracy, MT-DNNLARGE\text{MT-DNN}_{\text{LARGE}} achieves 91.6% (+1.5% absolute over the previous 90.1% state-of-the-art by Kim et al. and +0.6% over BERTLARGE\text{BERT}_{\text{LARGE}}). On SciTail test accuracy, MT-DNNLARGE\text{MT-DNN}_{\text{LARGE}} achieves 95.0% (+6.7% absolute over the previous 88.3% state-of-the-art by GPT and +0.6% over BERTLARGE\text{BERT}_{\text{LARGE}}).

  9. Knowl 9 — Experimental Setup and Training Hyperparameters for MT-DNN

    experimental setup

    The implementation of MT-DNN uses PyTorch and follows the optimization protocol of BERT:

    • Optimizer: Adamax optimizer.
    • Learning Rate: 5×10−55 \times 10^{-5} with a batch size of 32.
    • Schedule: Linear learning rate decay with a warmup ratio of 0.1 over the total training steps.
    • Epochs: Maximum number of epochs set to 5.
    • Dropout: Dropout rate of 0.1 for task-specific layers, except 0.3 for MNLI and 0.05 for CoLA.
    • Gradient Management: Gradient norm clipped within a threshold of 1.0 to prevent exploding gradients.
    • Tokenization: WordPiece tokenization with maximum sequence length capped at 512 tokens.

Coverage note — No substantial contributed material was omitted. All architectural components, loss functions, algorithms, benchmark evaluations on GLUE, SNLI, and SciTail, few-shot domain adaptation experiments, ablations with ST-DNN, and hyperparameter details are covered.

References

  1. 1.Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. 2015a. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642.
  2. 2.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015b. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics.
  3. 3.Chris Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, and Greg Hullender. 2005. Learning to rank using gradient descent. In Proceedings of the 22nd international conference on Machine learning, pages 89–96. ACM.
  4. 4.Rich Caruana. 1997. Multitask learning. Machine learning, 28(1):41–75.
  5. 5.Ronan Collobert, Jason Weston, Leon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. 2011. Natural language processing (almost) from scratch. Journal of Machine Learning Research, 12(Aug):2493–2537.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  7. 7.Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019. Unified language model pre-training for natural language understanding and generation. arXiv preprint arXiv:1905.03197.
  8. 8.J. Gao, M. Galley, and L. Li. 2018. Neural approaches to conversational AI. CoRR, abs/1809.08267.
  9. 9.Max Glockner, Vered Shwartz, and Yoav Goldberg. 2018. Breaking nli systems with sentences that require simple lexical inferences. In The 56th Annual Meeting of the Association for Computational Linguistics (ACL), Melbourne, Australia.
  10. 10.Han Guo, Ramakanth Pasunuru, and Mohit Bansal. 2018. Soft layer-specific multi-task summarization with entailment and question generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 687–697.
  11. 11.Po-Sen Huang, Xiaodong He, Jianfeng Gao, Li Deng, Alex Acero, and Larry Heck. 2013. Learning deep structured semantic models for web search using clickthrough data. In Proceedings of the 22nd ACM international conference on Conference on information & knowledge management, pages 2333–2338. ACM.
  12. 12.Tushar Khot, Ashish Sabharwal, and Peter Clark. 2018. SciTail: A textual entailment dataset from science question answering. In AAAI.
  13. 13.Seonhoon Kim, Jin-Hyuk Hong, Inho Kang, and Nojun Kwak. 2018. Semantic sentence matching with densely-connected recurrent and co-attentive information. arXiv preprint arXiv:1805.11360.
  14. 14.Diederik Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  15. 15.Xiaodong Liu, Kevin Duh, and Jianfeng Gao. 2018a. Stochastic answer networks for natural language inference. arXiv preprint arXiv:1804.07888.
  16. 16.Xiaodong Liu, Jianfeng Gao, Xiaodong He, Li Deng, Kevin Duh, and Ye-Yi Wang. 2015. Representation learning using multi-task deep neural networks for semantic classification and information retrieval. In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 912–921.
  17. 17.Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao. 2019. Improving multi-task deep neural networks via knowledge distillation for natural language understanding. arXiv preprint arXiv:1904.09482.
  18. 18.Xiaodong Liu, Yelong Shen, Kevin Duh, and Jianfeng Gao. 2018b. Stochastic answer networks for machine reading comprehension. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics.
  19. 19.Minh-Thang Luong, Quoc V Le, Ilya Sutskever, Oriol Vinyals, and Lukasz Kaiser. 2015. Multi-task sequence to sequence learning. arXiv preprint arXiv:1511.06114.
  20. 20.Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. arXiv preprint arXiv:1802.05365.
  21. 21.Jason Phang, Thibault Fevry, and Samuel R Bowman. 2018. Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks. arXiv preprint arXiv:1811.01088.
  22. 22.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training.
  23. 23.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100,000+ questions for machine comprehension of text. pages 2383–2392.
  24. 24.Sebastian Ruder12, Joachim Bingel, Isabelle Augenstein, and Anders Søgaard. 2019. Latent multi-task architecture learning.
  25. 25.Aarne Talman and Stergios Chatzikyriakidis. 2018. Testing the generalization power of neural network models across nli benchmarks. arXiv preprint arXiv:1810.09774.
  26. 26.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. arXiv preprint arXiv:1706.03762.
  27. 27.Alex Wang, Amapreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461.
  28. 28.Yichong Xu, Xiaodong Liu, Yelong Shen, Jingjing Liu, and Jianfeng Gao. 2018. Multi-task learning for machine reading comprehension. arXiv preprint arXiv:1809.06963.
  29. 29.Yu Zhang and Qiang Yang. 2017. A survey on multi-task learning. arXiv preprint arXiv:1707.08114.

Citation

MLA
Liu, X., et al. “Multi-Task Deep Neural Networks for Natural Language Understanding”. arXiv, 2019, http://arxiv.org/abs/1901.11504v2.
APA
Liu, X., He, P., Chen, W., & Gao, J. (2019). Multi-Task Deep Neural Networks for Natural Language Understanding. arXiv. http://arxiv.org/abs/1901.11504v2
Chicago
Liu, X., P. He, W. Chen, and J. Gao. 2019. “Multi-Task Deep Neural Networks for Natural Language Understanding”. arXiv. http://arxiv.org/abs/1901.11504v2.
Harvard
Liu, X. et al. (2019) “Multi-Task Deep Neural Networks for Natural Language Understanding”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1901.11504v2.
Vancouver
1. Liu X, He P, Chen W, Gao J (2019) Multi-Task Deep Neural Networks for Natural Language Understanding. arXiv

BibTeX

@article{liu2019multi,
  title = {Multi-Task Deep Neural Networks for Natural Language Understanding},
  author = {Liu, Xiaodong and He, Pengcheng and Chen, Weizhu and Gao, Jianfeng},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1901.11504v2},
  eprint = {1901.11504}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/