Deep Biaffine Attention for Neural Dependency Parsing

Timothy DozatChristopher D. Manning

article2016ICLR1,349 citations

Introduces a neural dependency parser using deep biaffine attention to predict syntactic arcs and labels, establishing state-of-the-art accuracy across multilingual treebanks and identifying critical hyperparameter choices for graph-based models.

Listen

Automated sentence analysis, known as dependency parsing, forms a critical foundation for modern language technologies and natural language understanding. However, parsing errors frequently undermine downstream system performance. While complex step-by-step parsers historically achieved higher accuracy, simpler holistic graph-based methods lagged behind. The article evaluates whether architectural refinements and targeted training configurations can enable a simpler, highly efficient graph-based model to match or exceed the accuracy of complex, state-of-the-art alternative systems.

To achieve this, the article introduces a deep biaffine attention framework within a bidirectional recurrent neural network. This design uses intermediate dimension-reducing layers to strip out irrelevant information before computing grammatical connections and relationship labels. The authors systematically evaluated this architecture across multiple standard multilingual benchmarks—including English, Chinese, Catalan, Czech, German, and Spanish—while analyzing the impact of network depth, cell designs, input regularization, and optimization hyperparameters.

Key findings show that the proposed design achieves leading performance across several benchmarks, reaching 95.7% unlabeled attachment score and 94.1% labeled attachment score on standard English data, outperforming previous graph-based baselines by about 1.8 to 2.2 percentage points. The model also established new performance benchmarks on multilingual datasets, particularly excelling on complex non-projective grammatical structures where step-by-step parsers struggle. Additionally, the deep biaffine mechanism processed over 410 sentences per second, proving significantly faster and more accurate than shallow or traditional multi-layer alternatives. Hyperparameter findings further highlighted that extensive dropout regularization across both words and part-of-speech tags, alongside adjusting the optimization decay parameter from standard defaults, was vital to prevent severe overfitting.

These results demonstrate that organizations can deploy simpler, faster graph-based parsing models without sacrificing structural accuracy, thereby reducing operational computational costs and latency. The findings also suggest that when adopting advanced neural networks, robust regularization across all input streams and careful optimizer tuning are just as critical as architectural complexity.

Decision-makers should consider adopting deep biaffine scoring architectures for language pipelines requiring both high throughput and strong structural accuracy. For future technical roadmaps, teams should explore richer pretrained word representations, improved handling of out-of-vocabulary words in morphologically diverse languages, and mechanisms to better capture complex phrase compositions to bridge the remaining gap in relationship label precision.

arXiv: 1611.01734
Cover for Deep Biaffine Attention for Neural Dependency Parsing

Abstract

This paper builds off recent work from Kiperwasser & Goldberg (2016) using neural attention in a simple graph-based dependency parser. We use a larger but more thoroughly regularized parser than other recent BiLSTM-based approaches, with biaffine classifiers to predict arcs and labels. Our parser gets state of the art or near state of the art performance on standard treebanks for six different languages, achieving 95.7% UAS and 94.1% LAS on the most popular English PTB dataset. This makes it the highest-performing graph-based parser on this benchmark---outperforming Kiperwasser Goldberg (2016) by 1.8% and 2.2%---and comparable to the highest performing transition-based parser (Kuncoro et al., 2016), which achieves 95.8% UAS and 94.6% LAS. We also show which hyperparameter choices had a significant effect on parsing accuracy, allowing us to achieve large gains over other graph-based approaches.

Table of Contents

  • 1 Introduction
  • 2 Background and Related Work
  • 3 Proposed Dependency Parser
  • 3.1 Deep biaffine attention
  • 3.2 Hyperparameter configuration
  • 4 Experiments & Results
  • 4.1 Datasets
  • 4.2 Hyperparameter choices
  • 4.2.1 Attention mechanism
  • 4.2.2 Network size
  • 4.2.3 Recurrent cell
  • 4.2.4 Embedding Dropout
  • 4.2.5 Optimizer
  • 4.3 Results
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Deep Biaffine Attention for Dependency Arc Scoring

    model/method

    In the deep biaffine neural dependency parser, the score for an arc connecting candidate head word jj to dependent word ii is computed using a biaffine attention function over lower-dimensional representations. Let ri∈R2dlstm\mathbf{r}_i \in \mathbb{R}^{2d_{lstm}} be the concatenated bidirectional LSTM output vector for word ii.

    First, non-linear multi-layer perceptron (MLP) layers project ri\mathbf{r}_i and rj\mathbf{r}_j into separate representation spaces for dependent and head roles: hi(arc−dep)=MLP(arc−dep)(ri)\mathbf{h}_i^{(arc-dep)} = \text{MLP}^{(arc-dep)}(\mathbf{r}_i) hj(arc−head)=MLP(arc−head)(rj)\mathbf{h}_j^{(arc-head)} = \text{MLP}^{(arc-head)}(\mathbf{r}_j) where hi(arc−dep),hj(arc−head)∈Rdarc\mathbf{h}_i^{(arc-dep)}, \mathbf{h}_j^{(arc-head)} \in \mathbb{R}^{d_{arc}}.

    For a sentence of length nn, let H(arc−head)∈Rn×darcH^{(arc-head)} \in \mathbb{R}^{n \times d_{arc}} be the matrix whose jj-th row is (hj(arc−head))⊤(\mathbf{h}_j^{(arc-head)})^\top. The vector of arc scores si(arc)∈Rn\mathbf{s}_i^{(arc)} \in \mathbb{R}^n assigning a score to each possible head j∈{1,…,n}j \in \{1, \dots, n\} for dependent ii is calculated via the biaffine transformation: si(arc)=H(arc−head)U(1)hi(arc−dep)+H(arc−head)u(2)\mathbf{s}_i^{(arc)} = H^{(arc-head)} U^{(1)} \mathbf{h}_i^{(arc-dep)} + H^{(arc-head)} \mathbf{u}^{(2)} where U(1)∈Rdarc×darcU^{(1)} \in \mathbb{R}^{d_{arc} \times d_{arc}} is a weight matrix modeling the bilinear interaction between head and dependent representations, and u(2)∈Rdarc\mathbf{u}^{(2)} \in \mathbb{R}^{d_{arc}} is a bias vector.

    Equivalently, the scalar score for an individual directed arc j→ij \to i is: sj,i(arc)=(hj(arc−head))⊤U(1)hi(arc−dep)+(hj(arc−head))⊤u(2)s_{j,i}^{(arc)} = (\mathbf{h}_j^{(arc-head)})^\top U^{(1)} \mathbf{h}_i^{(arc-dep)} + (\mathbf{h}_j^{(arc-head)})^\top \mathbf{u}^{(2)} This formulation directly models both the prior probability of word jj receiving any dependent via (hj(arc−head))⊤u(2)(\mathbf{h}_j^{(arc-head)})^\top \mathbf{u}^{(2)} and the specific affinity of head jj taking dependent ii via (hj(arc−head))⊤U(1)hi(arc−dep)(\mathbf{h}_j^{(arc-head)})^\top U^{(1)} \mathbf{h}_i^{(arc-dep)}.

  2. Knowl 2 — Biaffine Dependency Label Classifier

    model/method

    Once dependency arcs are determined, dependency relation labels are predicted using a multi-class fixed-class biaffine classifier. Given dependent word ii and its assigned (gold or predicted) head word yiy_i, separate MLP layers project the BiLSTM recurrent states ri\mathbf{r}_i and ryi\mathbf{r}_{y_i} into label-specific head and dependent representations: hi(label−dep)=MLP(label−dep)(ri)\mathbf{h}_i^{(label-dep)} = \text{MLP}^{(label-dep)}(\mathbf{r}_i) hyi(label−head)=MLP(label−head)(ryi)\mathbf{h}_{y_i}^{(label-head)} = \text{MLP}^{(label-head)}(\mathbf{r}_{y_i}) where hi(label−dep),hyi(label−head)∈Rdlabel\mathbf{h}_i^{(label-dep)}, \mathbf{h}_{y_i}^{(label-head)} \in \mathbb{R}^{d_{label}}.

    The vector of scores si(label)∈Rc\mathbf{s}_i^{(label)} \in \mathbb{R}^c across all cc possible dependency labels is given by: si(label)=(hyi(label−head))⊤U(1)hi(label−dep)+U(2)(hyi(label−head)⊕hi(label−dep))+b\mathbf{s}_i^{(label)} = (\mathbf{h}_{y_i}^{(label-head)})^\top \mathbf{U}^{(1)} \mathbf{h}_i^{(label-dep)} + U^{(2)} \left(\mathbf{h}_{y_i}^{(label-head)} \oplus \mathbf{h}_i^{(label-dep)}\right) + \mathbf{b} where ⊕\oplus denotes vector concatenation, U(1)∈Rdlabel×c×dlabel\mathbf{U}^{(1)} \in \mathbb{R}^{d_{label} \times c \times d_{label}} is a 3rd-order tensor for bilinear feature interactions per class, U(2)∈Rc×2dlabelU^{(2)} \in \mathbb{R}^{c \times 2d_{label}} is a weight matrix for linear feature interactions, and b∈Rc\mathbf{b} \in \mathbb{R}^c is a bias vector.

    This scoring function explicitly integrates four distinct linguistic components:

    1. b\mathbf{b}: The prior probability of each dependency label.
    2. Linear terms involving hi(label−dep)\mathbf{h}_i^{(label-dep)}: The likelihood of a label given only the dependent word.
    3. Linear terms involving hyi(label−head)\mathbf{h}_{y_i}^{(label-head)}: The likelihood of a label given only the head word.
    4. Bilinear term (hyi(label−head))⊤U(1)hi(label−dep)(\mathbf{h}_{y_i}^{(label-head)})^\top \mathbf{U}^{(1)} \mathbf{h}_i^{(label-dep)}: The likelihood of a label given the interaction between head and dependent.
  3. Knowl 3 — End-to-End Deep Biaffine Dependency Parsing Framework

    model/method

    The deep biaffine dependency parser processes an input sentence w1,…,wnw_1, \dots, w_n with corresponding part-of-speech tags t1,…,tnt_1, \dots, t_n through the following pipeline:

    1. Input Layer: For each token, a 100-dimensional word embedding (initialized with pretrained embeddings plus a learned embedding for words occurring ≥2\ge 2 times) is concatenated with a 100-dimensional learned POS tag embedding to produce input xi=[ewi;eti]∈R200\mathbf{x}_i = [\mathbf{e}_{w_i}; \mathbf{e}_{t_i}] \in \mathbb{R}^{200}.

    2. Encoder: A 3-layer bidirectional LSTM (400 dimensions in each direction, outputting ri∈R800\mathbf{r}_i \in \mathbb{R}^{800}) encodes context across the sequence.

    3. Role-Specific Projections: Four single-layer ReLU MLPs compute specialized representations:

      • Arc dependent: hi(arc−dep)∈R500\mathbf{h}_i^{(arc-dep)} \in \mathbb{R}^{500}
      • Arc head: hi(arc−head)∈R500\mathbf{h}_i^{(arc-head)} \in \mathbb{R}^{500}
      • Label dependent: hi(label−dep)∈R100\mathbf{h}_i^{(label-dep)} \in \mathbb{R}^{100}
      • Label head: hi(label−head)∈R100\mathbf{h}_i^{(label-head)} \in \mathbb{R}^{100}
    4. Scoring and Training: Arc scores are computed with biaffine attention; label scores are computed with a biaffine classifier. At training time, cross-entropy loss is computed independently for each word taking its gold head and gold label.

    5. Inference: At test time, head selection is subject to the Maximum Spanning Tree (MST) algorithm to guarantee that the output graph forms a valid, single-root directed dependency tree over the sentence.

  4. Knowl 4 — Dimensionality Reduction and Feature Stripping via Specialized MLPs

    model/method

    In neural graph-based parsers, recurrent states ri\mathbf{r}_i from the top BiLSTM layer must simultaneously encode information for identifying heads, identifying dependents, rejecting non-dependents, predicting labels, and transferring sequential context to surrounding tokens. Directly feeding ri\mathbf{r}_i into bilinear transformations (shallow bilinear attention) leads to:

    • Oversized parameter tensors (e.g., (801×c×801)(801 \times c \times 801) for label classification vs (101×c×101)(101 \times c \times 101) with MLPs).
    • Pronounced overfitting and slower parsing speeds.

    Applying specialized, dimension-reducing MLPs with non-linear activations prior to the biaffine transformations strips away information extraneous to that specific decision, projecting the BiLSTM outputs into lower-dimensional sub-spaces tailored specifically for dependent or head roles in arc and label prediction.

  5. Knowl 5 — Regularization Scheme with Independent Input Dropout and Timestep-Shared Dropout

    model/method

    To regularize the parser architecture, dropout is applied at multiple levels with specific retention policies:

    1. Input Word and Tag Dropout: During training, word embeddings and POS tag embeddings are dropped independently with a probability of 33%33\%. When either the word or tag is dropped, the remaining vector is scaled by a factor of 2. If both are dropped simultaneously, an all-zero vector is passed to the BiLSTM.

    2. Bayesian / Variational Recurrent Dropout: Dropout is applied with rate 33%33\% to both input connections and recurrent connections of the 3-layer BiLSTM. Crucially, the exact same dropout mask is reused across every recurrent timestep within a sentence.

    3. MLP and Classifier Dropout: Nodes in the feedforward MLP layers and the classifier heads are dropped with rate 33%33\%, with the dropout mask also held fixed across all timesteps for that sentence.

  6. Knowl 6 — Adam Optimization with Fast Gradient-Momentum Adaptation

    model/method

    The parser parameters are optimized using the Adam optimizer with a modified second-moment decay rate β2=0.9\beta_2 = 0.9 instead of the default β2=0.999\beta_2 = 0.999, while keeping β1=0.9\beta_1 = 0.9 and initial learning rate α=2×10−3\alpha = 2 \times 10^{-3}.

    The standard β2=0.999\beta_2 = 0.999 sets a very long moving average for gradient scale normalization, causing past gradient magnitudes to excessively influence current parameter updates and hindering fast adaptation to recent network changes. Lowering β2\beta_2 to 0.90.9 enables the optimizer to adapt rapidly to changes in gradient scale.

    The learning rate α\alpha is exponentially annealed during training according to: αt=α0⋅(0.75)t/5000\alpha_t = \alpha_0 \cdot (0.75)^{t / 5000} for approximately tmax⁡=50,000t_{\max} = 50,000 iterations (rounded up to the nearest full epoch).

  7. Knowl 7 — Dependency Parsing Performance Across Multilingual Benchmarks

    data/table

    The deep biaffine parser was evaluated against leading transition-based and graph-based dependency parsers on standard benchmarks: English Penn Treebank converted to Stanford Dependencies (PTB-SD 3.3.0), Chinese Penn Treebank (CTB 5.1), and the CoNLL 2009 shared task across six languages. Punctuation is excluded from evaluation on PTB-SD and CTB.

    English PTB-SD 3.3.0 Chinese PTB 5.1
    Type Model UAS LAS UAS LAS
    Transition Ballesteros et al. (2016) 93.56 91.42 87.65 86.21
    Transition Andor et al. (2016) 94.61 92.79 – –
    Transition Kuncoro et al. (2016) 95.8 94.6 – –
    Graph Kiperwasser Goldberg (2016) 93.9 91.9 87.6 86.1
    Graph Cheng et al. (2016) 94.10 91.49 88.1 85.7
    Graph Hashimoto et al. (2016) 94.67 92.90 – –
    Graph Deep Biaffine 95.74 94.08 89.30 88.23
    Catalan Chinese Czech
    Model UAS LAS UAS LAS UAS LAS
    Andor et al. (2016) 92.67 89.83 84.72 80.85 88.94 84.56
    Deep Biaffine 94.69 92.02 88.90 85.38 92.08 87.38
    English German Spanish
    Model UAS LAS UAS LAS UAS LAS
    Andor et al. (2016) 93.22 91.23 90.91 89.15 92.62 89.95
    Deep Biaffine 95.21 93.20 93.46 91.44 94.34 91.65

    The deep biaffine model establishes state-of-the-art results among graph-based parsers, outperforming Kiperwasser & Goldberg (2016) by 1.84%1.84\% UAS and 2.18%2.18\% LAS on PTB-SD 3.3.0. It also achieves state-of-the-art results on all six CoNLL 2009 languages, substantially surpassing Andor et al. (2016) due in part to the graph parser's native capacity to model non-projective dependencies present in CoNLL datasets.

  8. Knowl 8 — Architectural and Hyperparameter Ablation on PTB-SD 3.5.0

    data/table

    Ablation experiments evaluated on the English PTB-SD 3.5.0 test set assess the impact of classifier design, network depth/width, recurrent cell type, dropout scheme, and optimizer hyperparameters.

    Classifier Model UAS LAS Sents/s Network Size UAS LAS Sents/s
    Deep (Default) 95.75 94.22 410.91 3 layers, 400d 95.75 94.22 410.91
    Shallow 95.74 94.00* 298.99 3 layers, 300d 95.82 94.24 460.01
    Shallow, 50% drop 95.73 94.05* 300.04 3 layers, 200d 95.55* 93.89* 469.45
    Shallow, 300d 95.63* 93.86* 373.24 2 layers, 400d 95.62* 93.98* 497.99
    MLP 95.53* 93.91* 367.44 4 layers, 400d 95.83 94.22 362.09
    Recurrent Cell Input Dropout
    LSTM 95.75 94.22 410.91 Default 95.75 94.22 –
    GRU 93.18* 91.08* 435.32 No word dropout 95.74 94.08* –
    Cif-LSTM 95.67 94.06* 463.25 No tag dropout 95.28* 93.60* –
    Optimizer No tags 95.77 93.91* –
    β2=0.9\beta_2 = 0.9 95.75 94.22 –
    β2=0.999\beta_2 = 0.999 95.53* 93.91* –

    Asterisks denote statistically significant differences compared to the default model (p<0.05p < 0.05).

    Key takeaways from these ablations:

    1. Deep vs. Shallow / MLP: The deep biaffine architecture is both faster (410.91 vs. 298.99 and 367.44 sentences/sec) and more accurate than shallow bilinear or MLP classifiers.
    2. Recurrent Cell: GRU severely degrades accuracy (93.18%93.18\% UAS), likely because its lack of an independent output gate prevents maintaining sparse representations resilient to high dropout. Coupled input-forget gate LSTM (Cif-LSTM) achieves higher speed (463.25 sentences/sec) with minimal accuracy loss.
    3. Input Dropout: Dropping both words and POS tags during training is essential; removing tag dropout degrades UAS to 95.28%95.28\%, which is worse than training without POS tags entirely (95.77%95.77\% UAS).
    4. Optimization: Setting Adam's β2=0.9\beta_2 = 0.9 yields significant gains over the standard β2=0.999\beta_2 = 0.999 (94.22%94.22\% vs. 93.91%93.91\% LAS).
  9. Knowl 9 — Expressive and Structural Limitations of First-Order Graph Parsers

    limitation

    While the deep biaffine graph parser achieves competitive UAS with transition-based models (e.g., 95.74%95.74\% vs. 95.8%95.8\% UAS on PTB-SD 3.3.0), it exhibits a noticeable lag in Labeled Attachment Score (LAS) compared to phrasal compositional transition models (such as Kuncoro et al., 2016 at 94.6%94.6\% LAS vs. 94.08%94.08\% LAS).

    This gap stems from architectural and structural factors:

    1. Lack of Phrasal Compositionality: First-order graph-based parsers compute edge and label scores from independent word-pair representations rather than explicitly composing parsed phrases into subtree representations.
    2. Reduced Syntactic History: Graph-based edge scoring lacks access to previous parsing decisions during inference, unlike transition-based systems that maintain explicit stack states.
    3. Sensitivity to Tag and Embedding Inefficiencies: The model remains reliant on external POS tagger accuracy and pretrained word embeddings (e.g., GloVe) for fine-grained dependency relation labeling.

Coverage note — None was omitted; all key architectural components, equations, regularization techniques, optimization findings, benchmark evaluations, ablation studies, and stated limitations are fully represented.

References

  1. 1.Daniel Andor, Chris Alberti, David Weiss, Aliaksei Severyn, Alessandro Presta, Kuzman Ganchev, Slav Petrov, and Michael Collins. Globally normalized transition-based neural networks. In Association for Computational Linguistics, 2016. URL https://arxiv.org/abs/1603.06042.
  2. 2.Gabor Angeli, Melvin Johnson Premkumar, and Christopher D Manning. Leveraging linguistic structure for open domain information extraction. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics (ACL 2015), 2015.
  3. 3.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. International Conference on Learning Representations, 2014.
  4. 4.Miguel Ballesteros, Yoav Goldberg, Chris Dyer, and Noah A Smith. Training with exploration improves a greedy stack-LSTM parser. Proceedings of the conference on empirical methods in natural language processing, 2016.
  5. 5.Samuel R Bowman, Jon Gauthier, Abhinav Rastogi, Raghav Gupta, Christopher D Manning, and Christopher Potts. A fast unified model for parsing and sentence understanding. ACL 2016, 2016.
  6. 6.Danqi Chen and Christopher D Manning. A fast and accurate dependency parser using neural networks. In Proceedings of the conference on empirical methods in natural language processing, pp. 740–750, 2014.
  7. 7.Hao Cheng, Hao Fang, Xiaodong He, Jianfeng Gao, and Li Deng. Bi-directional attention with agreement for dependency parsing. arXiv preprint arXiv:1608.02076, 2016.
  8. 8.Chris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, and Noah A Smith. Transition-based dependency parsing with stack long short-term memory. Proceedings of the conference on empirical methods in natural language processing, 2015.
  9. 9.Yarin Gal and Zoubin Ghahramani. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. International Conference on Machine Learning, 2015.
  10. 10.Klaus Greff, Rupesh Kumar Srivastava, Jan Koutník, Bas R Steunebrink, and Jürgen Schmidhuber. LSTM: A search space odyssey. IEEE Transactions on Neural Networks and Learning Systems, 2015.
  11. 11.Kazuma Hashimoto, Caiming Xiong, Yoshimasa Tsuruoka, and Richard Socher. A joint many-task model: Growing a neural network for multiple nlp tasks. arXiv preprint arXiv:1611.01587, 2016.
  12. 12.Diederik Kingma and Jimmy Ba. Adam: A method for stochastic optimization. International Conference on Learning Representations, 2014.
  13. 13.Eliyahu Kiperwasser and Yoav Goldberg. Simple and accurate dependency parsing using bidirectional LSTM feature representations. Transactions of the Association for Computational Linguistics, 4:313–327, 2016.
  14. 14.Adhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Graham Neubig, and Noah A. Smith. What do recurrent neural network grammars learn about syntax? CoRR, abs/1611.05774, 2016. URL http://arxiv.org/abs/1611.05774.
  15. 15.Omer Levy and Yoav Goldberg. Dependency-based word embeddings. In ACL 2014, pp. 302–308, 2014.
  16. 16.Minh-Thang Luong, Hieu Pham, and Christopher D Manning. Effective approaches to attention-based neural machine translation. Empirical Methods in Natural Language Processing, 2015.
  17. 17.Ankur P Parikh, Hoifung Poon, and Kristina Toutanova. Grounded semantic parsing for complex knowledge extraction. In Proceedings of North American Chapter of the Association for Computational Linguistics, pp. 756–766, 2015.
  18. 18.Kristina Toutanova, Dan Klein, Christopher D Manning, and Yoram Singer. Feature-rich part-of-speech tagging with a cyclic dependency network. In Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology-Volume 1, pp. 173–180. Association for Computational Linguistics, 2003.
  19. 19.Kristina Toutanova, Xi Victoria Lin, and Wen-tau Yih. Compositional learning of embeddings for relation paths in knowledge bases and text. In ACL, 2016.
  20. 20.David Weiss, Chris Alberti, Michael Collins, and Slav Petrov. Structured training for neural network transition-based parsing. Annual Meeting of the Association for Computational Linguistics, 2015.

Citation

MLA
Dozat, T., and C. D. Manning. “Deep Biaffine Attention for Neural Dependency Parsing”. arXiv, 2016, http://arxiv.org/abs/1611.01734v3.
APA
Dozat, T., & Manning, C. D. (2016). Deep Biaffine Attention for Neural Dependency Parsing. arXiv. http://arxiv.org/abs/1611.01734v3
Chicago
Dozat, T., and C. D. Manning. 2016. “Deep Biaffine Attention for Neural Dependency Parsing”. arXiv. http://arxiv.org/abs/1611.01734v3.
Harvard
Dozat, T. and Manning, C.D. (2016) “Deep Biaffine Attention for Neural Dependency Parsing”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1611.01734v3.
Vancouver
1. Dozat T, Manning CD (2016) Deep Biaffine Attention for Neural Dependency Parsing. arXiv

BibTeX

@article{dozat2016deep,
  title = {Deep Biaffine Attention for Neural Dependency Parsing},
  author = {Dozat, Timothy and Manning, Christopher D.},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1611.01734v3},
  eprint = {1611.01734}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors