A Fast and Accurate Dependency Parser using Neural Networks
Danqi ChenChristopher D. Manning
Presents a greedy transition-based dependency parser powered by a neural network that replaces millions of sparse indicator features with compact dense representations, achieving high parsing accuracy while processing over 1,000 sentences per second.
Automated sentence analysis, known as dependency parsing, is essential for language-driven technology applications such as search, translation, and information extraction. Traditional transition-based parsers rely on millions of hand-crafted, sparse indicator features. This legacy design creates significant bottlenecks: extracting and querying these complex features consumes over 95% of total runtime, manual templates fail to capture all necessary word combinations, and models generalize poorly on unseen data.
The article evaluates whether replacing manual, sparse features with a compact neural network utilizing dense vector representations can simultaneously improve parsing accuracy and execution speed.
To demonstrate this, the researchers built a greedy transition-based parser powered by a single-hidden-layer neural network. The model maps words, grammatical tags (part-of-speech tags), and relationship labels into compact, continuous vector representations (embeddings). It employs a novel cubic activation function to naturally capture multi-word interactions and utilizes a pre-computation caching technique during runtime to eliminate repetitive mathematical calculations. The system was evaluated across standard English (Penn Treebank) and Chinese (Chinese Treebank) benchmarks covering tens of thousands of sentences.
The evaluation yielded several key findings. First, the neural parser achieved a parsing throughput of 654 to 1,013 sentences per second on English text, processing sentences roughly 20 times faster than standard baseline transition parsers and up to 100 times faster than graph-based alternatives like MSTParser. Second, it delivered an approximate 2% improvement in attachment accuracy across both English and Chinese benchmarks compared to standard greedy parsers, reaching 92.0% unlabeled accuracy on standard English tests. Third, the novel cubic activation function alone contributed a 0.8% to 1.2% accuracy improvement over standard neural activation functions such as tanh and sigmoid. Finally, incorporating part-of-speech tag embeddings drove substantial performance gains, boosting accuracy by 1.7% in English and nearly 10% in Chinese.
These findings prove that natural language processing pipelines do not need to trade accuracy for processing speed. By slashing feature computation overhead and learning rich compact features automatically, this neural approach delivers substantial operational cost and latency savings for high-volume text processing systems while outperforming manual feature engineering.
Organizations deploying large-scale text parsing systems should consider transitioning from sparse, template-based architectures to compact neural classifiers using dense representations. Engineering teams should also adopt runtime pre-computation caching for frequent vocabulary words to unlock maximum throughput. As an immediate next step, developers can explore integrating this neural classifier with search-based decoding methods (such as beam search) or incorporating additional positional features to further improve parsing quality.
Confidence in these findings is high given the rigorous evaluation across multiple languages, dependency frameworks, and standard benchmarks. However, leaders should note that greedy parsing can be vulnerable to early decision errors propagating down the sentence, and the study was conducted within controlled, standard treebank benchmarks where sentences are largely well-formed.
- Paper: Natural Language Processing (almost) from Scratch, Ronan Collobert et al. (2011). Read this account of neural NLP replacing hand-crafted features with learned dense representations first; it establishes the feature-learning approach that this parser adapts to dependency parsing.
- Paper: Deep Biaffine Attention for Neural Dependency Parsing, Timothy Dozat et al. (2016). This later dependency parser advances the neural parsing approach with deep biaffine attention, making it a direct next step from the source’s transition-based neural model.
