keyword
chart parser
A chart parser is a natural language processing algorithm that analyzes the grammatical structure of sentences by storing intermediate parsing results in a specialized data structure called a chart. By applying dynamic programming principles, it records both completed syntactic constituents and incomplete grammatical hypotheses as entries that span specific segments of the input text. This memoization technique prevents redundant recomputation of shared substructures during the parsing process, allowing efficient handling of grammatical ambiguity common in natural language. Chart parsers can operate using top-down, bottom-up, or mixed search strategies, and they serve as the foundation for prominent syntactic parsing methods such as the Earley algorithm and the Cocke-Younger-Kasami algorithm.
2 items

NLTK: The Natural Language Toolkit
Steven Bird
Why you should read this
Presents an open-source Python suite integrating symbolic and statistical natural language processing algorithms with annotated corpora to teach computational linguistics through hands-on model implementation.
NLTK, the Natural Language Toolkit, is a suite of open source program modules, tutorials and problem sets, providing ready-to-use computational linguistics courseware. NLTK covers symbolic and statistical natural language processing, and is interfaced to annotated corpora. Students augment and replace existing components, learn structured programming by example, and manipulate sophisticated models from the outset.
Added
2026-09-10

A Maximum-Entropy-Inspired Parser
Eugene Charniak
Why you should read this
Proposes a generative, maximum-entropy-inspired parser that achieves a state-of-the-art 13% reduction in error rate on the Penn Treebank by leveraging a lexicalized Markov grammar and the strategic prediction of a lexical head's pre-terminal before the head itself.
We present a new parser for parsing down to Penn tree-bank style parse trees that achieves 90.1% average precision/recall for sentences of length 40 and less, and 89.5% for sentences of length 100 and less when trained and tested on the previously established [5,9,10,15,17] “standard” sections of the Wall Street Journal tree-bank. This represents a 13% decrease in error rate over the best single-parser results on this corpus [9]. The major technical innovation is the use of a “maximum-entropy-inspired” model for conditioning and smoothing that let us successfully to test and combine many different conditioning events. We also present some partial results showing the effects of different conditioning information, including a surprising 2% improvement due to guessing the lexical head’s pre-terminal before guessing the lexical head.
Added
2026-02-21
