Toward an Architecture for Never-Ending Language Learning

Andrew CarlsonJustin BetteridgeBryan KisielBurr SettlesEstevam HruschkaTom Mitchell

article2010AAAI2,265 citations

Proposes an architecture and design principles for an autonomous agent that continuously reads the web to populate a structured knowledge base while improving its own information extraction methods over time through coupled semi-supervised learning.

Listen

Building automated artificial intelligence systems that can continuously acquire structured knowledge from the vast expanse of the web is a crucial challenge in computer science. Most conventional extraction systems rely on static, supervised training, which limits their adaptability and scope. To overcome these constraints, the article evaluates the design and long-term viability of a continuous learning architecture capable of extracting factual knowledge around the clock while systematically improving its own reading and extraction techniques over time.

The researchers developed and evaluated an autonomous prototype called the Never-Ending Language Learner (NELL). The system coordinates an ensemble of four complementary extraction components that process free text, HTML tables, word structure features, and relational inference rules. These subsystems propose facts to a central Knowledge Integrator, which verifies and promotes candidates based on multi-source confidence and ontological constraints such as mutual exclusivity. Starting with a minimal seed ontology of 123 categories and 55 relations, NELL executed across 66 automated iterations over a continuous 67-day deployment analyzing millions of web pages.

The evaluation yielded several key findings regarding system performance and sustainability. First, the system maintained a consistent pace of discovery, acquiring a cumulative total of 242,453 new beliefs, comprising 95 percent category instances and 5 percent relational facts. Second, an audit of promoted beliefs revealed an overall precision rate of approximately 74 percent across the full deployment. Third, the system demonstrated a progressive decline in precision over time, dropping from 90 percent in early iterations (1–22) to 71 percent in intermediate iterations (23–44), and finally to 57 percent in later rounds (45–66). Fourth, the multi-component architecture proved effective at scale, with more than half of all promoted beliefs requiring verification from multiple independent extraction methods to meet acceptance thresholds.

These findings indicate that continuous, semi-supervised extraction from web data is technically viable and capable of rapidly generating substantial knowledge bases with minimal initial labeling. However, the steady erosion in precision over extended runs highlights the risk of error propagation and conceptual drift. When extraction components share correlated misinterpretations—such as confusing web tracking cookies with baked goods—mistakes reinforce each other and gradually degrade data reliability. This dynamic confirms that autonomous learners must incorporate ongoing constraints to maintain output quality over multi-month operating timelines.

To maintain high precision without sacrificing autonomy, future deployments should integrate targeted, active human oversight. The authors recommend dedicating 10 to 15 minutes of human interaction per day, not to label raw facts manually, but to review uncertain high-level extraction patterns and relational inference rules before errors compound. Additional recommended priorities include transitioning from string-based representations to entity-level modeling, automating the discovery of new categories, and applying more rigorous probabilistic filtering across all sub-components.

Confidence in these findings is supported by a large-scale, 67-day empirical deployment using real-world web corpora and independent human audits. A primary limitation remains the lack of temporal modeling, meaning the current system cannot distinguish between past and present facts. In addition, domains heavily contaminated with web spam, such as online gaming, showed disproportionately steep accuracy drops. Stakeholders relying on similar continuous learning architectures should account for these error accumulation risks and validate specialized domains before fully automating critical data pipelines.

Cover for Toward an Architecture for Never-Ending Language Learning

Abstract

We consider here the problem of building a never-ending language learner; that is, an intelligent computer agent that runs forever and that each day must (1) extract, or read, information from the web to populate a growing structured knowledge base, and (2) learn to perform this task better than on the previous day. In particular, we propose an approach and a set of design principles for such an agent, describe a partial implementation of such a system that has already learned to extract a knowledge base containing over 242,000 beliefs with an estimated precision of 74% after running for 67 days, and discuss lessons learned from this preliminary attempt to build a never-ending learning agent.

Table of Contents

  • Introduction
  • Approach
  • Related Work
  • Implementation
  • Experimental Evaluation
  • Methodology
  • Results
  • Discussion
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Never-Ending Language Learner Architecture and Coupled Semi-Supervised Learning Framework

    model/method

    The Never-Ending Language Learner (NELL) architecture is designed for lifelong, continuous knowledge extraction from the web. The system operates 24/7 in an iterative loop performing two core tasks: extracting information from unstructured and semi-structured web sources to populate a structured knowledge base (KB), and continuously learning to improve extraction accuracy using accumulated beliefs.

    The architecture is governed by five foundational design principles:

    1. Uncorrelated Error Subsystems: Incorporating diverse extraction components with distinct error profiles so that the probability of joint erroneous extractions is minimized as the product of individual component error probabilities.
    2. Inter-related Knowledge Types: Concurrently learning category instances (entity types) and relation instances across categories, while also inferring new relation instances from existing KB facts.
    3. Coupled Semi-Supervised Learning: Constraining learning via an ontology that specifies subset hierarchies, mutual exclusion between categories (e.g., athlete is mutually exclusive with city), and argument type restrictions on relations.
    4. Candidate vs. Belief Separation: Maintaining a two-tier knowledge representation distinguishing unverified candidate facts (with provenance and confidence scores) from high-confidence promoted beliefs.
    5. Uniform Knowledge Representation: Using a shared frame-based repository across all learning, inference, and integration modules.

    Each iteration of NELL approximates an Expectation-Maximization (EM) process: the E-step estimates candidate facts and updates beliefs in the shared KB via a Knowledge Integrator, while the M-step retrains the extraction and classification subsystems on the updated belief set.

  2. Knowl 2 — Knowledge Integrator Promotion Algorithm

    algorithm

    The Knowledge Integrator (KI) evaluates candidate facts proposed by extraction subsystems and decides which candidates to promote to belief status in the knowledge base while enforcing ontological constraints.

    Input: Current knowledge base KB\mathcal{KB} containing promoted beliefs B\mathcal{B}, set of candidate facts C\mathcal{C} proposed by extractors with associated posterior probabilities Pe(c)P_e(c) and source justifications, ontology O\mathcal{O} specifying mutual exclusion pairs M\mathcal{M} and relation argument types T\mathcal{T}, promotion limit L=250L = 250 per predicate per iteration.
    Output: Updated belief set B′\mathcal{B}'.
    for each predicate p∈Op \in \mathcal{O} do
        promoted_count ←0\leftarrow 0
        for each candidate fact c∈Cc \in \mathcal{C} proposed for predicate pp do
            if promoted_count ≥L\ge L then
                break
            end if
            valid ←true\leftarrow \text{true}
            if pp is a category and cc asserts p(e)p(e) then
                for each category mm such that (p,m)∈M(p, m) \in \mathcal{M} do
                    if m(e)∈Bm(e) \in \mathcal{B} then
                        valid ←false\leftarrow \text{false}
                    end if
                end for
            else if pp is a relation and cc asserts p(e1,e2)p(e_1, e_2) with argument types (T1,T2)∈T(T_1, T_2) \in \mathcal{T} then
                if e1e_1 is not a candidate for T1T_1 or ∃m1∈M(T1)\exists m_1 \in \mathcal{M}(T_1) such that m1(e1)∈Bm_1(e_1) \in \mathcal{B} then
                    valid ←false\leftarrow \text{false}
                end if
                if e2e_2 is not a candidate for T2T_2 or ∃m2∈M(T2)\exists m_2 \in \mathcal{M}(T_2) such that m2(e2)∈Bm_2(e_2) \in \mathcal{B} then
                    valid ←false\leftarrow \text{false}
                end if
            end if
            if valid is true then
                has_high_single_confidence ←∃ extractor e such that Pe(c)>0.9\leftarrow \exists \text{ extractor } e \text{ such that } P_e(c) > 0.9
                has_multiple_sources ←(count of extractors proposing c≥2)\leftarrow (\text{count of extractors proposing } c \ge 2)
                if has_high_single_confidence or has_multiple_sources then
                    B←B∪{c}\mathcal{B} \leftarrow \mathcal{B} \cup \{c\}
                    promoted_count ←promoted_count+1\leftarrow \text{promoted\_count} + 1
                end if
            end if
        end for
    end for
    return B\mathcal{B}
  3. Knowl 3 — Extraction and Inference Subsystems in NELL

    model/method

    NELL employs four distinct subsystem components that extract, classify, and infer facts for categories and relations:

    1. Coupled Pattern Learner (CPL): An unstructured free-text extractor. CPL analyzes co-occurrence statistics between noun phrases and part-of-speech contextual patterns across a large parsed text corpus (e.g., 2 billion sentences). It induces text extraction patterns for categories and relations, filtering out overly general patterns using mutual exclusion constraints between predicates.
    2. Coupled SEAL (CSEAL): A semi-structured web extractor. CSEAL constructs search engine queries by sub-sampling known beliefs for each predicate. It crawls matching web pages, learns HTML table and list extraction wrappers, and extracts candidate entities and relation pairs, using mutual exclusion relations to filter out overly general wrappers.
    3. Coupled Morphological Classifier (CMC): An entity classifier composed of binary L2L_2-regularized logistic regression models (one per category). CMC extracts morphological features (including character n-grams, capitalization, affixes, and part-of-speech tags) from noun phrases. For categories with ≥100\ge 100 promoted beliefs, it inspects candidates proposed by other subsystems and classifies up to 30 new beliefs per predicate per iteration with posterior probability ≥0.75\ge 0.75, using mutually exclusive categories as negative training data.
    4. Rule Learner (RL): A first-order relational learning component based on FOIL. RL mines probabilistic Horn clauses over relation beliefs in the knowledge base to infer unstated relation instances from existing facts.
  4. Knowl 4 — Candidate Fact Probability Assignment in CPL and CSEAL

    equation

    In NELL's Coupled Pattern Learner (CPL) and Coupled SEAL (CSEAL) extraction modules, the heuristic probability P(x∈p)P(x \in p) assigned to an extracted candidate fact xx for a predicate pp is calculated as:

    P(x∈p)=1−0.5cP(x \in p) = 1 - 0.5^c

    where:

    • c∈N≥0c \in \mathbb{N}_{\ge 0} denotes the count of independent extraction mechanisms confirming the instance xx for predicate pp.
    • In CPL, cc is the number of distinct promoted textual patterns that extract candidate xx.
    • In CSEAL, cc is the number of distinct unfiltered HTML wrappers (mined from web lists and tables) that extract candidate xx.

    A single pattern or wrapper (c=1c = 1) yields P=0.5P = 0.5, two (c=2c = 2) yield P=0.75P = 0.75, three (c=3c = 3) yield P=0.875P = 0.875, and four or more (c≥4c \ge 4) yield P≥0.9375P \ge 0.9375, which exceeds the Knowledge Integrator's single-source high-confidence threshold of 0.90.9.

  5. Knowl 5 — Knowledge Accumulation and Empirical Precision Over 66 Iterations

    data/table

    NELL was evaluated during an uninterrupted 67-day run comprising 66 iterations, initialized with 123 categories and 55 relations (each provided with 10–15 seed instances). Precision was measured by human judges evaluating random samples of 100 promoted beliefs within each 22-iteration period.

    Iteration Span Estimated Precision (%) Promoted Beliefs Count
    Iterations 1–22 90 88,502
    Iterations 23–44 71 77,835
    Iterations 45–66 57 76,116
    Total / Weighted Average 74 242,453

    NELL accumulated 242,453 beliefs (95% category instances, 5% relation instances) with an overall weighted precision of 74%. While the volume of newly promoted facts remained steady across iterations (76,000–89,000 per 22-iteration period), precision declined over time due to the exhaustion of easily extractable patterns and the compounding of errors in bootstrap semi-supervised learning.

  6. Knowl 6 — Provenance and Multi-Source Promotion Distribution

    empirical result

    After 66 iterations and 242,453 promoted beliefs, the distribution of belief provenance across NELL's four subsystem extractors (Coupled Pattern Learner [CPL], Coupled SEAL [CSEAL], Coupled Morphological Classifier [CMC], and Rule Learner [RL]) demonstrated the central role of multi-source verification:

    • Beliefs promoted from single-source high confidence (P>0.9P > 0.9):
      • CSEAL alone: 51,987 beliefs
      • CPL alone: 48,786 beliefs
      • RL alone: 1,914 beliefs
      • CMC alone: 423 beliefs
    • Beliefs promoted from overlapping multi-source evidence:
      • CSEAL and CMC: 58,880 beliefs
      • CPL and CMC: 45,403 beliefs
      • CPL and CSEAL: 34,447 beliefs
      • CMC and RL: 509 beliefs
      • CSEAL and RL: 78 beliefs
      • CPL and RL: 25 beliefs

    More than half of all beliefs in the knowledge base were promoted because multiple extraction components independently proposed the candidate with lower individual confidence, validating the multi-view semi-supervised learning strategy.

  7. Knowl 7 — Predicate-Specific Precision and Extraction Dynamics

    data/table

    Human evaluation of sampled beliefs (25 instances per time period per predicate) across 7 representative categories and 7 relations highlights variation in precision and promotion volume across different semantic domains.

    Estimated Precision (%) Number of Promotions
    Predicate Iter 1–22 Iter 23–44 Iter 45–66 Iter 1–22 Iter 23–44 Iter 45–66
    Categories
    cardGame 40 20 0 584 552 2,472
    city 92 80 96 4,311 3,362 1,002
    magazine 96 68 80 1,235 788 664
    recordLabel 100 100 100 1,384 890 748
    restaurant 96 88 92 242 568 523
    scientist 96 100 100 768 1 404
    vertebrate 100 100 96 1,196 1,362 714
    Relations
    athletePlaysForTeam 100 100 100 113 304 39
    ceoOfCompany 100 100 100 82 8 9
    coachesTeam 100 100 100 196 121 12
    productType 28 44 20 35 156 195
    teamPlaysAgainstTeam 96 100 100 283 553 232
    teamPlaysSport 100 100 86 79 158 14
    teamWonTrophy 88 72 44 119 104 174

    Most predicates sustained precision >90%> 90\%. Extreme failure modes occurred in cardGame (dropping to 0% precision due to web spam causing parsing errors on promotional text strings) and productType (20--44% precision because an overly broad item type constraint allowed associative noun pairs like ("Photoshop", "graphics") to pass type checking).

  8. Knowl 8 — Probabilistic Horn Clause Induction in the Rule Learner

    model/method

    The Rule Learner (RL) applies inductive logic programming (FOIL) over relation instances in the knowledge base to discover probabilistic Horn clauses that infer new relation instances from existing beliefs. RL executes periodically (every 10 iterations).

    Learned rules specify a conditional probability P(Consequent∣Antecedents)P(\text{Consequent} \mid \text{Antecedents}) representing the empirical regularity over KB instances. Examples of induced rules include:

    • 0.95:athletePlaysSport(X,basketball)⇐athleteInLeague(X,NBA)0.95: \text{athletePlaysSport}(X, \text{basketball}) \Leftarrow \text{athleteInLeague}(X, \text{NBA})
    • 0.91:teamPlaysInLeague(X,NHL)⇐teamWonTrophy(X,Stanley Cup)0.91: \text{teamPlaysInLeague}(X, \text{NHL}) \Leftarrow \text{teamWonTrophy}(X, \text{Stanley Cup})
    • 0.90:athleteInLeague(X,Y)⇐athletePlaysForTeam(X,Z)∧teamPlaysInLeague(Z,Y)0.90: \text{athleteInLeague}(X, Y) \Leftarrow \text{athletePlaysForTeam}(X, Z) \land \text{teamPlaysInLeague}(Z, Y)
    • 0.88:cityInState(X,Y)⇐cityCapitalOfState(X,Y)∧cityInCountry(X,USA)0.88: \text{cityInState}(X, Y) \Leftarrow \text{cityCapitalOfState}(X, Y) \land \text{cityInCountry}(X, \text{USA})

    In experimental runs, RL produced an average of 66.5 novel rules per 10 iterations with a 92% human approval rate during spot-checking. Approximately 12% of approved rules inferred candidate facts not proposed by any other rule, generating an average of 69.5 novel instances per rule.

  9. Knowl 9 — Semantic Drift and Correlated Error Modes in Multi-Component Extraction

    limitation

    While NELL assumes that subsystem extractors make uncorrelated errors, semi-supervised bootstrapping is subject to semantic drift when extraction components develop correlated failure patterns:

    1. Correlated Extractor Errors: For the category bakedGood, the Coupled Pattern Learner (CPL) learned the context pattern "X are enabled in" from the seed instance "cookies". This caused CPL to extract "persistent cookies" as a candidate bakedGood. Concurrently, the Coupled Morphological Classifier (CMC) placed high positive weight on the suffix/word "cookies", assigning high posterior probability to "persistent cookies". Because both components produced false positives on the same phrase, the Knowledge Integrator promoted "persistent cookies" to a belief.
    2. Irreversible Belief Promotion: Once candidate facts are promoted to beliefs, NELL never demotes them. Erroneous promotions permanently contaminate the training pool for future iterations.
    3. String-Level Knowledge Representation: Storing knowledge as surface noun strings rather than disambiguated entity identifiers leads to polysemy errors and prevents entity resolution.
    4. Lack of Temporal Scoping: Facts lack temporal metadata, causing historically true assertions (e.g., former coaches or past affiliations) to be conflated with current truths.

Coverage note — The knowls comprehensively cover NELL's architectural principles, subsystem models (CPL, CSEAL, CMC, RL, KI), probability formulas, empirical precision trajectories, provenance distributions, predicate-level breakdowns, and error modes/limitations. Omitted are standard third-party software configuration details (such as OpenNLP and Tokyo Cabinet tool settings) and the qualitative list of all individual patterns/rules learned.

References

  1. 1.Anderson, J. R.; Byrne, M. D.; Douglass, S.; Lebiere, C.; and Qin, Y. 2004. An integrated theory of the mind. Psychological Review 111(4):1036–1050.
  2. 2.Banko, M., and Etzioni, O. 2007. Strategies for lifelong knowledge extraction from the web. In Proc. of K-CAP.
  3. 3.Blum, A., and Mitchell, T. 1998. Combining labeled and unlabeled data with co-training. In Proc. of COLT.
  4. 4.Callan, J., and Hoy, M. 2009. Clueweb09 data set. http://boston.lti.cs.cmu.edu/Data/clueweb09/.
  5. 5.Carbonell, J.; Etzioni, O.; Gil, Y.; Joseph, R.; Knoblock, C.; Minton, S.; and Veloso, M. 1991. PRODIGY: an integrated architecture for planning and learning. SIGART Bull. 2(4):51–55.
  6. 6.Carlson, A.; Betteridge, J.; Wang, R. C.; Jr., E. R. H.; and Mitchell, T. M. 2010. Coupled semi-supervised learning for information extraction. In Proc. of WSDM.
  7. 7.Caruana, R. 1997. Multitask learning. Machine Learning 28:41–75.
  8. 8.Chang, M.-W.; Ratinov, L.-A.; and Roth, D. 2007. Guiding semi-supervision with constraint-driven learning. In Proc. of ACL.
  9. 9.Collins, M., and Singer, Y. 1999. Unsupervised models for named entity classification. In Proc. of EMNLP.
  10. 10.Curran, J. R.; Murphy, T.; and Scholz, B. 2007. Minimising semantic drift with mutual exclusion bootstrapping. In Proc. of PACLING.
  11. 11.Downey, D.; Etzioni, O.; and Soderland, S. 2005. A probabilistic model of redundancy in information extraction. In Proc. of IJCAI.
  12. 12.Druck, G.; Settles, B.; and McCallum, A. 2009. Active learning by labeling features. In Proc. of EMNLP.
  13. 13.Erman, L.; Hayes-Roth, F.; Lesser, V.; and Reddy, D. 1980. The HEARSAY-II speech-understanding system: Integrating knowledge to resolve uncertainty. Computing Surveys 12(2):213–253.
  14. 14.Etzioni, O.; Cafarella, M.; Downey, D.; Popescu, A.-M.; Shaked, T.; Soderland, S.; Weld, D. S.; and Yates, A. 2004. Methods for domain-independent information extraction from the web: an experimental comparison. In Proc. of AAAI.
  15. 15.Hearst, M. A. 1992. Automatic acquisition of hyponyms from large text corpora. In Proc. of COLING.
  16. 16.Laird, J.; Newell, A.; and Rosenbloom, P. 1987. SOAR: An architecture for general intelligence. Artif. Intel. 33:1–64.
  17. 17.Langley, P.; McKusick, K. B.; Allen, J. A.; Iba, W. F.; and Thompson, K. 1991. A design for the ICARUS architecture. SIGART Bull. 2(4):104–109.
  18. 18.Lenat, D. B. 1983. Eurisko: A program that learns new heuristics and domain concepts. Artif. Intel. 21(1-2):61–98.
  19. 19.Mitchell, T. M.; Allen, J.; Chalasani, P.; Cheng, J.; Etzioni, O.; Ringuette, M. N.; and Schlimmer, J. C. 1991. Theo: A framework for self-improving systems. Arch. for Intelligence 323–356.
  20. 20.Nahm, U. Y., and Mooney, R. J. 2000. A mutually beneficial integration of data mining and information extraction. In Proc. of AAAI.
  21. 21.Paşca, M.; Lin, D.; Bigham, J.; Lifchits, A.; and Jain, A. 2006. Names and similarities on the web: fact extraction in the fast lane. In Proc. of ACL.
  22. 22.Pennacchiotti, M., and Pantel, P. 2009. Entity extraction via ensemble semantics. In Proc. of EMNLP.
  23. 23.Quinlan, J. R., and Cameron-Jones, R. M. 1993. Foil: A midterm report. In Proc. of ECML.
  24. 24.Riloff, E., and Jones, R. 1999. Learning dictionaries for information extraction by multi-level bootstrapping. In Proc. of AAAI.
  25. 25.Settles, B. 2009. Active learning literature survey. Computer Sciences Technical Report 1648, University of Wisconsin–Madison.
  26. 26.Thrun, S., and Mitchell, T. 1995. Lifelong robot learning. In Robotics and Autonomous Systems, volume 15, 25–46.
  27. 27.Wang, R. C., and Cohen, W. W. 2009. Character-level analysis of semi-structured documents for set expansion. In Proc. of EMNLP.
  28. 28.Yang, X.; Kim, S.; and Xing, E. 2009. Heterogeneous multitask learning with joint sparsity constraints. In NIPS 2009.
  29. 29.Yangarber, R. 2003. Counter-training in discovery of semantic patterns. In Proc. of ACL.
  30. 30.Yarowsky, D. 1995. Unsupervised word sense disambiguation rivaling supervised methods. In Proc. of ACL.

Citation

MLA
Carlson, A., et al. “Toward an Architecture for Never-Ending Language Learning”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 24, no. 1, 2010, pp. 1306–13, https://doi.org/10.1609/aaai.v24i1.7519.
APA
Carlson, A., Betteridge, J., Kisiel, B., Settles, B., Hruschka, E., & Mitchell, T. (2010). Toward an Architecture for Never-Ending Language Learning. Proceedings of the AAAI Conference on Artificial Intelligence, 24(1), 1306–1313. https://doi.org/10.1609/aaai.v24i1.7519
Chicago
Carlson, A., J. Betteridge, B. Kisiel, B. Settles, E. Hruschka, and T. Mitchell. 2010. “Toward an Architecture for Never-Ending Language Learning”. Proceedings of the AAAI Conference on Artificial Intelligence 24 (1): 1306–13. https://doi.org/10.1609/aaai.v24i1.7519.
Harvard
Carlson, A. et al. (2010) “Toward an Architecture for Never-Ending Language Learning”, Proceedings of the AAAI Conference on Artificial Intelligence, 24(1), pp. 1306–1313. Available at: https://doi.org/10.1609/aaai.v24i1.7519.
Vancouver
1. Carlson A, Betteridge J, Kisiel B, Settles B, Hruschka E, Mitchell T (2010) Toward an Architecture for Never-Ending Language Learning. Proceedings of the AAAI Conference on Artificial Intelligence 24:1306–1313

BibTeX

@article{Carlson_2010, title={Toward an Architecture for Never-Ending Language Learning}, volume={24}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/aaai.v24i1.7519}, DOI={10.1609/aaai.v24i1.7519}, number={1}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Carlson, Andrew and Betteridge, Justin and Kisiel, Bryan and Settles, Burr and Hruschka, Estevam and Mitchell, Tom}, year={2010}, month=July, pages={1306–1313} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF