Natural language processing: state of the art, current trends and challenges

Diksha KhuranaAditya C KoliKiran KhatterSukhdev Singh

article2017Multimedia tools and applications1,568 citations

Presents a comprehensive historical and technical overview of natural language processing by breaking down its architectural levels, natural language generation components, modern real-world applications, and open research challenges.

Listen

As digital textual information rapidly expands, organizations face substantial challenges in effectively interpreting, organizing, and utilizing human language computationally. Natural language processing bridges this gap by enabling computers to understand and generate human language, eliminating the need for human users to interact solely through complex programming syntax.

The article provides a broad review of the field, examining its foundational linguistic levels, historical evolution, algorithmic approaches, core applications, and recent commercial deployments.

The review evaluates established and emerging technologies across multiple domains. It traces historical milestones from early machine translation in the late 1940s to contemporary statistical and unsupervised machine learning methods, categorizing these tools into Natural Language Understanding and Natural Language Generation.

The key findings highlight that human language analysis spans seven interdependent linguistic levels: phonology, morphology, lexical analysis, syntax, semantics, discourse, and pragmatics. Processing methods have shifted from rule-based symbolic frameworks toward statistical, machine-learning, and deep neural approaches, with established methods like phrase chunking reaching high benchmark accuracy around 94.3% F1 score. Furthermore, practical implementations now span diverse sectors, including machine translation, automated document summarization, spam classification, medical record structuring, and enterprise dialogue systems.

These findings demonstrate that automated language systems can dramatically cut operational costs and labor requirements by automating high-volume document indexing, data classification, and regulatory compliance. Moreover, modern text-processing pipelines improve response timelines and user access across customer support, healthcare informatics, and cross-border communications, though ambiguity across syntax and informal web text remains an operational risk.

Organizations should adopt modular pipelines to allow flexible updates as component technologies improve, while pairing statistical methods with domain-specific knowledge to manage linguistic ambiguities. Enterprise teams must also evaluate privacy trade-offs when selecting conversational interfaces, as seen in proprietary messaging channels avoiding broad third-party platform access.

The article is a broad qualitative literature review rather than a single empirical benchmark study. Consequently, readers should exercise caution regarding specific performance claims in noisy real-world environments, as informal language, multi-language mixing, and cross-sentence discourse still require further development and evaluation.

arXiv: 1708.05148
Cover for Natural language processing: state of the art, current trends and challenges

Abstract

Natural language processing (NLP) has recently gained much attention for representing and analysing human language computationally. It has spread its applications in various fields such as machine translation, email spam detection, information extraction, summarization, medical, and question answering etc. The paper distinguishes four phases by discussing different levels of NLP and components of Natural Language Generation (NLG) followed by presenting the history and evolution of NLP, state of the art presenting the various applications of NLP and current trends and challenges.

Table of Contents

  • 1. Introduction
  • 2. Levels of NLP
  • 1. Phonology
  • 2. Morphology
  • 3. Natural Language Generation
  • 4. History of NLP
  • 6. Applications of NLP
  • 6.1 Machine Translation
  • 6.2 Text Categorization
  • 6.3 Spam Filtering
  • 6.4 Information Extraction
  • 7. Approaches
  • 7.1 Hidden Markov Model (HMM)
  • 7.2 Naive Bayes Classifiers
  • 8. NLP in Talk
  • 8.1 ACE Powered GDPR Robot Launched by RAVN Systems [104]
  • 8.2 Eno A Natural Language Chatbot Launched by Capital One [105]
  • 8.3 Future of BI in Natural Language Processing [106]
  • 8.4 Using Natural Language Processing and Network Analysis to Develop a Conceptual Framework for Medication Therapy Management Research [107]
  • 8.5 Meet the Pilot, world's first language translating earbuds [108]
  • References

Knowls

  1. Knowl 1 — Hierarchical Levels of Linguistic Analysis in Natural Language Processing

    model/method

    Natural language processing can be structured across seven hierarchical levels of linguistic analysis to represent, interpret, or generate human language:

    1. Phonology: The systematic arrangement and study of speech sounds used to encode meaning in spoken language.
    2. Morphology: The study of word formation and internal word structure by decomposing words into morphemes (the smallest meaningful linguistic units).
    3. Lexical Analysis: The interpretation of individual words, including assigning part-of-speech (POS) tags to each word based on local context and representing word-level senses.
    4. Syntactic Analysis: The examination of sentence structure using formal grammars and parsers to uncover structural dependency relationships and grammatical hierarchy between words.
    5. Semantic Analysis: The determination of literal meaning through interactions among word senses in a sentence, including word sense disambiguation within context.
    6. Discourse Analysis: The processing of multi-sentence text units focusing on inter-sentential connections, coreference, and document-level organization.
    7. Pragmatic Analysis: The inference of intended meaning beyond literal text by integrating real-world knowledge, situational context, conversational goals, and speaker intentions.
  2. Knowl 2 — Three-Stage Architecture for Natural Language Generation

    model/method

    Natural Language Generation (NLG) transforms non-linguistic internal representations into fluent, meaningful natural language text across three primary computational phases:

    1. Content Planning (Macroplanning): Selects the relevant subset of information from a knowledge base based on communicative goals, domain models, and user models, and establishes the overall discourse structure.
    2. Sentence Planning (Microplanning): Aggregates selected information units into sentence-sized segments, determines linguistic expressions, and performs lexicalisation (the selection of specific words and phrases).
    3. Surface Realization: Applies morphosyntactic and grammatical rules to map the structured sentence plans into final written natural language text or synthesized speech output.
  3. Knowl 3 — Taxonomy of Morphemes in Morphological Analysis

    definition

    In natural language morphology, words are composed of morphemes, which represent the smallest discrete units of grammatical and semantic meaning:

    • Lexical Morpheme: An independent base word that carries standalone semantic meaning (e.g., "table", "chair") and cannot be divided further.
    • Grammatical Morpheme: A morpheme that provides grammatical modification or relationship when combined with lexical morphemes (e.g., "-ed", "-ing", "-est", "-ly", "-ful").
    • Bound Morpheme: A grammatical morpheme that cannot occur in isolation and must be attached to another morpheme (e.g., inflectional affixes denoting verb tense or pluralization).
    • Derivational Morpheme: An affix combined with a base word that creates a new lexeme or changes the grammatical category of the base word.
  4. Knowl 4 — Generative versus Discriminative Modeling Paradigms in Statistical NLP

    model/method

    Statistical machine learning methods for natural language processing are categorized into generative and discriminative frameworks:

    • Generative Models: Estimate the joint probability distribution P(X,Y)P(X, Y) over observed inputs XX and targets or latent states YY. They can generate synthetic observations and encode structural domain assumptions (e.g., Naive Bayes classifiers, Hidden Markov Models). However, incorporating many overlapping or non-independent features makes joint probability estimation difficult.
    • Discriminative Models: Directly estimate the conditional posterior probability distribution P(Y∣X)P(Y \mid X) or determine decision boundaries based on observations without modeling the distribution of XX. They allow flexible use of large sets of overlapping, correlated features (e.g., Logistic Regression, Conditional Random Fields, Support Vector Machines).
  5. Knowl 5 — Multivariate Bernoulli versus Multinomial Event Models for Text Classification

    model/method

    In naive Bayes text categorization over a fixed vocabulary V={w1,w2,…,w∣V∣}V = \{w_1, w_2, \dots, w_{|V|}\}, documents are represented using two different event models:

    • Multi-variate Bernoulli Model: Documents are represented as binary vectors d=(x1,x2,…,x∣V∣)d = (x_1, x_2, \dots, x_{|V|}), where xi∈{0,1}x_i \in \{0, 1\} denotes the presence or absence of vocabulary word wiw_i in the document. Word frequency and word position are ignored.
    • Multinomial Model: Documents are represented as sequences or bags of word tokens, capturing both the presence of vocabulary words and their exact frequency counts within the document.
  6. Knowl 6 — Discourse-Level NLP Subtasks: Anaphora Resolution and Text Structure Recognition

    definition

    Discourse-level natural language processing operates across multi-sentence texts through two key subtasks:

    • Anaphora Resolution: The computational task of identifying referentially dependent words (such as pronouns or definite noun phrases) and replacing or linking them with the specific real-world entities (antecedents) to which they refer in preceding sentences.
    • Discourse and Text Structure Recognition: The identification of rhetorical and functional relationships between component sentences in a text (e.g., background, elaboration, contrast), enabling a unified structural and semantic representation of the document as a whole.
  7. Knowl 7 — Modular Data-Centric Pipeline Architecture for Multilingual Event Extraction

    model/method

    Cross-lingual event extraction across multiple source languages can be implemented using a modular, data-centric pipeline architecture structured like UNIX pipes:

    • Individual modules receive standardized data inputs, perform a specific linguistic analysis or extraction task, and generate standardized outputs that serve as inputs to downstream modules.
    • The pipeline cascades from low-level NLP operations (tokenization, morphological analysis, POS tagging) to high-level semantic components (cross-lingual named entity linking, semantic role labeling, time normalization).
    • Independent language pipelines feed their extracted outputs into a centralized system to construct event-centric knowledge graphs detailing events, participants, locations, and temporal relationships.

Coverage note — Historical summaries of early dialogue systems (e.g., BASEBALL, LUNAR, SHRDLU) and brief reviews of external commercial products (RAVN GDPR robot, Capital One Eno chatbot, Pilot translation earbuds) were omitted as they represent external background rather than original technical contributions.

References

  1. 1.Chomsky, Noam, 1965, Aspects of the Theory of Syntax, Cambridge, Massachusetts: MIT Press.
  2. 2.Rospocher, M., van Erp, M., Vossen, P., Fokkens, A., Aldabe,I., Rigau, G., Soroa, A., Ploeger, T., and Bogaard, T.(2016). Building event-centric knowledge graphs from news. Web Semantics: Science, Services and Agents on the World Wide Web, In Press.
  3. 3.Shemtov, H. (1997). Ambiguity management in natural language generation. Stanford University.
  4. 4.Emele, M. C., & Dorna, M. (1998, August). Ambiguity preserving machine translation using packed representations. In Proceedings of the 36th Annual Meeting of the Association for Computational Linguistics and 17th International Conference on Computational Linguistics-Volume 1 (pp. 365-371). Association for Computational Linguistics.
  5. 5.Knight, K., & Langkilde, I. (2000, July). Preserving ambiguities in generation via automata intersection. In AAAI/IAAI (pp. 697-702).
  6. 6.Nation, K., Snowling, M. J., & Clarke, P. (2007). Dissecting the relationship between language skills and learning to read: Semantic and phonological contributions to new vocabulary learning in children with poor reading comprehension. Advances in Speech Language Pathology, 9(2), 131-139.
  7. 7.Liddy, E. D. (2001). Natural language processing.
  8. 8.Feldman, S. (1999). NLP Meets the Jabberwocky: Natural Language Processing in Information Retrieval. ONLINE-WESTON THEN WILTON-, 23, 62-73.
  9. 9."Natural Language Processing." Natural Language Processing RSS. N.p., n.d. Web. 25 Mar. 2017
  10. 10.Hutchins, W. J. (1986). Machine translation: past, present, future (p. 66). Chichester: Ellis Horwood.
  11. 11.Hutchins, W. J. (Ed.). (2000). Early years in machine translation: memoirs and biographies of pioneers (Vol. 97). John Benjamins Publishing.
  12. 12.Green Jr, B. F., Wolf, A. K., Chomsky, C., & Laughery, K. (1961, May). Baseball: an automatic question-answerer. In Papers presented at the May 9-11, 1961, western joint IRE-AIEE-ACM computer conference (pp. 219-224). ACM.
  13. 13.Woods, W. A. (1978). Semantics and quantification in natural language question answering. Advances in computers, 17, 1-87.
  14. 14.Hendrix, G. G., Sacerdoti, E. D., Sagalowicz, D., & Slocum, J. (1978). Developing a natural language interface to complex data. ACM Transactions on Database Systems (TODS), 3(2), 105-147.
  15. 15.Alshawi, H. (1992). The core language engine. MIT press.
  16. 16.Kamp, H., & Reyle, U. (1993). Tense and Aspect. In From Discourse to Logic (pp. 483-689). Springer Netherlands.
  17. 17.Lea , W.A Trends in speech recognition , Englewoods Cliffs , NJ: Prentice Hall , 1980.
  18. 18.Young, S. J., & Chase, L. L. (1998). Speech recognition evaluation: a review of the US CSR and LVCSR programmes. Computer Speech & Language, 12(4), 263-279.
  19. 19.Sundheim, B. M., & Chinchor, N. A. (1993, March). Survey of the message understanding conferences. In Proceedings of the workshop on Human Language Technology (pp. 56-60). Association for Computational Linguistics.
  20. 20.Wahlster, W., & Kobsa, A. (1989). User models in dialog systems. In User models in dialog systems (pp. 4-34). Springer Berlin Heidelberg.
  21. 21.McKeown, K.R. Text generation , Cambridge: Cambridge University Press , 1985.
  22. 22.Small S.L., Cortell G.W., and Tanenhaus , M.K. Lexical Ambiguity Resolutions , San Mateo , CA : Morgan Kauffman, 1988.
  23. 23.Manning, C. D., & Schütze, H. (1999). Foundations of statistical natural language processing (Vol. 999). Cambridge: MIT press.
  24. 24.Mani, I., & Maybury, M. T. (Eds.). (1999). Advances in automatic text summarization (Vol. 293). Cambridge, MA: MIT press.
  25. 25.Yi, J., Nasukawa, T., Bunescu, R., & Niblack, W. (2003, November). Sentiment analyzer: Extracting sentiments about a given topic using natural language processing techniques. In Data Mining, 2003. ICDM 2003. Third IEEE International Conference on (pp. 427-434). IEEE.
  26. 26.Yi, J., Nasukawa, T., Bunescu, R., & Niblack, W. (2003, November). Sentiment analyzer: Extracting sentiments about a given topic using natural language processing techniques. In Data Mining, 2003. ICDM 2003. Third IEEE International Conference on (pp. 427-434). IEEE.
  27. 27.Tapaswi, N., & Jain, S. (2012, September). Treebank based deep grammar acquisition and Part-Of-Speech Tagging for Sanskrit sentences. In Software Engineering (CONSEG), 2012 CSI Sixth International Conference on (pp. 1-4). IEEE.
  28. 28.Ranjan, P., & Basu, H. V. S. S. A. (2003). Part of speech tagging and local word grouping techniques for natural language parsing in Hindi. In Proceedings of the 1st International Conference on Natural Language Processing (ICON 2003).
  29. 29.Diab, M., Hacioglu, K., & Jurafsky, D. (2004, May). Automatic tagging of Arabic text: From raw text to base phrase chunks. In Proceedings of HLT-NAACL 2004: Short papers (pp. 149-152). Association for Computational Linguistics.
  30. 30.Sha, F., & Pereira, F. (2003, May). Shallow parsing with conditional random fields. In Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology-Volume 1 (pp. 134-141). Association for Computational Linguistics.
  31. 31.McDonald, R., Crammer, K., & Pereira, F. (2005, October). Flexible text segmentation with structured multilabel classification. In Proceedings of the conference on Human Language Technology and Empirical Methods in Natural Language Processing (pp. 987-994). Association for Computational Linguistics.
  32. 32.Sun, X., Morency, L. P., Okanohara, D., & Tsujii, J. I. (2008, August). Modeling latent-dynamic in shallow parsing: a latent conditional model with improved inference. In Proceedings of the 22nd International Conference on Computational Linguistics-Volume 1 (pp. 841-848). Association for Computational Linguistics.
  33. 33.Ritter, A., Clark, S., & Etzioni, O. (2011, July). Named entity recognition in tweets: an experimental study. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (pp. 1524-1534). Association for Computational Linguistics.
  34. 34.Sharma, S., Srinivas, PYKL, & Balabantaray, RC (2016). Emotion Detection using Online Machine Learning Method and TLBO on Mixed Script. In Proceedings of Language Resources and Evaluation Conference 2016 (pp. 47-51).
  35. 35.Palmer, M., Gildea, D., & Kingsbury, P. (2005). The proposition bank: An annotated corpus of semantic roles. Computational linguistics, 31(1), 71-106.
  36. 36.Benson, E., Haghighi, A., & Barzilay, R. (2011, June). Event discovery in social media feeds. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume 1 (pp. 389-398). Association for Computational Linguistics.
  37. 37.Tillmann, C., Vogel, S., Ney, H., Zubiaga, A., & Sawaf, H. (1997, September). Accelerated DP based search for statistical translation. In Eurospeech.
  38. 38.Bangalore, S., Rambow, O., & Whittaker, S. (2000, June). Evaluation metrics for generation. In Proceedings of the first international conference on Natural language generation-Volume 14 (pp. 1-8). Association for Computational Linguistics
  39. 39.Nießen, S., Och, F. J., Leusch, G., & Ney, H. (2000, May). An Evaluation Tool for Machine Translation: Fast Evaluation for MT Research. In LREC
  40. 40.Papineni, K., Roukos, S., Ward, T., & Zhu, W. J. (2002, July). BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting on association for computational linguistics (pp. 311-318). Association for Computational Linguistics
  41. 41.Doddington, G. (2002, March). Automatic evaluation of machine translation quality using n-gram co-occurrence statistics. In Proceedings of the second international conference on Human Language Technology Research (pp. 138-145). Morgan Kaufmann Publishers Inc
  42. 42.Papineni, K., Roukos, S., Ward, T., & Zhu, W. J. (2002, July). BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting on association for computational linguistics (pp. 311-318). Association for Computational Linguistics
  43. 43.Doddington, G. (2002, March). Automatic evaluation of machine translation quality using n-gram co-occurrence statistics. In Proceedings of the second international conference on Human Language Technology Research (pp. 138-145). Morgan Kaufmann Publishers Inc
  44. 44.Hayes, P. J. (1992). Intelligent high-volume text processing using shallow, domain-specific techniques. Text-based intelligent systems: Current research and practice in information extraction and retrieval, 227-242.
  45. 45.Cohen, W. W. (1996, March). Learning rules that classify e-mail. In AAAI spring symposium on machine learning in information access (Vol. 18, p. 25).
  46. 46.Sahami, M., Dumais, S., Heckerman, D., & Horvitz, E. (1998, July). A Bayesian approach to filtering junk e-mail. In Learning for Text Categorization: Papers from the 1998 workshop (Vol. 62, pp. 98-105).
  47. 47.Androutsopoulos, I., Paliouras, G., Karkaletsis, V., Sakkis, G., Spyropoulos, C. D., & Stamatopoulos, P. (2000). Learning to filter spam e-mail: A comparison of a naive bayesian and a memory-based approach. arXiv preprint cs/0009009.
  48. 48.Rennie, J. (2000, August). ifile: An application of machine learning to e-mail filtering. In Proc. KDD 2000 Workshop on Text Mining, Boston, MA
  49. 49.Drucker, H., Wu, D., & Vapnik, V. N. (1999). Support vector machines for spam categorization. IEEE Transactions on Neural networks, 10(5), 1048-1054
  50. 50.Carreras, X., & Marquez, L. (2001). Boosting trees for anti-spam email filtering. arXiv preprint cs/0109015
  51. 51.BERGER, A. L., DELLA PIETRA, S. A., AND DELLA PIETRA, V. J. 1996. A maximum entropy approach to natural language processing. Computational Linguistics 22, 1, 39–71
  52. 52.Sakkis, G., Androutsopoulos, I., Paliouras, G., Karkaletsis, V., Spyropoulos, C. D., & Stamatopoulos, P. (2001). Stacking classifiers for anti-spam filtering of e-mail. arXiv preprint cs/0106040..
  53. 53.Lewis, D. D. (1998, April). Naive (Bayes) at forty: The independence assumption in information retrieval. In European conference on machine learning (pp. 4-15). Springer Berlin Heidelberg
  54. 54.McCallum, A., & Nigam, K. (1998, July). A comparison of event models for naive bayes text classification. In AAAI-98 workshop on learning for text categorization (Vol. 752, pp. 41-48).
  55. 55.McCallum, A., & Nigam, K. (1998, July). A comparison of event models for naive bayes text classification. In AAAI-98 workshop on learning for text categorization (Vol. 752, pp. 41-48).
  56. 56.Porter, M. F. (1980). An algorithm for suffix stripping. Program, 14(3), 130-137
  57. 57.Hayes, P. J. (1992). Intelligent high-volume text processing using shallow, domain-specific techniques. Text-based intelligent systems: Current research and practice in information extraction and retrieval, 227-242
  58. 58.Morin, E. (1999, August). Automatic acquisition of semantic relations between terms from technical corpora. In Proc. of the Fifth International Congress on Terminology and Knowledge Engineering-TKE’99.
  59. 59.Bondale, N., Maloor, P., Vaidyanathan, A., Sengupta, S., & Rao, P. V. (1999). Extraction of information from open-ended questionnaires using natural language processing techniques. Computer Science and Informatics, 29(2), 15-22
  60. 60.Glasgow, B., Mandell, A., Binney, D., Ghemri, L., & Fisher, D. (1998). MITA: An information-extraction approach to the analysis of free-form text in life insurance applications. AI magazine, 19(1), 59.
  61. 61.Ahonen, H., Heinonen, O., Klemettinen, M., & Verkamo, A. I. (1998, April). Applying data mining techniques for descriptive phrase extraction in digital document collections. In Research and Technology Advances in Digital Libraries, 1998. ADL 98. Proceedings. IEEE International Forum on (pp. 2-11). IEEE.
  62. 62.Zajic, D. M., Dorr, B. J., & Lin, J. (2008). Single-document and multi-document summarization techniques for email threads using sentence compression. Information Processing & Management, 44(4), 1600-1610.
  63. 63.Fattah, M. A., & Ren, F. (2009). GA, MR, FFNN, PNN and GMM based models for automatic text summarization. Computer Speech & Language, 23(1), 126-144.
  64. 64.Gong, Y., & Liu, X. (2001, September). Generic text summarization using relevance measure and latent semantic analysis. In Proceedings of the 24th annual international ACM SIGIR conference on Research and development in information retrieval (pp. 19-25). ACM.
  65. 65.Dunlavy, D. M., O’Leary, D. P., Conroy, J. M., & Schlesinger, J. D. (2007). QCS: A system for querying, clustering and summarizing documents. Information processing & management, 43(6), 1588-1605.
  66. 66.Wan, X. (2008). Using only cross-document relationships for both generic and topic-focused multi-document summarizations. Information Retrieval, 11(1), 25-49.
  67. 67.Ouyang, Y., Li, W., Li, S., & Lu, Q. (2011). Applying regression models to query-focused multi-document summarization. Information Processing & Management, 47(2), 227-237.
  68. 68.Mani, I., & Maybury, M. T. (Eds.). (1999). Advances in automatic text summarization (Vol. 293). Cambridge, MA: MIT press.
  69. 69.Riedhammer, K., Favre, B., & Hakkani-Tür, D. (2010). Long story short–global unsupervised models for keyphrase based meeting summarization. Speech Communication, 52(10), 801-815.
  70. 70.Wang, D., Zhu, S., Li, T., & Gong, Y. (2009, August). Multi-document summarization using sentence-based topic models. In Proceedings of the ACL-IJCNLP 2009 Conference Short Papers (pp. 297-300). Association for Computational Linguistics.
  71. 71.Wang, D., Zhu, S., Li, T., Chi, Y., & Gong, Y. (2011). Integrating document clustering and multidocument summarization. ACM Transactions on Knowledge Discovery from Data (TKDD), 5(3), 14.
  72. 72.Fang, H., Lu, W., Wu, F., Zhang, Y., Shang, X., Shao, J., & Zhuang, Y. (2015). Topic aspect-oriented summarization via group selection. Neurocomputing, 149, 1613-1619.
  73. 73.Sager, N., Lyman, M., Nhan, N. T., & Tick, L. J. (1995). Medical language processing: applications to patient data representation and automatic encoding. Methods of information in medicine, 34(1-2), 140-146.
  74. 74.Chi, E. C., Lyman, M. S., Sager, N., Friedman, C., & Macleod, C. (1985, November). A database of computer-structured narrative: methods of computing complex relations. In Proceedings of the Annual Symposium on Computer Application in Medical Care (p. 221). American Medical Informatics Association.
  75. 75.Grishman, R., Sager, N., Raze, C., & Bookchin, B. (1973, June). The linguistic string parser. In Proceedings of the June 4-8, 1973, national computer conference and exposition (pp. 427-434). ACM.
  76. 76.Hirschman, L., Grishman, R., & Sager, N. (1976, June). From text to structured information: automatic processing of medical reports. In Proceedings of the June 7-10, 1976, national computer conference and exposition (pp. 267-275). ACM.
  77. 77.Sager, N. (1981). Natural language information processing. Addison-Wesley Publishing Company, Advanced Book Program.
  78. 78.Lyman, M., Sager, N., Friedman, C., & Chi, E. (1985, November). Computer-structured narrative in ambulatory care: its use in longitudinal review of clinical data. In Proceedings of the Annual Symposium on Computer Application in Medical Care (p. 82). American Medical Informatics Association.
  79. 79.McCray, A. T., & Nelson, S. J. (1995). The representation of meaning in the UMLS. Methods of information in medicine, 34(1-2), 193-201.
  80. 80.McGray, A. T., Sponsler, J. L., Brylawski, B., & Browne, A. C. (1987, November). The role of lexical knowledge in biomedical text understanding. In Proceedings of the Annual Symposium on Computer Application in Medical Care (p. 103). American Medical Informatics Association.
  81. 81.McCray, A. T. (1991). Natural language processing for intelligent information retrieval. In Engineering in Medicine and Biology Society, 1991. Vol. 13: 1991., Proceedings of the Annual International Conference of the IEEE (pp. 1160-1161). IEEE.
  82. 82.McCray, A. T. (1991). Extending a natural language parser with UMLS knowledge. In Proceedings of the Annual Symposium on Computer Application in Medical Care (p. 194). American Medical Informatics Association.
  83. 83.McCray, A. T., Srinivasan, S., & Browne, A. C. (1994). Lexical methods for managing variation in biomedical terminologies. In Proceedings of the Annual Symposium on Computer Application in Medical Care (p. 235). American Medical Informatics Association.
  84. 84.McCray, A. T., & Razi, A. (1994). The UMLS Knowledge Source server. Medinfo. MEDINFO, 8, 144-147.
  85. 85.Scherrer, J. R., Revillard, C., Borst, F., Berthoud, M., & Lovis, C. (1994). Medical office automation integrated into the distributed architecture of a hospital information system. Methods of information in medicine, 33(2), 174-179.
  86. 86.Baud, R. H., Rassinoux, A. M., & Scherrer, J. R. (1992). Natural language processing and semantical representation of medical texts. Methods of information in medicine, 31(2), 117-125.
  87. 87.Lyman, M., Sager, N., Chi, E. C., Tick, L. J., Nhan, N. T., Su, Y., ... & Scherrer, J. (1989, November). Medical Language Processing for Knowledge Representation and Retrievals. In Proceedings. Symposium on Computer Applications in Medical Care (pp. 548-553). American Medical Informatics Association.
  88. 88.Nhàn, N. T., Sager, N., Lyman, M., Tick, L. J., Borst, F., & Su, Y. (1989, November). A Medical Language Processor for Two Indo-European Languages. In Proceedings. Symposium on Computer Applications in Medical Care (pp. 554-558). American Medical Informatics Association.
  89. 89.Sager, N., Lyman, M., Tick, L. J., Borst, F., Nhan, N. T., Revillard, C., ... & Scherrer, J. R. (1989). Adapting a medical language processor from English to French. Medinfo, 89, 795-799.
  90. 90.Borst, F., Sager, N., Nhàn, N. T., Su, Y., Lyman, M., Tick, L. J., ... & Scherrer, J. R. (1989). Analyse automatique de comptes rendus d'hospitalisation. In Degoulet P, Stephan JC, Venot A, Yvon PJ, rédacteurs. Informatique et Santé, Informatique et Gestion des Unités de Soins, Comptes Rendus du Colloque AIM-IF, Paris (pp. 246-56). [5]
  91. 91.Baud, R. H., Rassinoux, A. M., & Scherrer, J. R. (1991). Knowledge representation of discharge summaries. In AIME 91 (pp. 173-182). Springer Berlin Heidelberg.
  92. 92.Baud, R. H., Alpay, L., & Lovis, C. (1994). Let’s Meet the Users with Natural Language Understanding. Knowledge and Decisions in Health Telematics: The Next Decade, 12, 103.
  93. 93.Rassinoux, A. M., Baud, R. H., & Scherrer, J. R. (1992). Conceptual graphs model extension for knowledge representation of medical texts. MEDINFO, 92, 1368-1374.
  94. 94.Morel-Guillemaz, A. M., Baud, R. H., & Scherrer, J. R. (1990). Proximity Processing of Medical Text. In Medical Informatics Europe’90 (pp. 625-630). Springer Berlin Heidelberg.
  95. 95.Rassinoux, A. M., Michel, P. A., Juge, C., Baud, R., & Scherrer, J. R. (1994). Natural language processing of medical texts within the HELIOS environment. Computer methods and programs in biomedicine, 45, S79-96.
  96. 96.Rassinoux, A. M., Juge, C., Michel, P. A., Baud, R. H., Lemaitre, D., Jean, F. C., ... & Scherrer, J. R. (1995, June). Analysis of medical jargon: The RECIT system. In Conference on Artificial Intelligence in Medicine in Europe (pp. 42-52). Springer Berlin Heidelberg.
  97. 97.Friedman, C., Cimino, J. J., & Johnson, S. B. (1993). A conceptual model for clinical radiology reports. In Proceedings of the Annual Symposium on Computer Application in Medical Care (p. 829). American Medical Informatics Association.
  98. 98."Natural Language Processing." Natural Language Processing RSS. N.p., n.d. Web. 23 Mar. 2017.
  99. 99.[Srihari S. Machine Learning: Generative and Discriminative Models. 2010. http:// www.cedar.buffalo.edu/wsrihari/CSE574/Discriminative-Generative.pdf (accessed 31 May 2011).]
  100. 100.[Elkan C. Log-Linear Models and Conditional Random Fields. 2008. http://cseweb. ucsd.edu/welkan/250B/cikmtutorial.pdf (accessed 28 Jun 2011). 62. Hearst MA, Dumais ST, Osman E, et al. Support vector machines]
  101. 101.[Jurafsky D, Martin JH. Speech and Language Processing. 2nd edn. Englewood Cliffs, NJ: Prentice-Hall, 2008.]
  102. 102.[Sonnhammer ELL, Eddy SR, Birney E, et al. Pfam: Multiple sequence alignments and HMM-profiles of protein domains. Nucleic Acids Res 1998;26:320]
  103. 103.[Sonnhammer, E. L., Eddy, S. R., Birney, E., Bateman, A., & Durbin, R. (1998). Pfam: multiple sequence alignments and HMM-profiles of protein domains. Nucleic acids research, 26(1), 320-322]
  104. 104.Systems, RAVN. "RAVN Systems Launch the ACE Powered GDPR Robot - Artificial Intelligence to Expedite GDPR Compliance." Stock Market. PR Newswire, n.d. Web. 19 Mar. 2017.
  105. 105."Here's Why Natural Language Processing is the Future of BI." SmartData Collective. N.p., n.d. Web. 19 Mar. 2017
  106. 106."Using Natural Language Processing and Network Analysis to Develop a Conceptual Framework for Medication Therapy Management Research." AMIA ... Annual Symposium proceedings. AMIA Symposium. U.S. National Library of Medicine, n.d. Web. 19 Mar. 2017
  107. 107.Ogallo, W., & Kanter, A. S. (2017, February 10). Using Natural Language Processing and Network Analysis to Develop a Conceptual Framework for Medication Therapy Management Research. Retrieved April 10, 2017, from https://www.ncbi.nlm.nih.gov/pubmed/28269895?dopt=Abstract
  108. 108.Ochoa, A. (2016, May 25). Meet the Pilot: Smart Earpiece Language Translator. Retrieved April 10, 2017, from https://www.indiegogo.com/projects/meet-the-pilot-smart-earpiece-language-translator-headphones-travel

Citation

MLA
Khurana, D., et al. “Natural Language Processing: State of the Art, Current Trends and Challenges”. Multimedia Tools and Applications, vol. 82, no. 3, 2022, pp. 3713–44, https://doi.org/10.1007/s11042-022-13428-4.
APA
Khurana, D., Koli, A., Khatter, K., & Singh, S. (2022). Natural language processing: state of the art, current trends and challenges. Multimedia Tools and Applications, 82(3), 3713–3744. https://doi.org/10.1007/s11042-022-13428-4
Chicago
Khurana, D., A. Koli, K. Khatter, and S. Singh. 2022. “Natural Language Processing: State of the Art, Current Trends and Challenges”. Multimedia Tools and Applications 82 (3): 3713–44. https://doi.org/10.1007/s11042-022-13428-4.
Harvard
Khurana, D. et al. (2022) “Natural language processing: state of the art, current trends and challenges”, Multimedia Tools and Applications, 82(3), pp. 3713–3744. Available at: https://doi.org/10.1007/s11042-022-13428-4.
Vancouver
1. Khurana D, Koli A, Khatter K, Singh S (2022) Natural language processing: state of the art, current trends and challenges. Multimedia Tools and Applications 82:3713–3744

BibTeX

@article{Khurana_2022, title={Natural language processing: state of the art, current trends and challenges}, volume={82}, ISSN={1573-7721}, url={http://dx.doi.org/10.1007/s11042-022-13428-4}, DOI={10.1007/s11042-022-13428-4}, number={3}, journal={Multimedia Tools and Applications}, publisher={Springer Science and Business Media LLC}, author={Khurana, Diksha and Koli, Aditya and Khatter, Kiran and Singh, Sukhdev}, year={2022}, month=July, pages={3713–3744} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF