Is ChatGPT a General-Purpose Natural Language Processing Task Solver?

Chengwei QinAston ZhangZhuosheng ZhangJiaao ChenMichihiro YasunagaDiyi Yang

article2023EMNLP841 citations

Evaluates ChatGPT's zero-shot performance across twenty standard benchmarks spanning seven task categories, identifying clear strengths in reasoning alongside persistent weaknesses in structured predictions like sequence tagging.

Listen

Recent advancements in large language models have shown strong conversational abilities, most notably with the debut of ChatGPT. However, organizations and researchers face uncertainty regarding whether ChatGPT can function as a reliable, general-purpose solver for diverse language tasks without requiring task-specific training. Understanding its real-world baseline capabilities and limitations across standard workflows is essential for leaders evaluating where generative artificial intelligence can be safely and effectively deployed.

The article evaluates the zero-shot performance of ChatGPT—meaning its ability to execute tasks without prior fine-tuning on downstream training data—across 20 benchmark datasets spanning seven core categories: reasoning, natural language inference, question answering, dialogue, summarization, named entity recognition, and sentiment analysis. The study benchmarks ChatGPT against its predecessor (GPT-3.5) as well as specialized, fine-tuned models, employing both standard zero-shot prompting and step-by-step reasoning prompts.

The findings show that ChatGPT is an effective tool for selected workflows but falls short of being a universal task solver. First, ChatGPT outperforms previous models on arithmetic reasoning (achieving up to 95.8% accuracy with step-by-step reasoning prompts), dialogue reasoning (76.2% accuracy), and sentiment analysis (93.7% accuracy). Second, it demonstrates strong capabilities on natural language inference and reading comprehension tasks, showing a distinct strength in confirming factual entailments over non-entailments. Third, ChatGPT frequently underperforms GPT-3.5 on commonsense, symbolic, and logical reasoning benchmarks. Fourth, it struggles significantly with structured sequence tagging, achieving an overall score of only 53.2% on named entity recognition compared to around 94% achieved by fine-tuned models. Finally, ChatGPT exhibits verbosity in summarization, producing longer texts that yield lower standard overlap scores, with explicit word-limit constraints further degrading output quality.

These results indicate that while ChatGPT provides immediate value for dialogue generation, sentiment classification, and mathematical problem-solving, relying on it for specialized structured data extraction or strictly constrained summarization introduces operational risks and lower accuracy. Organizations should note that task-specific fine-tuned models consistently outperform zero-shot ChatGPT across nearly all benchmarks. When planning system architectures, decision-makers should deploy specialized models for high-precision information extraction and reserve ChatGPT for interactive dialogue and open-ended analysis. Further evaluations on larger datasets and few-shot in-context learning configurations should be conducted before replacing dedicated machine learning pipelines in mission-critical applications.

arXiv: 2302.06476
Cover for Is ChatGPT a General-Purpose Natural Language Processing Task Solver?

Abstract

Spurred by advancements in scale, large language models (LLMs) have demonstrated the ability to perform a variety of natural language processing (NLP) tasks zero-shot -- i.e., without adaptation on downstream data. Recently, the debut of ChatGPT has drawn a great deal of attention from the natural language processing (NLP) community due to the fact that it can generate high-quality responses to human input and self-correct previous mistakes based on subsequent conversations. However, it is not yet known whether ChatGPT can serve as a generalist model that can perform many NLP tasks zero-shot. In this work, we empirically analyze the zero-shot learning ability of ChatGPT by evaluating it on 20 popular NLP datasets covering 7 representative task categories. With extensive empirical studies, we demonstrate both the effectiveness and limitations of the current version of ChatGPT. We find that ChatGPT performs well on many tasks favoring reasoning capabilities (e.g., arithmetic reasoning) while it still faces challenges when solving specific tasks such as sequence tagging. We additionally provide in-depth analysis through qualitative case studies.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Large Language Models
  • 2.2 Zero-Shot Learning
  • 2.3 Chain-of-Thought Prompting
  • 3 Methodology
  • 4 Experiments
  • 4.1 Tasks and Datasets
  • 4.2 Experimental Results
  • 4.2.1 Arithmetic Reasoning
  • 4.2.2 Commonsense, Symbolic, and Logical Reasoning
  • 4.2.3 Natural Language Inference
  • 4.2.4 Question Answering
  • 4.2.5 Dialogue
  • 4.2.6 Summarization
  • 4.2.7 Named Entity Recognition
  • 4.2.8 Sentiment Analysis
  • 4.3 ChatGPT v.s. Full-Set or Few-Shot Fine-Tuning
  • 5 Conclusion
  • References
  • A Appendix
  • A.1 Example Input and Output Pairs of ChatGPT

Knowls

  1. Knowl 1 — Zero-Shot and Zero-Shot Chain-of-Thought Prompting Framework

    model/method

    To evaluate large language models such as ChatGPT (gpt-3.5-turbo) and GPT-3.5 (text-davinci-003) on downstream natural language processing tasks without gradient updates or fine-tuning, two zero-shot prompting paradigms are utilized:

    1. Standard Zero-Shot Prompting: Given a task instruction PP (which defines the task requirements and output format) and a test instance XX, both are concatenated as the input sequence to the model ff. The model generates the target output YY directly: Y=f(P,X)Y = f(P, X)

    2. Two-Stage Zero-Shot Chain-of-Thought (Zero-Shot-CoT): For multi-step reasoning tasks, a two-stage prompting mechanism is applied:

    • Stage 1 (Rationale Generation): A trigger prompt P1P_1 (specifically "Let's think step by step." ) is appended to the problem input XX. The model generates an intermediate reasoning chain (rationale) RR: R=f(X,P1)R = f(X, P_1)
    • Stage 2 (Answer Extraction): The original input XX, instruction P1P_1, generated rationale RR, and an answer extraction trigger P2P_2 (e.g., "Therefore, among A through E, the answer is") are concatenated to extract the final predicted answer YY: Y=f(X,P1,R,P2)Y = f(X, P_1, R, P_2)
  2. Knowl 2 — Comprehensive Zero-Shot Performance Comparison Across 20 NLP Benchmarks

    empirical result

    Evaluating ChatGPT (gpt-3.5-turbo) against GPT-3.5 (text-davinci-003) and task-specific fine-tuned models across 20 datasets covering 7 representative NLP task categories (reporting the higher score between standard zero-shot and zero-shot chain-of-thought for reasoning tasks) demonstrates that while ChatGPT exhibits strong capabilities as a generalist model, it underperforms dedicated fine-tuned models on most tasks.

    Task Category Dataset Metric ChatGPT GPT-3.5 Fine-Tuning
    Arithmetic Reasoning MultiArith Accuracy (%) 95.8 83.7 96.2
    Arithmetic Reasoning GSM8K Accuracy (%) 78.9 59.5 63.1
    Arithmetic Reasoning AddSub Accuracy (%) 88.6 87.3 93.9
    Arithmetic Reasoning AQUA-RAT Accuracy (%) 53.5 40.6 45.3
    Arithmetic Reasoning SingleEq Accuracy (%) 91.5 86.4 93.1
    Arithmetic Reasoning SVAMP Accuracy (%) 77.5 73.6 79.0
    Symbolic Reasoning Last Letter Accuracy (%) 70.2 54.4 99.4
    Symbolic Reasoning Coin Flip Accuracy (%) 65.8 97.8 100.0
    Logical Reasoning Date Understanding Accuracy (%) 72.6 77.0 65.3
    Logical Reasoning Tracking Shuffled Objects Accuracy (%) 58.7 39.7 23.9
    Commonsense Reasoning CSQA Accuracy (%) 73.7 74.9 82.3
    Commonsense Reasoning StrategyQA Accuracy (%) 61.1 61.1 77.8
    Commonsense Reasoning COPA Accuracy (%) 82.0 93.0 95.0
    Natural Language Inference RTE Accuracy (%) 85.9 80.1 95.8
    Natural Language Inference CB Accuracy (%) 89.3 83.9 100.0
    Question Answering BoolQ Accuracy (%) 87.3 84.7 91.2
    Dialogue MuTual Accuracy (%) 76.2 75.2 93.5
    Summarization SAMSum ROUGE Avg 31.0 32.4 40.5
    Named Entity Recognition CoNLL03 F1 53.2 53.5 94.6
    Sentiment Analysis SST2 Accuracy (%) 93.7 88.8 97.5

    ChatGPT outperforms GPT-3.5 on all arithmetic reasoning tasks, natural language inference, question answering, dialogue, and sentiment analysis, but trails GPT-3.5 on several commonsense, symbolic, and logical reasoning benchmarks as well as summarization.

  3. Knowl 3 — Arithmetic Reasoning Capabilities and Chain-of-Thought Gains

    empirical result

    In arithmetic reasoning benchmarks, ChatGPT (gpt-3.5-turbo) exhibits substantial performance advantages over GPT-3.5 (text-davinci-003) and earlier large language models both with and without Zero-Shot Chain-of-Thought (CoT).

    Model MultiArith GSM8K AddSub AQUA-RAT SingleEq SVAMP
    N/A CoT N/A CoT N/A CoT N/A CoT N/A CoT N/A CoT
    Zero-Shot
    text-davinci-002 22.7 78.7 12.5 40.7 77.0 74.7 22.4 33.5 78.7 78.7 58.8 63.7
    text-davinci-003 24.2 83.7 12.6 59.5 87.3 81.3 28.0 40.6 82.3 86.4 64.7 73.6
    ChatGPT 79.8 95.8 23.8 78.9 88.6 83.5 28.0 53.5 89.4 91.5 74.8 77.5
    Few-Shot
    UL2 (20B) 5.0 10.7 4.1 4.4 18.5 18.2 20.5 23.6 18.0 20.2 10.1 12.5
    LaMDA (137B) 7.6 44.9 6.5 14.3 43.0 51.9 25.5 20.6 48.8 58.7 29.5 37.5
    text-davinci-002 33.8 91.7 15.6 46.9 83.3 81.3 24.8 35.8 82.7 86.6 65.7 68.9
    Codex 44.0 96.2 19.7 63.1 90.9 90.9 29.5 45.3 86.8 93.1 69.9 76.4
    PaLM (540B) 42.2 94.7 17.9 56.9 93.9 91.9 25.2 35.8 86.5 92.3 69.4 79.0

    Without CoT, ChatGPT outperforms GPT-3.5 on 5 of the 6 datasets, notably on MultiArith (79.8%79.8\% vs. 24.2%24.2\%) and GSM8K (23.8%23.8\% vs. 12.6%12.6\%). When Zero-Shot-CoT is applied, ChatGPT achieves the highest accuracy across all arithmetic datasets, surpassing text-davinci-003 by +12.1%+12.1\% on MultiArith (95.8%95.8\%), +19.4%+19.4\% on GSM8K (78.9%78.9\%), and +12.9%+12.9\% on AQUA-RAT (53.5%53.5\%).

  4. Knowl 4 — Zero-Shot Performance on Commonsense, Symbolic, and Logical Reasoning

    empirical result

    Across commonsense, symbolic, and logical reasoning benchmarks, ChatGPT frequently underperforms GPT-3.5 (text-davinci-003), and Chain-of-Thought (CoT) prompting does not consistently yield accuracy improvements.

    Model Commonsense Reasoning Symbolic Reasoning Logical Reasoning
    CSQA StrategyQA COPA Last Letter Coin Flip Date Object
    N/A CoT N/A CoT N/A CoT N/A CoT N/A CoT N/A CoT N/A CoT
    Zero-Shot
    text-davinci-002 72.6 64.6 54.3 54.8 74.0 85.0 0.2 57.6 53.8 91.4 49.3 67.5 31.3 52.9
    text-davinci-003 74.9 70.0 57.2 61.1 93.0 64.0 0.0 54.4 49.0 97.8 56.6 77.0 27.1 39.7
    ChatGPT 73.7 71.5 61.1 55.5 78.0 82.0 0.4 70.2 21.8 65.8 48.0 72.6 31.6 58.7
    Few-Shot
    UL2 34.2 51.4 59.0 53.3 - - 0.6 18.8 70.4 67.1 13.5 14.0 - -
    LaMDA 53.6 57.9 62.4 65.4 - - 5.8 77.5 49.0 99.6 21.5 26.8 - -
    text-davinci-002 79.5 73.5 65.9 65.4 - - 0.2 59.0 57.2 97.2 43.8 52.1 - -
    Codex 82.3 77.9 67.1 73.2 - - - - - - 49.0 64.8 - -
    PaLM 78.1 79.9 68.6 77.8 95.0 - 7.6 99.4 98.1 100.0 49.0 65.3 23.9 -

    Key observations include:

    • Commonsense reasoning with CoT: In StrategyQA, zero-shot CoT reduces ChatGPT accuracy from 61.1%61.1\% to 55.5%55.5\%. Plausible rationales are generated, but the final predictions are frequently erroneous.
    • Symbolic reasoning gaps: On Coin Flip, ChatGPT attains only 21.8%21.8\% without CoT and 65.8%65.8\% with CoT, falling far behind GPT-3.5 (49.0%49.0\% and 97.8%97.8\%).
    • Logical reasoning: On Date Understanding, GPT-3.5 outperforms ChatGPT (77.0%77.0\% vs. 72.6%72.6\% with CoT). However, ChatGPT outperforms GPT-3.5 on Tracking Shuffled Objects (58.7%58.7\% vs. 39.7%39.7\% with CoT).
  5. Knowl 5 — Natural Language Inference Performance and Factual Entailment Preference

    empirical result

    In natural language inference (NLI) benchmarks, zero-shot ChatGPT achieves superior accuracy compared to zero-shot GPT-3.5 (text-davinci-003), FLAN, T0, and PaLM, driven by strong accuracy on entailed text pairs.

    Dataset ChatGPT (Zero-Shot) GPT-3.5 (Zero-Shot) FLAN (Zero-Shot) T0 (Zero-Shot) PaLM (Zero-Shot) PaLM-540B (Fine-Tuned)
    RTE 85.9% 80.1% 84.1% 80.8% 72.9% 95.8%
    CB 89.3% 83.9% 83.9% 70.1% 51.8% 100.0%

    Per-class accuracy on the RTE dataset demonstrates an asymmetric performance profile:

    Class ChatGPT GPT-3.5
    Entailment 92.5% 70.6%
    Not Entailment 78.6% 90.8%

    ChatGPT outperforms GPT-3.5 by +21.9%+21.9\% on true entailment examples (92.5%92.5\% vs. 70.6%70.6\%), but underperforms GPT-3.5 by −12.2%-12.2\% on non-entailment examples (78.6%78.6\% vs. 90.8%90.8\%). This indicates that ChatGPT is more effective at identifying factually consistent relationships, reflecting the human preference alignment embedded during reinforcement learning from human feedback (RLHF).

  6. Knowl 6 — Reading Comprehension and Question Answering on BoolQ

    empirical result

    On the BoolQ reading comprehension benchmark, zero-shot ChatGPT achieves 87.3%87.3\% accuracy, outperforming zero-shot GPT-3.5 (84.7%84.7\%), Chinchilla (83.7%83.7\%), FLAN (82.9%82.9\%), and Gopher (79.3%79.3\%), while approaching zero-shot PaLM (88.0%88.0\%).

    Zero-Shot Fine-Tuned
    ChatGPT GPT-3.5 PaLM Chinchilla FLAN Gopher CompassMTL T5-11B DeBERTa
    Accuracy 87.3% 84.7% 88.0% 83.7% 82.9% 79.3% 88.3% 91.2% 90.4%

    Per-class breakdown on BoolQ reveals that ChatGPT performs better on positive instances:

    • Class "Yes": ChatGPT attains 88.9%88.9\% accuracy (+7.8%+7.8\% compared to GPT-3.5's 81.1%81.1\%).
    • Class "No": ChatGPT attains 84.6%84.6\% accuracy (−6.0%-6.0\% compared to GPT-3.5's 90.6%90.6\%).

    Despite instructions requiring exact "yes" or "no" answers, ChatGPT occasionally outputs ambiguous responses (such as "It is unclear"), leading to classification failures.

  7. Knowl 7 — Dialogue Reasoning Evaluation on MuTual

    empirical result

    On the MuTual multi-turn dialogue reasoning benchmark, ChatGPT outperforms GPT-3.5 and earlier unsupervised methods, but remains behind specialized fine-tuned architectures.

    Model / Method Accuracy (%)
    Zero-Shot
    ChatGPT 76.2
    GPT-3.5 (text-davinci-003) 75.2
    Unsupervised
    TF-IDF 27.6
    Fine-Tuned
    Dual LSTM 26.6
    DAM 23.9
    SMN 27.4
    BERT 65.7
    RoBERTa 69.5
    GPT-2-FT 39.8
    MDFN 92.3
    BiDeN 93.5

    ChatGPT accurately tracks multi-turn conversational context without introducing unmentioned or extraneous information. However, fully fine-tuned models such as BiDeN (93.5%93.5\%) and MDFN (92.3%92.3\%) outperform zero-shot ChatGPT (76.2%76.2\%) by over 16 percentage points.

  8. Knowl 8 — Dialogue Summarization Verbosity and Length Restriction Effects on SAMSum

    empirical result

    On the SAMSum dialogue summarization benchmark, zero-shot ChatGPT achieves lower ROUGE scores than zero-shot GPT-3.5, driven by verbosity in its generated outputs.

    Metric Zero-Shot Fine-Tuned
    ChatGPT GPT-3.5 BART-large CODA
    ROUGE-1 42.4 44.0 49.1 50.1
    ROUGE-2 17.6 18.5 24.3 24.6
    ROUGE-L 33.0 34.7 45.8 46.9
    ROUGE Average 31.0 32.4 39.7 40.5

    The performance disparity is characterized by:

    • Output Length: Ground truth summaries average 20.020.0 words, GPT-3.5 responses average 23.323.3 words, and ChatGPT responses average 36.636.6 words, containing redundant conversational details.
    • Length Constrained Prompting: When prompted with "Please summarize the given conversation in no more than 25 words.", ChatGPT reduces its average length to 22.822.8 words; however, the average ROUGE-1/2/L score decreases from 31.031.0 to 30.630.6, showing that explicit zero-shot length constraints further degrade summary quality.
  9. Knowl 9 — Sequence Tagging and Named Entity Recognition Bottlenecks on CoNLL03

    empirical result

    Zero-shot evaluations on the CoNLL03 named entity recognition (NER) benchmark show that both ChatGPT and GPT-3.5 face substantial challenges with sequence tagging tasks compared to supervised models.

    Entity Type Zero-Shot Fine-Tuned
    ChatGPT GPT-3.5 Flair LUKE ACE
    All 53.2 53.5 93.0 93.9 94.6
    Location (Loc) 66.7 67.1 94.0 - -
    Person (Per) 87.2 78.0 97.4 - -
    Organization (Org) 51.4 50.0 91.9 - -
    Miscellaneous (Misc) 4.1 4.8 83.0 - -

    Key issues identified in zero-shot NER include:

    1. Substantial Performance Gap: Zero-shot LLMs achieve overall F1 scores of 53.253.2--53.553.5, trailing fine-tuned systems (94.694.6 F1 for ACE) by over 4040 F1 points.
    2. Failure on Miscellaneous Entities: Both models fail to capture the "Misc" category (4.14.1 F1 for ChatGPT, 4.84.8 F1 for GPT-3.5) due to discrepancies between the model's open-domain understanding and dataset-specific annotation boundaries.
    3. Decomposed Instruction Degradation: Designing separate prompt instructions to extract each entity type individually with GPT-3.5 reduces the overall F1 score to 34.834.8.
  10. Knowl 10 — Sentiment Classification on SST2 and Output Format Adherence

    empirical result

    On the SST2 sentiment analysis dataset, zero-shot ChatGPT attains 93.7%93.7\% overall accuracy, outperforming zero-shot GPT-3.5 (88.8%88.8\%), but underperforming zero-shot FLAN (94.6%94.6\%) and fine-tuned T5-11B (97.5%97.5\%).

    Subset Zero-Shot Fine-Tuned
    ChatGPT GPT-3.5 FLAN T5-11B
    All 93.7% 88.8% 94.6% 97.5%
    Positive 90.8% 88.1% - -
    Negative 96.7% 89.5% - -

    Key findings include:

    • Class Imbalance: ChatGPT's advantage over GPT-3.5 is driven primarily by negative sentiment classification (96.7%96.7\% vs. 89.5%89.5\%, a +7.2%+7.2\% margin), whereas positive sentiment classification shows smaller differences (90.8%90.8\% vs. 88.1%88.1\%).
    • Output Format Violations: Despite prompt instructions explicitly requiring exact "positive" or "negative" responses, both ChatGPT and GPT-3.5 occasionally generate unprompted labels such as "neutral" or "mixed", which lowers their accuracy compared to FLAN.
  11. Knowl 11 — Limitations of the Zero-Shot ChatGPT Evaluation

    limitation

    The empirical evaluation of ChatGPT across NLP tasks is bounded by three specific experimental limitations:

    1. Dataset and Task Breadth: Due to API usage costs, larger-scale datasets and several task categories (e.g., machine translation, semantic parsing) were excluded from evaluation.
    2. Prompt Template Coverage: Results for proprietary models not publicly queryable (such as PaLM) are sourced from published literature, and evaluated models were tested on a limited set of standard prompt templates rather than exhaustively optimized prompt variants.
    3. Absence of Few-Shot Evaluation: The empirical benchmark evaluates zero-shot and zero-shot Chain-of-Thought prompting; it does not evaluate how ChatGPT's few-shot in-context learning compares with zero-shot learning across these task suites.

Coverage note — Qualitative input/output sample tables in the appendix were omitted as they provide individual anecdotal model outputs rather than generalizable experimental findings.

References

  1. 1.Alan Akbik, Duncan Blythe, and Roland Vollgraf. 2018. Contextual string embeddings for sequence labeling. In Proceedings of the 27th international conference on computational linguistics, pages 1638–1649.
  2. 2.BIG-bench collaboration BIG-bench collaboration. 2021. Beyond the imitation game: Measuring and extrapolating the capabilities of language models. https://github.com/google/BIG-bench. "Accessed: 2022-05-07".
  3. 3.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  4. 4.Jiaao Chen and Diyi Yang. 2021. Simple conversational data augmentation for semi-supervised abstractive dialogue summarization. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6605–6616, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  5. 5.Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen. 2022. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. arXiv preprint arXiv:2211.12588.
  6. 6.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022a. Palm: Scaling language modeling with pathways.
  7. 7.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022b. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  8. 8.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30.
  9. 9.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  10. 10.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2924–2936, Minneapolis, Minnesota. Association for Computational Linguistics.
  11. 11.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems.
  12. 12.Leyang Cui, Yu Wu, Shujie Liu, Yue Zhang, and Ming Zhou. 2020. MuTual: A dataset for multi-turn dialogue reasoning. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1406–1416, Online. Association for Computational Linguistics.
  13. 13.Ido Dagan, Oren Glickman, and Bernardo Magnini. 2006. The pascal recognising textual entailment challenge. In Machine Learning Challenges. Evaluating Predictive Uncertainty, Visual Object Classification, and Recognising Tectual Entailment: First PASCAL Machine Learning Challenges Workshop, MLCW 2005, Southampton, UK, April 11-13, 2005, Revised Selected Papers, pages 177–190. Springer.
  14. 14.Marie-Catherine De Marneffe, Mandy Simons, and Judith Tonhauser. 2019. The commitmentbank: Investigating projection in naturally occurring discourse. In proceedings of Sinn und Bedeutung, volume 23, pages 107–124.
  15. 15.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  16. 16.Bosheng Ding, Chengwei Qin, Linlin Liu, Lidong Bing, Shafiq Joty, and Boyang Li. 2022. Is gpt-3 a good data annotator? arXiv preprint arXiv:2212.10450.
  17. 17.Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot. 2022. Complexity-based prompting for multi-step reasoning. arXiv preprint arXiv:2210.00720.
  18. 18.Zhenyong Fu, Tao Xiang, Elyor Kodirov, and Shaogang Gong. 2017. Zero-shot learning on semantic class prototype graph. IEEE transactions on pattern analysis and machine intelligence, 40(8):2009–2022.
  19. 19.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346–361.
  20. 20.Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer. 2019. Samsum corpus: A human-annotated dialogue dataset for abstractive summarization. arXiv preprint arXiv:1911.12237.
  21. 21.Ben Goertzel. 2014. Artificial general intelligence: concept, state of the art, and future prospects. Journal of Artificial General Intelligence, 5(1):1.
  22. 22.Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu. 2023. How close is chatgpt to human experts? comparison corpus, evaluation, and detection. arXiv preprint arXiv:2301.07597.
  23. 23.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020. Deberta: Decoding-enhanced bert with disentangled attention. arXiv preprint arXiv:2006.03654.
  24. 24.Namgyu Ho, Laura Schmid, and Se-Young Yun. 2022. Large language models are reasoning teachers. arXiv preprint arXiv:2212.10071.
  25. 25.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556.
  26. 26.Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kushman. 2014. Learning to solve arithmetic word problems with verb categorization. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 523–533, Doha, Qatar. Association for Computational Linguistics.
  27. 27.Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2022. Decomposed prompting: A modular approach for solving complex tasks. arXiv preprint arXiv:2210.02406.
  28. 28.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In Thirty-sixth Conference on Neural Information Processing Systems (NeurIPS 2022).
  29. 29.Rik Koncel-Kedziorski, Hannaneh Hajishirzi, Ashish Sabharwal, Oren Etzioni, and Siena Dumas Ang. 2015. Parsing algebraic word problems into equations. Transactions of the Association for Computational Linguistics, 3:585–597.
  30. 30.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension.
  31. 31.Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al. 2022. Solving quantitative reasoning problems with language models. arXiv preprint arXiv:2206.14858.
  32. 32.Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. 2022a. On the advance of making language models better reasoners. arXiv preprint arXiv:2206.02336.
  33. 33.Yiyang Li, Hai Zhao, and Zhuosheng Zhang. 2022b. Back to the future: Bidirectional information decoupling network for multi-turn dialogue modeling. In The 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP 2022).
  34. 34.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2022. Holistic evaluation of language models. ArXiv preprint, abs/2211.09110.
  35. 35.Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017. Program induction by rationale generation: Learning to solve and explain algebraic word problems. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 158–167, Vancouver, Canada. Association for Computational Linguistics.
  36. 36.Longxiang Liu, Zhuosheng Zhang, Hai Zhao, Xi Zhou, and Xiang Zhou. 2021. Filling the gap of utterance-aware and speaker-aware representation for multi-turn dialogue. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 13406–13414.
  37. 37.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  38. 38.Ryan Lowe, Nissan Pow, Iulian Vlad Serban, and Joelle Pineau. 2015. The ubuntu dialogue corpus: A large dataset for research in unstructured multi-turn dialogue systems. In Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 285–294.
  39. 39.Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022a. Learn to explain: Multimodal reasoning via thought chains for science question answering. ArXiv preprint, abs/2209.09513.
  40. 40.Pan Lu, Liang Qiu, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, Tanmay Rajpurohit, Peter Clark, and Ashwin Kalyan. 2022b. Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning. arXiv preprint arXiv:2209.14610.
  41. 41.Lucie Charlotte Magister, Jonathan Mallinson, Jakub Adamek, Eric Malmi, and Aliaksei Severyn. 2022. Teaching small language models to reason. ArXiv preprint, abs/2212.08410.
  42. 42.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.
  43. 43.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-task generalization via natural language crowdsourcing instructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3470–3487.
  44. 44.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  45. 45.Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are NLP models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2080–2094, Online. Association for Computational Linguistics.
  46. 46.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, page 9.
  47. 47.Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. 2021. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446.
  48. 48.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer.
  49. 49.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR, 21(140):1–67.
  50. 50.Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. 2011. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In AAAI spring symposium: logical formalizations of commonsense reasoning, pages 90–95.
  51. 51.Subhro Roy and Dan Roth. 2015. Solving general arithmetic word problems. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 1743–1752, Lisbon, Portugal. Association for Computational Linguistics.
  52. 52.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2022. Learning to retrieve prompts for in-context learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2655–2671.
  53. 53.Erik F Sang and Fien De Meulder. 2003. Introduction to the conll-2003 shared task: Language-independent named entity recognition. arXiv preprint cs/0306050.
  54. 54.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Tali Bers, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M. Rush. 2021a. Multitask prompted training enables zero-shot task generalization.
  55. 55.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2021b. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207.
  56. 56.Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, et al. 2022. Language models are multilingual chain-of-thought reasoners. arXiv preprint arXiv:2210.03057.
  57. 57.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al. 2022. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv preprint arXiv:2201.11990.
  58. 58.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, A. Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Conference on Empirical Methods in Natural Language Processing.
  59. 59.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615.
  60. 60.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. CommonsenseQA: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4149–4158, Minneapolis, Minnesota. Association for Computational Linguistics.
  61. 61.Yi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier García, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Neil Houlsby, and Donald Metzler. 2022. Unifying language learning paradigms. arXiv preprint arXiv:2205.05131.
  62. 62.Xiaolong Wang, Yufei Ye, and Abhinav Gupta. 2018. Zero-shot recognition via semantic embeddings and knowledge graphs. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6857–6866.
  63. 63.Xinyu Wang, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, and Kewei Tu. 2020. Automated concatenation of embeddings for structured prediction.
  64. 64.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022a. Rationale-augmented ensembles in language models. arXiv preprint arXiv:2207.00747.
  65. 65.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022b. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  66. 66.Albert Webson and Ellie Pavlick. 2022. Do prompt-based models really understand the meaning of their prompts? In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2300–2344, Seattle, United States. Association for Computational Linguistics.
  67. 67.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652.
  68. 68.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. In Thirty-sixth Conference on Neural Information Processing Systems (NeurIPS 2022).
  69. 69.Yu Wu, Wei Wu, Chen Xing, Ming Zhou, and Zhoujun Li. 2017. Sequential matching network: A new architecture for multi-turn response selection in retrieval-based chatbots. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 496–505.
  70. 70.Ikuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda, and Yuji Matsumoto. 2020. Luke: Deep contextualized entity representations with entity-aware self-attention.
  71. 71.Meng Ye and Yuhong Guo. 2017. Zero-shot classification with discriminative semantic representation learning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7140–7148.
  72. 72.Eric Zelikman, Yuhuai Wu, and Noah D Goodman. 2022. Star: Bootstrapping reasoning with reasoning. arXiv preprint arXiv:2203.14465.
  73. 73.Aston Zhang, Zachary C Lipton, Mu Li, and Alexander J Smola. 2021. Dive into deep learning. arXiv preprint arXiv:2106.11342.
  74. 74.Li Zhang, Tao Xiang, and Shaogang Gong. 2017. Learning a deep embedding model for zero-shot learning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2021–2030.
  75. 75.Zhuosheng Zhang, Shuohang Wang, Yichong Xu, Yuwei Fang, Wenhao Yu, Yang Liu, Hai Zhao, Chenguang Zhu, and Michael Zeng. 2022. Task compass: Scaling multi-task pre-training with task prefix. In Findings of The 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP 2022).
  76. 76.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2023a. Automatic chain of thought prompting in large language models. In The Eleventh International Conference on Learning Representations (ICLR 2023).
  77. 77.Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. 2023b. Multimodal chain-of-thought reasoning in language models. arXiv preprint arXiv:2302.00923.
  78. 78.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. 2022. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625.
  79. 79.Xiangyang Zhou, Lu Li, Daxiang Dong, Yi Liu, Ying Chen, Wayne Xin Zhao, Dianhai Yu, and Hua Wu. 2018. Multi-turn response selection for chatbots with deep attention matching network. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1118–1127.

Citation

MLA
Qin, C., et al. “Is ChatGPT a General-Purpose Natural Language Processing Task Solver?”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 1339–84, https://doi.org/10.18653/v1/2023.emnlp-main.85.
APA
Qin, C., Zhang, A., Zhang, Z., Chen, J., Yasunaga, M., & Yang, D. (2023). Is ChatGPT a General-Purpose Natural Language Processing Task Solver?. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1339–1384. https://doi.org/10.18653/v1/2023.emnlp-main.85
Chicago
Qin, C., A. Zhang, Z. Zhang, J. Chen, M. Yasunaga, and D. Yang. 2023. “Is ChatGPT a General-Purpose Natural Language Processing Task Solver?”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1339–84. https://doi.org/10.18653/v1/2023.emnlp-main.85.
Harvard
Qin, C. et al. (2023) “Is ChatGPT a General-Purpose Natural Language Processing Task Solver?”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1339–1384. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.85.
Vancouver
1. Qin C, Zhang A, Zhang Z, Chen J, Yasunaga M, Yang D (2023) Is ChatGPT a General-Purpose Natural Language Processing Task Solver?. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1339–1384

BibTeX

@inproceedings{qin-etal-2023-chatgpt,
    title = "Is {C}hat{GPT} a General-Purpose Natural Language Processing Task Solver?",
    author = "Qin, Chengwei  and
      Zhang, Aston  and
      Zhang, Zhuosheng  and
      Chen, Jiaao  and
      Yasunaga, Michihiro  and
      Yang, Diyi",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.85/",
    doi = "10.18653/v1/2023.emnlp-main.85",
    pages = "1339--1384"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/