A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity

Yejin BangSamuel CahyawijayaNayeon LeeWenliang DaiDan SuBryan WilieHoly LoveniaZiwei JiTiezheng YuWilly Chung

article2023International Joint Conference on Natural Language Processing1,824 citationsArea Chair Award (Language Modeling and Analysis)

Reveals critical limits in ChatGPT's reasoning, multilingual generation, and factual hallucination across twenty-three datasets while demonstrating how multi-turn interactive prompting significantly improves task performance.

Listen

The rapid adoption of conversational artificial intelligence has highlighted the need for rigorous, independent evaluations to understand what these models can reliably do. The article presents an extensive, zero-shot quantitative benchmark of the December 2022 release of ChatGPT. The objective of the study was to evaluate its multitask capabilities across eight core language tasks, assess its multilingual and multimodal competencies, and diagnose its reliability in terms of reasoning, factuality, hallucination, and multi-turn interactivity.

To establish these benchmarks without application programming interface access, the authors conducted single-run experiments across 23 datasets spanning tasks such as question answering, summarization, machine translation, sentiment analysis, dialogue tracking, and fact-checking. Credibility is supported through the use of standardized test collections and manual verification of responses. A specialized national flag-drawing dataset was also created to assess whether the model could generate visual output using scalable vector graphics code as an intermediate medium. For most tasks, sample sizes ranged from 30 to 200 instances.

The key findings reveal notable strengths alongside clear vulnerabilities. First, in zero-shot performance across standard language datasets, the model outperformed earlier foundational systems on 9 out of 13 benchmarks and matched or exceeded fully specialized, fine-tuned systems on 4 tasks. Second, the system demonstrated significant performance disparities across languages: while it performed strongly on high-resource, Latin-script languages, it degraded sharply on low-resource and non-Latin languages, showing better capability in understanding non-Latin text than generating it. Third, the model proved to be an inconsistent reasoner, scoring an average accuracy of only 63.41% across 10 reasoning categories. It performed well on commonsense, deductive, and analogical tasks but struggled severely on inductive, mathematical (scoring 23.33%), multi-hop (scoring 26.67%), and spatial tasks. Fourth, the system consistently exhibited extrinsic hallucinations by generating unverified information not contained in the source input. Finally, multi-turn interactivity served as an effective mechanism for improvement, where follow-up instructions increased summarization quality by about 8% on standard lexical overlap metrics and improved translation accuracy.

These results demonstrate that while interactive language models offer strong general-purpose capabilities and intuitive collaborative refinement, their propensity to hallucinate and fail at multi-step reasoning introduces operational, legal, and compliance risks. Deploying the system in high-stakes environments—such as task-oriented customer service, complex analytical workflows, or low-resource linguistic regions—without human oversight could lead to factual inaccuracies and flawed decision-making.

Organizations should treat conversational models as interactive collaborators rather than fully autonomous agents. Stakeholders should implement multi-turn prompting frameworks to refine and verify outputs, integrate external knowledge retrieval systems to ground factuality, and deploy automated hallucination detection safeguards. Strong caution is advised for tasks requiring complex logic, mathematical precision, or low-resource translation until targeted fine-tuning and further domain-specific validation are completed.

Cover for A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity

Abstract

This paper proposes a framework for quantitatively evaluating interactive LLMs such as ChatGPT using publicly available data sets. We carry out an extensive technical evaluation of ChatGPT using 23 data sets covering 8 different common NLP application tasks. We evaluate the multitask, multilingual and multi-modal aspects of ChatGPT based on these data sets and a newly designed multimodal dataset. We find that ChatGPT outperforms LLMs with zero-shot learning on most tasks and even outperforms fine-tuned models on some tasks. We find that it is better at understanding non-Latin script languages than generating them. It is able to generate multimodal content from textual prompts, via an intermediate code generation step. Moreover, we find that ChatGPT is 63.41% accurate on average in 10 different reasoning categories under logical reasoning, non-textual reasoning, and commonsense reasoning, hence making it an unreliable reasoner. It is, for example, better at deductive than inductive reasoning. ChatGPT suffers from hallucination problems like other LLMs and it generates more extrinsic hallucinations from its parametric memory as it does not have access to an external knowledge base. Finally, the interactive feature of ChatGPT enables human collaboration with the underlying LLM to improve its performance, i.e, 8% ROUGE-1 on summarization and 2% ChrF++ on machine translation, in a multi-turn "prompt engineering" fashion. We also release codebase for evaluation set extraction.

Table of Contents

  • 1 Introduction
  • 2 Multitask, Multilingual, and Multimodal Evaluations of ChatGPT
  • 2.1 Multitask Ability of ChatGPT
  • 2.2 Evaluating Multilinguality of ChatGPT
  • 2.2.1 Language Understanding
  • 2.2.2 Language Generation
  • 2.3 Evaluating Multimodality of ChatGPT
  • 3 Reasoning Evaluations of ChatGPT
  • 4 Factuality and Hallucination
  • 5 Evaluating Interactivity in ChatGPT
  • 6 Evaluation of GPT-4
  • 7 Conclusion and Discussion
  • References
  • A Background and Related Work
  • A.1 ChatGPT
  • A.2 LLM benchmark and evaluation
  • A.3 ChatGPT Evaluation
  • B General Experimental Details
  • C Multitask Evaluation of ChatGPT
  • C.1 Summarization
  • C.2 Machine Translation
  • C.3 Sentiment Analysis
  • C.4 Question Answering
  • C.5 Misinformation Detection
  • C.6 ChatGPT on Dialogue Tasks
  • C.6.1 Knowledge-Grounded Open-Domain Dialogue
  • C.6.2 Task-Oriented Dialogue Experimental Setups
  • D ChatGPT on Multilinguality
  • E Multimodality: Flag Drawing Task
  • F Details for Reasoning Evaluations
  • F.1 Logical Reasoning
  • F.1.1 Deductive vs. Inductive Reasoning
  • F.1.2 Abductive Reasoning
  • F.2 Non-textual Semantic Reasoning
  • F.3 Commonsense Reasoning
  • F.4 Causal, Multi-Hop, and Analogical Reasoning
  • G Details for Hallucination Evaluations
  • H Details for Interactivity Evaluation
  • H.1 Interactivity on Summarization
  • H.2 Interactivity on Machine Translation
  • H.2.1 Experiment 1: Multi-turn Post-Editting
  • H.2.2 Experiment 2: Automatic post-editing
  • H.3 Interactivity on Multimodal Generation
  • I Results for Evaluation of GPT-4
  • I.1 Results on Multitask Ability
  • I.2 Results on Mulilinguality
  • I.3 Results on Reasoning
  • J List of Evaluation Datasets
  • K Examples from Machine Translation and Post-Editing

Knowls

  1. Knowl 1 — Multitask Zero-Shot Benchmark Performance of ChatGPT Across NLP Tasks

    data/table

    ChatGPT (evaluated on the 15 December 2022 release) was benchmarked in a zero-shot setting against state-of-the-art fully fine-tuned models (Fine-Tuned SOTA) and zero-shot large language models (Zero-Shot SOTA) across 8 standard NLP tasks comprising 21 datasets. ChatGPT outperforms prior zero-shot LLMs on 9 out of 13 datasets with reported baselines and surpasses fine-tuned task-specific models on 4 datasets (EntailmentBank, Pep-3k, COVID-Scientific, and NusaX-Bug on sentiment analysis).

    Task Dataset Metric Fine-Tuned SOTA Zero-Shot SOTA ChatGPT
    Summarization CNN/DM ROUGE-1 44.47 35.27 35.29
    Summarization SAMSum ROUGE-1 47.28 - 35.29
    MT (XXX→\rightarrowEng) FLoRes-200 (HRL) ChrF++ 63.50 - 58.64
    MT (XXX→\rightarrowEng) FLoRes-200 (LRL) ChrF++ 54.90 - 27.75
    MT (Eng→\rightarrowXXX) FLoRes-200 (HRL) ChrF++ 54.40 - 51.12
    MT (Eng→\rightarrowXXX) FLoRes-200 (LRL) ChrF++ 41.90 - 21.57
    Sentiment Analysis NusaX - Eng Macro F1 92.60 61.50 83.24
    Sentiment Analysis NusaX - Ind Macro F1 91.60 59.30 82.13
    Sentiment Analysis NusaX - Jav Macro F1 84.20 55.70 79.64
    Sentiment Analysis NusaX - Bug Macro F1 70.00 55.90 55.84
    Question Answering bAbI task (15 | 16) Accuracy 100 | 100 - 93.3 | 66.7
    Question Answering EntailmentBank Accuracy 86.50 78.58 93.30
    Question Answering CLUTRR Accuracy 95.00 28.60 43.30
    Question Answering StepGame (k=9 | k=1) Accuracy 48.4 | 98.7 - 23.3 | 63.3
    Question Answering Pep-3k AUC 67.00 - 93.30
    Misinformation Detection COVID-Social Accuracy 77.70 50.00 73.30
    Misinformation Detection COVID-Scientific Accuracy 74.70 71.10 92.00
    Task-Oriented Dialogue MultiWOZ2.2 JGA 60.60 46.70 24.40
    Task-Oriented Dialogue MultiWOZ2.2 BLEU 19.10 - 5.65
    Task-Oriented Dialogue MultiWOZ2.2 Inform Rate 95.70 - 71.10
    Open-Domain KGD OpenDialKG BLEU | ROUGE-L 20.8 | 40.0 3.1 | 29.5 4.1 | 18.6
    Open-Domain KGD OpenDialKG FeQA 48.00 23.00 15.00

    Evaluation subsets for ChatGPT use 30 to 200 samples per dataset. High-resource languages (HRL) and low-resource languages (LRL) follow the NLLB classifications. Zero-shot SOTA models by task are: Summarization: InstructGPT; MT: NLLB-200; Sentiment Analysis: XLM-R LARGE; QA: ST-MoE-32B, ZeroQA, GPT-3; Misinformation: GPT-2; Task-Oriented Dialogue: D3ST; Open-Domain KGD: GPT-Jurassic-6B.

  2. Knowl 2 — Reasoning Performance Profile of ChatGPT Across Ten Subcategories

    data/table

    Across 10 distinct categories of reasoning evaluated over 634 total test instances, ChatGPT achieved an average accuracy of 63.41%, exhibiting significant variation across different reasoning modalities.

    Reasoning Category Dataset / Subtask Result (Correct / Total)
    Deductive EntailmentBank 28 / 30
    bAbI (task 15) 28 / 30 (as-is: 19 / 30)
    Inductive CLUTRR 13 / 30
    bAbI (task 16) 20 / 30 (as-is: 0 / 30)
    Abductive α\alphaNLI 26 / 30
    Mathematical MATH 13 / 30
    Temporal TimeDial 26 / 30
    Spatial SpartQA (hard | basic) 8 / 32 | 20 / 32
    StepGame (hard | basic) 7 / 30 | 19 / 30
    StepGame (cardinal) 17 / 20
    StepGame (diagonal) 11 / 20
    StepGame (clock) 5 / 20
    Commonsense CommonsenseQA 27 / 30
    PiQA 25 / 30
    Pep-3k (Hard) 28 / 30
    Causal E-CARE 24 / 30
    Multi-hop HotpotQA 8 / 30
    Analogical Letter string analogy 30 / 30

    ChatGPT excels at analogical reasoning (100%), commonsense reasoning (83.3%--93.3%), temporal reasoning (86.7%), and abductive reasoning (86.7%), but shows severe deficiencies in multi-hop reasoning (26.7%), mathematical reasoning (43.3% on MATH subset), and complex spatial reasoning (23.3%--25.0% on hard spatial subsets).

  3. Knowl 3 — Inductive versus Deductive Reasoning Disparity and Prompt Sensitivity

    empirical result

    ChatGPT exhibits a marked asymmetry between deductive and inductive logical reasoning, acting as a 'lazy reasoner' when induction is required without explicit instruction.

    Reasoning Type Task (Default Prompt) Task (Prompt Engineered) Advanced Benchmark
    Deductive bAbI task 15: 19 / 30 bAbI task 15: 28 / 30 EntailmentBank: 28 / 30
    Inductive bAbI task 16: 0 / 30 bAbI task 16: 20 / 30 CLUTRR: 13 / 30

    When evaluated on basic bAbI task 16 without prompt engineering, ChatGPT failed completely (0/30), defaulting to the response: "It is not specified what <attribute> <entity> is." However, when the prompt explicitly requests inductive inference ("Based on the given facts above, do a reasonable inference on this question using inductive reasoning:"), accuracy increased to 20/30 (66.7%). On deductive tasks (bAbI task 15), ChatGPT achieved 19/30 (63.3%) under the default prompt and 28/30 (93.3%) with prompt engineering.

    On the CLUTRR benchmark for inductive kinship relation reasoning, ChatGPT struggled to induce rules across complex combinations (13/30), frequently confusing kinship levels (such as failing to differentiate son from grandson, while successfully differentiating daughter from granddaughter).

  4. Knowl 4 — Non-Textual Semantic Reasoning in ChatGPT: Spatial, Mathematical, and Temporal

    empirical result

    Evaluation of non-textual semantic understanding reveals strong temporal reasoning but deficient spatial and mathematical reasoning:

    1. Temporal Reasoning: On TimeDial, ChatGPT scored 86.67% (26/30), outperforming reported zero-shot performance of large language models such as Chinchilla (68.8%) and Gopher (50.9%).
    2. Mathematical Reasoning: On the MATH dataset, ChatGPT scored 23.33% (7/30). While it generally understands the problem framing, it routinely fails to compute correct multi-step numerical or algebraic solutions.
    3. Spatial Reasoning: On multi-hop spatial datasets, ChatGPT achieved 43.33% (26/60) on StepGame and 43.75% (28/64) on SpartQA. On hard subsets involving long chains (k=9k=9 relations in StepGame or multi-relation compositions in SpartQA), performance dropped to 23.33% (7/30) and 25.00% (8/32), respectively.

    A breakdown across 20-sample subsets of elementary spatial relation descriptions in StepGame shows:

    • Basic cardinal directions (e.g., above, below): 85.0% accuracy (17/20).
    • Diagonal directions (e.g., lower-left, upper-right): 55.0% accuracy (11/20).
    • Clock-position directions (e.g., G is at Y's 6 o'clock): 25.0% accuracy (5/20).
  5. Knowl 5 — Multilingual Performance Disparity Across Resource Levels and Script Types

    empirical result

    ChatGPT's performance correlates directly with the pre-training data availability in CommonCrawl across high-resource (HRL: ≥1%\ge 1\%), medium-resource (MRL: ≥0.01%\ge 0.01\%), low-resource (LRL: ≥0.0001%\ge 0.0001\%), and extremely low-resource (X-LRL: <0.0001%< 0.0001\%) languages.

    Language CommonCrawl Size (%) Resource Tier SA Acc. (%) LID Acc. (%) MT (XXX→\rightarrowEng | Eng→\rightarrowXXX)
    English 46.320% HRL 84% 100% -
    Chinese 4.837% HRL - - 24 / 30 | 14 / 30
    French 4.604% HRL - - 29 / 30 | 25 / 30
    Indonesian 0.781% MRL 80% 100% 28 / 30 | 19 / 30
    Korean 0.679% MRL - - 22 / 30 | 12 / 30
    Javanese 0.002% LRL 78% 0% 7 / 30 | 6 / 30
    Sundanese 0.001% LRL - - 9 / 30 | 0 / 30
    Buginese 0.000% X-LRL 56% 12% -

    Key empirical findings include:

    • Understanding vs. Identification: ChatGPT achieves 78% sentiment analysis accuracy on Javanese despite 0% accuracy in language identification, demonstrating semantic comprehension without explicit meta-linguistic identification.
    • Translation Asymmetry: ChatGPT translates from non-English source languages into English (XXX→EngXXX\rightarrow\text{Eng}) far more accurately than from English into target languages (Eng→XXX\text{Eng}\rightarrow XXX).
    • Script Disparity: Non-Latin script generation (Chinese: 14/30, Korean: 12/30) is substantially worse than Latin script generation (French: 25/30, Indonesian: 19/30) even among high- and medium-resource languages.
  6. Knowl 6 — Multimodal Image Synthesis via Intermediate SVG Code Generation

    empirical result

    To evaluate visual representation capabilities in a text-only model, a national flag drawing task was conducted on 50 national flags across continents. ChatGPT generates Scalable Vector Graphics (SVG) code via multi-turn prompting.

    The generation protocol consists of: (1) describing the visual appearance of the flag, (2) generating SVG code from that description, and (3) iteratively requesting fixes for identified errors (up to 2 correction turns). Flag drawings are evaluated across four error types: layout (34% occurrence), color (20%), missing components (18%), and shape/size (68%). Flag quality is categorized into grades A (0 errors) through E (≥4\ge 4 errors).

    Grade (# Errors) Turn 1 (w/o description) Turn 1 (with description) Turn 2 Turn 3
    A (0 errors) 0% 4% 12% 24%
    B (1 error) 4% 22% 24% 24%
    C (2 errors) 16% 18% 12% 10%
    D (3 errors) 18% 24% 26% 20%
    E (≥4\ge 4 errors) 62% 32% 26% 22%

    An ablation removing the initial textual description step causes Grade E failures to spike from 32% to 62%, with 0% Grade A results. Generating an explicit natural language description first serves as a visual chain-of-thought mechanism.

  7. Knowl 7 — Multi-Turn Interactive Prompting for Performance Refinement

    empirical result

    Exploiting ChatGPT's multi-turn conversational interface enables substantial performance gains without retraining across three domains:

    1. Dialogue Summarization (SAMSum, 50 dialogues): ChatGPT's initial summaries are often overly verbose. Appending a second-turn instruction ("Please make the summary shorter") yields an increase of 7.99 points in ROUGE-1 (from 35.29 to 43.28) and 1.64 points in ROUGE-2.
    2. Machine Translation Post-Editing (WMT 2022 English→\rightarrowMarathi APE shared task, 50 samples): Prompting ChatGPT to perform automatic post-editing ("Could you perform a post-editing to ensure the meaning is equivalent to '[INPUT_SENTENCE]'?") improves back-translation alignment to the source English text across all metrics:
    Evaluation Target Metric Without APE With APE
    Post-Edited Marathi Text HTER (↓\downarrow) 88.14 88.79
    SacreBLEU (↑\uparrow) 4.81 4.20
    METEOR (↑\uparrow) 13.10 12.74
    Source English Text (Back-translation) HTER (↓\downarrow) 65.36 63.13
    SacreBLEU (↑\uparrow) 25.54 27.20
    METEOR (↑\uparrow) 43.71 47.51
    BERTScore (↑\uparrow) 92.30 92.59
    1. Multimodal SVG Refinement: Iterative multi-turn error correction reduced visual defects, with 34% of flag drawings improving from Turn 1 to Turn 2, and 36% improving from Turn 2 to Turn 3, increasing Grade A outputs from 4% to 24%.
  8. Knowl 8 — Intrinsic and Extrinsic Hallucination Profiles in ChatGPT

    empirical result

    ChatGPT exhibits distinct hallucination behaviors categorized into intrinsic (contradicting provided source information) and extrinsic (generating unverifiable content not present in the source):

    • Intrinsic Hallucinations: Rarely observed in summarization or open-domain knowledge-grounded dialogue, but present in multimodal code generation (e.g., correctly stating in text that the Mexican flag has "three vertical bands", yet generating SVG code rendering horizontal bands).
    • Extrinsic Hallucinations: Highly prevalent across multiple NLP tasks:
      • Factual Extrinsic: Adding accurate background knowledge absent from the source (e.g., expanding "Britain and five other countries" into "P5+1 (US, UK, France, China, Russia, and Germany)" in summarization).
      • Non-factual Extrinsic: Fabricating false details. In question answering, ChatGPT introduced unmentioned familial relationships (e.g., asserting "stepdaughter" instead of daughter). In machine translation into Korean, it hallucinated medical procedures ("저주파 치료", transcutaneous electrical nerve stimulation) absent in the source English text. In task-oriented dialogue, it fabricated train ticket prices (£10--£30) and travel durations (1h 30m).
    • TruthfulQA Benchmark: On 66 test questions from TruthfulQA designed to elicit human misconceptions, ChatGPT generated untruthful or false answers 35.38% of the time (23/65 errors).
    • Misinformation Fact-Checking: ChatGPT detected COVID-19 misinformation with 92.0% accuracy (46/50) on scientific claims and 73.33% accuracy (22/30, excluding refusal cases) on social claims.
  9. Knowl 9 — Failure Modes of ChatGPT in Task-Oriented Dialogue Modeling

    limitation

    When evaluated on task-oriented dialogue (TOD) using MultiWOZ 2.2 under both modular (Dialogue State Tracking and Action-based Response Generation) and unified (end-to-end multi-turn interaction over a structured database) setups, ChatGPT exhibits major structural failure modes:

    1. Poor Dialogue State Tracking: In the modular setup, ChatGPT achieved a Joint Goal Accuracy (JGA) of only 24.40%, trailing the fine-tuned SOTA (60.60%) and zero-shot D3ST (46.70%).
    2. Multi-Turn Context and Belief State Loss: In unified interactions, ChatGPT overwrites previous turn constraints upon receiving new criteria (e.g., searching for "rating 3 or higher" followed by "Italian food" leads ChatGPT to recommend a 2-star Italian restaurant, dropping the prior rating filter unless explicitly reminded by the user).
    3. Elementary Constraint Reasoning Failure: When queries require numerical or categorical filtering (e.g., "rating of 3 or higher" or "European food" mapped across countries), ChatGPT fails to answer correctly 66% of the time.
    4. Unconstrained Hallucination of Database Attributes: ChatGPT generates plausible but fabricated prices, availability states, and contact information not grounded in the provided database.
  10. Knowl 10 — Zero-Shot Evaluation Protocol for Interactive Large Language Models

    experimental setup

    The evaluation framework tests conversational LLMs across 23 public NLP datasets spanning 8 application tasks and 10 reasoning categories. The protocol evaluates the 15 December 2022 public release of ChatGPT under the following design:

    • Zero-Shot Prompting: Tasks are formulated as natural language instructions without few-shot examples or in-context demonstrations.
    • Representative Subsampling: Because evaluation was conducted via the web user interface prior to API availability, subsets of 30 to 200 diverse test samples were sampled per benchmark.
    • Multi-Turn Interaction Probing: Multi-turn prompt protocols were designed specifically for summarization (two-turn length control), machine translation (two-turn translation followed by automated post-editing), and multimodal flag generation (three-turn descriptive code generation and iterative bug fixing).
    • Evaluation Metrics: Standard quantitative metrics (ROUGE-1/2/L, BLEU, ChrF++, Macro F1, JGA, SacreBLEU, METEOR, HTER, BERTScore, FeQA) combined with manual validation by native speakers for multilingual translation quality, reasoning rationales, and visual code rendering accuracy.

Coverage note — None was omitted; all primary contributed benchmark results across multitask NLP, multilinguality, multimodal SVG generation, fine-grained reasoning categories, hallucination analyses, and interactive prompt engineering were converted into standalone knowls.

References

  1. 1.
    1. Chatgpt vs satya nadella over biryani: The chatbot is learning from its mistakes.
  2. 2.Alham Fikri Aji, Genta Indra Winata, Fajri Koto, Samuel Cahyawijaya, Ade Romadhony, Rahmad Mahendra, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Timothy Baldwin, Jey Han Lau, and Sebastian Ruder. 2022. One country, 700+ languages: NLP challenges for underrepresented languages and dialects in Indonesia. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7226–7249, Dublin, Ireland. Association for Computational Linguistics.
  3. 3.Ömer Aydın and Enis Karaarslan. 2022. Openai chatgpt generated literature review: Digital twin in healthcare. Available at SSRN 4308687.
  4. 4.Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan. Association for Computational Linguistics.
  5. 5.Paul Bartha. 2013. Analogy and analogical reasoning.
  6. 6.Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Wen tau Yih, and Yejin Choi. 2020. Abductive commonsense reasoning. In International Conference on Learning Representations.
  7. 7.Prajjwal Bhargava and Vincent Ng. 2022. Commonsense knowledge reasoning and generation with pretrained language models: a survey. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 12317–12325.
  8. 8.Pushpak Bhattacharyya, Rajen Chatterjee, Markus Freitag, Diptesh Kanojia, Matteo Negri, and Marco Turchi. 2022. Findings of the wmt 2022 shared task on automatic post-editing. In Proceedings of the Seventh Conference on Machine Translation, pages 109–117, Abu Dhabi.
  9. 9.David G.W. Birch. 2022. Chatgpt is a window into the real future of financial services.
  10. 10.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. 2020. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432–7439.
  11. 11.Alexandre Blanco-Gonzalez, Alfonso Cabezon, Alejandro Seco-Gonzalez, Daniel Conde-Torres, Paula Antelo-Riveiro, Angel Pineiro, and Rebeca Garcia-Fandino. 2022. The role of ai in drug discovery: Challenges, opportunities, and strategies. arXiv preprint arXiv:2212.08104.
  12. 12.Samuel Cahyawijaya, Holy Lovenia, Alham Fikri Aji, Genta Indra Winata, Bryan Wilie, Rahmad Mahendra, Christian Wibisono, Ade Romadhony, Karissa Vincentio, Fajri Koto, Jennifer Santoso, David Moeljadi, Cahya Wirawan, Frederikus Hudi, Ivan Halim Parmonangan, Ika Alfina, Muhammad Satrio Wicaksono, Ilham Firdausi Putra, Samsul Rahmadani, Yulianti Oenang, Ali Akbar Septiandri, James Jaya, Kaustubh D. Dhole, Arie Ardiyanti Suryani, Rifki Afina Putri, Dan Su, Keith Stevens, Made Nindyatama Nityasya, Muhammad Farid Adilazuarda, Ryan Ignatius, Ryandito Diandaru, Tiezheng Yu, Vito Ghifari, Wenliang Dai, Yan Xu, Dyah Damapuspita, Cuk Tho, Ichwanul Muslim Karo Karo, Tirana Noor Fatyanosa, Ziwei Ji, Pascale Fung, Graham Neubig, Timothy Baldwin, Sebastian Ruder, Herry Sujaini, Sakriani Sakti, and Ayu Purwarianti. 2022. Nusacrowd: Open source initiative for indonesian nlp resources.
  13. 13.Samuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio, Xiaohong Li, Adhiguna Kuncoro, Sebastian Ruder, Zhi Yuan Lim, Syafri Bahar, Masayu Khodra, Ayu Purwarianti, and Pascale Fung. 2021. IndoNLG: Benchmark and resources for evaluating Indonesian natural language generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8875–8898, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  14. 14.Ethan C. Chau and Noah A. Smith. 2021. Specializing multilingual language models: An empirical study. In Proceedings of the 1st Workshop on Multilingual Representation Learning, pages 51–61, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  15. 15.Jonathan H Choi, Kristin E Hickman, Amy Monahan, and Daniel Schwarcz. 2023. Chatgpt goes to law school. Available at SSRN.
  16. 16.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways.
  17. 17.Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences.
  18. 18.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge.
  19. 19.Cookup.ai. 2022. Chatgpt - where it lacks.
  20. 20.Wenliang Dai, Lu Hou, Lifeng Shang, Xin Jiang, Qun Liu, and Pascale Fung. 2022. Enabling multimodal generation on CLIP via vision-language knowledge distillation. In Findings of the Association for Computational Linguistics: ACL 2022, pages 2383–2395, Dublin, Ireland. Association for Computational Linguistics.
  21. 21.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023a. Instructblip: Towards general-purpose vision-language models with instruction tuning.
  22. 22.Wenliang Dai, Zihan Liu, Ziwei Ji, Dan Su, and Pascale Fung. 2023b. Plausible may not be faithful: Probing object hallucination in vision-language pre-training. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 2136–2148, Dubrovnik, Croatia. Association for Computational Linguistics.
  23. 23.Bhavana Dalvi, Peter Jansen, Oyvind Tafjord, Zhengnan Xie, Hannah Smith, Leighanna Pipatanangkura, and Peter Clark. 2021. Explaining answers with entailment trees. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7358–7370.
  24. 24.Ernest Davis. 2023a. Benchmarks for automated commonsense reasoning: A survey. arXiv preprint arXiv:2302.04752.
  25. 25.Ernest Davis. 2023b. Mathematics, word problems, common sense, and artificial intelligence. arXiv preprint arXiv:2301.09723.
  26. 26.Web Desk. 2023. Colombian judge uses chatgpt in ruling, triggers debate.
  27. 27.Igor Douven. 2017. Abduction.
  28. 28.Michael Dowling and Brian Lucey. 2023. Chatgpt for (finance) research: The bananarama conjecture. Finance Research Letters, page 103662.
  29. 29.Li Du, Xiao Ding, Kai Xiong, Ting Liu, and Bing Qin. 2022. e-CARE: a new dataset for exploring explainable causal reasoning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 432–446, Dublin, Ireland. Association for Computational Linguistics.
  30. 30.Simon Frieder, Luca Pinchetti, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Christian Petersen, Alexis Chevalier, and Julius Berner. 2023. Mathematical capabilities of chatgpt.
  31. 31.Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2021. A framework for few-shot language model evaluation.
  32. 32.Aidan Gilson, Conrad Safranek, Thomas Huang, Vimig Socrates, Ling Chi, Richard Andrew Taylor, and David Chartash. 2022. How well does chatgpt do when taking the medical licensing exams? the implications of large language models for medical education and knowledge assessment. medRxiv, pages 2022–12.
  33. 33.Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer. 2019. Samsum corpus: A human-annotated dialogue dataset for abstractive summarization. EMNLP-IJCNLP 2019, page 70.
  34. 34.Yoav Goldberg. 2023. Some remarks on large language models.
  35. 35.Cindy Gordon. 2023. Chatgpt is the fastest growing app in the history of web applications.
  36. 36.Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2021. The flores-101 evaluation benchmark for low-resource and multilingual machine translation.
  37. 37.Tanya Goyal, Junyi Jessy Li, and Greg Durrett. 2022. News summarization and evaluation in the era of gpt-3. arXiv preprint arXiv:2209.12356.
  38. 38.Roberto Gozalo-Brizuela and Eduardo C Garrido-Merchan. 2023. Chatgpt is not all you need. a state of the art review of large generative ai models. arXiv preprint arXiv:2301.04655.
  39. 39.Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu. 2023. How close is chatgpt to human experts? comparison corpus, evaluation, and detection. arXiv preprint arXiv:2301.07597.
  40. 40.James Hawthorne. 2021. Inductive Logic. In Edward N. Zalta, editor, The Stanford Encyclopedia of Philosophy, Spring 2021 edition. Metaphysics Research Lab, Stanford University.
  41. 41.Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc.
  42. 42.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. 2022. Training compute-optimal large language models.
  43. 43.Krystal Hu. 2023. Chatgpt sets record for fastest-growing user base - analyst note.
  44. 44.Jie Huang and Kevin Chen-Chuan Chang. 2022. Towards reasoning in large language models: A survey. arXiv preprint arXiv:2212.10403.
  45. 45.Arfinda Ilmania, Abdurrahman, Samuel Cahyawijaya, and Ayu Purwarianti. 2018. Aspect detection and sentiment classification using deep neural network for indonesian aspect-based sentiment analysis. In 2018 International Conference on Asian Language Processing (IALP), pages 62–67.
  46. 46.Hadar Yoana Jabotinsky and Roee Sarel. 2022. Co-authoring with an ai? ethical dilemmas and artificial intelligence. Ethical Dilemmas and Artificial Intelligence (December 15, 2022).
  47. 47.Katharina Jeblick, Balthasar Schachtner, Jakob Dexl, Andreas Mittermeier, Anna Theresa Stüber, Johanna Topalis, Tobias Weber, Philipp Wesp, Bastian Sabel, Jens Ricke, et al. 2022. Chatgpt makes medicine easy to swallow: An exploratory case study on simplified radiology reports. arXiv preprint arXiv:2212.14882.
  48. 48.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung. 2022a. Survey of hallucination in natural language generation. ACM Comput. Surv. Just Accepted.
  49. 49.Ziwei Ji, Zihan Liu, Nayeon Lee, Tiezheng Yu, Bryan Wilie, Min Zeng, and Pascale Fung. 2022b. Rho (ρ\rho): Reducing hallucination in open-domain dialogues with knowledge grounding. arXiv preprint arXiv:2212.01588.
  50. 50.Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang, and Zhaopeng Tu. 2023. Is chatgpt a good translator? a preliminary study.
  51. 51.Arianna Johnson. 2023. Is chatgpt partisan? poems about trump and biden raise questions about the ai bot’s bias-here’s what experts think.
  52. 52.Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6282–6293, Online. Association for Computational Linguistics.
  53. 53.Jennifer A. Kingson. 2023. Friend or foe? teachers debate chatgpt.
  54. 54.Jan Kocoń, Igor Cichecki, Oliwier Kaszyca, Mateusz Kochanek, Dominika Szydło, Joanna Baran, Julita Bielaniewicz, Marcin Gruza, Arkadiusz Janz, Kamil Kanclerz, et al. 2023. Chatgpt: Jack of all trades, master of none. Information Fusion, page 101861.
  55. 55.Escape Velocity Labs. 2022. Chatgpt imitates logical reasoning surprisingly well.
  56. 56.Viet Dac Lai, Nghia Trung Ngo, Amir Pouran Ben Veyseh, Hieu Man, Franck Dernoncourt, Trung Bui, and Thien Huu Nguyen. 2023. Chatgpt beyond english: Towards a comprehensive evaluation of large language models in multilingual learning. arXiv preprint arXiv:2304.05613.
  57. 57.Md Tahmid Rahman Laskar, M Saiful Bari, Mizanur Rahman, Md Amran Hossen Bhuiyan, Shafiq Joty, and Jimmy Huang. 2023a. A systematic study and comprehensive evaluation of ChatGPT on benchmark datasets. In Findings of the Association for Computational Linguistics: ACL 2023, pages 431–469, Toronto, Canada. Association for Computational Linguistics.
  58. 58.Md Tahmid Rahman Laskar, M Saiful Bari, Mizanur Rahman, Md Amran Hossen Bhuiyan, Shafiq Joty, and Jimmy Xiangji Huang. 2023b. A systematic study and comprehensive evaluation of chatgpt on benchmark datasets. arXiv preprint arXiv:2305.18486.
  59. 59.Anton E Lawson. 2005. What is the role of induction and deduction in reasoning and scientific inquiry? Journal of Research in Science Teaching, 42(6):716–740.
  60. 60.Nayeon Lee, Yejin Bang, Andrea Madotto, and Pascale Fung. 2021. Towards few-shot fact-checking via perplexity. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1971–1981, Online. Association for Computational Linguistics.
  61. 61.Nayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary, Pascale Fung, Mohammad Shoeybi, and Bryan Catanzaro. 2022. Factuality enhanced language models for open-ended text generation. In Advances in Neural Information Processing Systems.
  62. 62.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020a. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  63. 63.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020b. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880.
  64. 64.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas, Drew A. Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. 2022. Holistic evaluation of language models.
  65. 65.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland. Association for Computational Linguistics.
  66. 66.Zhaojiang Lin, Bing Liu, Andrea Madotto, Seungwhan Moon, Zhenpeng Zhou, Paul A Crook, Zhiguang Wang, Zhou Yu, Eunjoon Cho, Rajen Subba, et al. 2021. Zero-shot dialogue state tracking via cross-task transfer. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7890–7900.
  67. 67.Holy Lovenia, Bryan Wilie, Romain Barraud, Samuel Cahyawijaya, Willy Chung, and Pascale Fung. 2022. Every picture tells a story: Image-grounded controllable stylistic story generation. In Proceedings of the 6th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature, pages 40–52, Gyeongju, Republic of Korea. International Conference on Computational Linguistics.
  68. 68.Hongyuan Lu, Haoyang Huang, Shuming Ma, Dongdong Zhang, Wai Lam, and Furu Wei. 2022. Trip: Triangular document-level pre-training for multilingual language models. arXiv preprint arXiv:2212.07752.
  69. 69.Andrea Madotto, Zhaojiang Lin, Genta Indra Winata, and Pascale Fung. 2021. Few-shot bot: Prompt-based learning for dialogue systems. arXiv preprint arXiv:2110.08118.
  70. 70.Kyle Mahowald, Anna A Ivanova, Idan A Blank, Nancy Kanwisher, Joshua B Tenenbaum, and Evelina Fedorenko. 2023. Dissociating language and thought in large language models: a cognitive perspective. arXiv preprint arXiv:2301.06627.
  71. 71.Rui Mao, Guanyi Chen, Xulang Zhang, Frank Guerin, and Erik Cambria. 2023. Gpteval: A survey on assessments of chatgpt and gpt-4. arXiv preprint arXiv:2308.12488.
  72. 72.Bernard Marr. 2022. What does chatgpt really mean for businesses?
  73. 73.Vaibhav Mavi, Anubhav Jangra, and Adam Jatowt. 2022. A survey on multi-hop question answering and generation. arXiv preprint arXiv:2204.09140.
  74. 74.Pasquale Minervini, Sebastian Riedel, Pontus Stenetorp, Edward Grefenstette, and Tim Rocktäschel. 2020. Learning reasoning strategies in end-to-end differentiable proving. In Proceedings of the 37th International Conference on Machine Learning, ICML’20. JMLR.org.
  75. 75.Roshanak Mirzaee and Parisa Kordjamshidi. 2022. Transfer learning with synthetic corpora for spatial role labeling and reasoning. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 6148–6165, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  76. 76.Roshanak Mirzaee, Hossein Rajaby Faghihi, Qiang Ning, and Parisa Kordjamshidi. 2021. SPARTQA: A textual question answering benchmark for spatial reasoning. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4582–4598, Online. Association for Computational Linguistics.
  77. 77.Seungwhan Moon, Pararth Shah, Anuj Kumar, and Rajen Subba. 2019. Opendialkg: Explainable conversational reasoning with attention-based walks over knowledge graphs. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 845–854.
  78. 78.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2022. Crosslingual generalization through multitask finetuning. arXiv preprint arXiv:2211.01786.
  79. 79.Benjamin Muller, Antonios Anastasopoulos, Benoît Sagot, and Djamé Seddah. 2021. When being unseen from mBERT is just the beginning: Handling new languages with multilingual language models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 448–462, Online. Association for Computational Linguistics.
  80. 80.Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Çaglar Gulçehre, and Bing Xiang. 2016. Abstractive text summarization using sequence-to-sequence RNNs and beyond. In Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning, pages 280–290, Berlin, Germany. Association for Computational Linguistics.
  81. 81.Tomáš Nekvinda and Ondřej Dušek. 2021. Shades of BLEU, flavours of success: The case of MultiWOZ. In Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021), pages 34–46, Online. Association for Computational Linguistics.
  82. 82.Oded Nov, Nina Singh, and Devin M Mann. 2023. Putting chatgpt’s medical advice to the (turing) test. medRxiv, pages 2023–01.
  83. 83.OpenAI. 2023. Gpt-4 technical report.
  84. 84.Simon Ott, Konstantin Hebenstreit, Valentin Liévin, Christoffer Egeberg Hother, Milad Moradi, Maximilian Mayrhauser, Robert Praas, Ole Winther, and Matthias Samwald. 2023. Thoughtsource: A central hub for large language model reasoning data. arXiv preprint arXiv:2301.11596.
  85. 85.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback.
  86. 86.Maja Popović. 2015. chrF: character n-gram F-score for automatic MT evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392–395, Lisbon, Portugal. Association for Computational Linguistics.
  87. 87.Ian Porada, Kaheer Suleman, Adam Trischler, and Jackie Chi Kit Cheung. 2021. Modeling event plausibility with consistent conceptual abstraction. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1732–1743, Online. Association for Computational Linguistics.
  88. 88.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Brussels, Belgium. Association for Computational Linguistics.
  89. 89.Shuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, and Huajun Chen. 2022. Reasoning with language model prompting: A survey. arXiv preprint arXiv:2212.09597.
  90. 90.Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michihiro Yasunaga, and Diyi Yang. 2023. Is chatgpt a general-purpose natural language processing task solver? arXiv preprint arXiv:2302.06476.
  91. 91.Lianhui Qin, Aditya Gupta, Shyam Upadhyay, Luheng He, Yejin Choi, and Manaal Faruqui. 2021. TIMEDIAL: Temporal commonsense reasoning in dialog. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7066–7076, Online. Association for Computational Linguistics.
  92. 92.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR.
  93. 93.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  94. 94.Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. 2021. Scaling language models: Methods, analysis and insights from training gopher.
  95. 95.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2022. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(1).
  96. 96.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021. Zero-shot text-to-image generation. In International Conference on Machine Learning, pages 8821–8831. PMLR.
  97. 97.Fabin Rasheed. 2020. Gpt3 sees.
  98. 98.Partha Pratim Ray. 2023. Chatgpt: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope. Internet of Things and Cyber-Physical Systems.
  99. 99.Anna Rogers, Matt Gardner, and Isabelle Augenstein. 2022. Qa dataset explosion: A taxonomy of nlp resources for question answering and reading comprehension. ACM Comput. Surv. Just Accepted.
  100. 100.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684–10695.
  101. 101.David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli. 2019. Analysing mathematical reasoning abilities of neural models. In International Conference on Learning Representations.
  102. 102.Stephen Shankland. 2023. Why the chatgpt ai chatbot is blowing everyone’s mind.
  103. 103.Yiqiu Shen, Laura Heacock, Jonathan Elias, Keith D Hentel, Beatriu Reig, George Shih, and Linda Moy. 2023. Chatgpt and other large language models are double-edged swords.
  104. 104.Zhengxiang Shi, Qiang Zhang, and Aldo Lipani. 2022a. Stepgame: A new benchmark for robust multi-hop spatial reasoning in texts. Proceedings of the AAAI Conference on Artificial Intelligence, 36(10):11321–11329.
  105. 105.Zhengxiang Shi, Qiang Zhang, and Aldo Lipani. 2022b. StepGame: A new benchmark for robust multi-hop spatial reasoning in texts. Proceedings of the AAAI Conference on Artificial Intelligence, 36(10):11321–11329.
  106. 106.Denis Shiryaev. 2022. Drawing mona lisa with chatgpt.
  107. 107.Kurt Shuster, Jing Xu, Mojtaba Komeili, Da Ju, Eric Michael Smith, Stephen Roller, Megan Ung, Moya Chen, Kushal Arora, Joshua Lane, Morteza Behrooz, William Ngan, Spencer Poff, Naman Goyal, Arthur Szlam, Y-Lan Boureau, Melanie Kambadur, and Jason Weston. 2022. Blenderbot 3: a deployed conversational agent that continually learns to responsibly engage.
  108. 108.Koustuv Sinha, Shagun Sodhani, Jin Dong, Joelle Pineau, and William L Hamilton. 2019. Clutrr: A diagnostic benchmark for inductive reasoning from text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4506–4515.
  109. 109.Noah Smith. 2023. Why does chatgpt constantly lie?
  110. 110.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615.
  111. 111.Shane Storks, Qiaozi Gao, and Joyce Y Chai. 2019. Commonsense reasoning for natural language understanding: A survey of benchmarks, resources, and approaches. arXiv preprint arXiv:1904.01172, pages 1–60.
  112. 112.Dan Su, Xiaoguang Li, Jindi Zhang, Lifeng Shang, Xin Jiang, Qun Liu, and Pascale Fung. 2022. Read before generate! faithful long form question answering with machine reading. In Findings of the Association for Computational Linguistics: ACL 2022, pages 744–756.
  113. 113.Dan Su, Tiezheng Yu, and Pascale Fung. 2021. Improve query focused abstractive summarization by incorporating answer relevance. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3124–3131.
  114. 114.Zhiqing Sun, Xuezhi Wang, Yi Tay, Yiming Yang, and Denny Zhou. 2022. Recitation-augmented language models. In The Eleventh International Conference on Learning Representations.
  115. 115.Teo Susnjak. 2022. Chatgpt: The end of online exam integrity? arXiv preprint arXiv:2212.09292.
  116. 116.Alon Talmor, Yanai Elazar, Yoav Goldberg, and Jonathan Berant. 2020. olmpics-on what language model pre-training captures. Transactions of the Association for Computational Linguistics, 8:743–758.
  117. 117.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2018. Commonsenseqa: A question answering challenge targeting commonsense knowledge. arXiv preprint arXiv:1811.00937.
  118. 118.NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022. No language left behind: Scaling human-centered machine translation.
  119. 119.Richmond Thomason. 2018. Logic and artificial intelligence.
  120. 120.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. 2022. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239.
  121. 121.H Holden Thorp. 2023. Chatgpt is fun, but not an author.
  122. 122.Giuseppe Venuto. 2023. Giuven95/chatgpt-failures: Chatgpt failure archive.
  123. 123.Douglas Walton. 2014. Abductive reasoning. University of Alabama Press.
  124. 124.Ada Wan. 2022. Fairness in representation for multilingual NLP: Insights from controlled experiments on conditional language modeling. In International Conference on Learning Representations.
  125. 125.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018a. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 353–355, Brussels, Belgium. Association for Computational Linguistics.
  126. 126.Su Wang, Greg Durrett, and Katrin Erk. 2018b. Modeling semantic plausibility by injecting world knowledge. arXiv preprint arXiv:1804.00619.
  127. 127.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  128. 128.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. Self-instruct: Aligning language models with self-generated instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13484–13508, Toronto, Canada. Association for Computational Linguistics.
  129. 129.Peter Cathcart Wason and Philip Nicholas Johnson-Laird. 1972. Psychology of reasoning: Structure and content, volume 86. Harvard University Press.
  130. 130.Taylor Webb, Keith J. Holyoak, and Hongjing Lu. 2022a. Emergent analogical reasoning in large language models.
  131. 131.Taylor Webb, Keith J Holyoak, and Hongjing Lu. 2022b. Emergent analogical reasoning in large language models. arXiv preprint arXiv:2212.09196.
  132. 132.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022a. Emergent abilities of large language models. Transactions on Machine Learning Research.
  133. 133.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed H Chi, Quoc V Le, Denny Zhou, et al. 2022b. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems.
  134. 134.Jason Weston, Antoine Bordes, Sumit Chopra, and Tomás Mikolov. 2016a. Towards ai-complete question answering: A set of prerequisite toy tasks. In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings.
  135. 135.Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M Rush, Bart Van Merriënboer, Armand Joulin, and Tomas Mikolov. 2016b. Towards ai-complete question answering: A set of prerequisite toy tasks. In 4th International Conference on Learning Representations, ICLR 2016.
  136. 136.Bryan Wilie, Karissa Vincentio, Genta Indra Winata, Samuel Cahyawijaya, Xiaohong Li, Zhi Yuan Lim, Sidik Soleman, Rahmad Mahendra, Pascale Fung, Syafri Bahar, et al. 2020. Indonlu: Benchmark and resources for evaluating indonesian natural language understanding. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, pages 843–857.
  137. 137.Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra, Fajri Koto, Ade Romadhony, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Pascale Fung, Timothy Baldwin, Jey Han Lau, Rico Sennrich, and Sebastian Ruder. 2022. Nusax: Multilingual parallel sentiment dataset for 10 indonesian local languages.
  138. 138.Cameron R. Wolfe. 2023. Specialized llms: Chatgpt, lamda, galactica, codex, sparrow, and more.
  139. 139.BigScience Workshop, :, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major, Iz Beltagy, Huu Nguyen, Lucile Saulnier, Samson Tan, Pedro Ortiz Suarez, Victor Sanh, Hugo Laurençon, Yacine Jernite, Julien Launay, Margaret Mitchell, Colin Raffel, Aaron Gokaslan, Adi Simhi, Aitor Soroa, Alham Fikri Aji, Amit Alfassy, Anna Rogers, Ariel Kreisberg Nitzav, Canwen Xu, Chenghao Mou, Chris Emezue, Christopher Klamm, Colin Leong, Daniel van Strien, David Ifeoluwa Adelani, Dragomir Radev, Eduardo González Ponferrada, Efrat Levkovizh, Ethan Kim, Eyal Bar Natan, Francesco De Toni, Gérard Dupont, Germán Kruszewski, Giada Pistilli, Hady Elsahar, Hamza Benyamina, Hieu Tran, Ian Yu, Idris Abdulmumin, Isaac Johnson, Itziar Gonzalez-Dios, Javier de la Rosa, Jenny Chim, Jesse Dodge, Jian Zhu, Jonathan Chang, Jörg Frohberg, Joseph Tobing, Joydeep Bhattacharjee, Khalid Almubarak, Kimbo Chen, Kyle Lo, Leandro Von Werra, Leon Weber, Long Phan, Loubna Ben allal, Ludovic Tanguy, Manan Dey, Manuel Romero Muñoz, Maraim Masoud, María Grandury, Mario Šaško, Max Huang, Maximin Coavoux, Mayank Singh, Mike Tian-Jian Jiang, Minh Chien Vu, Mohammad A. Jauhar, Mustafa Ghaleb, Nishant Subramani, Nora Kassner, Nurulaqilla Khamis, Olivier Nguyen, Omar Espejel, Ona de Gibert, Paulo Villegas, Peter Henderson, Pierre Colombo, Priscilla Amuok, Quentin Lhoest, Rheza Harliman, Rishi Bommasani, Roberto Luis López, Rui Ribeiro, Salomey Osei, Sampo Pyysalo, Sebastian Nagel, Shamik Bose, Shamsuddeen Hassan Muhammad, Shanya Sharma, Shayne Longpre, Somaieh Nikpoor, Stanislav Silberberg, Suhas Pai, Sydney Zink, Tiago Timponi Torrent, Timo Schick, Tristan Thrush, Valentin Danchev, Vassilina Nikoulina, Veronika Laippala, Violette Lepercq, Vrinda Prabhu, Zaid Alyafeai, Zeerak Talat, Arun Raja, Benjamin Heinzerling, Chenglei Si, Davut Emre Taşar, Elizabeth Salesky, Sabrina J. Mielke, Wilson Y. Lee, Abheesht Sharma, Andrea Santilli, Antoine Chaffin, Arnaud Stiegler, Debajyoti Datta, Eliza Szczechla, Gunjan Chhablani, Han Wang, Harshit Pandey, Hendrik Strobelt, Jason Alan Fries, Jos Rozen, Leo Gao, Lintang Sutawika, M Saiful Bari, Maged S. Al-shaibani, Matteo Manica, Nihal Nayak, Ryan Teehan, Samuel Albanie, Sheng Shen, Srulik Ben-David, Stephen H. Bach, Taewoon Kim, Tali Bers, Thibault Fevry, Trishala Neeraj, Urmish Thakker, Vikas Raunak, Xiangru Tang, Zheng-Xin Yong, Zhiqing Sun, Shaked Brody, Yallow Uri, Hadar Tojarieh, Adam Roberts, Hyung Won Chung, Jaesung Tae, Jason Phang, Ofir Press, Conglong Li, Deepak Narayanan, Hatim Bourfoune, Jared Casper, Jeff Rasley, Max Ryabinin, Mayank Mishra, Minjia Zhang, Mohammad Shoeybi, Myriam Peyrounette, Nicolas Patry, Nouamane Tazi, Omar Sanseviero, Patrick von Platen, Pierre Cornette, Pierre François Lavallée, Rémi Lacroix, Samyam Rajbhandari, Sanchit Gandhi, Shaden Smith, Stéphane Requena, Suraj Patil, Tim Dettmers, Ahmed Baruwa, Amanpreet Singh, Anastasia Cheveleva, Anne-Laure Ligozat, Arjun Subramonian, Aurélie Névéol, Charles Lovering, Dan Garrette, Deepak Tunuguntla, Ehud Reiter, Ekaterina Taktasheva, Ekaterina Voloshina, Eli Bogdanov, Genta Indra Winata, Hailey Schoelkopf, Jan-Christoph Kalo, Jekaterina Novikova, Jessica Zosa Forde, Jordan Clive, Jungo Kasai, Ken Kawamura, Liam Hazan, Marine Carpuat, Miruna Clinciu, Najoung Kim, Newton Cheng, Oleg Serikov, Omer Antverg, Oskar van der Wal, Rui Zhang, Ruochen Zhang, Sebastian Gehrmann, Shachar Mirkin, Shani Pais, Tatiana Shavrina, Thomas Scialom, Tian Yun, Tomasz Limisiewicz, Verena Rieser, Vitaly Protasov, Vladislav Mikhailov, Yada Pruksachatkun, Yonatan Belinkov, Zachary Bamberger, Zdeněk Kasner, Alice Rueda, Amanda Pestana, Amir Feizpour, Ammar Khan, Amy Faranak, Ana Santos, Anthony Hevia, Antigona Unldreaj, Arash Aghagol, Arezoo Abdollahi, Aycha Tammour, Azadeh HajiHosseini, Bahareh Behroozi, Benjamin Ajibade, Bharat Saxena, Carlos Muñoz Ferrandis, Danish Contractor, David Lansky, Davis David, Douwe Kiela, Duong A. Nguyen, Edward Tan, Emi Baylor, Ezinwanne Ozoani, Fatima Mirza, Frankline Ononiwu, Habib Rezanejad, Hessie Jones, Indrani Bhattacharya, Irene Solaiman, Irina Sedenko, Isar Nejadgholi, Jesse Passmore, Josh Seltzer, Julio Bonis Sanz, Livia Dutra, Mairon Samagaio, Maraim Elbadri, Margot Mieskes, Marissa Gerchick, Martha Akinlolu, Michael McKenna, Mike Qiu, Muhammed Ghauri, Mykola Burynok, Nafis Abrar, Nazneen Rajani, Nour Elkott, Nour Fahmy, Olanrewaju Samuel, Ran An, Rasmus Kromann, Ryan Hao, Samira Alizadeh, Sarmad Shubber, Silas Wang, Sourav Roy, Sylvain Viguier, Thanh Le, Tobi Oyebade, Trieu Le, Yoyo Yang, Zach Nguyen, Abhinav Ramesh Kashyap, Alfredo Palasciano, Alison Callahan, Anima Shukla, Antonio Miranda-Escalada, Ayush Singh, Benjamin Beilharz, Bo Wang, Caio Brito, Chenxi Zhou, Chirag Jain, Chuxin Xu, Clémentine Fourrier, Daniel León Periñán, Daniel Molano, Dian Yu, Enrique Manjavacas, Fabio Barth, Florian Fuhrimann, Gabriel Altay, Giyaseddin Bayrak, Gully Burns, Helena U. Vrabec, Imane Bello, Ishani Dash, Jihyun Kang, John Giorgi, Jonas Golde, Jose David Posada, Karthik Rangasai Sivaraman, Lokesh Bulchandani, Lu Liu, Luisa Shinzato, Madeleine Hahn de Bykhovetz, Maiko Takeuchi, Marc Pàmies, Maria A Castillo, Marianna Nezhurina, Mario Sänger, Matthias Samwald, Michael Cullan, Michael Weinberg, Michiel De Wolf, Mina Mihaljcic, Minna Liu, Moritz Freidank, Myungsun Kang, Natasha Seelam, Nathan Dahlberg, Nicholas Michio Broad, Nikolaus Muellner, Pascale Fung, Patrick Haller, Ramya Chandrasekhar, Renata Eisenberg, Robert Martin, Rodrigo Canalli, Rosaline Su, Ruisi Su, Samuel Cahyawijaya, Samuele Garda, Shlok S Deshmukh, Shubhanshu Mishra, Sid Kiblawi, Simon Ott, Sinee Sang-aroonsiri, Srishti Kumar, Stefan Schweter, Sushil Bharati, Tanmay Laud, Théo Gigant, Tomoya Kainuma, Wojciech Kusa, Yanis Labrak, Yash Shailesh Bajaj, Yash Venkatraman, Yifan Xu, Yingxin Xu, Yu Xu, Zhe Tan, Zhongli Xie, Zifan Ye, Mathilde Bras, Younes Belkada, and Thomas Wolf. 2022. Bloom: A 176b-parameter open-access multilingual language model.
  140. 140.Yan Xu, Etsuko Ishii, Samuel Cahyawijaya, Zihan Liu, Genta Indra Winata, Andrea Madotto, Dan Su, and Pascale Fung. 2022. Retrieval-free knowledge-grounded dialogue response generation with adapters. In Proceedings of the Second DialDoc Workshop on Document-grounded Dialogue and Conversational Question Answering, pages 93–107.
  141. 141.Yan Xu, Deqian Kong, Dehong Xu, Ziwei Ji, Bo Pang, Pascale Fung, and Ying Nian Wu. 2023. Diverse and faithful knowledge-grounded dialogue generation via sequential posterior inference. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 38518–38534. PMLR.
  142. 142.Yunyi Yang, Yunhao Li, and Xiaojun Quan. 2021. Ubar: Towards fully end-to-end task-oriented dialog system with gpt-2. Proceedings of the AAAI Conference on Artificial Intelligence, 35(16):14230–14238.
  143. 143.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2369–2380.
  144. 144.Tiezheng Yu, Wenliang Dai, Zihan Liu, and Pascale Fung. 2021a. Vision guided generative pre-trained language models for multimodal abstractive summarization. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3995–4007, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  145. 145.Tiezheng Yu, Zihan Liu, and Pascale Fung. 2021b. Adaptsum: Towards low-resource domain adaptation for abstractive summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5892–5904.
  146. 146.Xiaoxue Zang, Abhinav Rastogi, Srinivas Sunkara, Raghav Gupta, Jianguo Zhang, and Jindong Chen. 2020. Multiwoz 2.2: A dialogue dataset with additional annotation corrections and state tracking baselines. ACL 2020, page 109.
  147. 147.Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. 2022. Star: Bootstrapping reasoning with reasoning. In Advances in Neural Information Processing Systems.
  148. 148.Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.
  149. 149.Jeffrey Zhao, Raghav Gupta, Yuanbin Cao, Dian Yu, Mingqiu Wang, Harrison Lee, Abhinav Rastogi, Izhak Shafran, and Yonghui Wu. 2022. Description-driven task-oriented dialog modeling. ArXiv, abs/2201.08904.
  150. 150.Xueliang Zhao, Wei Wu, Can Xu, Chongyang Tao, Dongyan Zhao, and Rui Yan. 2020. Knowledge-grounded dialogue generation with pre-trained language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3377–3390.
  151. 151.Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing. 2023a. Exploring ai ethics of chatgpt: A diagnostic analysis. arXiv preprint arXiv:2301.12867.
  152. 152.Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing. 2023b. Red teaming chatgpt via jailbreaking: Bias, robustness, reliability and toxicity.

Citation

MLA
Bang, Y., et al. “A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity”. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 675–718, https://doi.org/10.18653/v1/2023.ijcnlp-main.45.
APA
Bang, Y., Cahyawijaya, S., Lee, N., Dai, W., Su, D., Wilie, B., Lovenia, H., Ji, Z., Yu, T., Chung, W., Do, Q. V., Xu, Y., & Fung, P. (2023). A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), 675–718. https://doi.org/10.18653/v1/2023.ijcnlp-main.45
Chicago
Bang, Y., S. Cahyawijaya, N. Lee, et al. 2023. “A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity”. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), 675–718. https://doi.org/10.18653/v1/2023.ijcnlp-main.45.
Harvard
Bang, Y. et al. (2023) “A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity”, Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 675–718. Available at: https://doi.org/10.18653/v1/2023.ijcnlp-main.45.
Vancouver
1. Bang Y, Cahyawijaya S, Lee N, et al (2023) A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity. In: Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 675–718

BibTeX

@inproceedings{bang-etal-2023-multitask,
    title = "A Multitask, Multilingual, Multimodal Evaluation of {C}hat{GPT} on Reasoning, Hallucination, and Interactivity",
    author = "Bang, Yejin  and
      Cahyawijaya, Samuel  and
      Lee, Nayeon  and
      Dai, Wenliang  and
      Su, Dan  and
      Wilie, Bryan  and
      Lovenia, Holy  and
      Ji, Ziwei  and
      Yu, Tiezheng  and
      Chung, Willy  and
      Do, Quyet V.  and
      Xu, Yan  and
      Fung, Pascale",
    editor = "Park, Jong C.  and
      Arase, Yuki  and
      Hu, Baotian  and
      Lu, Wei  and
      Wijaya, Derry  and
      Purwarianti, Ayu  and
      Krisnadhi, Adila Alfa",
    booktitle = "Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = nov,
    year = "2023",
    address = "Nusa Dua, Bali",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.ijcnlp-main.45/",
    doi = "10.18653/v1/2023.ijcnlp-main.45",
    pages = "675--718"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF