EditEval: An Instruction-Based Benchmark for Text Improvements
Jane Dwivedi-Yu ♢^{\diamondsuit}♢ Timo Schick ♢^{\diamondsuit}♢ Zhengbao Jiang ♢,♡^{\diamondsuit,\heartsuit}♢,♡
Maria Lomeli ♢^{\diamondsuit}♢ Patrick Lewis ♢^{\diamondsuit}♢ Gautier Izacard ♢,♣^{\diamondsuit,\clubsuit}♢,♣
Edouard Grave ♢^{\diamondsuit}♢ Sebastian Riedel ♢,♠^{\diamondsuit,\spadesuit}♢,♠ Fabio Petroni ♢^{\diamondsuit}♢
♢^{\diamondsuit}♢ Meta AI Research, ♡^{\heartsuit}♡ Carnegie Mellon University,
♣^{\clubsuit}♣ Inria & ENS, PSL University, ♠^{\spadesuit}♠ University College London
{janeyu,schick,zhengbao,marialomeli,plewis,gizacard,
egrave,sriedel,fabiopetroni}@meta.com
Maria Lomeli ♢^{\diamondsuit}♢ Patrick Lewis ♢^{\diamondsuit}♢ Gautier Izacard ♢,♣^{\diamondsuit,\clubsuit}♢,♣
Edouard Grave ♢^{\diamondsuit}♢ Sebastian Riedel ♢,♠^{\diamondsuit,\spadesuit}♢,♠ Fabio Petroni ♢^{\diamondsuit}♢
♢^{\diamondsuit}♢ Meta AI Research, ♡^{\heartsuit}♡ Carnegie Mellon University,
♣^{\clubsuit}♣ Inria & ENS, PSL University, ♠^{\spadesuit}♠ University College London
{janeyu,schick,zhengbao,marialomeli,plewis,gizacard,
egrave,sriedel,fabiopetroni}@meta.com
Abstract
Evaluation of text generation to date has primarily focused on content created sequentially, rather than improvements on a piece of text. Writing, however, is naturally an iterative and incremental process that requires expertise in different modular skills such as fixing outdated information or making the style more consistent. Even so, comprehensive evaluation of a model's capacity to perform these skills and the ability to edit remains sparse. This work presents EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL: An instruction-based, benchmark and evaluation suite that leverages high-quality existing and new datasets for automatic evaluation of editing capabilities such as making text more cohesive and paraphrasing. We evaluate several pre-trained models, which shows that InstructGPT and PEER perform the best, but that most baselines fall below the supervised SOTA, particularly when neutralizing and updating information. Our analysis also shows that commonly used metrics for editing tasks do not always correlate well, and that optimization for prompts with the highest performance does not necessarily entail the strongest robustness to different models. Through the release of this benchmark,1^{1}1 and a publicly available leaderboard challenge,2^{2}2 we hope to unlock future research in developing models capable of iterative and more controllable editing.
1. Introduction
Large pre-trained language models have shown impressive text generation capabilities for a wide variety of tasks such as question answering, textual entailment, and summarization [1, 2, 3, 4, 5, 6]. However, to date, most work employing language models has focused on generating immutable text in a single pass. This is in stark contrast to the way in which humans develop articles of text, which is naturally an iterative process of small steps, each with a precise purpose [7]. This is a crucial process because it allows for analysis of "what’s working, what isn’t, and what it still needs" and adaptation to these needs along the way [8]. In many cases, a needed change may only become apparent after much of the text is created, such as in the case of a reorganization or fixing inconsistencies or contradictions [9]. In this way, the current paradigm of generating text passages in a single pass can be severely limiting.
Additionally, the current paradigm of continuous left-to-right generation is less controllable and not flexible to human-in-the-loop collaboration and feedback, and this absence of experienced human mediation in the writing process can be highly detrimental to the quality of the final product [10]. While there are some existing production tools geared towards working with humans to compose articles and emails, such as Smart Compose from Google1 and text predictions from Microsoft2, these mostly focus on sentence completion and are not developed to improve upon prior text. A more powerful editing assistant, however, would not only be capable of providing recommendations for continuations of the text but also improvements upon the already existing text, such as making the tone more consistent, making diction more precise, or adding more engaging information. These AI tools should also permit iterative and non-sequential development of the text [7], which is naturally unavoidable, for example, if new or missing information or external references are required to update the text or if a reshuffling/rebalancing of text is needed.
In this work, we alternatively promote iterative text generation and improvement—successive iterations of modular additions and modifications of the text that are relevant to text editing such as making text clearer and adding missing information. Many datasets for natural language tasks are actually annotated at the sentence or paragraph level, rather than document or article level, naturally lending well to evaluating iterative edits.
We create EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL, a benchmark and evaluation suite that leverages high-quality existing and new datasets for automatic evaluation of editing capabilities. Currently, many of these pertinent datasets live in separate packages and are often formatted in uniquely distinct ways. EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL downloads each dataset from their most recent version and standardizes each into a single format conducive to evaluation. Additionally, we include popular metrics for each task and a set of human-generated prompts to robustly measure a model's capability in executing the modular task when instructed. Figure 1 shows examples of such prompts and an example of a corresponding edit that we might expect for the given text. Using these prompts, we evaluate and compare several state-of-the-art language models, such as GPT-3 [4], OPT [5], and PEER [11]. In summary, our contributions are as follows:
- We identify a set of tasks and datasets relevant to iterative text improvement and provide a pipeline to download and process these datasets into a single format.
- We open-source a publicly available instruction-based benchmark for automatic evaluation according to metrics commonly used for each editing task.
- We introduce a new dataset, WAFER-INSERT{\mathchoice{\text{W{\scriptsize AFER}-I{\scriptsize NSERT}}}{\text{W{\scriptsize AFER}-I{\scriptsize NSERT}}}{\text{W{\scriptscriptstyle AFER}-I{\scriptscriptstyle NSERT}}}{\text{WAFER-INSERT}}}WAFER-INSERT, for evaluating a model's capability to update information, which is based on the WAFER dataset [12].
- We provide a comparison of various state-of-the-art baselines evaluated on EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL at the dataset and prompt level.
2. Related Work
Several multitask evaluation benchmarks have been open-sourced to the community to support progress in natural language understanding including GLUE [13], SuperGLUE [14], decaNLP [15], and GEM [16]. These datasets, however, focus on a broad set of tasks in NLP (e.g., question answering, reading comprehension, and natural language inference). While all of these tasks are critical to natural language understanding, EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL focuses on curating a benchmark for measuring a model's capability to improve and edit text.
There are several datasets which focus on iterative text revisions in the domain of Wikipedia [17, 18], academic essays [19], and news articles [20]. These works, however, focus on one particular domain and in some cases, a particular style like argumentative writing [19]. EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL, on the other hand, includes examples from multiple domains: Wikipedia, Wikinews, news articles, and arXiv. ITERATER{\mathchoice{\text{I{\scriptsize TERA}T{\scriptsize E}R}}{\text{I{\scriptsize TERA}T{\scriptsize E}R}}{\text{I{\scriptscriptstyle TERA}T{\scriptscriptstyle E}R}}{\text{ITERATER}}}ITERATER [21] is perhaps closet to EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL in that it provides iterative tasks from multiple domains, but it has a limited number of such tasks: fluency, coherence, clarity, style, and meaning-changed. Because this is a great starting point, we have included ITERATER{\mathchoice{\text{I{\scriptsize TERA}T{\scriptsize E}R}}{\text{I{\scriptsize TERA}T{\scriptsize E}R}}{\text{I{\scriptscriptstyle TERA}T{\scriptscriptstyle E}R}}{\text{ITERATER}}}ITERATER in EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL, and we additionally develop prompts for these tasks since ITERATER{\mathchoice{\text{I{\scriptsize TERA}T{\scriptsize E}R}}{\text{I{\scriptsize TERA}T{\scriptsize E}R}}{\text{I{\scriptscriptstyle TERA}T{\scriptscriptstyle E}R}}{\text{ITERATER}}}ITERATER is not instruction-based. Additionally, unlike ITERATER{\mathchoice{\text{I{\scriptsize TERA}T{\scriptsize E}R}}{\text{I{\scriptsize TERA}T{\scriptsize E}R}}{\text{I{\scriptscriptstyle TERA}T{\scriptscriptstyle E}R}}{\text{ITERATER}}}ITERATER, EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL includes novel datasets for tasks such as updating text using new information and neutralizing the text, which are core components of editing a factually-correct and unbiased article.
3. The EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL Benchmark
EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL is an instruction-based benchmark for iterative text generation/modification. EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL sources existing high-quality datasets—most with human annotations—containing tasks relevant to editing. These datasets are combined into a unified evaluation tool and can be evaluated with any metric provided in EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL. A task here refers to a type of edit (e.g., simplification or neutralization), and the specific task dictates which set of prompts to be used (e.g., simplify this text).
We consider seven editing tasks in EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL. The corresponding datasets for each task included in EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL are enumerated in Table 1, along with the size of the test set. For ease of evaluation, we define a consistent format for all datasets in the EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL benchmark. Each dataset of every task has five core fields: ID, input text, gold edits, task type, and reference documents. The input text is the original text before revision, and the gold edits are the target edits for that specific task type. Lastly, the reference documents provide textual information from external articles or documents that are relevant to the task. The task that requires reference documents is updating, and otherwise, the reference documents field is empty.
The datasets in EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL were selected if they test a capability relevant to the art of editing and contain human-annotated gold edits, if possible. We also endeavored to include datasets that are broadly used by the community. The datasets in EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL are by no means exhaustive, but the EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL framework is flexible such that it can easily extend to new datasets and metrics in future versions.
3.1 Fluency, Clarity, and Coherence
In this section, we describe the two datasets that compose this set of tasks: fluency (fixing grammatical or spelling errors), clarity (making the text clearer), and coherence (making the text more cohesive).
JFLEG JHU FLuency-Extended GUG [22] focuses solely on the first task of fluency. JFLEG is based on the GUG (Grammatical vs Un-Grammatical) dataset [23], which is a dataset of sentences originally annotated for how grammatical the sentence is on a scale of 1 to 4. JFLEG builds upon the ungrammatical sentences in GUG and annotates each sentence with four corresponding corrected versions.
IteraTeR This dataset introduced by [21] contains both automatically-mined and human-annotated edits at the sentence and document-level. For our benchmark, we only utilize the sentence-level examples with human annotations. Additionally, ITERATER{\mathchoice{\text{I{\scriptsize TERA}T{\scriptsize E}R}}{\text{I{\scriptsize TERA}T{\scriptsize E}R}}{\text{I{\scriptscriptstyle TERA}T{\scriptscriptstyle E}R}}{\text{ITERATER}}}ITERATER has labels for the intent—the type of edit that produces the targets, which can be one of six classes: Fluency, coherence, clarity, style (conveying the writer's writing preferences), meaning-changed (updating or adding new information), and other (none of the others). We included all classes except style, meaning-changed, and other. We excluded style and other because these tasks had roughly 100 or less test examples, and the definitions were comparatively under-specified. We excluded meaning-changed because the task does not use reference documents for updating. This dataset is the only one in EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL that encompasses multiple tasks, and we refer to each respective subset using the abbreviations ITR-F (fluency), ITR-L (clarity), and ITR-O (coherence).
3.2 Paraphrasing
STSB For paraphrasing, we use the STS benchmark from SemEval-2018 [24], which comprises English datasets used in the STS tasks of SemEval between 2012 and 2017. The selection of datasets includes text from image captions, news headlines and user forums. Each example contains an original sentence, a target sentence, and a similarity score indicating whether the target is a paraphrase of the original. This dataset is used for classification or regression, but for EditEval, we utilize all instances that we are confident are paraphrases, i.e., have the max similarity score of 5, as targets for generation evaluation. While other datasets such as ParaSCI [25] exist for paraphrase generation, these are automatically curated rather than human annotated, and EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL strives to utilize human-annotated datasets where possible.
3.3 Simplification
Simplification can be considered a very similar task to paraphrasing with the additional constraint that the output must be simpler than the input. The datasets we utilize for simplification are TurkCorpus [26] and ASSET [27].
TurkCorpus This dataset, like ASSET, builds upon the Parallel Wikipedia Simplification (PWKP) [28]. The PWKP dataset uses the Simple English Wikipedia and Standard English Wikipedia in parallel to create original-simplification pairs automatically. However, several works found PWKP to have a large proportion of targets that are not simplified or only partially aligned with the input [29, 30, 31, 32], leading to the creation of a human-annotated corpus, TurkCorpus. TurkCorpus was manually created with eight reference simplifications for each original sentence in PWKP, but only used simplifications that are possible without deleting content or splitting sentences.
ASSET Because TurkCorpus encompassed only specific kinds of simplifications, this led to the creation of ASSET, which provides manually-produced simplifications through a much broader set of transformations. We include both in EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL, for the sake of comprehensiveness.
3.4 Neutralization
The task of neutralization refers to making a text more neutral. For example, in the sentence "Obama was an excellent president who served two terms from 2008 to 2016" the term excellent violates Wikipedia's neutral point of view (POV) policy3. For information-intensive content like Wikipedia and news articles in particular, reducing bias is crucial because bias is the single largest source of distrust in the media [33].
3.
https://en.wikipedia.org/wiki/Wikipedia:Neutral point of view
WNC We use the Wiki Neutrality Corpus [34], a collection of original and de-biased sentence pairs mined from Wikipedia edits by carefully filtering based on the editor's comments. While ideally we would like to include a human-annotated dataset, to our knowledge there does not exist a dataset for de-biasing article content at the sentence level.
3.5 Updating
In this section we describe the task of updating information which requires references, text from external sources that are relevant to the particular task. Because of token-length restrictions, each external article is chunked into texts of fixed length. We limit the scope of the task to three chunks, and we refer to these selected chunks as our reference documents. These references documents are represented in the edits by their index in the reference documents field (e.g., the first would be demarcated as [0]), and we discuss below how these reference documents were selected.
WAFER-insert The first dataset for updating information that we use is the WAFER dataset [12], which is a dataset collected from Wikipedia inline citations. Each instance of the original WAFER dataset contains a claim, the text surrounding the claim, and a set of external references, where the task is to choose one of the references to be cited after the claim. While the original intention of WAFER was to measure a system's capability to choose the correct citation, EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL utilizes WAFER for the task of inserting new information using content from the reference documents. We create WAFER-INSERT{\mathchoice{\text{W{\scriptsize AFER}-I{\scriptsize NSERT}}}{\text{W{\scriptsize AFER}-I{\scriptsize NSERT}}}{\text{W{\scriptscriptstyle AFER}-I{\scriptscriptstyle NSERT}}}{\text{WAFER-INSERT}}}WAFER-INSERT, which differs from WAFER in that the claim is deleted from the input. The goal here is to derive the original claim from the references and insert it into the text. For the reference documents, we select the top three chunks from the inline citation chunks that have the highest scores, using results from the verification engine introduced in [12].
FRUIT In addition to WAFER-INSERT{\mathchoice{\text{W{\scriptsize AFER}-I{\scriptsize NSERT}}}{\text{W{\scriptsize AFER}-I{\scriptsize NSERT}}}{\text{W{\scriptscriptstyle AFER}-I{\scriptscriptstyle NSERT}}}{\text{WAFER-INSERT}}}WAFER-INSERT, we include the FRUIT dataset ([35]), a dataset collected by comparing two snapshots of a Wikipedia article where one contains updated or new information. The reference documents were identified by searching for other Wikipedia articles that provide evidence that supports the update. However, because there is no certainty that the identified evidentiary Wikipedia articles support the claim, the authors of FRUIT created a gold set by employing human annotation to filter out any new claims that are unsupported. We include this gold set in EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL, and only include reference documents if they actually appear in the output. Unlike WAFER-INSERT{\mathchoice{\text{W{\scriptsize AFER}-I{\scriptsize NSERT}}}{\text{W{\scriptsize AFER}-I{\scriptsize NSERT}}}{\text{W{\scriptscriptstyle AFER}-I{\scriptscriptstyle NSERT}}}{\text{WAFER-INSERT}}}WAFER-INSERT, the target edit contains not only the updated information but also the citation. For EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL, this is for verification purposes only, and the citation is removed when computing the metrics.
4. Metrics
The metrics we included in EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL are ones that are (1) shown to have significant correlation with human judgement for a task in EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL and (2) commonly used to benchmark one of the datasets in EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL. Below, we discuss each set of metrics in detail.
EM and EM-Diff
Exact match (EM) is the percentage of examples for which the performed edit exactly matches any of the targets. EM-Diff is a variant of EM that is computed on the diff level, where diffs are obtained using Python's
difflib library. For a model output OOO, we compute EM-Diff as follows:SARI
Introduced by [26], SARI is an n-gram based metric commonly used for measuring simplification [36, 37] and other editing tasks such as sentence fusion [38]. It has been demonstrated to correlate most closely with human judgement for simplification compared to many other n-gram based metrics [26]. The metric measures how simplified a candidate system output is relative to the original and to the simplification references by rewarding words added, kept, or deleted in both the target and the output. More specifically, this is done by computing the arithmetic mean of n-gram F1-scores for each of the three operations. We utilize the EASSE [39] implementation of SARI, which addresses inconsistencies in the original implementation4.
BLEU and iBLEU
BLEU [40] is another n-gram based metric that encourages a high proportion of n-gram matches between the output and the targets. BLEU, originally intended for machine translation, is very commonly used for many editing tasks such as simplification [32, 41] and improving fluency [34, 21], and is shown to correlate well with human judgement of tasks such as grammaticality and meaning preservation [26].
For some tasks like simplification and paraphrasing, however, we require not only that the output is similar to the target, but that the output is sufficiently different from the input. iBLEU, a metric introduced by [42] is a weighted average of the BLEU score computed between the output and the targets and the negated BLEU score computed between the output and the input. More specifically, for a candidate output sentence O, human targets R, and an input text I, iBLEU is defined as:
[26] demonstrated that for simplification, iBLEU correlates on par or better with human judgement than BLEU does, though not as well as SARI on average.
GLEU GLEU [43] is another variant of BLEU frequently used for grammatical error correction [44, 45, 46]. The issue with using BLEU for minimal edits can be attributed to the difference between analyzing machine translation and editing tasks. In the former, an untranslated word should always be penalized, but in the editing setting, an unmodified word in both the target and the output does not necessarily need to be penalized. Unlike BLEU, GLEU is customized to penalize n-grams changed in the targets but left unchanged by the system output. [43] not only demonstrated that GLEU correlates well with human rankings of corrections, but also that GLEU correlates much better than BLEU does.
ROUGE and UpdateROUGE
For the task of updating or adding new information, we follow [35] and use ROUGE and UpdateROUGE [35]. ROUGE [47] is a popular n-gram based metric that is commonly used for evaluating summarization systems [48, 49], but is also used in other tasks such as improving fluency [50] and simplification [51]. ROUGE essentially measures the overlap in n-grams between the system output and the targets. UpdateROUGE, a simple modification of ROUGE, computes ROUGE on the updated sentences rather than the full text. This is intended for tasks such as updating, because a majority of the target will remain unchanged. On the other hand, when evaluating using ROUGE, a system can often superficially achieve high scores by simply copying the input.
5. Baselines
For each baseline, we use greedy decoding, and we do not perform any task-specific fine-tuning or in-context learning. We evaluate on EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL using the following baselines:
- GPT-3 ([4]) is a 175B parameter pretrained decoder-only model. We evaluate GPT-3 through OpenAI's API.5
- InstructGPT ([52]) is a variant of GPT-3 that was fine-tuned on a large dataset of instructions and corresponding outputs written by humans. We evaluate the text-davinci-001 version described in ([52]) since, at the time of writing, details about the training process for text-davinci-002 were not publicly available.
- OPT ([5]) is an open-source replica of GPT-3. Like GPT-3, it is not fine-tuned on any labeled data.
- T0 [53] is a pretrained encoder-decoder model, which has demonstrated better performance than GPT-3 on several tasks despite being much smaller. It is initialized from the LM Adapt variant of T5 [3] and is fine-tuned on examples from 170 existing NLP datasets that are prompted using around 2000 crowdsourced prompts.
- T0++ ([53]) is similar to T0, but trained on a few additional datasets from SuperGLUE [14].
- Tkkk-Instruct ([54]) is similar to T0 and T0++ but instead fine-tuned on their dataset, Natural Instructions v2, a collection of instructions for more than 1,600 tasks, including grammatical error correction and text simplification.
- PEER ([11]) A collaborative language model trained to infill parts of the writing process by leveraging self-training techniques. It is also initialized from the LM Adapt variant of T5, and further fine-tuned on edit histories from Wikipedia. We use the 3B and 11B PEER (SP) model (shortened here as PEER-3 and PEER-11, respectively), where SP refers to augmenting the training data with synthetic instructions and was shown to perform the best in [11].
6. Instructions
We evaluate these baselines on their general capability to accomplish each task when prompted in natural language in a zero-shot fashion. Because there are a diverse set of ways in which to instruct for each task, we manually construct a set of 3–11 prompts in order to more robustly evaluate performance. For each task prompt ttt and input iii, the model is given a formatted input following the template:
with an additional field for references, should they be required. Figure 2 shows an example of an input including references. For tasks without references, we exclude this field. Some slight modification to this template were made. For example, Tkkk-Instruct expects the prompt to be prefixed by the string "Definition:" rather than "Task:").
7. Results
We summarize results in Table 2 with the aforementioned baselines averaged over all datasets and the breakdown for each dataset in Table 3. To visualize the variance according to each model, we show boxplots for each dataset and model according to the SARI metric in Figure 3. Similarly, to visualize the variance for a given prompt, we present boxplots averaged across models in Figure 4. We discuss several observations below.
InstructGPT and PEER perform the best overall. In Table 2, we show the mean SARI scores for each model averaged across all tasks using the average, maximum, and minimum scores across prompts. When using the average and minimum across prompts (third and fifth column, respectively) we see that InstructGPT performs the best overall, but when using the maximum score across prompts (fourth column), PEER-11 performs the best. Table 3 enumerates the breakdown of the third column according to each dataset. In general, we see that InstructGPT achieves the highest scores with the exception of the updating and neutralization datasets, as well as ITR-F and ITR-L. For these datasets, the PEER models clearly outperform InstructGPT by a large margin, despite being nearly 60×\times× smaller than InstructGPT and GPT-3. The substantially smaller models (T0, T0++, and Tkkk-Instruct) struggle the most overall, even falling behind the copy baseline at times, except on ITR-L where Tkkk-Instruct performs the best.
Most baselines lag substantially behind the supervised SOTA, especially in the task of updating and neutralization. We show the supervised state-of-the-art results in the final row of Table 3, which in almost all cases surpasses the performance of the best baseline. The gap is largest for the tasks of neutralization and updating (34–50% decrease from the supervised SOTA to the best baseline scores), whereas for other tasks, this decrease is only within 5–14%. It is conceivable that the difficulty with these two particular tasks is a consequence of the comparatively fewer datasets and research devoted to them compared to that of the more mainstream NLP tasks, such as paraphrasing.
Tasks that are the most challenging are not necessarily ones with the highest variance across models. In observing Figure 3 (left), we see that the tasks which have the largest variance across models (assessed using the interquartile range (IQR)) are fluency and updating. This is despite the fact that the fluency datasets are arguably easier (i.e., many of the models come close to the supervised SOTA) than the updating datasets, exemplifying that difficulty and robustness can be independent axes. JFLEG also appears to be easier than ITR-F (average SARI score of 45.1 versus 38.2). This is not surprising since JFLEG sources from the TOEFL exam, which has primarily simpler and conversational sentences, whereas ITERATER{\mathchoice{\text{I{\scriptsize TERA}T{\scriptsize E}R}}{\text{I{\scriptsize TERA}T{\scriptsize E}R}}{\text{I{\scriptscriptstyle TERA}T{\scriptscriptstyle E}R}}{\text{ITERATER}}}ITERATER is composed of technical sentences from Wikipedia, ArXiv, and Wikinews. Likewise, TurkCorpus seems on average to be slightly easier than ASSET, which is expected since it was created to be more diverse than TurkCorpus.
PEER has the highest variance across all tasks, but OPT and GPT-3 are the least robust to different prompts.
From Figure 3 (right), we observe that the PEER models have the largest range in performance from dataset to dataset. Within each task, however, GPT-3 and OPT have the highest coefficient of variation or standard deviation normalized by the mean (6.74% and 6.70%, respectively), whereas for the 3B and 11B PEER models, these values are smaller (6.36% and 5.75%), as enumerated in Table 2. This could be a consequence of the fact that GPT-3 and OPT are not trained explicitly to follow instructions, whereas the remaining baselines are.
Prompts chosen according to maximum performance and prompts chosen according to robustness across models can be different. Ideally, we would like to create prompts that are not only robust to different models, but achieve the highest performance using the best baseline. In assessing variance from Figure 4, we see that certain prompts stand out as less robust relative to others. For example, for neutralization, Prompts #1, 2, and 7 are less robust likely because they use uncommonly used language such as "Remove points of views" or "Neutralize this text". Some of the prompts which are less robust for simplification (Prompts #4, 7) and paraphrasing (Prompts #4, 6) are sometimes ones with less specific commands such as "Rewrite this text" versus "Rewrite this with different wording"—in the case of the former, an empirical assessment shows that the models seem to more often copy the original text and make fewer modifications. Unfortunately, choosing prompts that are the most robust, does not always entail prompts which achieve the maximum score—Prompt #5 for clarity achieves the maximum but has the largest IQR. Some of the tasks exhibit a great degree of outlier behavior (coherence, paraphrasing, or neutralization), which is either due to T0 performing exceedingly low or InstructGPT/PEER performing exceedingly well. Other tasks such as fluency and updating seem to have prompts with roughly a similar range of performance.
Different metrics do not always correlate well with each other. We measure the Pearson correlation between each pair of metrics using evaluation scores for all baselines, which is shown in Figure 5 as a heatmap. We exclude PEER in this analysis since it shows exceedingly strong performance in some cases, and we exclude the updating datasets since they are of a very different nature from the other datasets. We find that while families of variants like BLEU and iBLEU as well as ROUGE and UpdateROUGE show strong correlation within each set (>>> 0.97), the two sets are inversely correlated with one another (-0.29 to -0.1). ROUGE actually appears to be the metric that most conflicts with all other metrics, whereas GLEU seems to be the metric that is most in harmony with the rest (0.41–0.76). Though SARI is not correlated with ROUGE, it is the metric which shows the strongest correlation with EM-Diff (0.83) and UpdateROUGE (0.7).
8. Discussion
We present EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL, a benchmark composed of handcrafted, task-specific instructions for several editing datasets across multiple domains. EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL is a means of evaluating models for these tasks according to multiple popular metrics, all within a single, unified tool. We show that while state-of-the-art models such as InstructGPT and PEER have impressive performance, in general the baselines lag behind the supervised state-of-the-art, particularly for the task of updating and neutralization. Our analysis of metrics and prompts shows that several popular metrics are not well-correlated, even conflicting at times, and that small changes in the wording of a prompt can lead to substantial changes in performance and robustness across models. This suggests further work is needed to develop models comprehensively capable of executing editing tasks in addition to developing a standardized way of measuring editing capabilities and systematically selecting prompts. In releasing this work, we hope to bolster work in which language models are utilized for text generation that is iterative, and therefore potentially more controllable, collaborative, and capable of revising and correcting text.
Limitations
Our evaluation tool is by no means an exhaustive measurement of editing capabilities. Firstly, there are additional domains that could potentially be added to EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL, such as books and blogs; as it currently stands, EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL is primarily constructed from the domain of Wikipedia. Fortunately, EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL 's framework is flexible to the addition of datasets, provided that it has an input and target edit. In the same spirit, there are additional editing tasks such as verifying facts, citing, and reorganizing sentences/paragraphs which would be valuable to include in EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL. While we recognize these tasks as valuable to include in EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL, we consider these to be out of scope for the work at hand. Finally, our results demonstrate that many of the metrics give conflicting signal as to the rankings of the baselines, indicating further work is needed to identify better metrics for measuring overall editing capacity.
Appendix
A. Domains
In EDITEVAL{\mathchoice{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptsize DIT}E{\scriptsize VAL}}}{\text{E{\scriptscriptstyle DIT}E{\scriptscriptstyle VAL}}}{\text{EDITEVAL}}}EDITEVAL we strive to encompass datasets from many different domains, with an emphasis on factual content. Below in Table 4, we enumerate these domains.




