Understanding Iterative Revision from Human-Written Text
Wanyu DuVipul RahejaDhruv KumarZae Myung KimMelissa LopezDongyeop Kang
Presents ITERATER, a multi-domain corpus of iteratively revised text annotated with edit intentions across sentence and paragraph granularities, demonstrating that modeling human revision goals improves automatic text editing systems.
Writing is inherently an iterative and strategic process where human authors continuously refine their drafts across multiple cycles. However, existing automated text revision systems typically simplify this complex workflow into single-pass, sentence-level paraphrasing or confine their models to narrow, single-domain contexts. This gap between actual human writing behavior and automated tools restricts the practical utility of modern writing assistants.
The article introduces a comprehensive framework and a multi-domain corpus designed to model how humans iteratively revise text across various depths and granularities. The main objective is to evaluate how specific edit intentions affect document quality and demonstrate that incorporating these intentions significantly improves the performance of computational text revision models.
To accomplish this, the authors constructed ITERATER, a dataset comprising 31,631 iterative document revisions containing 196,987 edit actions across Wikipedia articles, ArXiv scientific paper abstracts, and Wikinews stories. The team developed an edit intention taxonomy covering fluency, clarity, coherence, style, and meaning changes. They established a reference set of 4,018 human edit actions annotated through crowdsourcing and expert linguists, trained a classifier to automatically label the remaining full dataset, and benchmarked both edit-based and generative machine learning models against human revisions.
Key findings show that human editors perform the vast majority of edits during initial revision cycles, with editing activity sharply decreasing in deeper iterations. Clarity and fluency edits dominate across all domains, while scientific papers also exhibit high proportions of substantive meaning changes. Evaluators found that clarity, fluency, and coherence edits consistently enhance document quality, whereas subjective style edits slightly degrade it. When automated models are provided with explicit edit intentions, text generation performance improves markedly across evaluation metrics. However, in head-to-head comparisons, human revisions still outperform the best model revisions in overall quality in roughly 83% of evaluated document revisions, though models achieve competitive parity in meaning preservation and basic fluency.
These results demonstrate that automated text revision cannot rely solely on generic paraphrasing; systems must actively incorporate domain context and explicit editing goals to generate high-quality text. The findings also reveal that existing automated evaluation metrics correlate poorly with human judgments of fluency and coherence, indicating a clear operational risk in relying on standard metrics for text quality assurance.
Organizations developing or deploying automated writing assistance should integrate explicit edit-intention controls into their generative architectures rather than relying on unguided revision models. Furthermore, development teams should invest in creating more reliable automated quality metrics tailored to iterative refinement before deploying such models into high-stakes communication workflows.
The findings are currently bounded by the dataset's focus on formal writing domains and the lower reliability observed when predicting nuanced edit categories such as style and coherence. Readers can place high confidence in the positive impact of intent-conditioned modeling, but caution is warranted when automating subjective style adjustments or applying these models to informal communication channels.
No sufficiently relevant recommendations were found.
- Paper: Self-Refine: Iterative Refinement with Self-Feedback, Aman Madaan et al. (2023). SELF-REFINE carries iterative revision into a general feedback-and-rewrite method, showing how an LLM can use explicit critiques to improve outputs across varied tasks.
- Paper: DiffusER: Discrete Diffusion via Edit-based Reconstruction, Machel Reid et al. (2023). DiffusER extends iterative editing into a discrete diffusion generator that revises text through successive insertions, deletions, and replacements.
- Paper: Deep Researcher with Test-time Diffusion, Rujun Han et al. (2025). TTD-DR applies iterative draft refinement to long-form research reports, repeatedly using the evolving text to guide retrieval and incorporate new evidence.
