LLMs Corrupt Your Documents When You Delegate

Philippe LabanTobias SchnabelJennifer Neville

article2026arXiv11 citations

Introduces the DELEGATE-52 benchmark to demonstrate that even frontier language models silently corrupt a quarter of document content during long delegated tasks across 52 professional domains, exposing critical reliability failures that current agentic tools fail to resolve.

Listen

As knowledge workers increasingly delegate complex document editing tasks to large language models (LLMs), users must trust that these systems can execute instructions faithfully without corrupting underlying files. Because users often lack the time or domain expertise to audit every change in multi-step workflows, silent errors introduced by AI systems pose serious operational risks. The article evaluates the readiness of current frontier and open-source models for long-horizon delegated knowledge work and quantifies the extent of document degradation that occurs during repeated interactions.

To conduct this evaluation without requiring human reference annotations, the article introduces a benchmark called DELEGATE-52 alongside a round-trip relay evaluation method. The benchmark encompasses 310 realistic work environments across 52 professional domains spanning Science and Engineering, Code and Configuration, Creative and Media, Structured Records, and Everyday tasks. Each environment pairs real-world textual documents (2,000 to 5,000 tokens) with domain-relevant distractor files (8,000 to 12,000 tokens) and complex, reversible editing tasks. By applying forward edits followed by inverse instructions across sequential rounds (simulating up to 20 or 100 interactions), the approach measures semantic preservation using custom, domain-specific programmatic parsers across 19 leading LLMs.

The investigation yields several critical findings. First, document degradation is widespread: across 20 interactions, tested models lose an average of 50% of document content, and even top frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, and GPT 5.4) corrupt roughly 25% of content. Second, catastrophic corruption is pervasive, occurring in over 80% of model-domain combinations; Python code was the only domain out of 52 where most models demonstrated reliable readiness. Third, degradation is driven by sparse, severe failures rather than gradual drift, with critical errors (single-round score drops of 10% or more) accounting for 80% to 98% of total content loss. While weaker models fail primarily through outright content deletion, frontier models degrade files by subtly corrupting preserved text. Fourth, degradation compounds sharply with longer interactions, larger document sizes, and the presence of distractor context, and equipping models with basic agentic tool harnesses (such as file-writing and code execution) fails to mitigate degradation, increasing token overhead and compounding errors by an additional 6% on average.

These findings demonstrate that current AI models are unreliable delegates for end-to-end, unsupervised knowledge work. High proficiency in coding tasks gives a misleading impression of broad competency; models struggle significantly with natural language prose, complex restructurings, and specialized document formats. Organizations deploying LLMs in operational pipelines face high risks of silent data corruption, which can compromise data integrity, compliance, and decision-making over extended workflows.

Moving forward, stakeholders should avoid delegating long-horizon, autonomous editing tasks without human-in-the-loop validation, particularly in non-code domains. Practitioners and developers must move beyond short-turn benchmarks and evaluate systems in extended, multi-session workflows across diverse professions. Furthermore, researchers should explore using reversible cycle consistency frameworks to train models and develop more sophisticated agentic harnesses capable of precise document manipulation.

These conclusions are bounded by specific conditions: the benchmark primarily evaluates single-turn, reversible text document editing within constrained token budgets. Because real-world workflows often involve multi-turn conversational ambiguity, irreversible operations, and larger document contexts—factors shown to worsen degradation—the reported results should be viewed with high confidence as a conservative lower bound on the true severity of document corruption in delegated AI workflows.

Cover for LLMs Corrupt Your Documents When You Delegate

Abstract

Large Language Models (LLMs) are poised to disrupt knowledge work, with the emergence of delegated work as a new interaction paradigm (e.g., vibe coding). Delegation requires trust - the expectation that the LLM will faithfully execute the task without introducing errors into documents. We introduce DELEGATE-52 to study the readiness of AI systems in delegated workflows. DELEGATE-52 simulates long delegated workflows that require in-depth document editing across 52 professional domains, such as coding, crystallography, and music notation. Our large-scale experiment with 19 LLMs reveals that current models degrade documents during delegation: even frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) corrupt an average of 25% of document content by the end of long workflows, with other models failing more severely. Additional experiments reveal that agentic tool use does not improve performance on DELEGATE-52, and that degradation severity is exacerbated by document size, length of interaction, or presence of distractor files. Our analysis shows that current LLMs are unreliable delegates: they introduce sparse but severe errors that silently corrupt documents, compounding over long interaction.

Table of Contents

  • 1 Introduction
  • 2 The DELEGATE-52 Benchmark
  • 2.1 Evaluating Without References
  • 2.2 Benchmark Construction
  • 2.2.1 Work Environments
  • 2.2.2 Domain-Specific Evaluation
  • 2.2.3 Quality Assurance.
  • 3 Experiments
  • 4 Results
  • 4.1 Main Results
  • 4.2 Agent (With Tools) vs. LLM (Without Tools)
  • 4.3 Document Size Effect
  • 4.4 Length of Interaction
  • 4.5 Distractor Effect
  • 4.6 Delegation Beyond Textual Documents
  • 5 Analysis
  • 6 Implications
  • 7 Related Work
  • 8 Limitations
  • 9 Conclusion
  • References
  • A Instruction Compliance Validation
  • A.1 Methodology
  • A.2 Findings
  • B Backtranslation Properties, Assumptions, and Limitations
  • B.1 Properties
  • B.2 Assumptions
  • B.3 Limitations
  • C Alternative Evaluation Methods
  • C.1 Evaluation Methodology
  • C.2 Evaluation Findings
  • D Round-Robin Design Validation
  • E Critical Error Analysis
  • F Deletion vs. Corruption Decomposition
  • F.1 Methodology
  • F.2 Findings
  • G Document Characteristics Analysis
  • G.1 Document Characteristic Metrics
  • G.2 Category-Level Overview
  • G.3 Findings
  • H Semantic Operation Analysis
  • I Context-Size Experiment
  • I.1 Domain and Size Selection
  • I.2 Work Environment Construction
  • J Image Domain
  • K Dataset Creation Process
  • K.1 Stage 1: Domain Brainstorm
  • K.2 Stage 2: Domain Creation
  • K.3 Stage 3: Domain Quality Assurance
  • K.4 Stage 4: Environment Scaling
  • K.5 Stage 5: Work Environment Completion
  • K.6 Stage 6: Edit Task Quality Assurance
  • K.7 Stage 7: Distractor Quality Assurance
  • K.8 Stage 8: Final Work Environment Validation
  • L Model Details
  • M Agentic Harness (Operating Models with Tools)
  • M.1 Tools and Execution Environment
  • M.2 Iteration Budget and Termination
  • M.3 Distractor Handling
  • M.4 Logging and Metadata
  • N Document Desiderata

Knowls

  1. Knowl 1 — Round-Trip Relay Simulation and Reference-Free Semantic Evaluation

    model/method

    The round-trip relay simulation evaluates the capability of Large Language Models (LLMs) to preserve document semantics across long-horizon delegated workflows without requiring annotated ground-truth intermediate references.

    Given a seed document ss, an invertible editing task consists of a forward instruction x→x^{\to} defining transformation σ\sigma and an inverse backward instruction x←x^{\leftarrow} defining σ−1\sigma^{-1}. In an independent single-turn call, the LLM executes the forward instruction to produce a transformed document t=σ(s)=LLM(s;x→)t = \sigma(s) = \text{LLM}(s; x^{\to}). In a subsequent independent call, the LLM executes the backward instruction to produce a reconstructed document s^=σ−1(t)=LLM(t;x←)\hat{s} = \sigma^{-1}(t) = \text{LLM}(t; x^{\leftarrow}).

    To simulate a workflow of length k=2nk = 2n interactions (nn full round-trips), nn invertible task pairs (σ1,σ1−1),…,(σn,σn−1)(\sigma_1, \sigma_1^{-1}), \dots, (\sigma_n, \sigma_n^{-1}) are chained sequentially:

    s^n=(σn−1∘σn∘⋯∘σ1−1∘σ1)(s)\hat{s}_n = \left(\sigma_n^{-1} \circ \sigma_n \circ \dots \circ \sigma_1^{-1} \circ \sigma_1\right)(s)

    The primary metric is the reconstruction score RS@k(s)∈[0,1]RS@k(s) \in [0, 1] after kk delegated interactions (k/2k/2 round-trips):

    RS@k(s)=sim(s,s^k/2)RS@k(s) = \text{sim}\left(s, \hat{s}_{k/2}\right)

    where sim(si,sj)∈[0,1]\text{sim}(s_i, s_j) \in [0, 1] is a domain-specific programmatic evaluation function. The scoring pipeline first parses candidate and reference documents into structured elements (e.g., ingredients, steps, and tips in a recipe; transactions in an accounting ledger) and computes a weighted similarity across structural components. This decouples surface-level text variations (such as ingredient reordering or unit aliases) from true semantic loss or corruption.

  2. Knowl 2 — The DELEGATE-52 Benchmark Specification

    experimental setup

    DELEGATE-52 is a benchmark designed to evaluate document preservation and editing accuracy across long-horizon delegated workflows in 52 professional domains grouped into five categories:

    1. Science & Engineering (11 domains: Aviation, Circuit, Crystal, MathLean, Molecule, Protein, Quantum, Robotics, Satellite, StarCatalog, Weather).
    2. Code & Configuration (11 domains: DBSchema, DNS, Docker, Filesystem, Graphviz, Infra, JSON, Makefile, Malware, Python, Translation).
    3. Creative & Media (11 domains: AudioSyn, Fiction, FontEng, LaTeX, MusicSheet, Obj3D, Screenplay, Slides, Subtitles, Vector, Weaving).
    4. Structured Records (11 domains: Accounting, Calendar, EDIFACT, Emails, Genealogy, Geodata, Geotrack, HamRadio, LibCatalog, Spreadsheet, Treebank).
    5. Everyday (8 domains: Chess, EarnCall, FoodMenu, JobBoard, Landmarks, Playlist, Recipe, Transit).

    The benchmark contains 310 work environments (6 per domain) and 2,125 unique editing tasks. Each work environment comprises:

    • Seed Document: A real-world, unencoded textual document of 2,000–5,000 tokens (tokenized via GPT-4 tiktoken) sourced with permissive licensing.
    • Invertible Edit Tasks: A set of 5–10 domain-specific forward/backward edit instruction pairs requiring non-trivial, in-place semantic transformations (e.g., splitting, classifying, converting units, or restructuring) rather than simple text append/crop operations.
    • Distractor Context: 1–5 topically related, non-interfering documents totaling 8,000–12,000 tokens to emulate realistic imperfect retrieval contexts.
  3. Knowl 3 — Document Content Degradation Across 19 LLMs on DELEGATE-52

    data/table

    Evaluation of 19 LLMs over a 10-round-trip relay (20 interactions) on DELEGATE-52 demonstrates that all models suffer cumulative document corruption. Even frontier models lose approximately 25% of content fidelity by interaction 20, with the average model degrading by 50%.

    Model 2 4 6 8 10 12 14 16 18 20
    GPT 5 Nano 30.3 17.4 12.8 12.2 11.4 11.1 10.5 10.3 10.1 10.0
    GPT 4o 45.6 29.6 23.9 19.9 18.8 17.3 16.5 16.2 15.6 14.7
    OSS 120B 73.1 52.4 36.5 28.3 25.0 22.2 20.3 20.2 19.8 19.2
    Large 3 82.4 71.8 59.8 53.8 46.1 43.4 40.7 37.4 35.9 35.5
    3 Flash 76.0 61.6 57.1 49.6 47.5 42.8 41.1 39.5 36.6 35.8
    GPT 5 Mini 86.3 75.1 66.2 60.4 55.2 50.5 48.1 47.0 45.6 45.1
    GPT 5 Chat 83.3 73.3 66.0 60.2 56.4 53.0 50.9 49.1 47.8 46.8
    o1 86.4 76.7 68.6 63.3 57.6 53.9 53.2 50.2 49.2 48.1
    o3 85.2 75.2 65.9 60.7 58.1 53.5 50.8 49.4 48.9 48.2
    GPT 5 91.5 80.9 71.6 66.3 62.1 58.5 55.9 53.3 51.4 48.3
    GPT 4.1 88.9 79.8 70.9 67.7 62.2 56.8 54.8 51.7 49.8 49.5
    Grok 4 91.7 85.4 78.5 74.0 69.0 67.2 65.4 62.1 61.4 59.3
    GPT 5.1 90.8 82.8 78.0 74.0 69.9 66.7 64.9 62.9 61.7 60.5
    Kimi K2.5 91.1 86.1 83.0 75.6 73.3 70.0 68.8 66.4 64.9 64.1
    Claude 4.6 Sonnet 92.2 85.7 81.8 78.2 74.9 71.7 70.2 69.1 66.9 66.0
    GPT 5.2 92.7 86.9 82.2 77.9 74.4 71.6 70.0 68.5 67.1 66.1
    GPT 5.4 94.3 89.3 85.4 82.0 79.4 76.4 74.6 73.1 72.1 71.5
    Claude 4.6 Opus 94.2 90.1 86.8 82.5 79.5 78.0 76.3 75.2 74.3 73.1
    Gemini 3.1 Pro 96.8 93.5 91.4 88.9 86.6 83.9 82.2 81.2 80.9 80.9

    A model is designated as "ready" for delegation in a domain if RS@20≥98%RS@20 \ge 98\%. Catastrophic degradation (RS@20≤80%RS@20 \le 80\%) occurs across >80% of model-domain pairs. Python is an outlier where 17 of 19 models achieve the ready status. The top overall model (Gemini 3.1 Pro) is ready in only 11 of 52 domains. Short-term performance after 2 interactions is not predictive of long-horizon trajectories (e.g., GPT 5 and Kimi K2.5 start at 91.5 vs. 91.1 at interaction 2, but diverge to 48.3 vs. 64.1 at interaction 20).

  4. Knowl 4 — Sparse Critical Failures Account for Majority of Document Degradation

    empirical result

    Analysis of individual relay simulation trajectories reveals that document decay does not occur through a steady accumulation of small errors ("death by a thousand cuts"). Instead, degradation is driven by sparse critical failures—defined as a single round-trip that causes a score drop ≥10\ge 10 percentage points.

    Across the 19 evaluated LLMs, single-step critical failures account for 80.9% to 97.0% of all observed score loss. Models preserve near-perfect fidelity during most round-trips, interspersed with occasional catastrophic rounds where 10–30+ points are lost in a single editing cycle. By interaction 20, the majority of relay runs for 18 out of 19 models have suffered at least one critical failure (e.g., 38.1% of runs for Gemini 3.1 Pro, 49.7% for Claude 4.6 Opus, 55.2% for GPT 5.4, up to 97.2% for GPT 5 Nano).

    Stronger frontier models do not reduce the magnitude of small execution errors; rather, they delay the onset of critical failures and reduce their frequency across interaction rounds.

  5. Knowl 5 — Error Mode Transition: Content Deletion in Weaker Models vs. Semantic Corruption in Frontier Models

    empirical result

    Decomposing document degradation into content deletion (loss of structural elements) versus semantic corruption (elements present but distorted, hallucinated, or mutated) reveals distinct failure profiles across model tiers.

    Let nrefn_{\text{ref}} and ngenn_{\text{gen}} denote the parsed structural element counts in the reference and generated documents, and let s∈[0,1]s \in [0, 1] be the reconstruction score. Defining coverage as coverage=min⁡(ngen/nref,1)\text{coverage} = \min(n_{\text{gen}} / n_{\text{ref}}, 1):

    • Deletion component: 1−coverage1 - \text{coverage}
    • Corruption component: coverage−s\text{coverage} - s
    • Total degradation: (1−coverage)+(coverage−s)=1−s(1 - \text{coverage}) + (\text{coverage} - s) = 1 - s

    In count-based analysis across 38 structured domains:

    • For weaker models (e.g., GPT 4o, GPT 5 Nano), content deletion constitutes 70% to 73% of total score degradation.
    • For frontier models (e.g., Claude 4.6 Opus, Claude 4.6 Sonnet), deletion accounts for only 22% to 27% of degradation, while semantic corruption accounts for 73% to 78%.

    In failure-mode tag analysis, deletion accounts for 35% of all tagged failure instances across models (ranging from 28% for Claude 4.6 Opus to 43% for GPT 5 Chat and GPT 4o).

  6. Knowl 6 — Impact of Agentic Tool Harness on Multi-Round Document Degradation

    data/table

    Providing LLMs with a basic agentic tool harness—including read_file, write_file, delete_file, run_python, and finish—worsens document degradation over 20 interactions compared to direct single-turn text output.

    Model Mode 2 4 6 8 10 12 14 16 18 20
    GPT 5.4 (Direct) 94.3 89.3 85.4 82.0 79.4 76.4 74.6 73.1 72.1 71.5
    GPT 5.4 (Agentic) 89.2 86.2 82.0 79.0 75.4 71.7 69.8 69.1 68.2 68.3
    GPT 5.2 (Direct) 92.7 86.9 82.2 77.9 74.4 71.6 70.0 68.5 67.1 66.1
    GPT 5.2 (Agentic) 90.7 85.1 77.9 74.5 69.4 67.5 65.0 63.6 63.2 63.4
    GPT 5.1 (Direct) 90.8 82.8 78.0 74.0 69.9 66.7 64.9 62.9 61.7 60.5
    GPT 5.1 (Agentic) 83.2 75.2 68.8 63.3 59.7 58.1 56.1 54.4 53.0 52.1
    GPT 4.1 (Direct) 88.9 79.8 70.9 67.7 62.2 56.8 54.8 51.7 49.8 49.5
    GPT 4.1 (Agentic) 84.4 71.9 63.5 57.1 52.8 48.8 46.1 43.6 42.0 40.4

    Across all four models, operating with tools induces an average additional 6% score penalty by interaction 20. Tool use incurs substantial operational overhead: models execute 8–12 tool calls per task, consuming 2.0×2.0\times to 4.6×4.6\times more input tokens than direct generation. Furthermore, models favor manual file overwriting (47%–81% of edits) over programmatic code execution (10%–45%), limiting the precision advantages of code execution while suffering context-length degradation.

  7. Knowl 7 — Scaling Effects of Document Size and Interaction Horizon on Degradation

    data/table

    Document size and interaction horizon compound multiplicatively to accelerate document degradation.

    In document size ablation with GPT 5.4 across 5 scalable domains without distractors:

    Doc. Size 2 4 6 8 10 12 14 16 18 20
    1k tokens 97.3 96.4 96.1 96.2 95.9 95.4 95.4 91.8 91.8 91.4
    2k tokens 98.4 96.7 95.3 92.6 91.9 91.3 90.9 90.7 90.0 89.9
    4k tokens 94.1 92.1 91.4 84.0 83.1 82.4 81.0 79.9 79.1 79.0
    6k tokens 98.7 93.4 90.5 87.2 84.0 80.7 79.3 79.9 77.7 72.3
    8k tokens 94.1 92.1 83.9 79.7 74.2 72.9 72.6 69.0 67.0 67.4
    10k tokens 90.1 86.4 83.0 77.7 72.2 70.7 67.1 63.6 63.7 59.9

    Each additional 1,000 document tokens degrades GPT 5.4 reconstruction by ≈0.7%\approx 0.7\% after 2 interactions, but by ≈3.6%\approx 3.6\% after 20 interactions (a ≈5×\approx 5\times magnification).

    Extending relays to 100 interactions (50 round-trips) demonstrates monotonic degradation with no plateau:

    Model 10 20 30 40 50 60 70 80 90 100
    GPT 5.4 79.7 72.9 69.7 66.8 66.2 62.9 62.2 62.0 60.6 58.7
    GPT 5.2 75.6 67.1 63.5 60.0 56.5 56.6 53.1 50.6 51.5 50.4
    GPT 5.1 68.5 60.9 56.7 52.4 49.5 47.0 45.3 45.2 43.4 42.6
    GPT 4.1 58.4 49.3 44.4 41.5 39.3 38.2 37.2 35.6 34.2 33.3

    The rate of degradation decelerates in later rounds (rounds 5–25 account for 2×–3×2\times\text{--}3\times more loss than rounds 25–50), but models continue introducing new errors even when edit tasks repeat.

  8. Knowl 8 — Compounding Effect of Distractor Context in Delegated Workflows

    data/table

    The inclusion of 8,000–12,000 tokens of non-interfering distractor documents causes performance harm that compounds over extended interactions.

    Model Setting 2 4 6 8 10 12 14 16 18 20
    GPT 5.4 (w/ Distractor) 94.3 89.3 85.4 82.0 79.4 76.4 74.6 73.1 72.1 71.5
    GPT 5.4 (w/o Distractor) 94.7 90.9 88.4 86.4 83.5 82.1 81.0 79.8 79.2 77.8
    GPT 5.2 (w/ Distractor) 92.7 86.9 82.2 77.9 74.4 71.6 70.0 68.5 67.1 66.1
    GPT 5.2 (w/o Distractor) 93.4 88.8 85.8 83.0 80.9 78.5 76.0 75.1 74.7 74.5
    GPT 5.1 (w/ Distractor) 90.8 82.8 78.0 74.0 69.9 66.7 64.9 62.9 61.7 60.5
    GPT 5.1 (w/o Distractor) 94.1 87.8 84.1 79.8 76.3 72.4 70.4 69.8 67.8 67.0
    GPT 4.1 (w/ Distractor) 88.9 79.8 70.9 67.7 62.2 56.8 54.8 51.7 49.8 49.5
    GPT 4.1 (w/o Distractor) 88.1 79.2 73.6 67.6 64.4 60.4 57.4 56.1 54.7 52.3

    At interaction 2, removing distractors yields a modest improvement of 0.4%–4.0%. By interaction 20, the benefit of omitting distractors widens to 2.8%–8.4% across models. This indicates that measuring distraction effects in short single-turn interactions substantially underestimates the compounding harm of irrelevant context in multi-session workflows.

  9. Knowl 9 — Document Properties and Semantic Operations Influencing Edit Difficulty

    empirical result

    Statistical analysis of single round trips with GPT 5.2 across document characteristics and 11 tagged semantic operations shows strong structural predictors of degradation:

    Document Characteristics:

    • Repetitiveness (fraction of repeated 5-grams) is the strongest positive predictor of preservation (d=+0.261,p<0.001d = +0.261, p < 0.001).
    • Numerical fraction (tokens containing digits) and structural density (non-alphanumeric, non-whitespace characters) also improve preservation (d=+0.159,p<0.001d = +0.159, p < 0.001 and d=+0.119,p<0.001d = +0.119, p < 0.001).
    • Naturalness (ratio of function words to total words) is the strongest negative predictor (d=−0.260,p<0.001d = -0.260, p < 0.001).
    • Category-level scores: Science & Engineering (97.3%) and Code & Configuration (95.6%) outperform Creative & Media (94.4%), Structured Records (91.5%), and Everyday domains (89.9%).

    Semantic Operations:

    • Global document restructuring operations correlate significantly with lower reconstruction scores: Split and Merge (rpb=−0.080,p<0.001r_{\text{pb}} = -0.080, p < 0.001), Classification (rpb=−0.076,p<0.001r_{\text{pb}} = -0.076, p < 0.001), and Format Knowledge (rpb=−0.060,p<0.001r_{\text{pb}} = -0.060, p < 0.001).
    • Local token/passage operations correlate with higher reconstruction scores: String Manipulation (rpb=+0.068,p<0.001r_{\text{pb}} = +0.068, p < 0.001), Referencing (rpb=+0.043,p<0.001r_{\text{pb}} = +0.043, p < 0.001), and Context Expansion (rpb=+0.023,p<0.01r_{\text{pb}} = +0.023, p < 0.01).
    • Task complexity: The number of simultaneous semantic operations per task negatively correlates with reconstruction (Spearman r=−0.043,p<0.001r = -0.043, p < 0.001), with mean scores dropping from 94.0% for single-operation tasks to 82.6% for tasks requiring five operations.
  10. Knowl 10 — Multi-Round Degradation in Visual Image Editing Delegation

    data/table

    Evaluating 9 image generation models across 6 visual work environments (Wikipedia images resized to 512×512512\times512 PNG, 6–7 reversible edit pairs per environment) using a composite perceptual similarity metric:

    Score=0.50⋅SSIM+0.25⋅HSV_Corr+0.25⋅max⁡(0,1−2.5⋅RMSE)\text{Score} = 0.50 \cdot \text{SSIM} + 0.25 \cdot \text{HSV\_Corr} + 0.25 \cdot \max\left(0, 1 - 2.5 \cdot \text{RMSE}\right)

    reveals severe degradation over 20 interactions (10 round-trips):

    Model 2 4 6 8 10 12 14 16 18 20
    Instruct Pix2Pix 34.3 24.3 21.1 15.9 17.7 16.1 11.2 9.6 12.1 11.4
    Flux2 Dev 31.3 25.4 23.4 20.6 19.7 19.2 25.7 23.1 23.2 19.5
    GPT Image 1 30.3 27.6 20.9 22.0 19.6 19.1 19.5 20.8 22.5 20.9
    Flux Kontext 54.0 35.5 35.0 27.2 26.6 25.1 22.6 22.1 21.0 22.0
    Flux2 Klein 4B 38.8 31.3 28.7 26.9 24.4 23.1 22.5 24.0 21.8 22.4
    Flux2 Klein 9B 54.6 37.0 31.6 29.0 25.8 23.7 26.2 25.1 27.2 26.2
    Gemini 2.5 Flash Image 55.1 43.8 40.0 35.1 34.8 32.7 29.5 28.9 29.0 27.7
    Gemini 3 Pro Image 63.2 48.1 36.7 28.5 27.9 27.6 29.1 28.7 28.7 30.0
    Gemini 3.1 Flash Image 55.6 44.4 35.5 32.7 32.3 32.0 30.4 31.5 33.2 30.4

    No image generation model scores above 65% after 2 interactions, and all models degrade to 11.4%–30.4% by interaction 20, demonstrating that current visual models corrupt artifacts substantially faster than text-based LLMs.

Coverage note — None was omitted; all key benchmark definitions, formal equations, empirical results (including agentic, context size, interaction length, distractor, image modality, error analysis, and task property studies) and full quantitative tables were captured.

References

  1. 1.Miltiadis Allamanis, Sheena Panthaplackel, and Pengcheng Yin. Unsupervised evaluation of code llms with round-trip correctness. pp. 1050–1066, 2024.
  2. 2.Anthropic. The claude 3 model family: Opus, sonnet, haiku, 2024.
  3. 3.Alexander Bick, A. Blandin, and David J. Deming. The rapid adoption of generative ai. SSRN Electronic Journal, 2024.
  4. 4.Michelle Brachman, Amina El-Ashry, Casey Dugan, and Werner Geyer. How knowledge workers use and want to use llms in an enterprise context. Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, 2024.
  5. 5.Tim Brooks, Aleksander Holynski, and Alexei A. Efros. Instructpix2pix: Learning to follow image editing instructions. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 18392–18402, 2022.
  6. 6.Federico Cassano, Luisa Li, Akul Sethi, Noah Shinn, Abby Brennan-Jones, Anton Lozhkov, C. Anderson, and Arjun Guha. Can it edit? evaluating the ability of large language models to follow code editing instructions. ArXiv, abs/2312.12450, 2023.
  7. 7.Tuhin Chakrabarty, Philippe Laban, and Chien-Sheng Wu. Can ai writing be salvaged? mitigating idiosyncrasies and improving human-ai alignment in the writing process through edits. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 2024.
  8. 8.Tuhin Chakrabarty, Philippe Laban, and Chien-Sheng Wu. Ai-slop to ai-polish? aligning language models through edit-based writing rewards and test-time computation. ArXiv, abs/2504.07532, 2025.
  9. 9.Aaron Chatterji, T. Cunningham, David Deming, Zoë Hitzig, Christopher Ong, Carl Shan, and Kevin Wadman. How people use chatgpt. SSRN Electronic Journal, 2025.
  10. 10.Kaiyuan Chen, Yixin Ren, Yang Liu, X. Hu, Haotong Tian, Tianbao Xie, Fangfu Liu, Haoye Zhang, Hongzhang Liu, Yuan Gong, Chen Sun, Han Hou, Hui Yang, James Pan, Jian-Guang Lou, Jiayi Mao, Jizheng Liu, Jinpeng Li, Kangyi Liu, Ke Liu, Rui Wang, Runhao Li, Tong Niu, Wenlong Zhang, Wenqi Yan, Xuanzheng Wang, Yuchen Zhang, Yi-Hsin Hung, Yuan Jiang, Zexuan Liu, Zihang Yin, Zi min Ma, and Zhi-Wei Mo. xbench: Tracking agents productivity scaling with profession-aligned real-world evaluations. ArXiv, abs/2506.13651, 2025a.
  11. 11.Siqi Chen, Xinyu Dong, Haolei Xu, Xingyu Wu, Fei Tang, Hang Zhang, Yuchen Yan, Linjuan Wu, Wenqi Zhang, Guiyang Hou, Yongliang Shen, Weiming Lu, and Yueting Zhuang. SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation. 2025b.
  12. 12.Zi-Jian Cheng, Weixin Wang, Yu Zhao, Ziyang Ren, Jiaxuan Chen, Ruiyang Xu, Shuai Huang, Yang Chen, Guowei Li, Mengshi Wang, Yichen Xie, Renchuan Zhu, Zeren Jiang, Keda Lu, Yihong Li, Xiaoliang Wang, Liwei Liu, and Cam-Tu Nguyen. Lifebench: A benchmark for long-horizon multi-source memory. 2026.
  13. 13.Gheorghe Comanici et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. ArXiv, abs/2507.06261, 2025.
  14. 14.Fabrizio Dell’Acqua, Edward McFowland, Ethan Mollick, Hila Lifshitz-Assaf, Katherine C. Kellogg, Saran Rajendran, Lisa A. Krayer, F. Candelon, and K. Lakhani. Navigating the jagged technological frontier: Field experimental evidence of the effects of ai on knowledge worker productivity and quality. SSRN Electronic Journal, 2023.
  15. 15.Alexandre Drouin, Maxime Gasse, Massimo Caccia, I. Laradji, Manuel Del Verme, Tom Marty, L’eo Boisvert, Megh Thakkar, Quentin Cappart, David Vázquez, Nicolas Chapados, and Alexandre Lacoste. Workarena: How capable are web agents at solving common knowledge work tasks? ArXiv, abs/2403.07718, 2024.
  16. 16.Yiming Du, Bingbing Wang, Yangfan He, Bin Liang, Baojun Wang, Zhongyang Li, Lin Gui, Jeff Z. Pan, Ruifeng Xu, and Kam-Fai Wong. Memguide: Intent-driven memory selection for goal-oriented multi-session llm agents. 2025.
  17. 17.Jane Dwivedi-Yu, Timo Schick, Zhengbao Jiang, M. Lomeli, Patrick Lewis, Gautier Izacard, Edouard Grave, Sebastian Riedel, and F. Petroni. Editeval: An instruction-based benchmark for text improvements. ArXiv, abs/2209.13331, 2022.
  18. 18.Ahmed El-Kishky. Openai o1 system card. 2024.
  19. 19.Tyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock. Gpts are gpts: An early look at the labor market impact potential of large language models. ArXiv, abs/2303.10130, 2023.
  20. 20.Jiawei Guo, Ziming Li, Xueling Liu, Kaijing Ma, Tianyu Zheng, Zhouliang Yu, Ding Pan, Yizhi Li, Ruibo Liu, Yue Wang, Shuyue Guo, Xingwei Qu, Xiang Yue, Ge Zhang, Wenhu Chen, and Jie Fu. Codeeditorbench: Evaluating code editing capability of large language models. ArXiv, abs/2404.03543, 2024.
  21. 21.Kunal Handa, Alex Tamkin, Miles McCain, Saffron Huang, Esin Durmus, Sarah Heck, J. Mueller, Jerry Hong, Stuart Ritchie, Tim Belonax, Kevin K. Troy, Dario Amodei, Jared Kaplan, Jack Clark, and Deep Ganguli. Which economic tasks are performed with ai? evidence from millions of claude conversations. ArXiv, abs/2503.04761, 2025.
  22. 22.Di He, Yingce Xia, Tao Qin, Liwei Wang, Nenghai Yu, Tie-Yan Liu, and Wei-Ying Ma. Dual learning for machine translation. pp. 820–828, 2016.
  23. 23.Zexue He, Yu Wang, Churan Zhi, Yuanzhe Hu, Tzu-Ping Chen, Lang Yin, Zeming Chen, Tong Wu, Siru Ouyang, Zihan Wang, Jiaxin Pei, Julian McAuley, Yejin Choi, and A. Pentland. Memoryarena: Benchmarking agent memory in interdependent multi-session agentic tasks. 2026.
  24. 24.Christine Herlihy, Jennifer Neville, Tobias Schnabel, and Adith Swaminathan. On overcoming miscalibrated conversational priors in llm-based chatbots. ArXiv, abs/2406.01633, 2024.
  25. 25.Cong Duy Vu Hoang, Philipp Koehn, Gholamreza Haffari, and Trevor Cohn. Iterative back-translation for neural machine translation. pp. 18–24, 2018.
  26. 26.Zhaochen Hong, Haofei Yu, and Jiaxuan You. Consistencychecker: Tree-based evaluation of llm generalization capabilities. pp. 33039–33075, 2025.
  27. 27.Chuanrui Hu, Tong Li, Xingze Gao, Hongda Chen, Yi Bai, Dannong Xu, Tianwei Lin, Xinda Zhao, Xiaohong Li, Yunyun Han, Jian Pei, and Yafeng Deng. Evermembench: Benchmarking long-term interactive memory in large language models. 2026.
  28. 28.Kung-Hsiang Huang, Akshara Prabhakar, Sidharth Dhawan, Yixin Mao, Huan Wang, Silvio Savarese, Caiming Xiong, Philippe Laban, and Chien-Sheng Wu. Crmarena: Understanding the capacity of llm agents to perform professional crm tasks in realistic environments. pp. 3830–3850, 2024.
  29. 29.Aaron Hurst et al. Gpt-4o system card. 2024.
  30. 30.Daesik Jang, Morgan Lindsay Heisler, Linzi Xing, Yifei Li, E. Wang, Ying Xiong, Yong Zhang, and Zhenan Fan. Deckbench: Benchmarking multi-agent frameworks for academic slide generation and editing. 2026.
  31. 31.Jihyoung Jang, Minseong Boo, and Hyounghun Kim. Conversation chronicles: Towards diverse temporal and relational dynamics in multi-session conversations. ArXiv, abs/2310.13420, 2023.
  32. 32.Saurabh Jha, Rohan R. Arora, Yuji Watanabe, Takumi Yanagawa, Yinfang Chen, Jackson Clark, Bhavya Bhavya, Mudit Verma, Harshit Kumar, Hirokuni Kitahara, Noah Zheutlin, Saki Takano, Divya Pathak, Felix George, Xinbo Wu, B. Turkkan, Gerard Vanloo, M. Nidd, Ting Dai, Oishik Chatterjee, Pranjal Gupta, Suranjana Samanta, Pooja Aggarwal, Rong Lee, Pavankumar Murali, Jae wook Ahn, Debanjana Kar, Ameet Rahane, Carlos Fonseca, Amit M. Paradkar, Yu Deng, Pratibha Moogi, P. Mohapatra, Naoki Abe, Chandra Narayanaswami, Tianyin Xu, Lav R. Varshney, R. Mahindru, A. Sailer, Laura Shwartz, Daby M. Sow, Nicholas C. Fuller, Ruchir Puri Ibm, and U. I. Urbana-Champaign. Itbench: Evaluating ai agents across diverse real-world it automation tasks. ArXiv, abs/2502.05352, 2025.
  33. 33.Albert Qiaochu Jiang et al. Mistral 7b. ArXiv, abs/2310.06825, 2023.
  34. 34.Bowen Jiang, Zhuoqun Hao, Young-Min Cho, Bryan Li, Yuan Yuan, Sihao Chen, Lyle Ungar, C. J. Taylor, and Dan Roth. Know me, respond to me: Benchmarking llms for dynamic user profiling and personalized responses at scale. ArXiv, abs/2504.14225, 2025.
  35. 35.Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe-bench: Can language models resolve real-world github issues? ArXiv, abs/2310.06770, 2023.
  36. 36.M. Kapadnis, Lawanya Baghel, Atharva Naik, and C. Ros’e. Charteditbench: Evaluating grounded multi-turn chart editing in multimodal language models. 2026.
  37. 37.Tae Soo Kim, Yoonjoo Lee, Jaesang Yu, John Joon Young Chung, and Juho Kim. Discoverllm: From executing intents to discovering them. arXiv preprint arXiv:2602.03429, 2026.
  38. 38.Philippe Laban, Jesse Vig, Wojciech Kryscinski, Shafiq R. Joty, Caiming Xiong, and Chien-Sheng Wu. Swipe: A dataset for document-level simplification of wikipedia pages. pp. 10674–10695, 2023.
  39. 39.Philippe Laban, A. R. Fabbri, Caiming Xiong, and Chien-Sheng Wu. Summary of a haystack: A challenge to long-context llms and rag systems. pp. 9885–9903, 2024.
  40. 40.Philippe Laban, Hiroaki Hayashi, Yingbo Zhou, and Jennifer Neville. Llms get lost in multi-turn conversation. ArXiv, abs/2505.06120, 2025.
  41. 41.Black Forest Labs. FLUX.2: Frontier Visual Intelligence. https://bfl.ai/blog/flux-2, 2025.
  42. 42.Black Forest Labs, Stephen Batifol, A. Blattmann, Frederic Boesel, Saksham Consul, Cyril Diagne, Tim Dockhorn, Jack English, Zion English, Patrick Esser, Sumith Kulal, Kyle Lacey, Yam Levi, Cheng Li, Dominik Lorenz, Jonas Muller, Dustin Podell, Robin Rombach, Harry Saini, Axel Sauer, and Luke Smith. Flux.1 kontext: Flow matching for in-context image generation and editing in latent space. ArXiv, abs/2506.15742, 2025.
  43. 43.M. Lachaux, Baptiste Rozière, Lowik Chanussot, and Guillaume Lample. Unsupervised translation of programming languages. ArXiv, abs/2006.03511, 2020.
  44. 44.Guillaume Lample, Ludovic Denoyer, and Marc’Aurelio Ranzato. Unsupervised machine translation using monolingual corpora only. ArXiv, abs/1711.00043, 2017.
  45. 45.V. Levenshtein. Binary codes capable of correcting deletions, insertions, and reversals. Soviet physics. Doklady, 10:707–710, 1965.
  46. 46.Shuo Li, Jiajun Sun, Zhekai Wang, Xiaoran Fan, Hui Li, Di Yang, Zhiheng Xi, Yijun Wang, Zifei Shan, Tao Gui, Qi Zhang, and Xuanjing Huang. Charte3: A comprehensive benchmark for end-to-end chart editing. ArXiv, abs/2601.21694, 2026.
  47. 47.Xian Li, Ping Yu, Chunting Zhou, Timo Schick, Luke Zettlemoyer, Omer Levy, J. Weston, and M. Lewis. Self-alignment with instruction backtranslation. ArXiv, abs/2308.06259, 2023.
  48. 48.Xintong Li, Jalend Bantupalli, Ria Dharmani, Yuwei Zhang, and Jingbo Shang. Toward multi-session personalized conversation: A large-scale dataset and hierarchical tree framework for implicit reasoning. pp. 11493–11506, 2025.
  49. 49.Zheng Li, Xiang Chen, and Xiaojun Wan. Wikitableedit: A benchmark for table editing by natural language instruction. ArXiv, abs/2403.02962, 2024.
  50. 50.Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Annual Meeting of the Association for Computational Linguistics, pp. 74–81, 2004.
  51. 51.Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, F. Petroni, and Percy Liang. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12:157–173, 2023.
  52. 52.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, M. Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. ArXiv, abs/1907.11692, 2019.
  53. 53.Adyasha Maharana, Dong-Ho Lee, S. Tulyakov, Mohit Bansal, Francesco Barbieri, and Yuwei Fang. Evaluating very long-term conversational memory of llm agents. ArXiv, abs/2402.17753, 2024.
  54. 54.Nickil Maveli, Antonio Vergari, and Shay B. Cohen. Can llms compress (and decompress)? evaluating code understanding and execution via invertibility. ArXiv, abs/2601.13398, 2026.
  55. 55.Mantas Mazeika, Alice Gatti, Cristina Menghini, Udari Madhushani Sehwag, Shivam Singhal, Yury Orlovskiy, Steven Basart, Manasi Sharma, Denis Peskoff, Elaine Lau, Jaehyuk Lim, Lachlan Carroll, Alice Blair, V. Sivakumar, Sumana Basu, Brad Kenstler, Yuntao Ma, Julian Michael, Xiaoke Li, Oliver Ingebretsen, Aditya Mehta, Jean Mottola, John Teichmann, Kevin Yu, Zaina Shaik, Adam Khoja, Richard Ren, J. Hausenloy, Long Phan, Ye Htet, Ankit Aich, Tahseen Rabbani, Vivswan Shah, Andriy Novykov, F. Binder, K. Chugunov, L. Ramírez, Matias Geralnik, Hern’an Mesura, Dean Lee, E. Cardona, A. Diamond, Summer Yue, Alexandr Wang, Bing Liu, Ernesto Hernandez, and Dan Hendrycks. Remote labor index: Measuring ai automation of remote work. ArXiv, abs/2510.26787, 2025.
  56. 56.Shuhaib Mehri, Priyanka Kargupta, Tal August, and Dilek Hakkani-Tur. Multisessioncollab: Learning user preferences with memory to improve long-term collaboration. 2026.
  57. 57.Marcus J. Min, Yangruibo Ding, Luca Buratti, Saurabh Pujar, Gail E. Kaiser, Suman Jana, and Baishakhi Ray. Beyond accuracy: Evaluating self-consistency of code large language models with identitychain. ArXiv, abs/2310.14053, 2023.
  58. 58.Tarek Naous, Philippe Laban, Wei Xu, and Jennifer Neville. Flipping the dialogue: Training and evaluating user language models. arXiv preprint arXiv:2510.06552, 2025.
  59. 59.Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, N. Tezak, Jong Wook Kim, Chris Hallacy, Johannes Heidecke, Pranav Shyam, Boris Power, Tyna Eloundou Nekoul, G. Sastry, Gretchen Krueger, D. Schnurr, F. Such, K. Hsu, Madeleine Thompson, Tabarak Khan, T. Sherbakov, Joanne Jang, Peter Welinder, and Lilian Weng. Text and code embeddings by contrastive pre-training. ArXiv, abs/2201.10005, 2022.
  60. 60.Thao Nguyen, Jeffrey Li, Sewoong Oh, Ludwig Schmidt, Jason Weston, Luke S. Zettlemoyer, and Xian Li. Better alignment with instruction back-and-forth translation. pp. 13289–13308, 2024.
  61. 61.Kunato Nishina and Yusuke Matsui. Svgeditbench: A benchmark dataset for quantitative assessment of llm’s svg editing capabilities. ArXiv, abs/2404.13710, 2024.
  62. 62.Zach Nussbaum, John X. Morris, Brandon Duderstadt, and Andriy Mulyar. Nomic embed: Training a reproducible long context text embedder. ArXiv, abs/2402.01613, 2024.
  63. 63.Michael Ofengenden, Yunze Man, Ziqi Pang, and Yu-Xiong Wang. Pptarena: A benchmark for agentic powerpoint editing. ArXiv, abs/2512.03042, 2025.
  64. 64.Tejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim, Michele Wang, Olivia Watkins, S. Fishman, Marwan Aljubeh, Phoebe Thacker, Laurance Fauconnet, Natalie S. Kim, Patrick Chao, Samuel Miserendino, Gildas Chabot, David Li, Michael Sharman, Alexandra Barr, Amelia Glaese, and Jerry Tworek. Gdpval: Evaluating ai model performance on real-world economically valuable tasks. ArXiv, abs/2510.04374, 2025.
  65. 65.NORMAN G. Peterson, Michael D. Mumford, W. C. Borman, P. Jeanneret, E. Fleishman, Kerry Y. Levin, MICHAEL A. Campion, M. S. Mayfield, F. Morgeson, Kenneth Pearlman, M. Gowing, Anita R. Lancaster, M. Silver, and D. Dye. Understanding work using the occupational information network (o*net): Implications for practice and research. Personnel Psychology, 54:451–492, 2001.
  66. 66.V. Pimenova, Sarah Fakhoury, Christian Bird, M. Storey, and Madeline Endres. Good vibrations? a qualitative study of co-creation, communication, flow, and trust in vibe coding. ArXiv, abs/2509.12491, 2025.
  67. 67.Vipul Raheja, Dhruv Kumar, Ryan Koo, and Dongyeop Kang. Coedit: Text editing by task-specific instruction tuning. ArXiv, abs/2305.09857, 2023.
  68. 68.Baptiste Rozière, J Zhang, François Charton, M. Harman, Gabriel Synnaeve, and Guillaume Lample. Leveraging automated unit tests for unsupervised code translation. ArXiv, abs/2110.06773, 2021.
  69. 69.Rico Sennrich, B. Haddow, and Alexandra Birch. Improving neural machine translation models with monolingual data. ArXiv, abs/1511.06709, 2015.
  70. 70.Yijia Shao, Humishka Zope, Yucheng Jiang, Jiaxin Pei, D. Nguyen, Erik Brynjolfsson, and Diyi Yang. Future of work with ai agents: Auditing automation and augmentation potential across the u.s. workforce. ArXiv, abs/2506.06576, 2025.
  71. 71.Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H. Chi, Nathanael Scharli, and Denny Zhou. Large language models can be easily distracted by irrelevant context. pp. 31210–31227, 2023.
  72. 72.Aaditya K. Singh et al. Openai gpt-5 system card. 2025.
  73. 73.Alexa Siu and Raymond Fok. Augmenting expert cognition in the age of generative ai: Insights from document-centric knowledge work. ArXiv, abs/2503.24334, 2025.
  74. 74.J. Skalse, Nikolaus H. R. Howe, D. Krasheninnikov, and David Krueger. Defining and characterizing reward hacking. ArXiv, abs/2209.13085, 2022.
  75. 75.H. Somers. Round-trip translation: What is it good for? pp. 127–133, 2005.
  76. 76.Alexander Spangher, Xiang Ren, Jonathan May, and Nanyun Peng. Newsedits: A news article revision dataset and a novel document-level reasoning challenge. ArXiv, abs/2206.07106, 2022.
  77. 77.Adam Suma and Sam Dauncey. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. ArXiv, abs/2501.12948, 2025.
  78. 78.Kimi Team et al. Kimi k1.5: Scaling reinforcement learning with llms. ArXiv, abs/2501.12599, 2025.
  79. 79.Kiran Tomlinson, Sonia Jaffe, W. Wang, Scott Counts, and Siddharth Suri. Working with ai: Measuring the applicability of generative ai to occupations. 2025.
  80. 80.Mara Ulloa, Jenna L. Butler, Sankeerti Haniyur, Courtney Miller, Barrett Amos, Advait Sarkar, and M. Storey. Product manager practices for delegating work to generative ai: "accountability must not be delegated to non-human actors". ArXiv, abs/2510.02504, 2025.
  81. 81.Z. Wang, S. Vijayvargiya, Aspen Chen, Han Zhang, Venu Arvind Arangarajan, Jett Chen, Valerie Chen, Diyi Yang, Daniel Fried, and Graham Neubig. How well does agent development reflect real-world work? 2026.
  82. 82.Bolin Wei, Ge Li, Xin Xia, Zhiyi Fu, and Zhi Jin. Code generation as a dual task of code summarization. pp. 6559–6569, 2019.
  83. 83.Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, and Dong Yu. Longmemeval: Benchmarking chat assistants on long-term interactive memory. ArXiv, abs/2410.10813, 2024.
  84. 84.xAI. Grok. https://x.ai/blog/grok, 2025.
  85. 85.Frank F. Xu, Yufan Song, Boxuan Li, Yuxuan Tang, Kritanjali Jain, Meng Bao, Z. Wang, Xuhui Zhou, Zhitong Guo, Murong Cao, Ming-Hsuan Yang, Hao Lu, Amaad Martin, Zhe Su, L. Maben, Raj Mehta, Wayne Chi, Lawrence Jang, Yiqing Xie, Shuyan Zhou, and Graham Neubig. Theagentcompany: Benchmarking llm agents on consequential real world tasks. ArXiv, abs/2412.14161, 2024.
  86. 86.Jing Xu, Arthur Szlam, and J. Weston. Beyond goldfish memory: Long-term open-domain conversation. ArXiv, abs/2107.07567, 2021.
  87. 87.Yisen Xu, Jinqiu Yang, and Tse-Hsun Chen. Swe-refactor: A repository-level benchmark for real-world llm-based code refactoring. 2026.
  88. 88.Jialin Yang, Dongfu Jiang, Lipeng He, Sherman Siu, Yuxuan Zhang, Disen Liao, Zhuofeng Li, Huaye Zeng, Yiming Jia, Haozhe Wang, Benjamin Schneider, Chi Ruan, Wentao Ma, Z. Lyu, Yifei Wang, Yi Lu, Quy Duc Do, Ziyan Jiang, Ping Nie, and Wenhu Chen. Structeval: Benchmarking llms’ capabilities to generate structural outputs. Trans. Mach. Learn. Res., 2026, 2025.
  89. 89.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. ArXiv, abs/2210.03629, 2022.
  90. 90.Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan. τ-bench: A benchmark for tool-agent-user interaction in real-world domains. ArXiv, abs/2406.12045, 2024.
  91. 91.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert. ArXiv, abs/1904.09675, 2019.
  92. 92.Junhao Zheng, Xidi Cai, Qiuke Li, Duzhen Zhang, Zhongzhi Li, Yingying Zhang, Le Song, and Qianli Ma. Lifelongagentbench: Evaluating llm agents as lifelong learners. ArXiv, abs/2505.11942, 2025.
  93. 93.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, E. Xing, Haotong Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena. ArXiv, abs/2306.05685, 2023.
  94. 94.Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. 2017 IEEE International Conference on Computer Vision (ICCV), pp. 2242–2251, 2017.
  95. 95.Terry Yue Zhuo, Qiongkai Xu, Xuanli He, and Trevor Cohn. Rethinking round-trip translation for machine translation evaluation. pp. 319–337, 2022.

Citation

MLA
Laban, P., et al. “LLMs Corrupt Your Documents When You Delegate”. arXiv, 2026, http://arxiv.org/abs/2604.15597v1.
APA
Laban, P., Schnabel, T., & Neville, J. (2026). LLMs Corrupt Your Documents When You Delegate. arXiv. http://arxiv.org/abs/2604.15597v1
Chicago
Laban, P., T. Schnabel, and J. Neville. 2026. “LLMs Corrupt Your Documents When You Delegate”. arXiv. http://arxiv.org/abs/2604.15597v1.
Harvard
Laban, P., Schnabel, T. and Neville, J. (2026) “LLMs Corrupt Your Documents When You Delegate”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2604.15597v1.
Vancouver
1. Laban P, Schnabel T, Neville J (2026) LLMs Corrupt Your Documents When You Delegate. arXiv

BibTeX

@article{laban2026llms,
  title = {LLMs Corrupt Your Documents When You Delegate},
  author = {Laban, Philippe and Schnabel, Tobias and Neville, Jennifer},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2604.15597v1},
  eprint = {2604.15597}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission