LLMs Corrupt Your Documents When You Delegate
Philippe LabanTobias SchnabelJennifer Neville
Introduces the DELEGATE-52 benchmark to demonstrate that even frontier language models silently corrupt a quarter of document content during long delegated tasks across 52 professional domains, exposing critical reliability failures that current agentic tools fail to resolve.
As knowledge workers increasingly delegate complex document editing tasks to large language models (LLMs), users must trust that these systems can execute instructions faithfully without corrupting underlying files. Because users often lack the time or domain expertise to audit every change in multi-step workflows, silent errors introduced by AI systems pose serious operational risks. The article evaluates the readiness of current frontier and open-source models for long-horizon delegated knowledge work and quantifies the extent of document degradation that occurs during repeated interactions.
To conduct this evaluation without requiring human reference annotations, the article introduces a benchmark called DELEGATE-52 alongside a round-trip relay evaluation method. The benchmark encompasses 310 realistic work environments across 52 professional domains spanning Science and Engineering, Code and Configuration, Creative and Media, Structured Records, and Everyday tasks. Each environment pairs real-world textual documents (2,000 to 5,000 tokens) with domain-relevant distractor files (8,000 to 12,000 tokens) and complex, reversible editing tasks. By applying forward edits followed by inverse instructions across sequential rounds (simulating up to 20 or 100 interactions), the approach measures semantic preservation using custom, domain-specific programmatic parsers across 19 leading LLMs.
The investigation yields several critical findings. First, document degradation is widespread: across 20 interactions, tested models lose an average of 50% of document content, and even top frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, and GPT 5.4) corrupt roughly 25% of content. Second, catastrophic corruption is pervasive, occurring in over 80% of model-domain combinations; Python code was the only domain out of 52 where most models demonstrated reliable readiness. Third, degradation is driven by sparse, severe failures rather than gradual drift, with critical errors (single-round score drops of 10% or more) accounting for 80% to 98% of total content loss. While weaker models fail primarily through outright content deletion, frontier models degrade files by subtly corrupting preserved text. Fourth, degradation compounds sharply with longer interactions, larger document sizes, and the presence of distractor context, and equipping models with basic agentic tool harnesses (such as file-writing and code execution) fails to mitigate degradation, increasing token overhead and compounding errors by an additional 6% on average.
These findings demonstrate that current AI models are unreliable delegates for end-to-end, unsupervised knowledge work. High proficiency in coding tasks gives a misleading impression of broad competency; models struggle significantly with natural language prose, complex restructurings, and specialized document formats. Organizations deploying LLMs in operational pipelines face high risks of silent data corruption, which can compromise data integrity, compliance, and decision-making over extended workflows.
Moving forward, stakeholders should avoid delegating long-horizon, autonomous editing tasks without human-in-the-loop validation, particularly in non-code domains. Practitioners and developers must move beyond short-turn benchmarks and evaluate systems in extended, multi-session workflows across diverse professions. Furthermore, researchers should explore using reversible cycle consistency frameworks to train models and develop more sophisticated agentic harnesses capable of precise document manipulation.
These conclusions are bounded by specific conditions: the benchmark primarily evaluates single-turn, reversible text document editing within constrained token budgets. Because real-world workflows often involve multi-turn conversational ambiguity, irreversible operations, and larger document contexts—factors shown to worsen degradation—the reported results should be viewed with high confidence as a conservative lower bound on the true severity of document corruption in delegated AI workflows.
- Paper: How Language Model Hallucinations Can Snowball, Muru Zhang et al. (2024). Provides fundamental insights into how early errors snowball and compound across multi-step LLM interactions, establishing the behavioral mechanics underlying document corruption during delegated workflows.
- Paper: Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models, Mosh Levy et al. (2024). Establishes empirical evidence on how expanding context length and distractor tokens degrade multi-step reasoning performance in LLMs, directly informing DELEGATE-52's findings on document size and interaction length.
- Paper: LooGLE: Can Long-Context Language Models Understand Long Contexts?, Jiaqi Li et al. (2024). Demonstrates the failure of long-context language models to maintain accurate long-range dependencies across complex documents, establishing a baseline for evaluating long-form document editing reliability.
- Paper: Position: TrustLLM: Trustworthiness in Large Language Models, Yue Huang et al. (2024). Formulates the foundational taxonomy and evaluation benchmarks for LLM trustworthiness and reliability across critical deployment domains.
- Paper: MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation, Qian Huang et al. (2024). Introduces a foundational evaluation suite for testing autonomous LLM agents executing complex, multi-step workflows in realistic workspace environments.
- Paper: SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?, Samuel Miserendino et al. (2025). Evaluates frontier LLM failure modes in realistic, high-stakes freelance software tasks, highlighting the practical unreliability of autonomous execution.
- Paper: Unvalidated Trust: Cross-Stage Vulnerabilities in Large Language Model Architectures, Dominik Schwarz (2025). Analyzes cross-stage vulnerabilities and unvalidated trust inheritance across multi-step LLM architectures, providing structural context for silent error propagation.
- Paper: ExpertQA: Expert-Curated Questions and Attributed Answers, Chaitanya Malaviya et al. (2024). Highlights LLM unreliability and factuality degradation across specialized professional domains, motivating multi-domain evaluation like DELEGATE-52.
- Paper: GLM-5: from Vibe Coding to Agentic Engineering, GLM-5-Team et al. (2026). Explores the architectural and algorithmic transition from casual 'vibe coding' to robust agentic engineering to mitigate the severe execution and document corruption failures revealed in delegated tasks.
- Paper: LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation, Dongge Han et al. (2026). Proposes a modular procedural memory framework (LEGOMem) to address stateless execution and prevent compounding errors in multi-agent document editing and workflow automation.
- Paper: Decentralized Multi-Agent Systems with Shared Context, Yuzhen Mao et al. (2026). Develops a decentralized multi-agent system with shared, verified context to resolve coordination bottlenecks and error propagation in long-context document reasoning.
- Paper: Learning to Orchestrate Agents in Natural Language with the Conductor, Stefan Nielsen et al. (2026). Investigates learned meta-orchestration of heterogeneous LLM agents to improve task decomposition and execution fidelity in complex multi-step workflows.
- Paper: Control Illusion: The Failure of Instruction Hierarchies in Large Language Models, Yilin Geng et al. (2026). Investigates how LLMs fail to enforce system-level constraint hierarchies, explaining why delegated agents violate core operational rules during extended interactions.
- Paper: Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context, Keivan Alizadeh et al. (2026). Introduces self-reflective program search over extended inputs to overcome the long-context degradation and detail loss identified during delegated document processing.
- Paper: Toward Efficient Agents: Memory, Tool learning, and Planning, Xiaofang Yang et al. (2026). Surveys architectural strategies in agent memory, tool learning, and planning to overcome the compounding inefficiencies and unreliability present in long-horizon agent workflows.
