Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning
Ben Kereopa-YorkeGuillermo DíazHolly WrightReagan JohnstonRon F. Del RosarioTimothy Lynar
Reveals a critical vulnerability in AI agent architectures by demonstrating that corrupting structured knowledge graphs queried via tool-use reliably causes nine leading models to reach false conclusions through valid reasoning on a 42-million-node production system.
Artificial intelligence agents increasingly rely on structured knowledge graphs to understand large, complex codebases and perform tasks such as dependency management, vulnerability discovery, and automated code review. Connected through standard protocols like the Model Context Protocol, these knowledge graphs act as authoritative "oracles" that agents treat as absolute ground truth. However, because these systems often use shared, unsegmented database connections that permit write access without validating data origins, they introduce an unaddressed security threat. The article investigates "Oracle Poisoning," a newly defined attack class where an adversary alters structured graph data—rather than injecting instructions into prompts or altering model weights—causing artificial intelligence agents to reach incorrect, dangerous conclusions through flawless internal reasoning.
To demonstrate this threat, the researchers tested six attack scenarios against a live, production-scale software knowledge graph comprising 42 million nodes. They evaluated nine leading commercial artificial intelligence models from OpenAI, Anthropic, and Google across hundreds of trials using genuine agent software development kits where models autonomously invoked graph query tools. The investigation tested various attacker skill levels, delivery channels, and candidate defenses to assess systemic vulnerabilities and practical mitigations.
Key findings show that under directed queries, every tested model accepted fabricated security claims 100% of the time (269 out of 269 valid trials) when presented with moderately sophisticated poisoned data. An attacker required only a minimal footprint—modifying as few as one to two nodes or changing properties on an existing node without creating new ones—to manipulate agent outputs across multiple workflows. Furthermore, evaluating models via inline text instead of real agent tool calls masked vulnerabilities, in one case showing 0% trust inline but 100% trust during actual tool-use. When presented with open-ended security analysis prompts rather than direct yes-or-no queries, model trust dropped to between 3% and 55% as models tended to hedge rather than issue firm conclusions.
These results demonstrate that agent reasoning capability cannot overcome falsified underlying data; higher reasoning performance merely produces more articulate justifications for false premises. This introduces severe operational and compliance risks, as automated agents may falsely certify that vulnerabilities are patched or dependencies are safe. Structural analysis indicates that several major code intelligence platforms share these architectural preconditions, prioritizing developer convenience over data provenance.
To counter Oracle Poisoning, the article recommends a layered defense strategy. Enforcing read-only access control on tool database integrations is the most effective immediate mitigation, eliminating the direct mutation vector at minimal engineering cost. Deploying multi-tool cross-verification—enabling agents to cross-check graph claims against authoritative primary source code—dramatically reduces blind trust down to 0% to 25%. In contrast, textual defenses such as system prompt hardening proved entirely ineffective, and generic critical reviews generated false alarms at the same rate as true detections. Organizations must treat data store integrity as a primary security boundary and re-evaluate tool vulnerability whenever deploying updated model versions.
- Paper: Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection, Kai Greshake et al. (2023). Provides the foundational framework and threat model for indirect prompt injection in LLMs integrated with external tools and runtime data retrieval.
- Paper: Can Knowledge Graphs Reduce Hallucinations in LLMs? : A Survey, Garima Agrawal et al. (2024). Surveys the core architectures and mechanics of augmenting LLM reasoning with structured external knowledge graphs.
- Paper: Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs, Liyi Chen et al. (2024). Examines how agentic LLMs dynamically plan, query, and reason over structured graph facts during multi-hop problem solving.
- Paper: Evaluating the Instruction-Following Robustness of Large Language Models to Prompt Injection, Zekun Li et al. (2024). Benchmarks LLM robustness when distinguishing legitimate queries from adversarial content injected into retrieved reference text.
- Paper: BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents, Yifei Wang et al. (2024). Analyzes how adversarial triggers embedded in an agent's operating environment can compromise tool-use workflows.
- Paper: Neural-Symbolic Models for Logical Queries on Knowledge Graphs, Zhaocheng Zhu et al. (2022). Details the fundamentals of executing logical, multi-hop reasoning over structured knowledge graphs.
- Paper: Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use, Aradhye Agarwal et al. (2026). Develops a guardrail framework for multi-step tool-use agents to verify risks and safely refuse corrupted tool responses during reasoning.
