3DGen: AI-Assisted Generation of Provably Correct Binary Format Parsers
Sarah FakhouryMarkus KuppeShuvendu LahiriTahina RamananandroNikhil Swamy
Presents 3DGen, an automated framework that pairs AI agents with symbolic test generation to translate informal RFC specifications into provably correct, memory-safe C binary parsers.
Incorrect parsing of binary network inputs is a leading source of severe software security vulnerabilities. Manually converting ambiguous, natural language specifications from Request for Comments (RFC) documents into low-level, memory-unsafe languages like C is notoriously error-prone, leaving systems exposed to exploits for decades. Formal domain-specific languages (DSLs) like 3D offer provable correctness and memory safety via code generators like EverParse, but authoring these specifications by hand requires steep learning curves and substantial engineering effort.
The article introduces 3DGen, an automated framework designed to evaluate whether artificial intelligence agents can reliably translate natural language RFC documents into formal 3D specifications, which then automatically compile into verified, secure C code.
The framework combines a multi-agent system powered by GPT-4 with automated formal methods. The agents—split into planning, domain expertise, and 3D programming roles—interactively draft candidate specifications. To validate these candidates, the framework introduces 3DTestGen, a symbolic test generator that produces test cases and performs differential testing. The overall approach was evaluated across 20 standardized Internet protocols, using Wireshark as an initial external oracle to label valid and invalid network packets.
The evaluation revealed several critical findings. First, while 3DGen initially achieved a 45% pass rate (9 of 20 protocols) against standard Wireshark labeling, investigation showed that all 11 failing cases were due to Wireshark being overly permissive and ignoring RFC-mandated constraints. Once labels were corrected to match RFC standards, 3DGen achieved a 100% success rate across all 20 protocols. Second, when comparing generated specifications to existing expert-written specifications across 7 protocols, 3DGen matched equivalent behavior and uncovered three human errors in existing specifications for UDP, ICMP, and VXLAN. Third, symbolic differential analysis successfully distinguished subtle semantic divergences among candidate specifications, proving its value in grouping equivalent programs and surfacing missing constraints.
These findings demonstrate that targeting an analyzable domain-specific language rather than prompting an AI model for raw C code drastically mitigates software security risks. This intermediate representation allows symbolic tools to systematically check AI outputs, eliminate undefined behaviors, and generate provably secure C parsers at scale. Furthermore, the framework serves as an auditing mechanism to expose discrepancies in legacy tools and human-authored specifications.
Organizations developing security-critical software should consider adopting DSL-based AI pipelines with symbolic validation loops rather than directly deploying AI-generated imperative code. Before strong autonomous adoption, development teams should maintain human review over synthesized tests and conduct user studies to determine how seamlessly non-expert engineers can operate the intent-refinement loop.
The study's primary limitation is the framework's dependency on the quality and strictness of the labeling oracle, as shown by Wireshark's permissiveness and occasional under-constrained test sets. Users should exercise caution to ensure test suites are comprehensive, but confidence remains high that candidate specifications passing well-aligned symbolic test suites will compile into provably memory-safe and correct parsing code.
- Paper: Teaching Large Language Models to Self-Debug, Xinyun Chen et al. (2023). Introduces foundational techniques for iterative LLM self-debugging and refinement against execution feedback, which underpin the repeated repair cycles in 3DGen.
- Paper: Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning, Saibo Geng et al. (2023). Demonstrates grammar-constrained decoding to force neural models to adhere strictly to formal syntax, providing context for restricting LLM generation to formal domain-specific languages.
- Paper: Instruction Tuning for Secure Code Generation, Jingxuan He et al. (2024). Explores tuning language models to prevent memory and security vulnerabilities during code generation, addressing the core security motivation of safe parser synthesis.
- Paper: CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis, Erik Nijkamp et al. (2022). Establishes multi-turn conversational program synthesis frameworks that decompose complex specification translation tasks into iterative sub-goals.
- Paper: Competition-level code generation with AlphaCode, Yujia Li et al. (2022). Pioneers the methodology of generating candidate programs at scale and pruning them via automated test execution and behavioral clustering.
- Paper: VeruSAGE: A Study of Agent-Based Verification for Rust Systems, Chenyuan Yang et al. (2025). Extends the concept of AI-assisted formal systems verification by evaluating multi-agent LLM frameworks on end-to-end correctness proofs for real-world systems code.
- Paper: LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks, Po-Nien Kung et al. (2026). Applies agentic informal-to-formal decomposition and compiler-in-the-loop validation to construct mechanically verified mathematical proofs.
- Paper: RustAssistant: Using LLMs to Fix Compilation Errors in Rust Code, Pantazis Deligiannis et al. (2025). Investigates interactive LLM repair guided by compiler and static analysis diagnostics to resolve complex memory safety constraints in systems software.
- Paper: AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security, Dongrui Liu et al. (2026). Develops diagnostic guardrails and execution path attribution to monitor and secure autonomous multi-step AI agent workflows.
