CornerCase: Automated Extremal Testing of Protocol Implementations using LLMs
Rathin SinghaKuan QianSrinath SaikrishnanTracy ZhaoSoheil AbbaslooRyan BeckettSiva Kesava Reddy KakarlaTodd MillsteinGeorge Varghese
Introduces CornerCase, an automated testing framework that uses large language models to extract boundary constraints from protocol specifications and generate targeted edge-case tests, uncovering dozens of previously unknown bugs across widely used implementations of HTTP, DNS, BGP, SMTP, and QUIC.
Modern internet infrastructure relies on complex network protocol implementations across web servers, routing suites, and email engines. Software defects frequently emerge at the boundaries of technical specifications—such as inputs that sit just inside or outside permissible ranges, or commands that violate protocol state machines. While automated testing tools like fuzzers excel at detecting generic parsing crashes through random inputs, they struggle to generate semantically rich edge cases. Consequently, subtle boundary-handling flaws often escape detection, exposing systems to denial-of-service vulnerabilities, routing loops, and security bypasses.
The article demonstrates an automated extremal testing framework named CornerCase. Its primary objective is to evaluate how effectively large language models can convert natural-language specification documents into boundary-focused test cases to discover semantic and state-machine defects across diverse protocol implementations.
The authors develop a structured, black-box pipeline that operates without access to source code. The framework first ingests official specification documents (Requests for Comments, or RFCs) and uses a large language model to extract explicit validity constraints section by section, resolving any cross-referenced rules. The model then generates extremal test cases designed specifically at and near the boundaries of each constraint (both barely valid and barely invalid inputs). These tests are executed in parallel across multiple independent software implementations in isolated container environments. A differential testing engine flags behavioral discrepancies, and a secondary language-model stage analyzes, scores, and triages these anomalies to prioritize probable bugs for human review. The approach was evaluated across 38 implementations spanning five major network protocols: HTTP, DNS, BGP, SMTP, and QUIC.
Extremal testing identified 42 distinct bugs and specification inconsistencies across the evaluated systems. Developers have already acknowledged 26 of these issues and fully patched 18, with the remainder currently under investigation. Key discoveries include an HTTP server (h2o) that creates infinite redirect loops when processing URLs containing encoded null bytes, multiple web servers serving unauthorized files despite malformed or empty Host headers, routing software (GoBGP) failing to reject Autonomous System loops, and mail servers (Mailpit) permitting invalid nested transaction sequences. Crucially, an ablation study revealed that decomposing the testing process into section-wise processing, constraint extraction, and reference expansion generates up to 22 times more behavioral anomalies than attempting to generate tests in a single unconstrained step.
These findings demonstrate that specification-driven extremal testing effectively uncovers high-impact semantic, state-machine, and security flaws that traditional testing techniques miss. For organizations maintaining critical software infrastructure, the framework significantly reduces testing blind spots and improves compliance with open standards. By shifting the primary development bottleneck from test generation to result validation and triaging, the structured methodology makes systematic specification coverage feasible and scalable.
Engineering teams should adopt multi-stage specification analysis over naive direct test generation when integrating automated language models into quality assurance pipelines. Furthermore, organizations managing network protocol stacks should integrate differential boundary testing into continuous integration workflows to prevent edge-case regressions. To handle the high volume of discrepancies that automated testing produces, teams should implement automated scoring and grouping mechanisms to prioritize actionable flaws.
The framework currently focuses on explicitly stated boundaries and short transaction sequences, meaning implicit design assumptions, multi-step workflows, and interactions spanning multiple cross-referenced specifications remain areas for expansion. Nevertheless, the high rate of developer-confirmed and patched bugs provides strong confidence in the framework's effectiveness as a complementary tool alongside fuzzing and formal verification.
- Paper: 3DGen: AI-Assisted Generation of Provably Correct Binary Format Parsers, Sarah Fakhoury et al. (2025). This paper establishes how LLM agents can parse natural language RFC specifications to extract protocol constraints and drive automated differential testing, providing the foundational conceptual framework that CornerCase specializes for extremal boundary testing.
No sufficiently relevant recommendations were found.
