CornerCase: Automated Extremal Testing of Protocol Implementations using LLMs

Rathin SinghaKuan QianSrinath SaikrishnanTracy ZhaoSoheil AbbaslooRyan BeckettSiva Kesava Reddy KakarlaTodd MillsteinGeorge Varghese

article2026arXiv0 citations

Introduces CornerCase, an automated testing framework that uses large language models to extract boundary constraints from protocol specifications and generate targeted edge-case tests, uncovering dozens of previously unknown bugs across widely used implementations of HTTP, DNS, BGP, SMTP, and QUIC.

Listen

Modern internet infrastructure relies on complex network protocol implementations across web servers, routing suites, and email engines. Software defects frequently emerge at the boundaries of technical specifications—such as inputs that sit just inside or outside permissible ranges, or commands that violate protocol state machines. While automated testing tools like fuzzers excel at detecting generic parsing crashes through random inputs, they struggle to generate semantically rich edge cases. Consequently, subtle boundary-handling flaws often escape detection, exposing systems to denial-of-service vulnerabilities, routing loops, and security bypasses.

The article demonstrates an automated extremal testing framework named CornerCase. Its primary objective is to evaluate how effectively large language models can convert natural-language specification documents into boundary-focused test cases to discover semantic and state-machine defects across diverse protocol implementations.

The authors develop a structured, black-box pipeline that operates without access to source code. The framework first ingests official specification documents (Requests for Comments, or RFCs) and uses a large language model to extract explicit validity constraints section by section, resolving any cross-referenced rules. The model then generates extremal test cases designed specifically at and near the boundaries of each constraint (both barely valid and barely invalid inputs). These tests are executed in parallel across multiple independent software implementations in isolated container environments. A differential testing engine flags behavioral discrepancies, and a secondary language-model stage analyzes, scores, and triages these anomalies to prioritize probable bugs for human review. The approach was evaluated across 38 implementations spanning five major network protocols: HTTP, DNS, BGP, SMTP, and QUIC.

Extremal testing identified 42 distinct bugs and specification inconsistencies across the evaluated systems. Developers have already acknowledged 26 of these issues and fully patched 18, with the remainder currently under investigation. Key discoveries include an HTTP server (h2o) that creates infinite redirect loops when processing URLs containing encoded null bytes, multiple web servers serving unauthorized files despite malformed or empty Host headers, routing software (GoBGP) failing to reject Autonomous System loops, and mail servers (Mailpit) permitting invalid nested transaction sequences. Crucially, an ablation study revealed that decomposing the testing process into section-wise processing, constraint extraction, and reference expansion generates up to 22 times more behavioral anomalies than attempting to generate tests in a single unconstrained step.

These findings demonstrate that specification-driven extremal testing effectively uncovers high-impact semantic, state-machine, and security flaws that traditional testing techniques miss. For organizations maintaining critical software infrastructure, the framework significantly reduces testing blind spots and improves compliance with open standards. By shifting the primary development bottleneck from test generation to result validation and triaging, the structured methodology makes systematic specification coverage feasible and scalable.

Engineering teams should adopt multi-stage specification analysis over naive direct test generation when integrating automated language models into quality assurance pipelines. Furthermore, organizations managing network protocol stacks should integrate differential boundary testing into continuous integration workflows to prevent edge-case regressions. To handle the high volume of discrepancies that automated testing produces, teams should implement automated scoring and grouping mechanisms to prioritize actionable flaws.

The framework currently focuses on explicitly stated boundaries and short transaction sequences, meaning implicit design assumptions, multi-step workflows, and interactions spanning multiple cross-referenced specifications remain areas for expansion. Nevertheless, the high rate of developer-confirmed and patched bugs provides strong confidence in the framework's effectiveness as a complementary tool alongside fuzzing and formal verification.

arXiv: 2606.29124
  • Paper: 3DGen: AI-Assisted Generation of Provably Correct Binary Format Parsers, Sarah Fakhoury et al. (2025). This paper establishes how LLM agents can parse natural language RFC specifications to extract protocol constraints and drive automated differential testing, providing the foundational conceptual framework that CornerCase specializes for extremal boundary testing.

No sufficiently relevant recommendations were found.

Cover for CornerCase: Automated Extremal Testing of Protocol Implementations using LLMs

Abstract

Many software bugs in network protocol implementations arise near specification boundaries, such as inputs just within or outside allowed ranges, or messages that are valid in isolation but invalid in a given state. From the SSL Heartbleed exploit to TCP Christmas Tree packets, boundary inputs have repeatedly exposed critical weaknesses, yet remain under-tested by existing techniques such as fuzzing and model-based testing. We present CornerCase, an automated extremal testing approach that systematically targets such boundary behaviors. Our key idea is to decompose test generation into two stages: first, large language models (LLMs) extract explicit validity constraints from protocol specifications (e.g., RFCs) in a structured, section-by-section manner; second, extremal test cases are generated at or near the boundary of each constraint. These tests are executed across multiple implementations, and differential testing identifies inconsistencies. We evaluate CornerCase on widely used implementations of HTTP, DNS, BGP, SMTP, and QUIC, uncovering many previously unknown bugs. For example, the HTTP server h2o enters a redirect loop when processing URLs containing encoded null bytes. Overall, we used CornerCase to identify and file 42 anomalies; to date 26 have been acknowledged as bugs and 18 fixed, with others under active investigation

Table of Contents

  • 1 Introduction
  • 2 Motivating Examples
  • 2.1 HTTP Null Byte Loops (h2o)
  • 2.2 BGP Confederation Loops (GoBGP)
  • 2.3 SMTP Nesting Failures (Mailpit)
  • 3 Methodology
  • 3.1 Input
  • 3.2 Stage 1: Constraint Generation
  • 3.3 Stage 2: Test Generation
  • 3.4 Stage 3: Test Execution
  • 3.5 Stage 4: Differential Testing
  • 3.6 Stage 5: Result Analysis
  • 3.7 Stage 6: Triaging
  • 3.8 Summary
  • 4 Evaluation
  • 4.1 Experimental Setup
  • 4.2 Differential Anomalies
  • 4.3 Ablation Study
  • 4.4 Lessons Learned
  • 5 Related Work
  • 6 Limitations and Future Work
  • 7 Conclusion
  • References
  • A Test Format
  • A.1 HTTP
  • A.2 DNS
  • A.3 SMTP
  • A.4 BGP (Confederation)
  • A.5 QUIC
  • B LLM Prompts
  • B.1 Constraint Generation Prompt
  • B.2 Test Generation Prompt
  • B.3 Result Analysis Prompt
  • C Implementation Matrix

Knowls

  1. Knowl 1 — CornerCase Extremal Testing Architecture

    model/method

    CornerCase is an automated, protocol-agnostic, and implementation-agnostic extremal testing framework that targets boundary behaviors in network protocol implementations using Large Language Models (LLMs) and differential testing.

    Rather than asking an LLM to generate tests end-to-end from an entire specification, CornerCase decomposes test generation into a multi-stage pipeline:

    1. Specification Input & Formatting: Takes as input a natural language protocol specification (e.g., an RFC), a user-defined JSON schema specifying controllable test input and observable output fields, and an executable black-box test harness.
    2. Constraint Extraction: Decomposes the RFC into individual sections and uses an LLM to extract verbatim sentences containing testable normative rules and validity constraints that map to the provided test format.
    3. Extremal Test Generation: Prompts the LLM on small batches of extracted constraints (along with cross-referenced RFC context) to construct extremal test inputs that lie just at, just below, and just above the validity boundary (both barely valid positive cases and barely invalid negative cases).
    4. Differential Execution: Runs generated test cases across multiple independent, black-box implementations in containerized environments and detects behavioral divergence.
    5. Automated Analysis & Triaging: Employs an LLM to assess behavioral discrepancies, assign confidence scores indicating the probability of an RFC violation, group anomalies by constraint and boundary polarity tags, and prioritize unique candidates for manual validation.
  2. Knowl 2 — Section-Wise RFC Constraint Extraction and Context Resolution

    algorithm

    Constraint generation converts unstructured specification prose into an enumerable set of testable boundaries without semantic alteration.

    Input: RFC document text DD, JSON test input schema FF
    Output: Registry of extracted constraints C={(k,s)}C = \{(k, s)\} where kk is section number and ss is verbatim constraint sentence
    1. Segment DD line-by-line into sections using header regex pattern ^\s*(\d+(?:\.\d+)*)\.\s+(.+)$
    2. Initialize global constraint registry C←∅C \leftarrow \emptyset
    3. for each section with number kk and text TkT_k in DD:
    4. Prompt LLM with TkT_k and input schema FF to identify testable normative statements (MUST, MUST NOT, SHOULD, SHOULD NOT) and non-normative rules
    5. Receive candidate list of tuples [(k,s1),(k,s2),… ][(k, s_1), (k, s_2), \dots]
    6. for each returned tuple (k,s)(k, s):
    7. Verify that sentence ss matches the verbatim text in TkT_k
    8. Detect cross-references in ss matching regex patterns for section citations (e.g., "Section X.Y")
    9. Extract referenced section indices Rs={r1,r2,… }R_s = \{r_1, r_2, \dots\}
    10. if (k,s)∉C(k, s) \notin C then
    11. C←C∪{(k,s,Rs)}C \leftarrow C \cup \{(k, s, R_s)\}
    12. return CC

    Extracted constraints are categorized into eight classes:

    • Range constraints: Numeric bounds on values.
    • Size constraints: Bounds on strings, packets, lists, or field lengths.
    • Format constraints: Syntactic and encoding rules.
    • Dependency constraints: Inter-field relationship rules.
    • Presence constraints: Mandatory vs. forbidden conditions for fields/commands.
    • Enumeration constraints: Selection from fixed allowed value sets.
    • Ordering/state constraints: Sequencing rules for messages across protocol states.
    • Cross-field semantic constraints: Rules requiring evaluation of field values against protocol configuration or contextual state.
  3. Knowl 3 — Constraint-Driven Extremal Test Case Generation

    algorithm

    Extremal test generation produces minimal boundary inputs targeting exact specification constraints.

    Input: Constraint registry CC, JSON test input format FF, batch size B=5B = 5, RFC section text lookup TT
    Output: Array of structured test cases T\mathcal{T}
    1. Initialize test suite T←[]\mathcal{T} \leftarrow []
    2. Partition CC into batches {C1,C2,… }\{C_1, C_2, \dots\} of size ≤B\le B
    3. for each batch CbC_b:
    4. Collect referenced section texts Tref=⋃(k,s,Rs)∈Cb{Tr∣r∈Rs}T_{\text{ref}} = \bigcup_{(k, s, R_s) \in C_b} \{T_r \mid r \in R_s\}
    5. Construct prompt containing schema FF, constraints in CbC_b, and referenced section texts TrefT_{\text{ref}}
    6. Prompt LLM to generate extremal inputs targeting boundaries for each constraint:
    7. - "Almost valid" (barely satisfying boundary, e.g., length =Lmax⁡= L_{\max}, value =xmin⁡= x_{\min})
    8. - "Almost invalid" (barely violating boundary, e.g., length =Lmax⁡+1= L_{\max} + 1, single invalid character)
    9. - Generate tag G∈{ConstraintID_Positive,ConstraintID_Negative}G \in \{\text{ConstraintID\_Positive}, \text{ConstraintID\_Negative}\} for each test
    10. Receive JSON array of test objects OO
    11. for each test object o∈Oo \in O:
    12. Validate that oo strictly complies with schema FF
    13. Assign globally unique identifier test_id to oo
    14. Ensure field constraint contains the exact verbatim RFC sentence
    15. Append oo to T\mathcal{T}
    16. return T\mathcal{T}
  4. Knowl 4 — Differential Execution, Anomaly Scoring, and Tag-Based Triaging

    algorithm

    Differential analysis and LLM-guided triaging identify, score, and group protocol implementation discrepancies.

    Input: Test suite T\mathcal{T}, set of implementations I={I1,I2,…,Im}\mathcal{I} = \{I_1, I_2, \dots, I_m\}, test harness HH, confidence threshold θ\theta
    Output: Ranked and triaged bug candidates B\mathcal{B}
    1. Initialize raw anomalies A←[]\mathcal{A} \leftarrow []
    2. for each test case t∈Tt \in \mathcal{T}:
    3. for each implementation Ij∈II_j \in \mathcal{I}:
    4. Execute tt against IjI_j using harness HH to obtain structured JSON response R(Ij,t)R(I_j, t)
    5. if ∃j1,j2\exists j_1, j_2 such that R(Ij1,t)≠R(Ij2,t)R(I_{j_1}, t) \neq R(I_{j_2}, t) then
    6. Record anomaly record a=(t,{R(Ij,t)∣Ij∈I})a = (t, \{R(I_j, t) \mid I_j \in \mathcal{I}\})
    7. Append aa to A\mathcal{A}
    8. Initialize scored anomalies S←[]\mathcal{S} \leftarrow []
    9. for each batch of anomalies in A\mathcal{A}:
    10. Prompt LLM with implementation setups, test inputs, and disparate outputs
    11. Receive explanatory analysis and confidence score c∈[0,10]c \in [0, 10] per anomaly
    12. Append (a,c)(a, c) to S\mathcal{S}
    13. Group scored anomalies S\mathcal{S} by test tag t.tagt.\text{tag}
    14. Initialize triaged candidate set B←[]\mathcal{B} \leftarrow []
    15. for each tag group GG:
    16. Sort items in GG descending by confidence score cc
    17. Select highest-confidence test case a∗∈Ga^* \in G
    18. if a∗.c≥θa^*.c \ge \theta then
    19. Append a∗a^* to B\mathcal{B}
    20. return B\mathcal{B} for manual human validation
  5. Knowl 5 — Cross-Protocol Extremal Testing Evaluation and Yield

    data/table

    CornerCase was evaluated across five network protocols targeting standard RFC specifications: SMTP (RFC 5321), DNS (RFC 2181), BGP Confederation (RFC 5065), HTTP URI parsing (RFC 3986), and QUIC TLS handshake (RFC 8446).

    Protocol RFC Constraints Tests Anomalies Prioritized Triaged
    SMTP 5321 177 630 526 84 54
    DNS 2181 65 213 156 28 23
    BGP (Confederation) 5065 28 98 64 25 20
    HTTP 3986 239 757 297 75 42
    QUIC 8446 136 328 16 6 5

    Each extracted constraint yielded between 3 and 11 test cases covering boundary, sub-boundary, and super-boundary values. Following automated differential execution, confidence scoring (using an initial threshold θ=8\theta = 8), and tag-based deduplication, candidate anomalies were manually inspected.

    In total, 42 unique bugs and specification violations were discovered across all tested implementations: 26 were acknowledged by maintainers and 18 were patched/fixed.

  6. Knowl 6 — Ablation Analysis of Pipeline Structural Components

    data/table

    An ablation study evaluated the necessity of three structural components in the CornerCase pipeline:

    1. Section-wise RFC chunking (S) vs. full-RFC one-shot prompting.
    2. Explicit constraint extraction (C) vs. direct test generation from text.
    3. Targeted cross-reference expansion (R) vs. local section context only.
    Protocol Variant #Cons. #Tests #Anom.
    SMTP Base X 25 20
    S X 234 124
    SC 154 613 504
    SR X 284 168
    CR 41 169 95
    SCR 177 630 526
    DNS Base X 20 11
    S X 168 71
    SC 63 209 100
    SR X 162 86
    CR 27 96 66
    SCR 65 213 156
    BGP Confed Base X 12 2
    S X 68 45
    SC 23 85 49
    SR X 64 33
    CR 20 78 25
    SCR 28 98 64
    HTTP Base X 27 13
    S X 437 131
    SC 165 660 207
    SR X 430 139
    CR 28 124 47
    SCR 239 757 297
    QUIC Base X 20 0
    S X 230 8
    SC 168 344 20
    SR X 336 10
    CR 21 52 1
    SCR 136 328 16

    Key takeaways from the ablation data:

    • Section-by-section processing (S) yields a 10×10\times anomaly increase over Base for HTTP (13→13113 \to 131) and a 22×22\times increase for BGP (2→452 \to 45).
    • Adding constraint extraction (C) substantially expands test volume and anomaly detection across all protocols (e.g., finding 526 anomalies in SMTP under SCR vs. 124 under S).
    • Reference expansion (R) improves yield when references contain concrete testable parameters (such as in DNS: 156 under SCR vs. 100 under SC), but can introduce noise if cross-references consist primarily of descriptive commentary without controllable input parameters (e.g., in QUIC: 16 under SCR vs. 20 under SC).
  7. Knowl 7 — Protocol Implementation Bug Patterns and Representative Failures

    empirical result

    The 42 implementation bugs uncovered by CornerCase fall into three principal failure modes:

    1. Missing Input Validation:

      • HTTP Servers (Caddy, H2O, Nginx): Caddy and H2O served files when the Host header was entirely missing or empty instead of rejecting requests with 400 Bad Request. Caddy and Nginx accepted invalid IP-literal values in Host headers. In Caddy, requesting a percent-encoded null byte (GET /%00) propagated to a filesystem stat call, triggering an EINVAL error that leaked as a 500 Internal Server Error.
      • SMTP (aiosmtpd, Mailpit): aiosmtpd accepted syntactically malformed MAIL FROM and RCPT TO commands lacking angle brackets (e.g., RCPT TO:Postmaster), responding with 250 OK instead of returning a 501 syntax error. Mailpit accepted malformed source routes lacking colons in RCPT TO.
    2. State Machine Violations:

      • SMTP (Mailpit, OpenSMTPD): Mailpit accepted MAIL FROM before receiving an initial EHLO/HELO greeting. Mailpit also accepted nested MAIL FROM commands while a mail transaction was actively open (after RCPT TO but before DATA), returning 250 OK instead of rejecting the bad sequence with 503. OpenSMTPD returned a 503 error on a second EHLO command sent after MAIL FROM, violating the RFC 5321 requirement that a mid-transaction EHLO resets the session state.
    3. Semantic Misinterpretation:

      • HTTP Server (h2o): Processing a URI path containing a percent-encoded null byte (GET /%00) caused h2o to issue a 301 Moved Permanently redirect to /%00/, inducing an infinite redirect loop.
      • BGP (GoBGP, FRR, Batfish): GoBGP accepted BGP routes whose AS_PATH contained the router's own confederation identifier instead of treating them as routing loops. FRR incorrectly rejected valid routes containing its own Member-AS in a regular (non-confederation) AS_PATH as confederation loops. GoBGP established eBGP sessions and propagated routes between peers residing in the same Member-AS.
      • QUIC (kwik, quic-go): kwik permitted client handshakes offering exclusively obsolete TLS 1.2 versions in supported_versions to succeed rather than terminating the connection. quic-go rejected valid TLS 1.3 handshakes when obsolete TLS 1.2 was additionally included in supported_versions.
  8. Knowl 8 — Multi-Protocol Experimental Setup and Target Implementations

    experimental setup

    CornerCase was evaluated across 38 diverse open- and closed-source protocol implementations written in C, C++, Go, Rust, Python, Java, and C#, executed within isolated Docker containers to guarantee reproducibility.

    The evaluated suites comprised:

    • HTTP (5 servers): NGINX (C), Apache httpd (C), Caddy (Go), H2O (C), lighttpd (C). Evaluated on RFC 3986 URI parsing against synthesized filesystem directory/symlink structures.
    • DNS (10 servers): BIND (C), NSD (C), Knot DNS (C), PowerDNS (C++), CoreDNS (Go), YADIFA (C), HickoryDNS (Rust), gdnsd (C), TwistedNames (Python), Technitium (C#). Evaluated on RFC 2181 resource record sets (RRSet duplicate suppression, TTL consistency, zone cuts).
    • BGP Confederation (3 implementations): FRRouting (C), GoBGP (Go), Batfish (Java). Evaluated on RFC 5065 confederation handling and RFC 4456 route reflection under multi-router topologies.
    • SMTP (5 servers): smtpd (Python), aiosmtpd (Python), OpenSMTPD (C), Mailpit (Go), Stalwart (Rust). Evaluated on RFC 5321 command syntax and multi-command sequencing.
    • QUIC / TLS Handshake (15 implementations): quic-go (Go), go-x-net (Go), picoquic (C), HAProxy (C), MsQuic (C), mvfst (C++), nginx (C), quinn (Rust), neqo (Rust), quiche (Rust), lsquic (C), aioquic (Python, modified to support client mutation), ngtcp2 (C), kwik (Java), Chrome image for QuicInteropRunner. Evaluated on RFC 8446 TLS 1.3 ClientHello mutations.

    All prompt-based stages utilized GPT-5 through the OpenAI API.

  9. Knowl 9 — Limitations of Extremal Protocol Testing

    limitation

    CornerCase exhibits four principal limitations:

    1. Implicit Specification Constraints: The framework extracts constraints explicitly stated via normative or descriptive rules in RFC text. Implicit constraints—rules that are unstated in the RFC prose but fundamentally implied (such as payload lengths matching length fields in TLS Heartbleed, or invalid TCP flag combinations in Christmas Tree packets)—are not captured by standard text-scanning prompts.
    2. Short-Sequence Interaction Depth: Test generation currently targets isolated messages or short command sequences. It does not comprehensively explore long multi-step interaction workflows, asynchronous timing dependencies, or deep protocol state spaces.
    3. Single-RFC Processing Scope: Constraint extraction processes single specification documents in isolation. It does not automatically resolve or traverse multi-RFC dependency hierarchies where protocol semantics span multiple interdependent standards (e.g., HTTP/2 RFC 9113 depending on HTTP Semantics RFC 9110, which references URI Syntax RFC 3986).
    4. LLM Dependency and Context Dilution: While the two-stage decomposition is model-agnostic, LLM performance depends on context density. Broad context containing non-testable prose commentary (such as TLS 1.3 RFC 8446 cross-references) can dilute prompt focus and slightly reduce anomaly yield relative to unexpanded section-by-section extraction.

Coverage note — None was omitted; all contributed methodology, algorithms, experimental setup parameters, cross-protocol empirical findings, ablation study results, and limitations are fully covered.

References

  1. 1.aiosmtpd community. aiosmtpd - An asyncio based SMTP server. https://aiosmtpd.aio-libs.org/en/latest/, 2026.
  2. 2.American Fuzzing Lop AFL. AFL 2018. https://lcamtuf.coredump.cx/afl/.
  3. 3.Anthropic. Assessing Claude Mythos Preview’s Cybersecurity Capabilities. https://red.anthropic.com/2026/mythos-preview/, 2026. Accessed: 2026-04-19.
  4. 4.R. Can Aygun, Yehuda Afek, Anat Bremler-Barr, and Leonard Kleinrock. LAPRAD: LLM-Assisted PRotocol Attack Discovery. In IFIP Networking 2025 Proceedings, 2025. Also available as arXiv:2510.19264.
  5. 5.Asma Bhat and S. M. K. Quadri. Equivalence class partitioning and boundary value analysis - A review. In 2015 2nd International Conference on Computing for Sustainable Global Development (INDIACom), pages 1557–1562, 2015.
  6. 6.Brandon L Black and Community. gdnsd. https://gdnsd.org/, 2023. Github: https://github.com/gdnsd/gdnsd.
  7. 7.Marcel Böhme, Van-Thuan Pham, and Abhik Roychoudhury. Coverage-based greybox fuzzing as Markov chain. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pages 1032–1043, 2016.
  8. 8.Josip Bozic, Lina Marsso, Radu Mateescu, and Franz Wotawa. A formal TLS handshake model in LNT. arXiv preprint arXiv:1803.10319, 2018.
  9. 9.Cristian Cadar, Daniel Dunbar, Dawson R Engler, et al. KLEE: Unassisted and automatic generation of high-coverage tests for complex systems programs. In OSDI, volume 8, pages 209–224, 2008.
  10. 10.Cloudflare, Inc. quiche QUIC Implementation. https://github.com/cloudflare/quiche, 2018. (Accessed: 2026-04-10).
  11. 11.CoreDNS community. CoreDNS. https://coredns.io/, 2026. Github: https://github.com/coredns/coredns.
  12. 12.FRR community. The FRRouting protocol suite. https://frrouting.org/, 2026. Github: https://github.com/FRRouting/frr.
  13. 13.GoBGP community. GoBGP. https://github.com/osrg/gobgp, 2026.
  14. 14.PowerDNS Community. PowerDNS. https://www.powerdns.com/, 2026. Github: https://github.com/PowerDNS/pdns.
  15. 15.Internet Systems Consortium. BIND 9. https://www.isc.org/bind/, 2026. GitLab: https://gitlab.isc.org/isc-projects/bind9.
  16. 16.CZ.NIC. Knot. https://www.knot-dns.cz/, 2025. GitLab: https://gitlab.nic.cz/knot/knot-dns.
  17. 17.Stanislav Dashevskyi. A simple BGP fuzzer based on boofuzz. Github, 2023. https://github.com/Forescout/bgp_boofuzzer.
  18. 18.Joeri de Ruiter and Erik Poll. Protocol State Fuzzing of TLS Implementations. In USENIX Security Symposium, 2015.
  19. 19.Gelei Deng, Yi Liu, Victor Mayoral-Vilches, Peng Liu, Yuekang Li, Yuan Xu, Tianwei Zhang, Yang Liu, Martin Pinzger, and Stefan Rass. Pentestgpt: An llm-empowered automatic penetration testing tool. arXiv preprint arXiv:2308.06782, 2024.
  20. 20.Python developer community. SMTPD Python library. https://docs.python.org/3.10/library/smtpd.html, 2024.
  21. 21.Mailpit developers. Mailpit Email Testing Tool. https://github.com/axllent/mailpit, 2026.
  22. 22.OpenSMTPD developers. OpenSMTPD Mail Server. https://github.com/OpenSMTPD/OpenSMTPD, 2026.
  23. 23.Donatas Abraitis Donald Sharp and et al. Fuzzing targets and supported fuzzers available in FRR. Github, 2023. https://docs.frrouting.org/projects/dev-guide/en/latest/fuzzing.html.
  24. 24.Peter Doornbosch. Kwik QUIC Implementation. https://github.com/ptrd/kwik, 2018. (Accessed: 2026-04-10).
  25. 25.EURid.eu. Yadifa. https://www.yadifa.eu/, 2026. Github: https://github.com/yadifa/yadifa.
  26. 26.Angela Fan, Beliz Gokkaya, Mark Harman, Mitya Lyubarskiy, Shubho Sengupta, Shin Yoo, and Jie M. Zhang. Large language models for software engineering: Survey and open problems. arXiv preprint arXiv:2310.03533, 2023.
  27. 27.Ari Fogel, Stanley Fung, Luis Pedrosa, Meg Walraed-Sullivan, Ramesh Govindan, Ratul Mahajan, and Todd Millstein. A general approach to network configuration analysis. In Proceedings of the 12th USENIX Conference on Networked Systems Design and Implementation, NSDI'15, page 469–483, USA, 2015. USENIX Association.
  28. 28.Apache Software Foundation. Apache HTTP Server. https://httpd.apache.org, 1995. Source code: https://github.com/apache/httpd (accessed 2026-04-10).
  29. 29.Frederic Cambus. Fuzzing DNS zone parsers. https://www.cambus.net/fuzzing-dns-zone-parsers/.
  30. 30.Benjamin Fry and Community. Hickory-DNS. https://github.com/hickory-dns/hickory-dns, 2026. Github: https://github.com/hickory-dns/hickory-dns/.
  31. 31.Patrice Godefroid, Nils Klarlund, and Koushik Sen. DART: Directed automated random testing. In Proceedings of the 2005 ACM SIGPLAN conference on Programming language design and implementation, pages 213–223, 2005.
  32. 32.Xiujing Guo, Chen Li, and Tatsuhiro Tsuchiya. Boundary Value Test Input Generation using Prompt Engineering with LLMs: Fault Detection and Coverage analysis, 2025.
  33. 33.Matthew Holt. Caddy Web Server. https://caddyserver.com, 2015. Source code: https://github.com/caddyserver/caddy (accessed 2026-04-10).
  34. 34.Christian Huitema. picoquic QUIC Implementation. https://github.com/private-octopus/picoquic, 2017. (Accessed: 2026-04-10).
  35. 35.Siva Kesava Reddy Kakarla, Ryan Beckett, Todd Millstein, and George Varghese. SCALE: Automatically finding RFC compliance bugs in DNS nameservers. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22), pages 307–323, 2022.
  36. 36.Jan Kneschke. lighttpd Web Server. https://www.lighttpd.net, 2003. Source code: https://github.com/lighttpd/lighttpd1.4 (accessed 2026-04-10).
  37. 37.NLnet Labs. NSD. https://nlnetlabs.nl/projects/nsd/about/, 2026. Github: https://github.com/NLnetLabs/nsd.
  38. 38.Stalwart Labs. Stalwart Mail Server. https://github.com/stalwartlabs/stalwart, 2026.
  39. 39.Twisted Matrix Labs. TwistedNames. https://twisted.org/, 2026. Github: https://github.com/twisted/twisted.
  40. 40.Jeremy Lainé. aioquic QUIC Implementation. https://github.com/aiortc/aioquic, 2019. (Accessed: 2026-04-10).
  41. 41.Hyojeong Lee, Jeff Seibert, Dylan Fistrovic, Charles Killian, and Cristina Nita-Rotaru. Gatling: Automatic performance attack discovery in Large-scale Distributed systems. ACM Trans. Inf. Syst. Secur., 17(4), apr 2015.
  42. 42.LiteSpeed Technologies. LSQUIC QUIC Implementation. https://github.com/litespeedtech/lsquic, 2017. (Accessed: 2026-04-10).
  43. 43.Gordon Lyon. NMAP Network Scanning: The Official NMAP Project Guide to Network Discovery and Security Scanning. Insecure, 2009.
  44. 44.Meta Platforms, Inc. mvfst QUIC Implementation. https://github.com/facebook/mvfst, 2019. (Accessed: 2026-04-10).
  45. 45.Microsoft Corporation. MsQuic QUIC Implementation. https://github.com/microsoft/msquic, 2019. (Accessed: 2026-04-10).
  46. 46.MITRE Corporation. OpenSSL TLS Heartbeat Extension Read Overrun (CVE-2014-0160). https://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2014-0160, 2014. Accessed: 2026-04-17.
  47. 47.Rajdeep Mondal, Rathin Singha, Todd Millstein, George Varghese, Ryan Beckett, and Siva Kesava Reddy Kakarla. Eywa: Automating model based testing using llms. arXiv preprint arXiv:2312.06875, 2023.
  48. 48.Mozilla Corporation. Neqo QUIC Implementation. https://github.com/mozilla/neqo, 2019. (Accessed: 2026-04-10).
  49. 49.NGINX, Inc. NGINX QUIC Implementation. https://github.com/nginx/nginx, 2020. Project page: https://quic.nginx.org/ (Accessed: 2026-04-10).
  50. 50.NMAP Organization. Dns-fuzz. https://nmap.org/nsedoc/scripts/dns-fuzz.html.
  51. 51.Kazuho Oku. H2O HTTP Server. https://h2o.examp1e.net, 2014. Source Code: https://github.com/h2o/h2o(accessed 2026-04-10).
  52. 52.Peach Fuzzer. https://peachtech.gitlab.io/peach-fuzzer-community/.
  53. 53.quic-go contributors. quic-go QUIC Implementation. https://github.com/quic-go/quic-go, 2016. (Accessed: 2026-04-10).
  54. 54.QUIC Interop Working Group. Chrome Image for the QUIC Interop Runner. https://github.com/quic-interop/chrome-quic-interop-runner, 2020. (Accessed: 2026-04-10).
  55. 55.quinn-rs contributors. Quinn: QUIC Implementation in Rust. https://github.com/quinn-rs/quinn, 2018. (Accessed: 2026-04-10).
  56. 56.Muthu Ramachandran. Testing software components using boundary value analysis. In 2003 Proceedings 29th Euromicro Conference, pages 94–98. IEEE, 2003.
  57. 57.Marten Seemann and Jana Iyengar. Automating QUIC Interoperability Testing. In Proceedings of the Workshop on the Evolution, Performance, and Interoperability of QUIC, EPIQ'20, pages 8–13, New York, NY, USA, 2020. ACM. Co-located with SIGCOMM 2020, Virtual Event, USA.
  58. 58.Muhammad Sholeh, Irmah Gisfas, Muhammad Anwar Fauzi, et al. Black Box testing with Boundary Value Analysis and Equivalence Partitioning Methods. In Journal of Physics: Conference Series, volume 1823, page 012029. IOP Publishing, 2021.
  59. 59.Rathin Singha, Rajdeep Mondal, Ryan Beckett, Siva Kesava Reddy Kakarla, Todd Millstein, and George Varghese. MESSI: Behavioral Testing of BGP Implementations. In 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24), pages 1009–1023, 2024.
  60. 60.Rathin Singha, Harry Qian, Srinath Saikrishnan, Tracy Zhao, Ryan Beckett, Siva Kesava Reddy Kakarla, and George Varghese. Extremal testing for network software using llms, 2025.
  61. 61.Robert Swiecki. Honggfuzz - Security oriented software fuzzer. https://github.com/google/honggfuzz/tree/master/examples/bind.
  62. 62.Igor Sysoev. NGINX HTTP Server. https://nginx.org, 2004. Source code: https://github.com/nginx/nginx (accessed 2026-04-10).
  63. 63.Willy Tarreau. HAProxy QUIC Implementation. https://github.com/haproxy/haproxy, 2022. QUIC support added in v2.6. Canonical source: https://git.haproxy.org/ (Accessed: 2026-04-10).
  64. 64.The Go Authors. golang.org/x/net: QUIC Package. https://pkg.go.dev/golang.org/x/net/internal/quic, 2022. Source code: https://github.com/golang/net (Accessed: 2026-04-10).
  65. 65.Tatsuhiro Tsujikawa. ngtcp2 QUIC Implementation. https://github.com/ngtcp2/ngtcp2, 2017. (Accessed: 2026-04-10).
  66. 66.Junjie Wang, Yuchao Huang, Chunyang Chen, Zhe Liu, Song Wang, and Qing Wang. Software testing with large language models: Survey, landscape, and vision. IEEE Transactions on Software Engineering, 2024. Also available as arXiv:2307.07221.
  67. 67.Michal Zalewski. American Fuzzy Lop (AFL). https://lcamtuf.coredump.cx/afl/, 2014. Accessed: 2026-04-17.
  68. 68.Shreyas Zare and Community. Technitium DNS server. https://technitium.com/dns/, 2026. Github: https://github.com/TechnitiumSoftware/DnsServer.
  69. 69.Zhiqiang Zhang, Tianyong Wu, and Jian Zhang. Boundary value analysis in automatic white-box test generation. In 2015 IEEE 26th International Symposium on Software Reliability Engineering (ISSRE), pages 239–249. IEEE, 2015.
  70. 70.Xiaogang Zhou, Tianyi Zhang, and David Lo. Large language model for vulnerability detection: Emerging results and future directions. In Proceedings of the 2024 ACM/IEEE 44th International Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER), 2024.
  71. 71.Sam Hocevar. zzuf: multi-purpose fuzzer. https://caca.zoy.org/wiki/zzuf.

Citation

MLA
Singha, R., et al. “CornerCase: Automated Extremal Testing of Protocol Implementations Using LLMs”. arXiv, 2026, https://doi.org/10.48550/arxiv.2606.29124.
APA
Singha, R., Qian, K., Saikrishnan, S., Zhao, T., Abbasloo, S., Beckett, R., Kakarla, S. K. R., Millstein, T., & Varghese, G. (2026). CornerCase: Automated Extremal Testing of Protocol Implementations using LLMs. arXiv. https://doi.org/10.48550/arxiv.2606.29124
Chicago
Singha, R., K. Qian, S. Saikrishnan, et al. 2026. “CornerCase: Automated Extremal Testing of Protocol Implementations Using LLMs”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2606.29124.
Harvard
Singha, R. et al. (2026) “CornerCase: Automated Extremal Testing of Protocol Implementations using LLMs”. arXiv. Available at: https://doi.org/10.48550/arxiv.2606.29124.
Vancouver
1. Singha R, Qian K, Saikrishnan S, Zhao T, Abbasloo S, Beckett R, Kakarla SKR, Millstein T, Varghese G (2026) CornerCase: Automated Extremal Testing of Protocol Implementations using LLMs. https://doi.org/10.48550/arxiv.2606.29124

BibTeX

@misc{https://doi.org/10.48550/arxiv.2606.29124,
  doi = {10.48550/ARXIV.2606.29124},
  url = {https://arxiv.org/abs/2606.29124},
  author = {Singha, Rathin and Qian, Kuan and Saikrishnan, Srinath and Zhao, Tracy and Abbasloo, Soheil and Beckett, Ryan and Kakarla, Siva Kesava Reddy and Millstein, Todd and Varghese, George},
  keywords = {Networking and Internet Architecture (cs.NI), FOS: Computer and information sciences},
  title = {CornerCase: Automated Extremal Testing of Protocol Implementations using LLMs},
  publisher = {arXiv},
  year = {2026},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/