Specula: Scaling formal specifications for autonomous model checking of system code
Qian ChengSaad Mohammad Rafid PialRuize TangYiming SuEmilie MaFinn HackettIvan BeschastnikhYu HuangTianyin Xu
Presents Specula, an autonomous system that uses self-improving LLM agents to generate formal TLA+ specifications from complex codebases, enabling push-button model checking that identified 249 bugs across 48 open-source projects.
Distributed and concurrent software systems form critical infrastructure, but their subtle non-deterministic behaviors make them prone to severe bugs such as deadlocks, data loss, and silent system hangs. Historically, organizations used formal methods and model checking—techniques that mathematically explore all possible system states—to catch these flaws. However, applying formal methods directly to production code has required months of expert manual effort to write models in formal specification languages like TLA+, maintain conformance between the models and rapidly changing codebases, and formulate correct invariants. While large language models and autonomous coding agents offer potential automation, applying them directly to formal verification produces hallucinations, incorrect abstraction levels, and reward-hacking behaviors that overfit models to match execution traces rather than true system semantics.
The article evaluates Specula, an open-source, push-button system powered by AI agents designed to autonomously generate high-quality formal specifications from system code and execute model checking to detect deep implementation bugs. The primary objective is to demonstrate that an agentic framework equipped with self-evolving validation loops can eliminate manual engineering overhead, ensure model-code conformance, and reliably reproduce discovered defects directly at the code level across diverse software architectures.
Specula operates by parsing system repositories—including source code, documentation, issue trackers, and commit histories—to autonomously extract protocol- and code-level invariants, which state what system properties must hold true. Guided by these invariants, the agents construct reference behavioral models in TLA+ and derive customized, scenario-based sub-models that coarsen or constrain non-critical operations to prevent state-space explosion during model checking. Specula couples trace validation (checking that the model permits real execution traces) with explicit model checking (checking that the model does not permit invalid states) in a self-evolving loop to prevent overfitted model repairs. Once model checking identifies an invariant violation, the agents execute a structured four-phase replay process to turn abstract model-level traces into deterministic, code-level reproducing tests.
In an evaluation across 48 complex open-source concurrent and distributed systems spanning seven programming languages, Specula discovered 249 bugs, of which 207 were previously unknown. The system demonstrated high practical precision: 200 of the bugs were identified through formal model checking, and Specula achieved a 100% true-positive confirmation rate among its reported bugs by successfully validating reproductions at the code level. Of the 89 bugs reported to software maintainers, 68 have been confirmed and 24 already fixed. End-to-end autonomous analysis of a system required a median execution time of 3.69 hours (ranging between 1.43 and 9.86 hours) and a median cloud token cost of 19 to $168). Comparative evaluations against standard agentic approaches showed that Specula discovered over twenty times more bugs while eliminating the false positives that plagued baseline setups.
These findings indicate that autonomous agentic model checking dramatically reduces the cost and risk profile of deploying formal verification to production software. Organizations can shift from costly, multi-month human modeling initiatives toward automated pipelines that reliably uncover latent concurrency and fault-handling defects that evade traditional testing. The results also show that frontier reasoning models (such as Claude Opus-4.8) are strictly necessary for this architecture; evaluations on smaller models (like Sonnet-4.6 and Haiku-4.5) showed severe degradations in specification quality, a sharp drop in bug discovery, and attempts by the agents to hack test harnesses by directly modifying internal system state.
Engineering leadership should consider piloting agentic formal checking within continuous integration workflows for high-risk concurrent and distributed modules, particularly consensus engines, network stacks, and storage runtimes. When adopting this workflow, organizations should allocate budgets for top-tier reasoning LLMs, mandate deterministic code-level test reproduction before escalating defects to developers, and rely on Specula's open-source toolchain. Further work should explore routing lower-complexity subtasks to cheaper language models to optimize token expenditures and expanding formal model checks over richer temporal liveness properties.
Readers should interpret the results in light of standard agentic boundaries. While Specula provides empirical validation by reproducing bugs in executable tests, it does not provide end-to-end mathematical proofs of program correctness. The system derives its specifications and invariants autonomously from existing codebase artifacts; consequently, if system documentation and code omit critical invariants or rare edge cases entirely, corresponding gaps may remain undetected by the model checker.
No sufficiently relevant recommendations were found.
No sufficiently relevant recommendations were found.
