SentinelBench: A Benchmark for Long-Running Monitoring Agents

Matheus Kunzler MaldanerAdam FourneyAmanda SwearnginHussein MozannarGagan BansalMaya MuradRafah HosnSaleema Amershi

article2026arXiv4 citations

Introduces SentinelBench, an open-source benchmark of 100 dynamic web tasks that evaluates how effectively autonomous agents monitor changing interfaces and balance reaction speed against resource consumption over extended periods.

Listen

Most modern artificial intelligence (AI) benchmarks assume that an agent operates in a reactive setting where the environment changes only in response to direct actions. However, many practical long-running tasks require sustained attention and patient monitoring rather than continuous tool execution. Forcing active agents into continuous polling loops can lead to excessive token costs and task failure due to long trajectory lengths. The article introduces SentinelBench, an open-source evaluation suite designed to assess how well AI agents monitor dynamic web environments over time and act promptly when external events occur.

To evaluate monitoring capabilities under controlled and reproducible conditions, the authors built 10 lightweight, synthetic web applications mirroring popular platforms such as email, calendars, and stock trading. SentinelBench includes 100 scripted task scenarios spanning 10-minute baseline windows, with options to extend task durations. The tasks cover passive monitoring, active interventions, and no-operation controls where trigger conditions never occur. The authors benchmarked three AI models (GPT-5.4 with low reasoning, GPT-4o, and Qwen 3.5:9B) paired with either a traditional fixed-interval sleep tool or a purpose-built wait_for tool that inspects incremental web page text diffs using an integrated language model.

Key findings show that frontier model capability and tooling choices critically determine success and cost efficiency. First, GPT-5.4 achieved the highest overall task completion rates (75% with the wait_for tool and 68% with sleep), substantially outperforming GPT-4o and Qwen 3.5:9B (both below 50%). Second, utilizing the wait_for tool reduced median API costs by 2 to 5 times across standard 10-minute tasks by drastically cutting unnecessary agent tool calls. Third, when tasks were stretched to 40-minute durations, the cost disparity grew to nearly 10 times (4.65medianpertaskwithsleepversus4.65 median per task with sleep versus 0.48 with wait_for), while the wait_for configuration achieved a 13-point higher completion rate (69% versus 56%). Finally, agents relying on sleep tools frequently failed long tasks by terminating early before conditions were met.

These results demonstrate that agent harness design is just as vital as underlying model intelligence for long-horizon monitoring. Using specialized waiting tools prevents resource waste and improves task reliability without introducing substantial reaction time delays. Organizations deploying long-running autonomous workflows should avoid continuous polling loops and adopt event-driven waiting primitives to contain costs and prevent premature task termination. Future development should focus on testing more complex, ephemeral event conditions, subjective task triggers, and compressed time simulations to support reinforcement learning.

No sufficiently relevant recommendations were found.

Cover for SentinelBench: A Benchmark for Long-Running Monitoring Agents

Abstract

AI agents are increasingly asked to carry out work that spans minutes, hours, or longer. Yet the default model of agent behavior is continuous action: issuing tool calls, refreshing pages, searching for alternatives, or otherwise trying to force progress. This is the wrong approach for many long-running tasks, which are better served by a strategy of sustained attention. Instead, agents should monitor an environment, notice when an external event makes progress possible, then respond promptly without wasting resources while waiting. To measure progress on this class of tasks, we introduce SentinelBench, an open-source benchmark for time-evolving monitoring tasks.

SentinelBench contains 100 tasks across 10 synthetic web environments, including email, calendars, finance, professional networking, and entertainment. Each environment exposes a live web interface and replays a scripted sequence of events, requiring agents to navigate and reason about web pages whose state shifts underfoot. SentinelBench measures task completion, reaction time, and resource use, exposing the tradeoff between responsiveness and cost. We report results across three models and two browser-agent harnesses, establishing performance baselines for future comparison and demonstrating how agent design choices can dramatically impact key metrics. Together, these results show that SentinelBench distinguishes meaningful differences in agent behavior.

Citation

MLA
Maldaner, M. K., et al. “SentinelBench: A Benchmark for Long-Running Monitoring Agents”. arXiv, 2026, http://arxiv.org/abs/2606.05342v2.
APA
Maldaner, M. K., Fourney, A., Swearngin, A., Mozannar, H., Bansal, G., Murad, M., Hosn, R., & Amershi, S. (2026). SentinelBench: A Benchmark for Long-Running Monitoring Agents. arXiv. http://arxiv.org/abs/2606.05342v2
Chicago
Maldaner, M. K., A. Fourney, A. Swearngin, et al. 2026. “SentinelBench: A Benchmark for Long-Running Monitoring Agents”. arXiv. http://arxiv.org/abs/2606.05342v2.
Harvard
Maldaner, M.K. et al. (2026) “SentinelBench: A Benchmark for Long-Running Monitoring Agents”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2606.05342v2.
Vancouver
1. Maldaner MK, Fourney A, Swearngin A, Mozannar H, Bansal G, Murad M, Hosn R, Amershi S (2026) SentinelBench: A Benchmark for Long-Running Monitoring Agents. arXiv

BibTeX

@article{maldaner2026sentinelbench,
  title = {SentinelBench: A Benchmark for Long-Running Monitoring Agents},
  author = {Maldaner, Matheus Kunzler and Fourney, Adam and Swearngin, Amanda and Mozannar, Hussein and Bansal, Gagan and Murad, Maya and Hosn, Rafah and Amershi, Saleema},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2606.05342v2},
  eprint = {2606.05342}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/