SentinelBench: A Benchmark for Long-Running Monitoring Agents
Matheus Kunzler MaldanerAdam FourneyAmanda SwearnginHussein MozannarGagan BansalMaya MuradRafah HosnSaleema Amershi
Introduces SentinelBench, an open-source benchmark of 100 dynamic web tasks that evaluates how effectively autonomous agents monitor changing interfaces and balance reaction speed against resource consumption over extended periods.
Most modern artificial intelligence (AI) benchmarks assume that an agent operates in a reactive setting where the environment changes only in response to direct actions. However, many practical long-running tasks require sustained attention and patient monitoring rather than continuous tool execution. Forcing active agents into continuous polling loops can lead to excessive token costs and task failure due to long trajectory lengths. The article introduces SentinelBench, an open-source evaluation suite designed to assess how well AI agents monitor dynamic web environments over time and act promptly when external events occur.
To evaluate monitoring capabilities under controlled and reproducible conditions, the authors built 10 lightweight, synthetic web applications mirroring popular platforms such as email, calendars, and stock trading. SentinelBench includes 100 scripted task scenarios spanning 10-minute baseline windows, with options to extend task durations. The tasks cover passive monitoring, active interventions, and no-operation controls where trigger conditions never occur. The authors benchmarked three AI models (GPT-5.4 with low reasoning, GPT-4o, and Qwen 3.5:9B) paired with either a traditional fixed-interval sleep tool or a purpose-built wait_for tool that inspects incremental web page text diffs using an integrated language model.
Key findings show that frontier model capability and tooling choices critically determine success and cost efficiency. First, GPT-5.4 achieved the highest overall task completion rates (75% with the wait_for tool and 68% with sleep), substantially outperforming GPT-4o and Qwen 3.5:9B (both below 50%). Second, utilizing the wait_for tool reduced median API costs by 2 to 5 times across standard 10-minute tasks by drastically cutting unnecessary agent tool calls. Third, when tasks were stretched to 40-minute durations, the cost disparity grew to nearly 10 times (0.48 with wait_for), while the wait_for configuration achieved a 13-point higher completion rate (69% versus 56%). Finally, agents relying on sleep tools frequently failed long tasks by terminating early before conditions were met.
These results demonstrate that agent harness design is just as vital as underlying model intelligence for long-horizon monitoring. Using specialized waiting tools prevents resource waste and improves task reliability without introducing substantial reaction time delays. Organizations deploying long-running autonomous workflows should avoid continuous polling loops and adopt event-driven waiting primitives to contain costs and prevent premature task termination. Future development should focus on testing more complex, ephemeral event conditions, subjective task triggers, and compressed time simulations to support reinforcement learning.
- Paper: WebArena: A Realistic Web Environment for Building Autonomous Agents, Shuyan Zhou et al. (2023). WebArena establishes the foundational paradigm for evaluating autonomous agents via functional execution on simulated web platforms, which SentinelBench adapts to dynamic, time-evolving monitoring tasks.
- Paper: AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?, Ori Yoran et al. (2024). AssistantBench introduces benchmarking for realistic, time-consuming web navigation tasks across complex websites, laying essential groundwork for SentinelBench's focus on long-running web agents.
- Paper: WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models, Hongliang He et al. (2024). WebVoyager demonstrates end-to-end multimodal agent browsing across dynamic real-world web environments, providing key context for the browser-agent harnesses evaluated in SentinelBench.
- Paper: Mind2Web: Towards a Generalist Agent for the Web, Xiang Deng et al. (2023). Mind2Web outlines the core challenges of generalist web navigation and dynamic interface element interaction that underpin agent architectures tested in SentinelBench.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). This survey provides a comprehensive taxonomy of large language model agent architectures, planning, and evaluation methods that frame the continuous versus sustained-attention paradigms analyzed in SentinelBench.
No sufficiently relevant recommendations were found.
