SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations
Shuaiqi WangAadyaa MaddiZinan LinGiulia Fanti
Presents SynAE, an evaluation framework that systematically assesses the validity, fidelity, and diversity of synthetic execution traces across multi-turn tool calls and outputs to ensure synthetic datasets can reliably test tool-calling agents.
Organizations deploying artificial intelligence systems increasingly rely on automated agents that call external tools and execute multi-step workflows. However, testing these systems before deployment is challenging because real user interaction logs are often too small for comprehensive evaluation or contain sensitive personal data subject to privacy rules. To solve this, developers use synthetic datasets to augment or replace real user data. Despite this growing practice, teams currently lack systematic, quantitative methods to verify whether synthetic test data accurately reflects real-world operational conditions or preserves the validity of multi-step tool interactions.
The article introduces SynAE, a comprehensive evaluation framework designed to measure the quality of synthetic test datasets for multi-turn, tool-calling agents. Its primary objective is to evaluate how well synthetic benchmarks replicate and expand upon real data across three core dimensions: validity, fidelity, and diversity.
To establish credibility without relying on simple surface-level comparisons, the approach evaluates synthetic trajectories across four levels: conversational instructions, tool calls, final text outputs, and downstream agent performance. SynAE assesses validity by checking whether tool calls and final outputs correctly accomplish tasks, fidelity by measuring distributional and structural similarity against real data, and diversity through entropy-based metrics across dataset features. The authors evaluated SynAE on three established benchmarks (T1, BFCL, and ACP) covering conversational planning, function calling, and discrete action domains, while testing synthetic generation techniques including industry-standard tools (such as NVIDIA NeMo) and controlled modification schemes.
The analysis yielded several key findings. First, no single metric can capture synthetic benchmark quality. Interventions that improve data diversity, such as masking and re-generating instructions, frequently degrade data fidelity. Second, traditional text-similarity metrics can produce false confidence; for example, adding demonstrations to generation prompts improved surface vocabulary overlap while degrading deeper semantic recall and attribute alignment. Third, naive synthetic data generation methods often damage validity; directly replacing topic keywords in instructions improved diversity but caused task-completion validity to fall from 100% to 82% due to inconsistencies between prompts and subsequent tool calls. Finally, using higher-capacity model backends resolved these trade-offs, enabling simultaneous improvements in diversity and task validity without sacrificing fidelity.
These findings have direct operational and risk implications for artificial intelligence deployments. Evaluating tool-calling agents on unvalidated synthetic data creates hidden blind spots, risking deployment of faulty models that fail during execution or miscalculate user intent. Over-relying on superficial metrics or basic data manipulation increases compliance and safety risks while potentially distorting performance rankings among candidate models.
Organizations should adopt multi-axis evaluation frameworks to systematically audit synthetic benchmarks before using them in pre-deployment agent testing. Teams should use targeted diagnostic metrics to pinpoint whether a benchmark suffers from limited diversity, low fidelity, or invalid logic, and then apply suitable synthetic generation strategies rather than naive keyword substitutions. Where high fidelity and diversity are required, practitioners must budget for sufficiently capable generator models to avoid subtle execution failures.
The findings are well supported by systematic experiments across diverse domain benchmarks and validate strong alignment between automated evaluations and human judgment. However, readers should note that the evaluation is currently bounded by static, text-based multi-turn interactions. Confidence is highest for discrete tool-calling environments, while cautious interpretation is recommended when extending these conclusions to fully dynamic real-time environments or multi-agent collaborative workflows.
- Paper: ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs, Yujia Qin et al. (2023). Introduces ToolBench and automated multi-turn synthetic trajectory generation for tool-calling agents, establishing the primary data paradigms evaluated by SynAE.
- Paper: API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs, Minghao Li et al. (2023). Provides foundational benchmark protocols and dialogue-action trajectory evaluation standards for tool-augmented language models.
- Paper: Toolformer: Language Models Can Teach Themselves to Use Tools, Timo Schick et al. (2023). Establishes core methodologies for synthesizing self-supervised tool-use calls and execution demonstrations in language models.
- Paper: Agent-as-a-Judge: Evaluate Agents with Agents, Mingchen Zhuge et al. (2025). Pioneers multi-step trajectory and intermediate execution state evaluation for autonomous agents using automated judges.
- Paper: Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic Environments, Hongjin Su et al. (2025). Presents techniques for synthesizing multi-step agent trajectories and backwards task construction in realistic digital environments.
- Paper: Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents, Sourabrata Mukherjee et al. (2026). Examines intermediate tool-calling action trace fidelity and policy retention across multilingual agent evaluation benchmarks.
- Paper: LLMs Get Lost in Evolving User Intent, Jihoon Tack et al. (2026). Applies synthetic multi-turn trajectory generation to benchmark tool-using and reasoning agents under dynamically changing user specifications.
- Paper: Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking, Jiwan Chung et al. (2026). Extends trajectory-level quality assessment by introducing semantic state tracking to diagnose intermediate step failures in web agents.
- Paper: AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation, Priyam Sahoo et al. (2026). Analyzes process quality and structural alignment across agent execution trajectories to identify flawed workflows behind nominal benchmark passes.
