SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations

Shuaiqi WangAadyaa MaddiZinan LinGiulia Fanti

article2026arXiv1 citations

Presents SynAE, an evaluation framework that systematically assesses the validity, fidelity, and diversity of synthetic execution traces across multi-turn tool calls and outputs to ensure synthetic datasets can reliably test tool-calling agents.

Listen

Organizations deploying artificial intelligence systems increasingly rely on automated agents that call external tools and execute multi-step workflows. However, testing these systems before deployment is challenging because real user interaction logs are often too small for comprehensive evaluation or contain sensitive personal data subject to privacy rules. To solve this, developers use synthetic datasets to augment or replace real user data. Despite this growing practice, teams currently lack systematic, quantitative methods to verify whether synthetic test data accurately reflects real-world operational conditions or preserves the validity of multi-step tool interactions.

The article introduces SynAE, a comprehensive evaluation framework designed to measure the quality of synthetic test datasets for multi-turn, tool-calling agents. Its primary objective is to evaluate how well synthetic benchmarks replicate and expand upon real data across three core dimensions: validity, fidelity, and diversity.

To establish credibility without relying on simple surface-level comparisons, the approach evaluates synthetic trajectories across four levels: conversational instructions, tool calls, final text outputs, and downstream agent performance. SynAE assesses validity by checking whether tool calls and final outputs correctly accomplish tasks, fidelity by measuring distributional and structural similarity against real data, and diversity through entropy-based metrics across dataset features. The authors evaluated SynAE on three established benchmarks (T1, BFCL, and ACP) covering conversational planning, function calling, and discrete action domains, while testing synthetic generation techniques including industry-standard tools (such as NVIDIA NeMo) and controlled modification schemes.

The analysis yielded several key findings. First, no single metric can capture synthetic benchmark quality. Interventions that improve data diversity, such as masking and re-generating instructions, frequently degrade data fidelity. Second, traditional text-similarity metrics can produce false confidence; for example, adding demonstrations to generation prompts improved surface vocabulary overlap while degrading deeper semantic recall and attribute alignment. Third, naive synthetic data generation methods often damage validity; directly replacing topic keywords in instructions improved diversity but caused task-completion validity to fall from 100% to 82% due to inconsistencies between prompts and subsequent tool calls. Finally, using higher-capacity model backends resolved these trade-offs, enabling simultaneous improvements in diversity and task validity without sacrificing fidelity.

These findings have direct operational and risk implications for artificial intelligence deployments. Evaluating tool-calling agents on unvalidated synthetic data creates hidden blind spots, risking deployment of faulty models that fail during execution or miscalculate user intent. Over-relying on superficial metrics or basic data manipulation increases compliance and safety risks while potentially distorting performance rankings among candidate models.

Organizations should adopt multi-axis evaluation frameworks to systematically audit synthetic benchmarks before using them in pre-deployment agent testing. Teams should use targeted diagnostic metrics to pinpoint whether a benchmark suffers from limited diversity, low fidelity, or invalid logic, and then apply suitable synthetic generation strategies rather than naive keyword substitutions. Where high fidelity and diversity are required, practitioners must budget for sufficiently capable generator models to avoid subtle execution failures.

The findings are well supported by systematic experiments across diverse domain benchmarks and validate strong alignment between automated evaluations and human judgment. However, readers should note that the evaluation is currently bounded by static, text-based multi-turn interactions. Confidence is highest for discrete tool-calling environments, while cautious interpretation is recommended when extending these conclusions to fully dynamic real-time environments or multi-agent collaborative workflows.

arXiv: 2605.22564wsqwsq/SynAE
Cover for SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations

Abstract

Today, tool-calling agents are commonly evaluated or tested on static datasets of execution traces, including input commands, agent responses, and associated tool calls. However, internal production datasets are often insufficient or unusable for testing; for example, they may contain sensitive or proprietary data, or they may be too sparse to support comprehensive testing (especially pre-deployment). In these settings, practitioners are increasingly replacing or augmenting real datasets with synthetic ones for evaluation purposes. A key challenge is quantifying the relation between these synthetic datasets and the real data. We introduce SynAE, an evaluation framework for assessing how well synthetic benchmarks for multi-turn, tool-calling agents replicate and augment the characteristics of real data trajectories. SynAE assesses the validity, fidelity, and diversity of synthetic data across four metric categories: (i) task instructions and intermediate responses, (ii) tool calls, (iii) final outputs, and (iv) downstream evaluation. We evaluate SynAE using recent agent benchmarks and test common synthetic data failure modes via realistic and controlled generation schemes. SynAE detects fine-grained variations in data validity, fidelity and diversity, and shows that no single metric is sufficient to fully characterize synthetic data quality, motivating a multi-axis evaluation of synthetic data for agent testing. A demo of SynAE is available at this https URL, with code at this https URL.

Table of Contents

  • 1 Introduction
  • 1.1 Related Work
  • 2 SynAE Framework
  • 2.1 Evaluation Metrics
  • 2.1.1 Validity Metrics
  • 2.1.2 Fidelity Metrics
  • 2.1.3 Diversity Metrics
  • 3 Experiments
  • 3.1 Experimental Results
  • 3.1.1 SynAE can capture fine-grained variations in fidelity, diversity, and validity of synthetic data.
  • 3.1.2 No single metric can fully characterize synthetic data performance.
  • 3.1.3 Case study: Practitioners can leverage SynAE to iteratively diagnose and improve synthetic data generation.
  • 4 Conclusions and Limitations
  • References
  • A Related Work
  • B Prompts for LLM-as-a-Judge on Validity
  • B.1 Prompts for T1
  • B.1.1 Tool Call Evaluation
  • B.1.2 Output Evaluation
  • B.2 Prompts for BFCL
  • B.2.1 Tool Call Evaluation
  • B.3 Prompts for ACP
  • B.3.1 Output Evaluation
  • B.4 Agreement Rate with Human Annotation
  • C Prompts for LLM-as-a-Judge in Downstream Evaluation
  • C.1 Prompts for T1
  • C.1.1 Tool Call Evaluation
  • C.1.2 Output Evaluation
  • C.2 Prompts for BFCL
  • C.2.1 Tool Call Evaluation
  • D Computational Cost of SynAE Across Datasets
  • E Prompts for LLM-Based Synthetic Data Generation Methods
  • E.1 Prompts for T1
  • E.1.1 Blank Filling
  • E.1.2 In-Context Generation
  • E.2 Prompts for BFCL
  • E.2.1 Blank Filling
  • E.2.2 In-Context Generation
  • E.3 Prompts for ACP
  • E.3.1 Blank Filling
  • E.3.2 In-Context Generation
  • F Detailed Evaluation Results
  • G Evaluation Results of Nvidia NeMo
  • H No single baseline metric can fully characterize synthetic data performance
  • I Attribute Distribution of the Skewed and Augmented Datasets

Knowls

  1. Knowl 1 — SynAE Framework for Synthetic Agent Benchmark Evaluation

    model/method

    SynAE is an evaluation framework designed to quantify how well synthetic benchmarks for multi-turn, tool-calling agents replicate and augment real agent execution traces.

    An agent execution trajectory DiD_i within a dataset D={Di}i=1m\mathcal{D} = \{D_i\}_{i=1}^m consists of three components:

    1. An instruction-response sequence Ri=(ri,1,ri,2,…,ri,ℓi)R_i = (r_{i,1}, r_{i,2}, \dots, r_{i,\ell_i}) containing ℓi\ell_i alternating turns of user instructions and agent responses.
    2. A reference tool-call sequence Fi=(fi,1(ϕi,1),fi,2(ϕi,2),…,fi,qi(ϕi,qi))F_i = \big(f_{i,1}(\phi_{i,1}), f_{i,2}(\phi_{i,2}), \dots, f_{i,q_i}(\phi_{i,q_i})\big), where each fi,j∈Ff_{i,j} \in \mathcal{F} represents an executable function from the toolset F\mathcal{F}, ϕi,j\phi_{i,j} represents its argument inputs, and qiq_i is the total count of tool calls.
    3. A textual final output OiO_i.

    SynAE takes as input a real baseline dataset D\mathcal{D}, a synthetic dataset D′\mathcal{D}', and optional metric configurations (such as user-specified domain attributes or evaluator LLM agents A1,…,AhA_1, \dots, A_h). It assesses the synthetic dataset along three core pillars:

    • Validity: whether synthetic tool calls and outputs successfully fulfill the instructions.
    • Fidelity: distributional similarity between real and synthetic data across instructions/responses, tool calls, final outputs, and downstream agent task performance.
    • Diversity: structural and semantic spread of the synthetic trajectories using reference-free metrics.
  2. Knowl 2 — Tool Call Fidelity Metrics in SynAE

    model/method

    SynAE measures the structural and planning fidelity of synthetic tool-calling sequences relative to real baseline sequences using three distributional distance metrics:

    1. Tool Usage Match (TUM): Compares the marginal distributions of tool usage across all tool calls: TUM=TV(ωf,ωf′)\text{TUM} = \text{TV}(\omega_f, \omega'_f) where ωf\omega_f and ωf′\omega'_f represent the categorical tool selection probability distributions in the real and synthetic datasets respectively, and TV\text{TV} denotes the Total Variation distance.

    2. Tool Call Number Match (TCNM): Evaluates whether the number of tool calls per trajectory matches the real data distribution: TCNM=W2(ωq,ωq′)\text{TCNM} = W_2(\omega_q, \omega'_q) where ωq\omega_q and ωq′\omega'_q are the empirical distributions over the tool-call count qiq_i per trajectory sample DiD_i, and W2W_2 denotes the Wasserstein-2 distance.

    3. kk-Step Tool Planning Match: Quantifies tool-transition and planning fidelity by computing the weighted Total Variation distance between the conditional distributions of the next tool given the prefix sequence of the previous k−1k-1 tool calls: $$\text{$k$-Step Planning} = \sum_{f_1 \dots f_{k-1} \in \mathcal{F}^{k-1}} n_{f_1 \dots f_{k-1}} \cdot \text{TV}\left(\omega_{f \mid f_1 \dots f_{k-1}}, \omega'{f \mid f_1 \dots f{k-1}}\right) where F\mathcal{F} is the toolset, nf1…fk−1n_{f_1 \dots f_{k-1}} is the count of kk-step prefixes matching f1…fk−1f_1 \dots f_{k-1} in the real dataset, and ωf∣…\omega_{f \mid \dots} and ωf∣…′\omega'_{f \mid \dots} denote the empirical conditional next-tool distributions in the real and synthetic data, respectively (k∈{2,3}k \in \{2, 3\} by default).

  3. Knowl 3 — Downstream Task Fidelity Metrics in SynAE

    model/method

    To evaluate the downstream utility of a synthetic benchmark, SynAE executes a panel of evaluation agents A1,…,AhA_1, \dots, A_h on both the real benchmark D\mathcal{D} and the synthetic benchmark D′\mathcal{D}', assessing two downstream fidelity properties:

    1. Task Difficulty Difference (TDD): Measures the average absolute gap in agent task-completion success rates between real and synthetic benchmarks: TDD=1h∑j=1h∣Acc(Aj,D)−Acc(Aj,D′)∣\text{TDD} = \frac{1}{h} \sum_{j=1}^h \left| \text{Acc}(A_j, \mathcal{D}) - \text{Acc}(A_j, \mathcal{D}') \right| where Acc(Aj,D)\text{Acc}(A_j, \mathcal{D}) denotes the accuracy of agent AjA_j on dataset D\mathcal{D}. An agent-generated trajectory D~i,Aj\tilde{D}_{i, A_j} is judged for functional task success against the reference trajectory DiD_i using an LLM-as-a-judge (or automated ground-truth checkers).

    2. Ranking Divergence (RD): Measures whether the relative ranking of agent capabilities is preserved between the real and synthetic benchmarks. It is computed as the Spearman rank correlation ρs\rho_s between the agent performance rankings obtained on D\mathcal{D} versus D′\mathcal{D}' based on tool-call correctness or final output accuracy.

    A lower TDD indicates that the synthetic benchmark closely matches the real dataset's intrinsic difficulty, while a higher RD indicates that model evaluations on synthetic data accurately mirror real-world relative agent comparisons.

  4. Knowl 4 — Text and Attribute Fidelity Metrics in SynAE

    model/method

    SynAE measures instruction, response, and output fidelity using both structural and semantic distribution-matching metrics:

    • Key Node Dependency (KND): Captures sequential dependencies across dialogue turns by computing cosine similarities between text embeddings of adjacent elements: (i) an instruction ri,2t−1r_{i, 2t-1} and its immediate response ri,2tr_{i, 2t}, and (ii) a response ri,2tr_{i, 2t} and the subsequent instruction ri,2t+1r_{i, 2t+1}. KND computes the distributional distance between the resulting cosine-similarity distributions of the real and synthetic datasets.
    • Attribute Match (AM): Evaluates alignment across structural, statistical, and semantic metadata attributes (e.g., number of instruction turns, token lengths, task domains, geographic locations). AM applies Wasserstein-2 distance for continuous/numerical attributes and Total Variation distance for categorical attributes.
    • kk-NN Precision and kk-NN Recall: Evaluates semantic representation coverage and quality using embedding space neighborhoods. kk-NN Precision is the fraction of synthetic samples whose embedding distance to their nearest real neighbor is smaller than the distance to their kk-th nearest synthetic neighbor. kk-NN Recall is the fraction of real samples whose distance to their nearest synthetic neighbor is smaller than the distance to their kk-th nearest real neighbor.
    • Fréchet Inception Distance (FID): Measures the Wasserstein-2 distance between Gaussian approximations fitted to the text embedding distributions of the real and synthetic datasets.
  5. Knowl 5 — Reference-Free Diversity Metrics in SynAE

    model/method

    SynAE provides two reference-free diversity metrics that quantify dataset spread without requiring real baseline data:

    1. Vendi Score (Vendi): Given a dataset of mm trajectories and a pairwise positive semidefinite similarity matrix K∈Rm×mK \in \mathbb{R}^{m \times m} normalized such that Ki,i=1K_{i,i} = 1, the Vendi Score is computed from the eigenvalues λ1,…,λm\lambda_1, \dots, \lambda_m of the matrix K/mK/m as: Vendi(K)=exp⁡(−∑i=1mλilog⁡λi)\text{Vendi}(K) = \exp\left( -\sum_{i=1}^m \lambda_i \log \lambda_i \right) where 0log⁡0≜00 \log 0 \triangleq 0. The score is upper-bounded by mm (achieved when all samples are mutually orthogonal).
    • For instruction-response sequences RR and final outputs OO, Ki,jK_{i,j} is the cosine similarity between text embeddings (e.g., text-embedding-3-small).
    • For tool-call sequences FiF_i and FjF_j with lengths qiq_i and qjq_j, similarity is defined using normalized Levenshtein edit distance: Ki,j=1−Levenshtein(Fi,Fj)max⁡(qi,qj)K_{i,j} = 1 - \frac{\text{Levenshtein}(F_i, F_j)}{\max(q_i, q_j)}
    1. Attribute Diversity (AD): Measures diversity across user-specified discrete attribute categories (e.g., task intent, planning domains, entity types). If there are CC possible attribute combinations and pip_i is the empirical fraction of trajectories belonging to combination ii, AD is the Shannon entropy: AD=−∑i=1Cpilog⁡pi\text{AD} = -\sum_{i=1}^C p_i \log p_i Upper-bounded by log⁡C\log C, achieved when trajectories are uniformly distributed across all attribute combinations.
  6. Knowl 6 — Validity Evaluation and LLM-as-a-Judge Calibration in SynAE

    model/method

    Validity Rate (VR) in SynAE measures the proportion of synthetic task trajectories where the reference tool-call sequences and textual outputs correctly fulfill the task instructions without hallucinated parameters, nonexistent tool names, or inconsistent answers.

    By default, SynAE assesses validity using an LLM-as-a-judge (Mistral-7B-Instruct) prompted to perform binary (yes/no) verification of whether the tool call sequence or final output completely accomplishes the conversation goal (partial correctness is penalized as invalid).

    When calibrated against human expert annotations on 100 randomly sampled synthetic trajectories from the T1 benchmark, the LLM-as-a-judge validity checker achieved an F1F_1 score of 0.86 and a Cohen's κ\kappa of 0.61, indicating substantial agreement with human judgment.

  7. Knowl 7 — Multi-Metric Trade-offs Under Controlled Synthetic Data Perturbations

    empirical result

    Experiments using controlled synthetic generation schemes on the T1, BFCL, and ACP benchmarks demonstrate that no single metric can fully evaluate synthetic benchmark quality due to trade-offs across validity, fidelity, and diversity:

    • Blank Filling: Randomly masking instruction tokens with probability p∈[0.1,0.9]p \in [0.1, 0.9] and refilling them via an LLM lowers instruction and output kk-NN Precision (from ~1.0 down to <0.25 on T1) and degrades tool-planning match, but monotonically increases diversity (instruction Vendi Score increases from 9.768 at p=0p=0 to 25.029 at p=0.9p=0.9).
    • Oversampling: Duplicating a subset of trajectories at rate r∈[0.1,1.0]r \in [0.1, 1.0] maintains high kk-NN Precision (~0.99) but collapses kk-NN Recall (from 0.923 to 0.004 on T1 outputs) and diversity (instruction Vendi Score drops from 8.966 to 1.000).
    • In-Context Demonstration Count (kk): Increasing few-shot demonstrations k∈{0,1,3,5}k \in \{0, 1, 3, 5\} improves superficial n-gram overlap and FID, but does not reliably improve structural fidelity metrics (such as kk-NN Recall, KND, or Attribute Match).
    • Ranking Divergence Decoupling: Across perturbations, downstream Ranking Divergence (RD) often remains high (e.g., ρs≥0.5\rho_s \ge 0.5 or 1.01.0) even when tool-planning fidelity severely degrades, because higher-capacity models consistently outperform smaller models despite distribution shifts.
  8. Knowl 8 — Evaluation of NVIDIA NeMo Synthetic Datasets Across LLM Backends and Temperatures

    data/table

    Synthetic agent trajectories generated for the T1 benchmark using NVIDIA NeMo Data Designer across different backbone models (Mistral-24B-Instruct, Nemotron-Nano-9B-v2, Llama3.1-8B-Instruct) and sampling temperatures T∈{0.1,0.3,0.5,0.7,0.9}T \in \{0.1, 0.3, 0.5, 0.7, 0.9\} demonstrate systematic shifts in quality metrics:

    Model Temp VR (Output) ↑\uparrow kk-NN Prec ↑\uparrow kk-NN Rec ↑\uparrow Vendi ↑\uparrow AD ↑\uparrow
    Mistral-24B 0.1 0.798 0.382 0.089 4.692 2.449
    Mistral-24B 0.5 0.898 0.222 0.338 5.472 2.899
    Mistral-24B 0.9 0.893 0.013 0.533 7.472 3.077
    Nemotron-9B 0.1 0.764 0.209 0.271 5.145 2.347
    Nemotron-9B 0.5 0.820 0.237 0.302 7.619 3.125
    Nemotron-9B 0.9 0.856 0.049 0.667 9.524 3.487
    Llama3.1-8B 0.1 0.809 0.120 0.662 7.480 3.184
    Llama3.1-8B 0.5 0.881 0.259 0.160 9.236 3.257
    Llama3.1-8B 0.9 0.859 0.100 0.427 9.397 3.481

    Across all model backends, increasing the generation temperature increases diversity (Vendi Score and Attribute Diversity) and semantic coverage (kk-NN Recall), but degrades semantic quality (kk-NN Precision). Nemotron and Llama backends generate broader category diversity than Mistral, while maintaining comparable fidelity scores.

  9. Knowl 9 — Iterative Benchmark Refinement with SynAE Diagnostic Feedback

    data/table

    SynAE enables structured diagnosis and iterative correction of flaws in synthetic agent benchmarks. In a controlled case study starting from a skewed T1 dataset where four attraction types (culture, sport, culinary, guide) were subsampled to 10% representation:

    Metric Skewed Real Relabeling NeMo (Llama) NeMo (Nemotron) NeMo (GPT-4o-mini)
    Validity 1.00 0.82 0.98 0.99 0.99
    Fidelity 1.00 0.95 0.71 0.79 0.94
    Diversity 0.48 0.65 0.67 0.61 0.70
    • Diagnosis: SynAE identifies low diversity (0.48 vs. unskewed baseline of 0.72) and pinpoints specific under-represented attribute slices.
    • Heuristic Relabeling: Naive keyword replacement increases diversity to 0.65, but drops Validity Rate to 0.82 because modified instructions conflict with multi-turn conversation context and tool parameters.
    • Model-Based Synthesis: NVIDIA NeMo with smaller models restores validity (≥0.98\ge 0.98) but sacrifices fidelity (0.71–0.790.71\text{--}0.79). Scaling the generation backend to GPT-4o-mini achieves near-optimal validity (0.99), high fidelity (0.94), and restores unskewed diversity (0.70).
  10. Knowl 10 — Scope and Limitations of SynAE

    limitation

    SynAE is designed specifically for evaluating static execution trajectories (instructions, responses, tool calls, and textual outputs) in multi-turn tool-calling benchmarks.

    It does not natively evaluate:

    1. Interactive, dynamic environments where agent actions modify an external, stateful world model (such as interactive web navigation or bash environments) beyond static textual outputs.
    2. Multi-agent collaborative, communication, or competitive dynamics involving coordination across multiple simultaneous actors.

Coverage note — Omitted exact text prompt strings for specific synthetic generators (Appendices E.1-E.3) and raw evaluation runtime scripts, as their key algorithmic mechanisms and prompt behaviors are fully described in the methods and empirical knowls.

References

  1. 1.Omar Alonso and Kenneth Church. Evaluating the evaluations: A perspective on benchmarks. In ACM SIGIR Forum, volume 58, pages 1–27. ACM New York, NY, USA, 2025.
  2. 2.Anthropic. Demystifying evals for ai agents. https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents, 2026.
  3. 3.CapitalOne. Synthetic data matters for machine learning innovation. https://www.capitalone.com/tech/machine-learning/synthetic-data-research/, 2022.
  4. 4.Amartya Chakraborty, Paresh Dashore, Nadia Bathaee, Anmol Jain, Anirban Das, Shi-Xiong Zhang, Sambit Sahu, Milind Naphade, and Genta Indra Winata. T1: A tool-oriented conversational dataset for multi-turn agentic planning. arXiv preprint arXiv:2505.16986, 2025.
  5. 5.Enkrypt AI. What are specialized task ai agents? benefits, features & use cases explained. Enkrypt AI Blog (Guest Post), March 2024.
  6. 6.Dan Friedman and Adji Bousso Dieng. The vendi score: A diversity evaluation metric for machine learning. arXiv preprint arXiv:2210.02410, 2022.
  7. 7.Alexander Gill, Abhilasha Ravichander, and Ana Marasović. What has been lost with synthetic evaluation? arXiv preprint arXiv:2505.22830, 2025.
  8. 8.Ben Hutchinson, Negar Rostamzadeh, Christina Greer, Katherine Heller, and Vinodkumar Prabhakaran. Evaluation gaps in machine learning practice. In Proceedings of the 2022 ACM conference on fairness, accountability, and transparency, pages 1859–1876, 2022.
  9. 9.Shadi Iskander, Sofia Tolmach, Ori Shapira, Nachshon Cohen, and Zohar Karnin. Quality matters: Evaluating synthetic data for tool-using llms. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 4958–4976, 2024.
  10. 10.Dongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong, Hao Li, Le Zhuo, Shilin Yan, Pheng-Ann Heng, and Hongsheng Li. T2i-r1: Reinforcing image generation with collaborative semantic-level and token-level cot. arXiv preprint arXiv:2505.00703, 2025.
  11. 11.Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe-bench: Can language models resolve real-world github issues? arXiv preprint arXiv:2310.06770, 2023.
  12. 12.Sayash Kapoor, Benedikt Stroebl, Peter Kirgis, Nitya Nadgir, Zachary S Siegel, Boyi Wei, Tianci Xue, Ziru Chen, Felix Chen, Saiteja Utpala, et al. Holistic agent leaderboard: The missing infrastructure for ai agent evaluation. arXiv preprint arXiv:2510.11977, 2025.
  13. 13.Harsha Kokel, Michael Katz, Kavitha Srinivas, and Shirin Sohrabi. Acpbench: Reasoning about action, change, and planning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 26559–26568, 2025.
  14. 14.Renhao Li, Jianhong Tu, Yang Su, Yantao Liu, Fei Huang, Hamid Alinejad-Rokny, Derek F Wong, Junyang Lin, and Min Yang. Toolrm: Towards agentic tool-use reward modeling. arXiv preprint arXiv:2510.26167, 2025.
  15. 15.Jiachang Liu, Dinghan Shen, Yizhe Zhang, William B Dolan, Lawrence Carin, and Weizhu Chen. What makes good in-context examples for gpt-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd workshop on knowledge extraction and integration for deep learning architectures, pages 100–114, 2022.
  16. 16.Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al. Agentbench: Evaluating llms as agents. arXiv preprint arXiv:2308.03688, 2023.
  17. 17.Xing Han Lù, Amirhossein Kazemnejad, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J Pal, and Siva Reddy. Agentrewardbench: Evaluating automatic evaluations of web agent trajectories. arXiv preprint arXiv:2504.08942, 2025.
  18. 18.Gaurav Maheshwari, Dmitry Ivanov, and Kevin El Haddad. Efficacy of synthetic data as a benchmark. arXiv preprint arXiv:2409.11968, 2024.
  19. 19.Michael Majurski and Cynthia Matuszek. Grounding synthetic data evaluations of language models in unsupervised document corpora. arXiv preprint arXiv:2505.08905, 2025.
  20. 20.Amanda McGrath and Amanda Downie. What are vertical ai agents? IBM Think, n.d.
  21. 21.Sohum Mehta and Saaketh Bhojanam. Prompt genotyping: Quantifying the evaluation gap between synthetic benchmarks and real llm performance. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling.
  22. 22.Mahmoud Mohammadi, Yipeng Li, Jane Lo, and Wendy Yip. Evaluation and benchmarking of llm agents: A survey. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2, pages 6129–6139, 2025.
  23. 23.NVIDIA. NVIDIA NeMo. https://www.nvidia.com/en-us/ai-data-science/products/nemo/.
  24. 24.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730–27744, 2022.
  25. 25.Jiayi Pan, Yichi Zhang, Nicholas Tomlin, Yifei Zhou, Sergey Levine, and Alane Suhr. Autonomous evaluation and refinement of digital agents. arXiv preprint arXiv:2404.06474, 2024.
  26. 26.Melissa Z Pan, Negar Arabzadeh, Riccardo Cogo, Yuxuan Zhu, Alexander Xiong, Lakshya A Agrawal, Huanzhi Mao, Emma Shen, Sid Pallerla, Liana Patel, et al. Measuring agents in production. arXiv preprint arXiv:2512.04123, 2025.
  27. 27.Shishir G Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E Gonzalez. The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models. In Forty-second International Conference on Machine Learning.
  28. 28.Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, et al. Toolllm: Facilitating large language models to master 16000+ real-world apis. arXiv preprint arXiv:2307.16789, 2023.
  29. 29.Yijia Shao, Yucheng Jiang, Theodore Kanell, Peter Xu, Omar Khattab, and Monica Lam. Assisting in writing wikipedia-like articles from scratch with large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 6252–6278, 2024.
  30. 30.Yongliang Shen, Kaitao Song, Xu Tan, Wenqi Zhang, Kan Ren, Siyu Yuan, Weiming Lu, Dongsheng Li, and Yueting Zhuang. Taskbench: Benchmarking large language models for task automation. Advances in Neural Information Processing Systems, 37:4540–4574, 2024.
  31. 31.Alex Tamkin, Miles McCain, Kunal Handa, Esin Durmus, Liane Lovitt, Ankur Rathi, Saffron Huang, Alfred Mountfield, Jerry Hong, Stuart Ritchie, et al. Clio: Privacy-preserving insights into real-world ai use. arXiv preprint arXiv:2412.13678, 2024.
  32. 32.Qiaoyu Tang, Ziliang Deng, Hongyu Lin, Xianpei Han, Qiao Liang, Boxi Cao, and Le Sun. Toolalpaca: Generalized tool learning for language models with 3000 simulated cases. arXiv preprint arXiv:2306.05301, 2023.
  33. 33.B. Wang, W. Chen, H. Pei, C. Xie, M. Kang, C. Zhang, C. Xu, Z. Xiong, R. Dutta, R. Schaeffer, et al. DecodingTrust: A comprehensive assessment of trustworthiness in GPT models. In Conference on Neural Information Processing Systems, 2023.
  34. 34.Shuaiqi Wang, Vikas Raunak, Arturs Backurs, Victor Reis, Pei Zhou, Sihao Chen, Longqi Yang, Zinan Lin, Sergey Yekhanin, and Giulia Fanti. Struct-bench: A benchmark for differentially private structured text generation. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  35. 35.Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Ben Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, et al. Livebench: A challenging, contamination-limited llm benchmark. arXiv preprint arXiv:2406.19314, 2025.
  36. 36.Shengguang Wu, Keming Lu, Benfeng Xu, Junyang Lin, Qi Su, and Chang Zhou. Self-evolved diverse data sampling for efficient instruction tuning. arXiv preprint arXiv:2311.08182, 2023.
  37. 37.Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh J Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Information Processing Systems, 37:52040–52094, 2024.
  38. 38.Lang Xiong, Nishant Bhargava, Jianhang Hong, Jeremy Chang, Haihao Liu, Vasu Sharma, and Kevin Zhu. Probe-rewrite-evaluate: A workflow for reliable benchmarks and quantifying evaluation awareness. arXiv preprint arXiv:2509.00591, 2025.
  39. 39.Tianci Xue, Weijian Qi, Tianneng Shi, Chan Hee Song, Boyu Gou, Dawn Song, Huan Sun, and Yu Su. An illusion of progress? assessing the current state of web agents. arXiv preprint arXiv:2504.01382, 2025.
  40. 40.John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. Swe-agent: Agent-computer interfaces enable automated software engineering. Advances in Neural Information Processing Systems, 37:50528–50652, 2024.
  41. 41.John Yang, Akshara Prabhakar, Karthik Narasimhan, and Shunyu Yao. Intercode: Standardizing and benchmarking interactive coding with execution feedback. Advances in Neural Information Processing Systems, 36:23826–23854, 2023.
  42. 42.Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan. τ -bench: A benchmark for tool-agent-user interaction in real-world domains. arXiv preprint arXiv:2406.12045, 2024.
  43. 43.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. In The eleventh international conference on learning representations, 2022.
  44. 44.Asaf Yehudai, Lilach Eden, Alan Li, Guy Uziel, Yilun Zhao, Roy Bar-Haim, Arman Cohan, and Michal Shmueli-Scheuer. Survey on evaluation of llm-based agents. arXiv preprint arXiv:2503.16416, 2025.
  45. 45.Zeyu Zhang, Guohao Li, Zhenchang Xing, Alexandros Apostolopoulos, Yu Lin Lee, and Liang Zheng. Gecko: A simulation environment to ground agent tool calls with stateful feedback for refinement.
  46. 46.Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al. Webarena: A realistic web environment for building autonomous agents. arXiv preprint arXiv:2307.13854, 2023.
  47. 47.Kaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong, Diyi Yang, and Xing Xie. Dyval: Dynamic evaluation of large language models for reasoning tasks. arXiv preprint arXiv:2309.17167, 2023.
  48. 48.Yuxuan Zhu, Tengjun Jin, Yada Pruksachatkun, Andy Zhang, Shu Liu, Sasha Cui, Sayash Kapoor, Shayne Longpre, Kevin Meng, Rebecca Weiss, et al. Establishing best practices for building rigorous agentic benchmarks. arXiv preprint arXiv:2507.02825, 2025.
  49. 49.Kaijian Zou, Muhammad Khalifa, and Lu Wang. On many-shot in-context learning for long-context evaluation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 25605–25639, 2025.

Citation

MLA
Wang, S., et al. “SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations”. arXiv, 2026, https://doi.org/10.48550/arxiv.2605.22564.
APA
Wang, S., Maddi, A., Lin, Z., & Fanti, G. (2026). SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations. arXiv. https://doi.org/10.48550/arxiv.2605.22564
Chicago
Wang, S., A. Maddi, Z. Lin, and G. Fanti. 2026. “SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2605.22564.
Harvard
Wang, S. et al. (2026) “SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations”. arXiv. Available at: https://doi.org/10.48550/arxiv.2605.22564.
Vancouver
1. Wang S, Maddi A, Lin Z, Fanti G (2026) SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations. https://doi.org/10.48550/arxiv.2605.22564

BibTeX

@misc{https://doi.org/10.48550/arxiv.2605.22564,
  doi = {10.48550/ARXIV.2605.22564},
  url = {https://arxiv.org/abs/2605.22564},
  author = {Wang, Shuaiqi and Maddi, Aadyaa and Lin, Zinan and Fanti, Giulia},
  keywords = {Computation and Language (cs.CL), Machine Learning (cs.LG), Software Engineering (cs.SE), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations},
  publisher = {arXiv},
  year = {2026},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/