Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs
Ananya SinghaHarshita SahijwaniWalt WilliamsEmmanuel Aboah BoatengNick HausmanMiguel Di LucaKeegan ChoudhuryChaya BinetVu-Anh LeTianwei Chen
Presents a scalable, execution-verified data generation pipeline and benchmark dataset for automated Excel formula repair, providing standardized baselines for evaluating how well large language models fix semantic spreadsheet errors.
Spreadsheets are among the most widely used computing platforms globally, yet non-technical users frequently encounter complex formula errors. Most automated repair tools have historically targeted isolated syntax issues, leaving semantic and runtime errors—which make up the majority of real-world spreadsheet problems—unaddressed. Advancing automated repair systems has been severely constrained by the lack of high-quality benchmark datasets containing both faulty formulas and their broader spreadsheet contexts.
The article demonstrates an automated, cost-effective pipeline to generate a benchmark dataset for repairing Excel runtime errors and evaluates how effectively modern language models can resolve these errors when supplied with spreadsheet context.
The researchers collected and manually validated 59 real-world seed cases from online Excel user forums across five standard runtime error categories: #DIV/0!, #N/A, #NAME?, #REF!, and #VALUE!. Using these seeds, they implemented a bootstrap data generation pipeline that used few-shot prompting with GPT-4o to synthetically expand the dataset. Generated cases were subjected to dual-layer quality control comprising automated formula execution checks and an automated evaluation model that assessed semantic intent and difficulty. The resulting dataset, named FoRepBench, contains 618 validated benchmark examples. The authors then tested a context-aware formula repair baseline across proprietary and open-source language models, including GPT-4.1, GPT-4o, Phi-3, and Mistral.
The evaluation revealed several critical findings. First, synthetic dataset creation is highly economical; generating FoRepBench required an average of approximately 2.09 model calls per accepted sample, translating to an estimated generation cost of only $0.026 per sample. Second, the automated generation pipeline produced greater functional diversity than the original seed data, introducing common functions such as AVERAGE and CONCATENATE. Third, modern language models solved synthetic repairs with high accuracy, with GPT-4.1 achieving an 80% exact execution match and GPT-4o reaching 73%. However, repair performance dropped steeply when models faced real-world seed data, with GPT-4.1 dropping to 41% accuracy and GPT-4o to 35%. Open-source models lagged further behind, achieving execution match rates between 19% and 24% on real-world examples.
These results indicate that automated synthetic pipelines can efficiently scale benchmark datasets, but synthetic examples skew significantly toward simpler, localized formula corrections. Real-world user errors frequently involve multi-step logic rewrites and deeply nested functions that current generation techniques underrepresent. Consequently, organizations relying solely on synthetic benchmarks risk overestimating how well automated assistants will perform for actual end users in operational settings.
To bridge the gap between synthetic data and real-world complexity, the article recommends incorporating human-in-the-loop review or iterative feedback agents into synthetic generation pipelines. Technical teams developing spreadsheet assistants should also implement structure-aware context retrieval to feed relevant table areas into repair models, rather than relying on raw formula text alone.
Readers should interpret the synthetic performance metrics with caution. While execution correctness is guaranteed via programmatic checks, the automated evaluation filter showed only moderate agreement with human raters on the contextual realism of table data (Cohen's Kappa of 0.42). Additionally, the benchmark remains limited to single-sheet, English-language spreadsheets, meaning performance on multi-sheet workbooks and collaborative enterprise environments requires further evaluation.
- Paper: Repair Is Nearly Generation: Multilingual Program Repair with LLMs, Harshit Joshi et al. (2023). It establishes the foundational paradigm of using LLMs for multi-language program repair—specifically highlighting Excel formula repair—which directly informs the target work's repair framing.
- Paper: Teaching Large Language Models to Self-Debug, Xinyun Chen et al. (2023). It introduces iterative prompt-driven debugging and execution-feedback repair loops that underpin automated code and formula repair methodologies.
- Paper: Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation, Jiawei Liu et al. (2023). It pioneers rigorous execution-based validation and test synthesis to evaluate LLM code correctness, foundational to the execution-based checks used in the source benchmark.
- Paper: Fault-Aware Neural Code Rankers, Jeevana Priya Inala et al. (2022). It demonstrates how execution errors and failure modes can guide the ranking and repair of generated programs without manual rule engineering.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). It provides critical insights into the capabilities and failure modes of LLM-as-a-judge evaluators, which the source relies upon to ensure synthetic data fidelity.
- Paper: Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?, Wang Bill Zhu et al. (2026). It advances beyond general repair benchmarks to distinguish precise, minimal fault localization from complete code regeneration in LLM debugging.
- Paper: LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code, Naman Jain et al. (2025). It extends execution-based code repair and evaluation by creating dynamic, contamination-free benchmarks across practical programming scenarios.
- Paper: TableBench: A Comprehensive and Complex Benchmark for Table Question Answering, Xianjie Wu et al. (2025). It broadens spreadsheet and tabular evaluation from formula error repair to comprehensive multi-step analytical reasoning and question answering over tables.
- Paper: SWE-Exp: Experience-Driven Software Issue Resolution, Silin Chen et al. (2025). It explores experience-driven automated issue resolution to build repair agents that accumulate reusable diagnostic strategies over repeated repair tasks.
