Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs

Ananya SinghaHarshita SahijwaniWalt WilliamsEmmanuel Aboah BoatengNick HausmanMiguel Di LucaKeegan ChoudhuryChaya BinetVu-Anh LeTianwei Chen

article2025arXiv0 citations

Presents a scalable, execution-verified data generation pipeline and benchmark dataset for automated Excel formula repair, providing standardized baselines for evaluating how well large language models fix semantic spreadsheet errors.

Listen

Spreadsheets are among the most widely used computing platforms globally, yet non-technical users frequently encounter complex formula errors. Most automated repair tools have historically targeted isolated syntax issues, leaving semantic and runtime errors—which make up the majority of real-world spreadsheet problems—unaddressed. Advancing automated repair systems has been severely constrained by the lack of high-quality benchmark datasets containing both faulty formulas and their broader spreadsheet contexts.

The article demonstrates an automated, cost-effective pipeline to generate a benchmark dataset for repairing Excel runtime errors and evaluates how effectively modern language models can resolve these errors when supplied with spreadsheet context.

The researchers collected and manually validated 59 real-world seed cases from online Excel user forums across five standard runtime error categories: #DIV/0!, #N/A, #NAME?, #REF!, and #VALUE!. Using these seeds, they implemented a bootstrap data generation pipeline that used few-shot prompting with GPT-4o to synthetically expand the dataset. Generated cases were subjected to dual-layer quality control comprising automated formula execution checks and an automated evaluation model that assessed semantic intent and difficulty. The resulting dataset, named FoRepBench, contains 618 validated benchmark examples. The authors then tested a context-aware formula repair baseline across proprietary and open-source language models, including GPT-4.1, GPT-4o, Phi-3, and Mistral.

The evaluation revealed several critical findings. First, synthetic dataset creation is highly economical; generating FoRepBench required an average of approximately 2.09 model calls per accepted sample, translating to an estimated generation cost of only $0.026 per sample. Second, the automated generation pipeline produced greater functional diversity than the original seed data, introducing common functions such as AVERAGE and CONCATENATE. Third, modern language models solved synthetic repairs with high accuracy, with GPT-4.1 achieving an 80% exact execution match and GPT-4o reaching 73%. However, repair performance dropped steeply when models faced real-world seed data, with GPT-4.1 dropping to 41% accuracy and GPT-4o to 35%. Open-source models lagged further behind, achieving execution match rates between 19% and 24% on real-world examples.

These results indicate that automated synthetic pipelines can efficiently scale benchmark datasets, but synthetic examples skew significantly toward simpler, localized formula corrections. Real-world user errors frequently involve multi-step logic rewrites and deeply nested functions that current generation techniques underrepresent. Consequently, organizations relying solely on synthetic benchmarks risk overestimating how well automated assistants will perform for actual end users in operational settings.

To bridge the gap between synthetic data and real-world complexity, the article recommends incorporating human-in-the-loop review or iterative feedback agents into synthetic generation pipelines. Technical teams developing spreadsheet assistants should also implement structure-aware context retrieval to feed relevant table areas into repair models, rather than relying on raw formula text alone.

Readers should interpret the synthetic performance metrics with caution. While execution correctness is guaranteed via programmatic checks, the automated evaluation filter showed only moderate agreement with human raters on the contextual realism of table data (Cohen's Kappa of 0.42). Additionally, the benchmark remains limited to single-sheet, English-language spreadsheets, meaning performance on multi-sheet workbooks and collaborative enterprise environments requires further evaluation.

Cover for Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs

Abstract

Excel is a pervasive yet often complex tool, particularly for novice users, where runtime errors arising from logical mistakes or misinterpretations of functions pose a significant challenge. While large language models (LLMs) offer promising assistance by explaining formula errors, the automated correction of these semantic runtime errors remains an open problem. A primary challenge to advancing models for such scenarios is the severe lack of high-quality, comprehensive datasets for training and rigorous evaluation. This paper addresses this gap by introducing a novel approach for constructing a benchmark dataset specifically designed for Excel formula repair. We propose a data generation pipeline, which leverages a small set of curated seed samples from online forums to synthetically expand the dataset. Our pipeline integrates few-shot prompting with LLMs and employs a robust \textit{LLM-as-a-Judge} validation framework, combined with execution-based checks to ensure the correctness and semantic fidelity of the generated data. This process produced a benchmark dataset of 618 high-quality samples, covering common runtime errors. Furthermore, we propose a context-aware baseline technique for Excel formula repair that utilizes LLMs to leverage both the faulty formula, and relevant spreadsheet context. We evaluate the performance of various LLMs (GPT-4o, GPT-4.1, Phi-3, Mistral) on our newly generated benchmark using execution-based metrics. Our analysis demonstrates the dataset's quality through manual annotation and provides insights into error and function distributions. The proposed generation methodology is highly scalable and can be readily adapted to create evaluation benchmarks for similar code repair tasks in other low-resource programming languages.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 LLMs for Code Generation
  • 2.2 Excel Formula Generation and Repair
  • 3 Methodology
  • 3.1 Seed Data Curation
  • 3.1.1 Dataset Creation
  • 3.1.2 Manual Verification and Correction
  • 3.2 Bootstrap Generation
  • 3.2.1 Data Generation with Few-Shot Prompting
  • 3.2.2 Validating Generations executing Excel formulas
  • 3.2.3 Validating Generations with LLM-as-a-Judge Approach
  • 4 Formula Repair
  • 4.1 Baseline Repair Technique
  • 5 Experimental Setup
  • 5.1 Metrics
  • 6 Results
  • 6.1 RQ1: Data Distribution and Comparison with Seed Dataset
  • 6.2 RQ2: Synthetic Dataset Quality Based on Human Evaluation
  • 6.3 RQ3: Performance of Repair Task on Synthetic and Seed Data
  • 6.4 RQ4: Cost of Dataset Generation
  • 7 Discussion and Conclusion
  • References

Knowls

  1. Knowl 1 — FoRepBench: Benchmark Dataset for Context-Aware Excel Formula Repair

    definition

    FoRepBench (Formula Repair Benchmark) is a dataset designed to evaluate and train automated program repair models on semantic runtime errors in Microsoft Excel formulas. Each example in the dataset includes four core components:

    1. Tabular Data / Spreadsheet Context: Cell values, column headers, and tabular grid structures representing the context in which the formula error occurred.
    2. Faulty Formula: An Excel formula that produces one of five specific runtime errors: #DIV/0!, #N/A, #NAME?, #REF!, or #VALUE!.
    3. Correct Formula: A ground-truth repaired formula that resolves the runtime error, executes successfully against the spreadsheet context, and fulfills the user's intent.
    4. User Utterance: A natural language query describing the user's intended calculation or goal.

    The benchmark comprises 618 verified samples generated via few-shot bootstrapping and multi-stage execution and LLM-as-a-judge validation.

  2. Knowl 2 — Bootstrap Generator Pipeline for Excel Formula Repair Data

    model/method

    The Bootstrap Generator is a synthetic data generation and filtering pipeline that expands a small set of curated real-world seed examples into a large-scale formula repair benchmark:

    1. Few-Shot Prompting: Each curated seed data point is injected as a 1-shot example into a generation prompt evaluated using an LLM (GPT-4o at temperature 0.640.64, generating candidate batches of size N=25N=25). The model generates new tabular contexts, faulty formulas, repaired formulas, corresponding runtime error types, and user utterances.
    2. Execution-Based Verification: The generated formula pairs are evaluated against their generated spreadsheet tables using Calc.ts (a standalone Excel formula evaluation engine). Candidates are filtered out unless the faulty formula produces the specified runtime error (#DIV/0!, #N/A, #NAME?, #REF!, or #VALUE!) and the repaired formula executes without runtime errors.
    3. Chain-of-Thought LLM Validation: Candidates passing execution checks are evaluated by an LLM Validator using Chain-of-Thought (CoT) reasoning. The validator checks if the repair resolves the runtime error, verifies semantic alignment with the user's utterance and spreadsheet context, and assigns a difficulty rating (easy, medium, hard). Samples passing both stages are added to the final benchmark.
  3. Knowl 3 — Context-Aware Baseline Formula Repair Technique and Evaluation Metrics

    model/method

    The baseline Excel formula repair technique provides a single-call LLM framework for repairing semantic runtime errors by incorporating localized tabular context:

    1. Context Extraction: Because full spreadsheets can exceed LLM context windows, the nearest table associated with the faulty formula is identified, and only its column headers along with a small set of representative sample rows are extracted.
    2. Prompt Assembly and Repair Generation: A standardized prompt is constructed containing the extracted table context, the faulty formula, the runtime error type, and the user's natural language utterance. The LLM processes this input and generates a repaired formula alongside an explanatory natural language rationale.
    3. Evaluation Metrics:
      • Syntax Validity: Binary metric indicating whether the repaired formula successfully parses and compiles without syntax errors.
      • Can Execute: Binary metric indicating whether the repaired formula executes on the spreadsheet without triggering runtime errors (e.g., #VALUE!, #REF!, #DIV/0!).
      • Execution Match: Exact equality check between the output produced by executing the model's repaired formula and the output of the ground-truth correct formula on the spreadsheet context.
  4. Knowl 4 — Repair Performance Across LLMs on FoRepBench and Seed Datasets

    data/table

    Evaluating the context-aware baseline repair approach across proprietary (GPT-4.1, GPT-4o) and open-weight (Phi-3, Mistral) LLMs on both FoRepBench (N=618N=618) and the seed dataset (N=59N=59) yields the following execution-based performance metrics:

    LLM Dataset Origin Syntax Valid Can Execute Execution Match
    GPT-4.1 FoRepBench 1.00 0.96 0.80
    GPT-4.1 Seed Dataset 0.98 0.65 0.41
    GPT-4o FoRepBench 1.00 0.93 0.73
    GPT-4o Seed Dataset 0.96 0.63 0.35
    Phi-3 FoRepBench 0.81 0.77 0.58
    Phi-3 Seed Dataset 0.73 0.41 0.24
    Mistral FoRepBench 0.78 0.76 0.51
    Mistral Seed Dataset 0.67 0.37 0.19

    All models achieve higher accuracy on FoRepBench than on the seed dataset. GPT-4.1 achieves the highest performance (0.80 Execution Match on FoRepBench, 0.41 on Seed), and larger frontier models consistently outperform open-weight models across all three metrics.

  5. Knowl 5 — Seed Dataset Curation and Manual Verification Process

    model/method

    The seed dataset is constructed by scraping user posts from the MrExcel community forum to obtain real-world spreadsheet repair problems:

    1. Extraction and Workbook Reconstruction: Forum threads containing a faulty formula, table context, and an "accepted answer" formula are extracted. Cell addresses, values, formulas, and reported Excel error messages are parsed into reconstructed workbook structures.
    2. Execution Simulation: Formulas are executed using Calc.ts outside the Excel environment to confirm that the faulty formula triggers one of five targeted runtime errors (#N/A, #REF!, #VALUE!, #NAME?, #DIV/0!) and that the accepted answer resolves the error.
    3. Two-Stage Human Review: In Round 1, each sample is evaluated by an annotator on four criteria: requirement satisfaction by the correct formula, correct error generation by the faulty formula, table extraction accuracy, and context-utterance consistency. Because only one-third of samples met all criteria in Round 1, a second round involving three or more annotators per sample edited utterances for clarity, replaced incorrect formulas, and removed data leakage (such as pre-existing target output columns), producing N=59N=59 validated seed samples.
  6. Knowl 6 — LLM-as-a-Judge vs. Human Agreement on Synthetic Formula Quality

    empirical result

    Pairwise agreement measured via Cohen's Kappa (kappa\\kappa) between two expert human annotators and LLM Validator on a 24-sample synthetic subset indicates moderate inter-human agreement and lower human-LLM agreement:

    Annotator Pair Cohen's Kappa (κ\kappa)
    Annotator 1 vs. LLM Validator 0.42
    Annotator 2 vs. LLM Validator 0.42
    Annotator 1 vs. Annotator 2 0.60

    Disagreements between human annotators and LLM Validator primarily occur because LLM Validator verifies logical and computational consistency but is insensitive to the contextual plausibility of spreadsheet content (for instance, accepting tables where numeric columns like unit prices contain text strings such as "Three").

  7. Knowl 7 — Filter Yield of Bootstrap Generator by Runtime Error Type

    data/table

    Out of 1,095 synthetic samples that satisfied Calc.ts execution requirements, LLM Validator passed 618 samples (56.4%56.4\% overall acceptance rate), broken down by runtime error type as follows:

    Error Type Samples Generated (Calc.ts Executable) Samples Passed by LLM Validator
    #VALUE! 241 64
    #N/A 329 140
    #REF! 16 12
    #NAME? 222 158
    #DIV/0! 287 244
    Total 1,095 618

    #DIV/0! and #NAME? achieved high validation pass rates (85.0%85.0\% and 71.2%71.2\% respectively), while #VALUE! had the lowest pass rate (26.6%26.6\%) due to frequent semantic inconsistencies between tabular data and formula logic.

  8. Knowl 8 — Excel Function Frequencies in Seed Dataset vs. FoRepBench

    data/table

    The distribution of the top 10 most frequent Excel functions across error types in the seed dataset (N=59N=59) and FoRepBench (N=618N=618) illustrates function coverage:

    Seed Dataset (N=59N=59):

    Error Type AND FIND IF INDEX LEFT MATCH MID MIN SUM VLOOKUP
    #DIV/0! 0 0 1 0 0 0 0 2 1 0
    #N/A 1 0 2 8 2 10 1 0 1 7
    #NAME? 0 0 1 0 1 0 0 0 0 0
    #REF! 0 0 0 0 0 0 0 0 0 0
    #VALUE! 2 4 12 2 2 3 3 1 2 0

    FoRepBench (N=618N=618):

    Error Type AVERAGE AVRG CONCATENATE COUNT IF INDEX MATCH SUM VALUE VLOOKUP
    #DIV/0! 3 0 0 4 1 0 0 9 0 0
    #N/A 1 0 0 0 1 25 28 1 0 66
    #NAME? 2 4 0 0 2 0 0 1 0 0
    #REF! 0 0 0 0 0 0 0 1 0 2
    #VALUE! 9 0 16 0 10 2 3 11 4 0

    While #N/A errors remain dominated by lookup functions (INDEX, MATCH, VLOOKUP) in both datasets, FoRepBench introduces diverse functions not present in the seed data, such as AVERAGE and CONCATENATE.

  9. Knowl 9 — Computational and Financial Cost of FoRepBench Generation

    empirical result

    Generating FoRepBench using GPT-4o required a total of 1,154 LLM API calls:

    • Generation Calls: 59 calls (1 prompt per seed sample producing candidate batches of N=25N=25).
    • Validation Calls: 1,095 calls (1 CoT evaluation call per executable candidate).
    • Average Calls per Validated Sample: 2.092.09 calls per accepted benchmark sample.

    With an average token consumption of approximately 2,000 tokens per call, the generation pipeline achieved an average cost of approximately $0.02\$0.02 per validated sample.

  10. Knowl 10 — Limitations of Synthetic Excel Formula Repair Data

    limitation

    The synthetic data generation pipeline and resulting FoRepBench dataset have three main limitations:

    1. Complexity Gap: Synthetic instances feature shallower formula nesting and require fewer, more localized token edits (e.g., adding an IF guard) than real-world forum problems, which often require multi-step logic modifications and deeper function nesting.
    2. Judge Misalignment with Semantic Plausibility: The LLM-as-a-judge validator evaluates execution and logical consistency but struggles to detect implausible spreadsheet values (such as text entered in numeric columns).
    3. Scope Constraints: The dataset is restricted to single-sheet, English-language spreadsheets, omitting multi-sheet formulas, external workbook dependencies, and localized function names common in production environments.

Coverage note — None was omitted; all contributed methodologies, dataset construction steps, benchmark tables, empirical results, cost analyses, and limitations are fully represented.

References

  1. 1.Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al. 2024. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219 (2024).
  2. 2.Daniel W Barowy, Shan Gao, Alvin Cheung, and Brad A Myers. 2014. ExceLint: Automatically detecting spreadsheet formula errors. In Proceedings of the 36th International Conference on Software Engineering. ACM, 460–470.
  3. 3.Rohan Bavishi, Harshit Joshi, José Cambronero, Anna Fariha, Sumit Gulwani, Vu Le, Ivan Radiček, and Ashish Tiwari. 2022. Neurosymbolic repair for low-code formula languages. Proc. ACM Program. Lang. 6, OOPSLA2, Article 164 (Oct. 2022), 30 pages. doi:10.1145/3563327
  4. 4.Rohan Bavishi, Harshita Joshi, Jorge Cambronero, Ayesha Fariha, Sumit Gulwani, Vu Le, Ivan Radiček, and Aditya Tiwari. 2022. Neurosymbolic Repair for Low-Code Formula Languages. Proceedings of the ACM on Programming Languages 6, OOPSLA2 (2022).
  5. 5.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems 33 (2020).
  6. 6.Binyuan Chen, Qian Liu, Jinjie Jiang, and et al. 2023. CodeLLM: Evaluating Large Language Models on Code Generation. arXiv preprint arXiv:2305.14335 (2023).
  7. 7.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating Large Language Models Trained on Code. arXiv preprint arXiv:2107.03374 (2021).
  8. 8.Xinyun Chen, Petros Maniatis, Rishabh Singh, Charles Sutton, Hanjun Dai, Max Lin, and Denny Zhou. 2021. SpreadsheetCoder: Formula Prediction from Semi-structured Context. arXiv:2106.15339 [cs.SE] https://arxiv.org/abs/2106.15339
  9. 9.Georgios Gousios, Andy Zaidman, Margaret-Anne Storey, and Arie Van Deursen. 2015. Work practices and challenges in pull-based development: The integrator’s perspective. In Proceedings of the 37th IEEE/ACM International Conference on Software Engineering, Volume 1. IEEE, 358–368.
  10. 10.Aniruddh Gudibande, Xisen Li, Ethan Chi, Percy Liang, and Yuxin Wu. 2023. False sense of security: Evaluation misalignment in language models. arXiv preprint arXiv:2304.09106 (2023).
  11. 11.Sumit Gulwani. 2011. Automating string processing in spreadsheets using input-output examples. In Proceedings of the 38th Annual ACM SIGPLAN-SIGACT Symposium on Principles of Programming Languages (Austin, Texas, USA) (POPL ’11). Association for Computing Machinery, New York, NY, USA, 317–330. doi:10.1145/1926385.1926423
  12. 12.Felienne Hermans, Martin Pinzger, and Arie van Deursen. 2016. Detecting errors in spreadsheets. In Proceedings of the 38th International Conference on Software Engineering (ICSE). ACM, 818–828.
  13. 13.Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276 (2024).
  14. 14.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7B. arXiv:2310.06825 [cs.CL] https://arxiv.org/abs/2310.06825
  15. 15.Harshit Joshi, Abishai Ebenezer, José Cambronero Sanchez, Sumit Gulwani, Aditya Kanade, Vu Le, Ivan Radiček, and Gust Verbruggen. 2024. Flame: A small language model for spreadsheet formulas. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 12995–13003.
  16. 16.Sean Kandel, Andreas Paepcke, Joseph M Hellerstein, and Jeffrey Heer. 2011. Wrangler: Interactive visual specification of data transformation scripts. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. ACM, 3363–3372.
  17. 17.Xi Li et al. 2022. Competition-Level Code Generation with AlphaCode. In Proceedings of the International Conference on Machine Learning (ICML).
  18. 18.A. Liu et al. 2022. InCoder: A Generative Model for Code Infill. arXiv:2210.00745 [cs.CL]
  19. 19.Fangyu Liu, Yuxian Wu, Yixuan Liu, and et al. 2023. GPTEval: NLG evaluation using GPT-4 as the reference-free evaluator. arXiv preprint arXiv:2305.04648 (2023).
  20. 20.Martin Monperrus. 2018. Automatic Software Repair: A Bibliography. ACM Comput. Surv. 51, 1, Article 17 (Jan. 2018), 24 pages. doi:10.1145/3105906
  21. 21.Eric Nijkamp, Christopher Rosin, Antonio Martins, Adam Rogers, Thomas Wolf, Mikel Artetxe, Victor Costa, Sudheer Banerjee, Binh Shih, Emelie Siktberg, et al. 2022. CodeGen: An Open Large Language Model for Code Generation. arXiv preprint arXiv:2203.13474 (2022).
  22. 22.Usneek Singh, José Cambronero, Sumit Gulwani, Aditya Kanade, Anirudh Khatry, Vu Le, Mukul Singh, and Gust Verbruggen. 2024. An Empirical Study of Validating Synthetic Data for Formula Generation. arXiv:2407.10657 [cs.CL] https://arxiv.org/abs/2407.10657
  23. 23.Alex Wang and Ellie Pavlick. 2023. ChatGPT as a judge: Linguistic acceptability judgments. arXiv preprint arXiv:2304.03442 (2023).
  24. 24.Shuai Wang, Feng Li, Jia Zhou, Ruo Yan, and Jie Chen. 2021. CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. arXiv preprint arXiv:2109.00859 (2021).
  25. 25.Pengcheng Yin and Graham Neubig. 2018. Learning to Represent Programs with Graphs. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 195–206.
  26. 26.Wei Zhao, Zhitao Hou, Siyuan Wu, Yan Gao, Haoyu Dong, Yao Wan, Hongyu Zhang, Yulei Sui, and Haidong Zhang. 2024. NL2Formula: Generating Spreadsheet Formulas from Natural Language Queries. arXiv:2402.14853 [cs.CL] https://arxiv.org/abs/2402.14853
  27. 27.Lei Zheng, Xiaowei Wang, Baoxu Peng, Xin Wang, and Minlie Huang. 2023. Judging code generation with large language models: A comparative study. arXiv preprint arXiv:2305.17951 (2023).

Citation

MLA
Singha, A., et al. “Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs”. arXiv, 2025, http://arxiv.org/abs/2508.11715v1.
APA
Singha, A., Sahijwani, H., Williams, W., Boateng, E. A., Hausman, N., Luca, M. D., Choudhury, K., Binet, C., Le, V., Chen, T., Chen, O. R., Vesal, S., & Hasan, S. (2025). Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs. arXiv. http://arxiv.org/abs/2508.11715v1
Chicago
Singha, A., H. Sahijwani, W. Williams, et al. 2025. “Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs”. arXiv. http://arxiv.org/abs/2508.11715v1.
Harvard
Singha, A. et al. (2025) “Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2508.11715v1.
Vancouver
1. Singha A, Sahijwani H, Williams W, et al (2025) Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs. arXiv

BibTeX

@article{singha2025benchmark,
  title = {Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs},
  author = {Singha, Ananya and Sahijwani, Harshita and Williams, Walt and Boateng, Emmanuel Aboah and Hausman, Nick and Luca, Miguel Di and Choudhury, Keegan and Binet, Chaya and Le, Vu and Chen, Tianwei and Chen, Oryan Rokeah and Vesal, Sulaiman and Hasan, Sadid},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2508.11715v1},
  eprint = {2508.11715}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/