VERT: Reliable LLM Judges for Radiology Report Evaluation

Federica BolognaJean-Philippe CorbeilMatthew WilkensAsma Ben Abacha

article2026arXiv2 citations

Introduces VERT, an automated radiology report evaluation metric that achieves superior correlation with expert radiologist judgments across diverse imaging modalities while showing that lightweight fine-tuning can boost evaluation performance and reduce inference time by up to 37 times.

Listen

Automated evaluation of radiology reports is essential for scaling quality control in AI-assisted medical imaging, but existing automated metrics have largely been restricted to chest radiographs. This creates a significant blind spot when systems encounter varied clinical settings across different imaging modalities (such as CT and MRI) and diverse anatomical regions. Without robust and broadly applicable evaluation methods, clinical decision-makers cannot reliably benchmark generative AI tools across real-world hospital workflows.

The article systematically assesses how large language models (LLMs) can act as automated judges to evaluate radiology reports across diverse modalities and body regions. It also introduces and validates VERT, a new LLM-based evaluation metric designed to score reports with higher clinical reliability and alignment to expert radiologists.

The authors conducted empirical evaluations across two multi-modality, expert-annotated radiology benchmarks: RadEval (148 chest X-ray cases annotated by error counts) and RaTE-Eval (1,856 report pairs covering 9 imaging modalities and 22 anatomies). The study compared several prompting strategies, proprietary and open-source models, reasoning modes, few-shot variations, model ensembling, parameter-efficient fine-tuning (LoRA), and controlled clinical error injections.

Direct continuous accuracy scoring via VERT outperformed prior automated metrics, improving correlation with expert judgments by up to 11.7% relative to the standard GREEN metric. Parameter-efficient fine-tuning of an open-source model (Qwen3 30B) on roughly 1,300 samples yielded correlation gains of up to 25% while speeding up inference time by 37.2 times compared to proprietary API baselines. However, longer extended reasoning traces generally failed to improve correlation with human radiologists, and automated judges frequently underestimated error counts when reports contained high error density or nuanced errors such as incorrect anatomical locations and omitted comparisons.

These findings demonstrate that lightweight adaptation of open-source models offers a practical, high-throughput path for automated radiology quality assurance while cutting operational costs and API latency. However, because current automated judges struggle to detect complex clinical nuances such as location errors or temporal changes, organizations cannot yet rely on fully autonomous evaluation in safety-critical clinical deployments.

Healthcare technology leaders should prioritize direct-scoring prompts and domain-adapted open-source models for offline benchmarking and automated triage. For production deployments, leaders should maintain human-in-the-loop validation for reports with high error potential or complex comparative findings. Further research is recommended to improve automated detection of multi-category errors without requiring matched ground-truth reference reports.

While the study provides strong empirical evidence across multiple modalities, results are bounded by the reliance on two reference datasets and synthetic error validation steps. Decision-makers should maintain high confidence in using these approaches for large-scale comparative benchmarking, but exercise caution before replacing radiologist review in live clinical oversight.

arXiv: 2604.03376

No sufficiently relevant recommendations were found.

Cover for VERT: Reliable LLM Judges for Radiology Report Evaluation

Abstract

Current literature on radiology report evaluation has focused primarily on designing LLM-based metrics and fine-tuning small models for chest X-rays. However, it remains unclear whether these approaches are robust when applied to reports from other modalities and anatomies. Which model and prompt configurations are best suited to serve as LLM judges for radiology evaluation? We conduct a thorough correlation analysis between expert and LLM-based ratings. We compare three existing LLM-as-a-judge metrics (RadFact, GREEN, and FineRadScore) alongside VERT, our proposed LLM-based metric, using open- and closed-source models (reasoning and non-reasoning) of different sizes across two expert-annotated datasets, RadEval and RaTE-Eval, spanning multiple modalities and anatomies. We further evaluate few-shot approaches, ensembling, and parameter-efficient fine-tuning using RaTE-Eval. To better understand metric behavior, we perform a systematic error detection and categorization study to assess alignment of these metrics against expert judgments and identify areas of lower and higher agreement. Our results show that VERT improves correlation with radiologist judgments by up to 11.7% relative to GREEN. Furthermore, fine-tuning Qwen3 30B yield gains of up to 25% using only 1,300 training samples. The fine-tuned model also reduces inference time up to 37.2 times. These findings highlight the effectiveness of LLM-based judges and demonstrate that reliable evaluation can be achieved with lightweight adaptation.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Traditional and Hybrid Metrics for Radiology
  • 2.2 LLM-based Metrics for Radiology
  • 2.3 Fine-Tuning Models as Evaluators in Radiology
  • 3 Datasets
  • 4 Methodology
  • 4.1 Correlation With Human Judgments
  • 4.1.1 Zero-shot Prompting
  • 4.1.2 Ensembling
  • 4.1.3 Few-shot Prompting
  • 4.1.4 LoRA Fine-Tuning
  • 4.2 Error Detection and Categorization Analysis
  • 5 Results
  • 5.1 Correlation With Human Judgments
  • 5.1.1 Model-generated scores correlate with expert judgments
  • 5.1.2 Different zero-shot judges for different datasets
  • 5.1.3 Best LLMs as Radiology Evaluators
  • 5.1.4 Thinking does not improve correlation with expert judgments
  • 5.1.5 Regression-based ensemblings improve correlation with expert judgments
  • 5.1.6 Few-shot Prompting
  • 5.1.7 LoRA Fine-Tuning
  • 5.2 Error Detection and Categorization
  • 6 Conclusion
  • References
  • A Dataset Statistics
  • B Comparison of GREEN, F1, and Weighted Error Metrics
  • B.1 Notation
  • B.2 GREEN Score
  • B.3 F1 Score
  • B.4 Formula Variant as Metric
  • B.5 Unified Formulation
  • B.6 Summary
  • B.7 Metric Simulations
  • C F1-Score for Error Type Categorization
  • D Prompts for Radiology Report Evaluation
  • D.1 GREEN
  • D.2 VERT
  • D.3 Formula
  • D.4 Rubric
  • D.5 Few-shot Variants
  • D.6 rad-err: RadEval Error-Type Examples
  • D.7 rate-err: RaTE-Eval Error-Type Examples
  • D.8 rad-err-10 human: RadEval 10-Shot with Human Annotations
  • D.9 rate-err-10 VERT: RaTE-Eval 10-Shot with VERT Assessments
  • E Prompts for Radiology Report Evaluation
  • F Prompts for Error Injection
  • F.1 Error Type (a): False Prediction of Finding
  • F.2 Error Type (b): Omission of Finding
  • F.3 Validation Prompt

Knowls

  1. Knowl 1 — The VERT Metric for Radiology Report Evaluation

    model/method

    VERT is an LLM-as-a-judge metric designed for evaluating candidate radiology reports against expert reference reports across diverse imaging modalities and anatomies. Unlike previous approaches that derive quality scores post hoc via deterministic formulas over counted errors (such as GREEN) or enforce discrete scoring rubrics, VERT instructs the large language model to identify clinical errors and directly predict a continuous overall accuracy score.

    The VERT evaluation process operates as follows:

    1. The model is provided with the reference radiology report and the generated candidate report.
    2. The model identifies and enumerates clinically significant and clinically insignificant errors across six defined categories:
      • (a) False report of a finding in the candidate report.
      • (b) Missing a finding present in the reference report.
      • (c) Misidentification of anatomical location or position.
      • (d) Misassessment of finding severity.
      • (e) Mentioning a comparison absent from the reference.
      • (f) Omitting a comparison detailing a change relative to a prior study.
    3. The model extracts and counts matched findings present in both reports.
    4. The model directly outputs an overall accuracy score formulated as a continuous value in the range [0.00,1.00][0.00, 1.00] with two decimal places, reflecting the candidate report's clinical accuracy in light of the identified errors and matched findings.
  2. Knowl 2 — Unified Formulation and Comparative Properties of Radiology Report Evaluation Metrics

    theoretical result

    Radiology report evaluation metrics based on matched findings and error counts can be formalized under a single unified equation. Let TP\text{TP} denote the number of correctly matched findings between a candidate report and a reference report, FPS\text{FP}_S and FNS\text{FN}_S denote the false positive and false negative clinically significant errors respectively with S=FPS+FNSS = \text{FP}_S + \text{FN}_S, and FPI\text{FP}_I and FNI\text{FN}_I denote clinically insignificant errors with I=FPI+FNII = \text{FP}_I + \text{FN}_I.

    The unified score formulation is: Score=αTPαTP+βSS+βII\text{Score} = \frac{\alpha \text{TP}}{\alpha \text{TP} + \beta_S S + \beta_I I}

    Special cases of this formulation include:

    • GREEN Score: α=1\alpha = 1, βS=1\beta_S = 1, βI=0\beta_I = 0, yielding: G=TPTP+SG = \frac{\text{TP}}{\text{TP} + S} GREEN represents a normalized accuracy over findings. It assigns a score of 0.50.5 whenever TP=S\text{TP} = S, but can be sensitive to report length because variations in TP\text{TP} shift the score even when the error count SS is constant.
    • F1 Score: α=2\alpha = 2, βS=1\beta_S = 1, βI=0\beta_I = 0, yielding: F1=2TP2TP+S\text{F1} = \frac{2\text{TP}}{2\text{TP} + S} F1 corresponds to the harmonic mean of precision and recall over significant errors. By weighting TP\text{TP} by a factor of 2, it reduces sensitivity to total report length and stabilizes comparisons across reports with varying finding counts.
    • Formula Variant: α=1\alpha = 1, βS=2\beta_S = 2, βI=0.5\beta_I = 0.5, yielding: Gw=TPTP+2S+0.5IG_w = \frac{\text{TP}}{\text{TP} + 2S + 0.5I} This variant introduces cost-sensitive weighting penalizing high-risk significant errors while retaining partial penalties for minor errors.
  3. Knowl 3 — Comparison of Zero-Shot Scoring Prompt Instructions

    empirical result

    Evaluating different prompting formulations for LLM judges shows that instructing an LLM to directly predict a continuous accuracy score (VERT) achieves higher rank correlation with radiologist error annotations than computing post-hoc ratios (GREEN), providing mathematical scoring formulas, or defining discrete rubrics.

    Evaluation on the chest X-ray RadEval benchmark measures absolute Kendall's τ\tau correlation between automatic scores and human-annotated error counts:

    Model GREEN VERT Formula Rubric
    GPT-4.1-mini 0.332 0.371 0.242 0.322
    Gemini 2.5 flash 0.305 0.308 0.242 0.254

    All reported correlations are statistically significant (p<0.01p < 0.01). Constraining the model with explicit mathematical equations or discrete rubrics consistently degrades alignment with expert radiologist judgments.

  4. Knowl 4 — Benchmark Evaluation of LLM Judges Across Modalities and Anatomies

    empirical result

    Evaluations across RadEval (148 chest X-ray studies with candidate error counts) and RaTE-Eval (1,856 test paragraph findings spanning 9 modalities and 22 anatomical regions with expert quality ratings on a 0–5 scale) establish the comparative correlation of zero-shot LLM judges using absolute Kendall's τ\tau:

    Model RadFact GREEN GREEN-EC GREEN F1 FineRadScore VERT
    RadEval — Error Annotations
    GPT-4.1 mini 0.2946 0.3322 0.5606 0.4441 0.1725 0.3710
    Gemini 3 Flash (low) – 0.1675 0.0206* 0.4103 0.1810 0.1937
    Claude Sonnet 4.6 (low) – 0.2931 0.5103 0.3924 0.2333 0.3291
    RaTE-Eval — Expert Scores (0–5)
    GPT-4.1 mini 0.3197 0.4195 0.1966 0.3369 0.1372 0.4474
    Gemini 3 Flash (low) – 0.0655* 0.0061* 0.4641 0.1148 0.1617
    Claude Sonnet 4.6 (low) – 0.4322 0.2461 0.3782 0.1954 0.4526

    Asterisks () indicate results lacking statistical significance (p≥0.01p \ge 0.01).

    When evaluating on RaTE-Eval expert scores, Claude Opus 4.6 (low) achieves τ=0.4633\tau = 0.4633 with VERT and τ=0.4493\tau = 0.4493 with GREEN, while GPT-4.1 nano achieves τ=0.5040\tau = 0.5040 with VERT. VERT improves correlation with human judgments by up to 11.7% relative to GREEN across standard model backbones.

  5. Knowl 5 — Impact of Reasoning Tokens and Thinking Effort on LLM Judges

    empirical result

    Varying the reasoning budget (thinking effort) in reasoning LLMs does not consistently improve evaluation alignment with expert radiologist ratings on RaTE-Eval (0–5 scale) using the GREEN prompt formulation:

    Thinking Level GPT-5 GPT-5 mini Gemini 3 Flash Claude Opus Claude Sonnet Qwen3 30B Qwen3 4B
    None – – 0.0795* 0.4409 0.4343 0.0411* 0.0434*
    Low 0.3333 0.3495 0.0655* 0.4493 0.4322 – –
    Medium (default) 0.3674 0.3482 0.1183 0.3590 0.2988 0.2851 0.0475*
    High 0.3508 0.2750 0.1207 0.2004 0.2054 – –

    Asterisks () indicate p≥0.01p \ge 0.01.

    For Claude and GPT models, high thinking levels lead to substantial drops in Kendall's τ\tau correlation (e.g., Claude Opus drops from 0.4493 at Low to 0.2004 at High; Claude Sonnet drops from 0.4343 at None to 0.2054 at High). Small open-source models (Qwen3 4B) fail to produce statistically significant correlations under any thinking configuration.

  6. Knowl 6 — Ensembling LLM Judges via Linear Regression

    empirical result

    Ensembling predictions from multiple diverse models—specifically GPT-4.1, Claude Opus 4.6 (low), and Gemini 3 Flash (low)—improves correlation with radiologist scores on the RaTE-Eval test set when aggregated using linear regression fitted on 500 training instances:

    GREEN VERT
    Ensemble Method SCORE F1 EC All SCORE All
    Single-Model Baseline 0.4493 0.4641 0.2461 – 0.4633 –
    Simple Average 0.4149 0.4249 0.2649 – 0.4671 –
    Linear Regression 0.4478 0.4908 0.3486 0.5126 0.5683 0.5573

    Simple score averaging underperforms individual top-tier baselines in half of the settings. In contrast, linear regression aggregation on VERT scores achieves Kendall's τ=0.5683\tau = 0.5683, representing a 22.7% relative improvement over the best single-model baseline (τ=0.4633\tau = 0.4633).

  7. Knowl 7 — Parameter-Efficient Fine-Tuning of Qwen3-30B for Radiology Report Evaluation

    model/method

    Parameter-efficient fine-tuning of an open-source mixture-of-experts model on expert-annotated radiology reports yields superior alignment and substantial inference acceleration over frontier proprietary LLM APIs.

    Adaptation Setup

    • Base Model: Qwen3-30B-A3B-Instruct-2507 (128 total experts, 8 active per token).
    • Quantization & LoRA: Base weights frozen in 4-bit NormalFloat (computation in bfloat16); LoRA adapters of rank r=32r = 32 and α=64\alpha = 64 injected into all linear layers except the MoE router layers.
    • Optimization: AdamW optimizer with a cosine learning rate scheduler, peak learning rate 2×10−42 \times 10^{-4}, trained for up to 5 epochs.
    • Data: RaTE-Eval training split comprising 1,486 report paragraphs (90% training, 10% validation) formatted with the first two lines of the VERT prompt and wrapping target scores in [0.0,1.0][0.0, 1.0] rounded to one decimal place.

    Results and Efficiency

    • Correlation Progression: Kendall's τ\tau correlation on the RaTE-Eval test set improves from 0.4223 (epoch 1) to 0.486 (epoch 2), 0.521 (epoch 3), and 0.578 (epoch 5).
    • Performance: The 5-epoch fine-tuned model (τ=0.578\tau = 0.578) outperforms the best zero-shot proprietary model (Claude Opus 4.6 low with VERT, τ=0.4633\tau = 0.4633) by 24.8% relative and surpasses the multi-model linear regression ensemble (τ=0.5683\tau = 0.5683).
    • Inference Speed: Deployed via vLLM on a single A100 GPU, the fine-tuned model processes the entire RaTE-Eval test benchmark in 36 seconds, compared to 22.3 minutes for Claude Opus 4.6 via API, achieving a 37.2×37.2\times reduction in runtime.
  8. Knowl 8 — Count-Level Precision, Recall, and F1 Formulation for Multi-Category Radiology Error Evaluation

    equation

    To evaluate error detection across specific error categories c∈{(a),(b),(c),(d),(e),(f)}c \in \{(a), (b), (c), (d), (e), (f)\}, count-level matches and discrepancies are calculated for each report ii. Let hi,c∈N0h_{i,c} \in \mathbb{N}_0 represent the ground-truth human-annotated error count for category cc, and let gi,c∈N0g_{i,c} \in \mathbb{N}_0 represent the model-predicted error count for category cc.

    Per-report true positives, false positives, and false negatives are defined as: TPi,c=min⁡(hi,c,gi,c),FPi,c=max⁡(0,gi,c−hi,c),FNi,c=max⁡(0,hi,c−gi,c)\text{TP}_{i,c} = \min(h_{i,c}, g_{i,c}), \quad \text{FP}_{i,c} = \max(0, g_{i,c} - h_{i,c}), \quad \text{FN}_{i,c} = \max(0, h_{i,c} - g_{i,c})

    Aggregating across all reports in the evaluation set: TPc=∑iTPi,c,FPc=∑iFPi,c,FNc=∑iFNi,c\text{TP}_c = \sum_i \text{TP}_{i,c}, \quad \text{FP}_c = \sum_i \text{FP}_{i,c}, \quad \text{FN}_c = \sum_i \text{FN}_{i,c}

    Category-level Precision, Recall, and F1 are then computed as: Precisionc=TPcTPc+FPc,Recallc=TPcTPc+FNc,F1c=2⋅Precisionc⋅RecallcPrecisionc+Recallc\text{Precision}_c = \frac{\text{TP}_c}{\text{TP}_c + \text{FP}_c}, \quad \text{Recall}_c = \frac{\text{TP}_c}{\text{TP}_c + \text{FN}_c}, \quad \text{F1}_c = \frac{2 \cdot \text{Precision}_c \cdot \text{Recall}_c}{\text{Precision}_c + \text{Recall}_c} If a denominator evaluates to zero, the corresponding score is set to 0.

  9. Knowl 9 — Systematic Error Detection Biases and Category Failure Modes of LLM Judges

    limitation

    Granular error analysis on RadEval and controlled error-injection experiments on RaTE-Eval reveal distinct systematic failure modes in LLM-based radiology evaluators:

    1. Error Count Over- and Under-estimation Bias: When evaluating reports with fewer than 6 expert-annotated significant errors, LLM judges consistently overestimate total error counts. Conversely, for reports containing more than 6 errors, models consistently underestimate the total number of errors.
    2. Severe Category Performance Disparity: LLM evaluators successfully identify coarse presence/absence errors but struggle with nuanced semantic and longitudinal discrepancies:
      • On RadEval, models achieve F1=0.66–0.69F_1 = 0.66\text{--}0.69 on (a) false predictions and F1=0.77–0.78F_1 = 0.77\text{--}0.78 on (b) missing findings, but drop to F1=0.39–0.46F_1 = 0.39\text{--}0.46 on (c) incorrect location, (d) incorrect severity, (e) false comparison, and (f) missing comparison.
      • On RaTE-Eval error injection tests, models achieve F1=0.90–0.92F_1 = 0.90\text{--}0.92 on (a) false findings, but only F1=0.69–0.72F_1 = 0.69\text{--}0.72 on (b) missing findings.
    3. Underestimation of Multi-Error Omissions: In reports where multiple clinically significant missing findings (k∈{2,3}k \in \{2, 3\}) are introduced, LLM judges consistently detect only a subset of omitted findings.
  10. Knowl 10 — Curated vs. Random Few-Shot Prompting for LLM Radiology Evaluation

    empirical result

    Evaluating few-shot prompting variations on 100-report subsets using GPT-4.1-mini shows that randomly sampled in-context examples outperform carefully curated error-specific and rule-described prompt demonstrations:

    Rad-Eval Shots RaTE-Eval Shots
    Dataset 0-shot 3 5 10 rad-err rad-err 10h 3 5 10 rate-err rate-err 10v
    RadEval 0.379 0.499 0.372 0.397 0.390 0.431 0.289 0.367 0.341 0.423 0.398
    RaTE-Eval 0.465 0.468 0.495 0.528 0.476 0.489 0.484 0.504 0.497 0.405 0.492

    On RadEval, 3 random RadEval shots achieve τ=0.499\tau = 0.499 (+12.0% over zero-shot), outperforming 10 human curated shots (τ=0.431\tau = 0.431). On RaTE-Eval, 10 random RadEval shots achieve τ=0.528\tau = 0.528 (+6.3% over zero-shot), outperforming 10 curated VERT shots (τ=0.492\tau = 0.492). Providing larger quantities of examples or hand-crafted category explanations does not monotonically improve evaluator correlation.

Coverage note — Detailed prompt text formatting examples and the multi-turn GPT-4.1 error injection verification subroutines were omitted in favor of the core evaluation metrics, theoretical formulations, and empirical findings.

References

  1. 1.Shruthi Bannur, Kenza Bouzid, Daniel C. Castro, Anton Schwaighofer, Sam Bond-Taylor, Maximilian Ilse, Fernando P’erez-Garc’ia, Valentina Salvatelli, Harshita Sharma, Felix Meissen, Mercy Prasanna Ranjit, Shaury Srivastav, Julia Gong, Fabian Falck, Ozan Oktay, Anja Thieme, Matthew P. Lungren, Maria T. A. Wetscherek, Javier Alvarez-Valle, and Stephanie L. Hyland. Maira-2: Grounded radiology report generation. arXiv, abs/2406.04449, 2024. URL https://arxiv.org/abs/2406.04449.
  2. 2.Gihun Cho, Seunghyun Jang, Hanbin Ko, Inhyeok Baek, and Chang Min Park. Crepe: Rapid chest x-ray report evaluation by predicting multi-category error counts. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 21749–21766, 2025.
  3. 3.Amin Dada, Osman Alperen Koraş, Marie Bauer, Jean-Philippe Corbeil, Amanda Butler Contreras, Constantin Marc Seibold, Kaleb E Smith, Julian Friedrich, and Jens Kleesiek. Does biomedical training lead to better medical performance? In Ofir Arviv, Miruna Clinciu, Kaustubh Dhole, Rotem Dror, Sebastian Gehrmann, Eliya Habba, Itay Itzhak, Simon Mille, Yotam Perlitz, Enrico Santus, Joao Sedoc, Michal Shmueli Scheuer, Gabriel Stanovsky, and Oyvind Tafjord (eds.), Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM²), pp. 46–59, Vienna, Austria and virtual meeting, July 2025. Association for Computational Linguistics. ISBN 979-8-89176-261-9. URL https://aclanthology.org/2025.gem-1.5/.
  4. 4.Jean-Benoit Delbrouck, Justin Xu, Johannes Moll, Alois Thomas, Zhihong Chen, Sophie Ostmeier, Asfandyar Azhar, Kelvin Zhenghao Li, Andrew Johnston, Christian Bluethgen, et al. Automated structured radiology report generation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 26813–26829, 2025.
  5. 5.Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. Qlora: Efficient finetuning of quantized llms. Advances in neural information processing systems, 36:10088–10115, 2023.
  6. 6.Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al. A survey on llm-as-a-judge. The Innovation, 2024.
  7. 7.Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. In International Conference on Learning Representations (ICLR), 2022. URL https://openreview.net/forum?id=nZeVKeeFYf9.
  8. 8.Alyssa Huang, Oishi Banerjee, Kay Wu, Eduardo Pontes Reis, and Pranav Rajpurkar. Fineradscore: A radiology report line-by-line evaluation technique generating corrections with severity scores. In Machine Learning for Healthcare Conference. PMLR, 2024.
  9. 9.Mohamed Sobhi Jabal, Pranav Warman, Jikai Zhang, Kartikeye Gupta, Ayush Jain, Maciej Mazurowski, Walter Wiggins, Kirti Magudia, and Evan Calabrese. Language models and retrieval augmented generation for automated structured data extraction from diagnostic reports. arXiv preprint arXiv:2409.10576, 2024.
  10. 10.Saahil Jain, Ashwin Agrawal, Adriel Saporta, Steven Truong, Du Nguyen Duong, Tan Bui, Pierre Chambon, Yuhao Zhang, Matthew P Lungren, Andrew Y Ng, et al. Radgraph: Extracting clinical entities and relations from radiology reports. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1), 2021.
  11. 11.Woosuk Kwon. vLLM: An Efficient Inference Engine for Large Language Models. PhD thesis, UC Berkeley, 2025.
  12. 12.Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pp. 74–81, 2004.
  13. 13.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. G-eval: Nlg evaluation using gpt-4 with better human alignment. In Proceedings of the 2023 conference on empirical methods in natural language processing, pp. 2511–2522, 2023.
  14. 14.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019.
  15. 15.Harsha Nori, Yin Tat Lee, Sheng Zhang, Dean Carignan, Richard Edgar, Nicolo Fusi, Nicholas King, Jonathan Larson, Yuanzhi Li, Weishung Liu, et al. Can generalist foundation models outcompete special-purpose tuning? case study in medicine. Medicine, 84 (88.3):77–3, 2023.
  16. 16.Sophie Ostmeier, Justin Xu, Zhihong Chen, Maya Varma, Louis Blankemeier, Christian Bluethgen, Arne Edward Michalson Md, Michael Moseley, Curtis Langlotz, Akshay S Chaudhari, et al. Green: Generative radiology report evaluation and error notation. In Findings of the association for computational linguistics: EMNLP 2024, pp. 374–390, 2024.
  17. 17.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp. 311–318, 2002.
  18. 18.Inwoo Seo, Eunkyoung Bae, Joo-Young Jeon, Young-Sang Yoon, and Jiho Cha. The era of foundation models in medical imaging is approaching: A scoping review of the clinical value of large-scale generative ai applications in radiology. arXiv preprint arXiv:2409.12973, 2024.
  19. 19.Akshay Smit, Saahil Jain, Pranav Rajpurkar, Anuj Pareek, Andrew Y Ng, and Matthew Lungren. Combining automatic labelers and expert annotations for accurate radiology report labeling using bert. In Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP), pp. 1500–1519, 2020.
  20. 20.Annalisa Szymanski, Noah Ziems, Heather A Eicher-Miller, Toby Jia-Jun Li, Meng Jiang, and Ronald A Metoyer. Limitations of the llm-as-a-judge approach for evaluating llm outputs in expert knowledge tasks. In Proceedings of the 30th international conference on intelligent user interfaces, pp. 952–966, 2025.
  21. 21.Ming-Liang Wang, Rui-Peng Zhang, Wen-Juan Wu, Yu Lu, Xiao-Er Wei, Zheng Sun, Bao-Hui Guan, Jun-Jie Zhang, Xue Wu, Lei Zhang, Tian-Le Wang, and Yue-Hua Li. Evaluation of large language models for diagnostic impression generation from brain mri report findings: a multicenter benchmark and reader study. npj Digital Medicine, 9:187, 2026. doi: 10.1038/s41746-026-02380-4.
  22. 22.Justin Xu, Xi Zhang, Javid Abderezaei, Julie Bauml, Roger Boodoo, Fatemeh Haghighi, Ali Ganjizadeh, Eric Brattain, Dave Van Veen, Zaiqiao Meng, et al. Radeval: A framework for radiology text evaluation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 546–557, 2025.
  23. 23.An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025.
  24. 24.Feiyang Yu, Mark Endo, Rayan Krishnan, Ian Pan, Andy Tsai, Eduardo Pontes Reis, Eduardo Kaiser Ururahy Nunes Fonseca, Henrique Min Ho Lee, Zahra Shakeri Hossein Abad, Andrew Y Ng, et al. Evaluating progress in automatic chest x-ray radiology report generation. Patterns, 4(9), 2023a.
  25. 25.Feiyang Yu, Mark Endo, Rayan Krishnan, Ian Pan, Andy Tsai, Eduardo Pontes Reis, Eduardo Kaiser Ururahy Nunes Fonseca, Henrique Lee, Zahra Shakeri, Andrew Ng, Curtis Langlotz, Vasantha Kumar Venugopal, and Pranav Rajpurkar. Radiology Report Expert Evaluation (ReXVal) Dataset. PhysioNet, June 2023b. doi: 10.13026/2fp8-qr71. URL https://doi.org/10.13026/2fp8-qr71. Version 1.0.0.
  26. 26.Juan Manuel Zambrano Chaves, Shih-Cheng Huang, Yanbo Xu, Hanwen Xu, Naoto Usuyama, Sheng Zhang, Fei Wang, Yujia Xie, Mahmoud Khademi, Ziyi Yang, et al. A clinically accessible small multimodal radiology model and evaluation metric for chest x-ray findings. Nature Communications, 16(1):3108, 2025.
  27. 27.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations, 2020.
  28. 28.Weike Zhao, Chaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. Ratescore: A metric for radiology report generation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 15004–15019, 2024.
  29. 29.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems, 36:46595–46623, 2023.

Citation

MLA
Bologna, F., et al. “VERT: Reliable LLM Judges for Radiology Report Evaluation”. arXiv, 2026, http://arxiv.org/abs/2604.03376v1.
APA
Bologna, F., Corbeil, J.-P., Wilkens, M., & Abacha, A. B. (2026). VERT: Reliable LLM Judges for Radiology Report Evaluation. arXiv. http://arxiv.org/abs/2604.03376v1
Chicago
Bologna, F., J.-P. Corbeil, M. Wilkens, and A. B. Abacha. 2026. “VERT: Reliable LLM Judges for Radiology Report Evaluation”. arXiv. http://arxiv.org/abs/2604.03376v1.
Harvard
Bologna, F. et al. (2026) “VERT: Reliable LLM Judges for Radiology Report Evaluation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2604.03376v1.
Vancouver
1. Bologna F, Corbeil J-P, Wilkens M, Abacha AB (2026) VERT: Reliable LLM Judges for Radiology Report Evaluation. arXiv

BibTeX

@article{bologna2026vert,
  title = {VERT: Reliable LLM Judges for Radiology Report Evaluation},
  author = {Bologna, Federica and Corbeil, Jean-Philippe and Wilkens, Matthew and Abacha, Asma Ben},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2604.03376v1},
  eprint = {2604.03376}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/