VERT: Reliable LLM Judges for Radiology Report Evaluation
Federica BolognaJean-Philippe CorbeilMatthew WilkensAsma Ben Abacha
Introduces VERT, an automated radiology report evaluation metric that achieves superior correlation with expert radiologist judgments across diverse imaging modalities while showing that lightweight fine-tuning can boost evaluation performance and reduce inference time by up to 37 times.
Automated evaluation of radiology reports is essential for scaling quality control in AI-assisted medical imaging, but existing automated metrics have largely been restricted to chest radiographs. This creates a significant blind spot when systems encounter varied clinical settings across different imaging modalities (such as CT and MRI) and diverse anatomical regions. Without robust and broadly applicable evaluation methods, clinical decision-makers cannot reliably benchmark generative AI tools across real-world hospital workflows.
The article systematically assesses how large language models (LLMs) can act as automated judges to evaluate radiology reports across diverse modalities and body regions. It also introduces and validates VERT, a new LLM-based evaluation metric designed to score reports with higher clinical reliability and alignment to expert radiologists.
The authors conducted empirical evaluations across two multi-modality, expert-annotated radiology benchmarks: RadEval (148 chest X-ray cases annotated by error counts) and RaTE-Eval (1,856 report pairs covering 9 imaging modalities and 22 anatomies). The study compared several prompting strategies, proprietary and open-source models, reasoning modes, few-shot variations, model ensembling, parameter-efficient fine-tuning (LoRA), and controlled clinical error injections.
Direct continuous accuracy scoring via VERT outperformed prior automated metrics, improving correlation with expert judgments by up to 11.7% relative to the standard GREEN metric. Parameter-efficient fine-tuning of an open-source model (Qwen3 30B) on roughly 1,300 samples yielded correlation gains of up to 25% while speeding up inference time by 37.2 times compared to proprietary API baselines. However, longer extended reasoning traces generally failed to improve correlation with human radiologists, and automated judges frequently underestimated error counts when reports contained high error density or nuanced errors such as incorrect anatomical locations and omitted comparisons.
These findings demonstrate that lightweight adaptation of open-source models offers a practical, high-throughput path for automated radiology quality assurance while cutting operational costs and API latency. However, because current automated judges struggle to detect complex clinical nuances such as location errors or temporal changes, organizations cannot yet rely on fully autonomous evaluation in safety-critical clinical deployments.
Healthcare technology leaders should prioritize direct-scoring prompts and domain-adapted open-source models for offline benchmarking and automated triage. For production deployments, leaders should maintain human-in-the-loop validation for reports with high error potential or complex comparative findings. Further research is recommended to improve automated detection of multi-category errors without requiring matched ground-truth reference reports.
While the study provides strong empirical evidence across multiple modalities, results are bounded by the reliance on two reference datasets and synthetic error validation steps. Decision-makers should maintain high confidence in using these approaches for large-scale comparative benchmarking, but exercise caution before replacing radiologist review in live clinical oversight.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). Provides a comprehensive foundational survey on the LLM-as-a-Judge paradigm, framing the methodologies, biases, and alignment evaluations critical for automated judging systems.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). Establishes a systematic taxonomy of techniques, attributes, and prompting/tuning strategies for deploying LLMs as judges across complex evaluation scenarios.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). Introduces form-based prompting and chain-of-thought scoring with LLM evaluators to improve alignment with human judgment, establishing a baseline methodology for LLM-based metrics.
- Paper: Large Language Models are not Fair Evaluators, Peiyi Wang et al. (2024). Analyzes severe positional biases in LLM evaluators and outlines essential calibration strategies for reliable comparative scoring.
- Paper: TRUE: Re-evaluating Factual Consistency Evaluation, Or Honovich et al. (2022). Formulates standardized benchmarking for factual consistency evaluation across text generation systems, directly motivating domain-specific factual evaluation in radiology.
- Paper: CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison, Jeremy Irvin et al. (2019). Introduces foundational chest radiograph dataset curation and expert-labeled clinical observation extraction protocols for evaluating radiology interpretations.
No sufficiently relevant recommendations were found.
