RaTEScore: A Metric for Radiology Report Generation
Weike ZhaoChaoyi WuXiaoman ZhangYa ZhangYanfeng WangWeidi Xie
Proposes an entity-aware evaluation metric and comprehensive medical named entity dataset that measure the clinical accuracy of AI-generated radiology reports by handling medical synonyms and negation expressions far better than traditional text-generation metrics.
The rapid development of generative artificial intelligence for interpreting medical imaging requires accurate automated evaluation methods to verify the clinical quality and safety of generated reports. Standard text evaluation tools rely on surface word overlap or general language embeddings, failing to recognize medical synonyms, detect clinical negations, or prioritize critical diagnostic findings. Meanwhile, large language model evaluators are computationally expensive and prone to subjective bias, and existing domain-specific metrics remain narrow, focusing primarily on chest X-rays. To address this bottleneck, the article introduces RaTEScore, a lightweight, entity-aware evaluation metric designed to measure clinical consistency across diverse medical imaging modalities and body regions.
The framework evaluates reports through three core stages: extracting clinical entities using a dedicated medical named-entity recognition model across five categories (anatomy, abnormality, disease, non-abnormality, and non-disease); embedding these entities using a medical synonym disambiguation module; and calculating an entity-level similarity score that penalizes polarity mismatches and weights entity pairs by their clinical significance. To support this framework, the authors created two public resources: RaTE-NER, a dataset of over 40,000 sentences across 9 imaging modalities and 22 anatomical regions drawn from MIMIC-IV and Radiopaedia, and RaTE-Eval, a multi-task benchmark containing sentence-level, paragraph-level, and synthetic test suites annotated by experienced radiologists.
Key findings show that RaTEScore consistently outperforms existing evaluation metrics in aligning with expert clinical judgment. On the public ReXVal chest X-ray benchmark, RaTEScore achieved a leading Kendall correlation of 0.527 against baseline metrics. On the multi-modal RaTE-Eval benchmark, it achieved the highest correlation on both sentence-level ratings (Pearson correlation of 0.54) and paragraph-level ratings (Pearson correlation of 0.653 and Kendall correlation of 0.462). Furthermore, on synthetic test sets designed to assess semantic robustness, RaTEScore correctly distinguished between synonymous rewrites and negated reports with 67.0% accuracy, substantially surpassing standard natural language processing baselines such as BERTScore (14.0%) and BLEU (11.9%).
These findings provide immediate practical value for deploying clinical artificial intelligence. By reliably penalizing critical diagnostic contradictions and handling diverse medical vocabularies, RaTEScore reduces clinical risk and offers a low-cost, explainable alternative to expensive human grading or closed commercial language models. Organizations developing medical report generation systems should adopt RaTEScore and its open benchmark to monitor model safety and performance. However, because the system was evaluated strictly within radiological text and utilizes an off-the-shelf embedding encoder without task-specific fine-tuning, decision-makers should exercise caution before extending it to broader healthcare domains, such as clinical summarization or medical question answering, until further domain-specific validation is conducted.
- Paper: On the Blind Spots of Model-Based Evaluation Metrics for Text Generation, Tianxing He et al. (2023). Its stress tests expose how text-generation metrics can miss serious semantic errors, motivating RaTEScore’s explicit checks for clinical contradictions and robustness.
- Paper: TRUE: Re-evaluating Factual Consistency Evaluation, Or Honovich et al. (2022). TRUE establishes factual consistency as a distinct evaluation problem, a key prerequisite for understanding RaTEScore’s emphasis on whether clinical findings and their polarity agree.
- Paper: Towards a Unified Multi-Dimensional Evaluator for Text Generation, Ming Zhong et al. (2022). UNIEVAL shows how automated evaluation can separate quality dimensions rather than rely on surface similarity, framing RaTEScore’s entity-level clinical scoring.
- Paper: Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text, Sebastian Gehrmann et al. (2023). This survey maps the shortcomings of text-generation metrics and evaluation practices that RaTEScore addresses with clinically grounded scoring and expert validation.
- Paper: VERT: Reliable LLM Judges for Radiology Report Evaluation, Federica Bologna et al. (2026). VERT directly continues RaTEScore’s radiology-evaluation work by using the RaTE-Eval benchmark to develop and assess an LLM-based alternative against expert judgments.
