RaTEScore: A Metric for Radiology Report Generation

Weike ZhaoChaoyi WuXiaoman ZhangYa ZhangYanfeng WangWeidi Xie

article2024EMNLP89 citations

Proposes an entity-aware evaluation metric and comprehensive medical named entity dataset that measure the clinical accuracy of AI-generated radiology reports by handling medical synonyms and negation expressions far better than traditional text-generation metrics.

Listen

The rapid development of generative artificial intelligence for interpreting medical imaging requires accurate automated evaluation methods to verify the clinical quality and safety of generated reports. Standard text evaluation tools rely on surface word overlap or general language embeddings, failing to recognize medical synonyms, detect clinical negations, or prioritize critical diagnostic findings. Meanwhile, large language model evaluators are computationally expensive and prone to subjective bias, and existing domain-specific metrics remain narrow, focusing primarily on chest X-rays. To address this bottleneck, the article introduces RaTEScore, a lightweight, entity-aware evaluation metric designed to measure clinical consistency across diverse medical imaging modalities and body regions.

The framework evaluates reports through three core stages: extracting clinical entities using a dedicated medical named-entity recognition model across five categories (anatomy, abnormality, disease, non-abnormality, and non-disease); embedding these entities using a medical synonym disambiguation module; and calculating an entity-level similarity score that penalizes polarity mismatches and weights entity pairs by their clinical significance. To support this framework, the authors created two public resources: RaTE-NER, a dataset of over 40,000 sentences across 9 imaging modalities and 22 anatomical regions drawn from MIMIC-IV and Radiopaedia, and RaTE-Eval, a multi-task benchmark containing sentence-level, paragraph-level, and synthetic test suites annotated by experienced radiologists.

Key findings show that RaTEScore consistently outperforms existing evaluation metrics in aligning with expert clinical judgment. On the public ReXVal chest X-ray benchmark, RaTEScore achieved a leading Kendall correlation of 0.527 against baseline metrics. On the multi-modal RaTE-Eval benchmark, it achieved the highest correlation on both sentence-level ratings (Pearson correlation of 0.54) and paragraph-level ratings (Pearson correlation of 0.653 and Kendall correlation of 0.462). Furthermore, on synthetic test sets designed to assess semantic robustness, RaTEScore correctly distinguished between synonymous rewrites and negated reports with 67.0% accuracy, substantially surpassing standard natural language processing baselines such as BERTScore (14.0%) and BLEU (11.9%).

These findings provide immediate practical value for deploying clinical artificial intelligence. By reliably penalizing critical diagnostic contradictions and handling diverse medical vocabularies, RaTEScore reduces clinical risk and offers a low-cost, explainable alternative to expensive human grading or closed commercial language models. Organizations developing medical report generation systems should adopt RaTEScore and its open benchmark to monitor model safety and performance. However, because the system was evaluated strictly within radiological text and utilizes an off-the-shelf embedding encoder without task-specific fine-tuning, decision-makers should exercise caution before extending it to broader healthcare domains, such as clinical summarization or medical question answering, until further domain-specific validation is conducted.

arXiv: 2406.16845
Cover for RaTEScore: A Metric for Radiology Report Generation

Table of Contents

  • 1 Introduction
  • 2 Methods
  • 2.1 General Pipeline
  • 2.2 Medical Named Entity Recognition
  • 2.3 Synonym Disambiguation Encoding
  • 2.4 Scoring Procedure
  • 2.5 Implementation Details
  • 3 RaTE-Eval Benchmark
  • 4 Experiments
  • 4.1 Baselines
  • 4.2 Results in ReXVal dataset
  • 4.3 Results in RaTE-Eval benchmark
  • 4.4 Ablation Study
  • 4.4.1 NER Module Discussion
  • 4.4.2 Entity Encoding Module Discussion
  • 5 Related Work
  • 5.1 General Text Evaluation Metric
  • 5.2 Radiological Text Evaluation Metric
  • 5.3 Medical Named-Entity Recognition
  • 6 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Scoring Example
  • A.2 Automatic Annotation Approach
  • A.3 Involving Anatomies and Modalities in MIMIC-IV Data
  • A.4 Guidelines for Radiologists
  • A.5 Example for Simulation Reports
  • A.6 Baselines
  • A.7 Failure Cases in ReXVal Dataset
  • A.8 Pretrained BERT Model Introduction
  • A.9 NER Module Implementation Details

Knowls

  1. Knowl 1 — RaTEScore matches and weights report entities in both directions

    equation

    RaTEScore compares a reference radiology report xx with a candidate report x^\hat{x} by matching each candidate entity to its most similar reference entity. An entity is a name–type pair; let the reference contain MM entities (ni,ti)(n_i,t_i) with embedding vectors fif_i, and the candidate contain NN entities (n^j,t^j)(\hat n_j,\hat t_j) with embedding vectors f^j\hat f_j. Cosine similarity is denoted by cos⁡\operatorname{cos}, and i∗(j)i^*(j) is the index of the reference entity whose embedding has the highest cosine similarity to candidate entity jj:

    i∗(j)=arg⁡max⁡1≤i≤Mcos⁡(fi,f^j).i^*(j)=\arg\max_{1\leq i\leq M}\operatorname{cos}(f_i,\hat f_j).

    A learnable 5×55\times5 affinity matrix WW weights pairs of entity types. A learnable mismatch factor pp penalizes matches whose entity types differ. The candidate-to-reference score is the affinity-weighted mean of the matched entity similarities:

    S(x,x^)=∑j=1NW(ti∗(j),t^j) sim⁡(ei∗(j),e^j)∑j=1NW(ti∗(j),t^j),sim⁡(ei,e^j)={p cos⁡(fi,f^j),ti≠t^j,cos⁡(fi,f^j),ti=t^j.S(x,\hat{x})= \frac{\displaystyle\sum_{j=1}^{N}W(t_{i^*(j)},\hat t_j)\,\operatorname{sim}(e_{i^*(j)},\hat e_j)} {\displaystyle\sum_{j=1}^{N}W(t_{i^*(j)},\hat t_j)}, \qquad \operatorname{sim}(e_i,\hat e_j)= \begin{cases} p\,\operatorname{cos}(f_i,\hat f_j), & t_i\ne\hat t_j,\\ \operatorname{cos}(f_i,\hat f_j), & t_i=\hat t_j. \end{cases}

    Because the matching direction changes which report supplies the entities being averaged, S(x,x^)S(x,\hat{x}) and S(x^,x)S(\hat{x},x) need not be equal. RaTEScore combines them as an F1-style harmonic mean, with score zero when both directional scores sum to zero:

    RaTEScore⁡(x,x^)={0,S(x,x^)+S(x^,x)=0,2S(x,x^)S(x^,x)S(x,x^)+S(x^,x),otherwise.\operatorname{RaTEScore}(x,\hat{x})= \begin{cases} 0, & S(x,\hat{x})+S(\hat{x},x)=0,\\ \displaystyle\frac{2S(x,\hat{x})S(\hat{x},x)}{S(x,\hat{x})+S(\hat{x},x)}, & \text{otherwise}. \end{cases}

    The type-affinity matrix and mismatch factor are fitted using a small amount of human rating data with tree-structured Parzen estimator hyperparameter search. The type weights are intended to reflect clinical importance, while the mismatch factor reduces the credit for semantically similar entities whose types differ, including positive versus negated findings.

  2. Knowl 2 — Five entity types distinguish anatomy, findings, diagnoses, and negation

    definition

    RaTEScore represents each extracted medical entity as a pair consisting of its name and one of five types: Anatomy, Abnormality, Non-Abnormality, Disease, or Non-Disease. Anatomy labels identify body structures. Abnormality labels cover notable imaging findings such as masses, effusions, and edema; Non-Abnormality marks an abnormality expressed as absent or negated, as in “no evidence of pleural effusion.” Disease denotes higher-level diagnostic conclusions, such as pneumonia or lymphadenopathy, while Non-Disease marks a negated disease entity. This scheme makes negation part of the entity representation rather than treating a positive finding and its negated form as interchangeable.

  3. Knowl 3 — The metric uses token-level NER and BioLORD entity embeddings

    model/method

    RaTEScore first applies a medical named-entity recognition model to a report and extracts typed entities, then encodes each entity name into a vector before comparing reports. The selected NER system is an IOB token-classification model initialized with DeBERTa-v3: each token receives a beginning-of-entity, inside-of-entity, or outside-of-entity label, predicted using a classification layer on the model's output representations. It was trained for 10 epochs with batch size 96 and learning rate 10−510^{-5} on one NVIDIA GeForce GTX 3090 GPU. For entity-name encoding, the method uses the off-the-shelf BioLORD-2023-C model, which was trained on medical entity–concept information. The resulting vectors allow semantically related medical names to receive high cosine similarity even when their surface forms differ.

  4. Knowl 4 — RaTE-NER combines manually annotated and automatically enriched radiology text

    data/table

    RaTE-NER supplies the typed entity annotations used to train and evaluate the metric's NER component. Its described sources are 13,235 manually annotated sentences from 1,816 MIMIC-IV reports, covering 9 imaging modalities and 23 anatomical regions, and 33,605 sentences from 17,432 Radiopaedia reports used to broaden coverage of rarer conditions. The authors also report manually labeling 3,529 sentences for testing. The dataset's entity-level split totals are:

    Split Entity annotations Non-duplicate entities
    Train 107,909 60,682
    Dev 13,065 9,800
    Test 13,159 9,434

    For the Radiopaedia-derived material, GPT-4 was prompted to extract entities from reports, and the extracted names were checked against medical knowledge resources including UMLS, SNOMED CT, and ICD-10 using MedCPT similarity. Candidate matches with cosine similarity below 0.83 were filtered out, as were sentences whose entity-annotation density was below 0.7. Negation and polarity were identified using medspaCy and negative-context terms such as “no,” “without,” “unremarkable,” and “intact.” This combination of sources and filtering was designed to expand beyond common findings while retaining a manually annotated test component.

  5. Knowl 5 — RaTE-Eval sentence ratings normalize errors by opportunities for error

    experimental setup

    The sentence-level component of RaTE-Eval contains 2,215 reference–candidate sentence pairs from MIMIC-IV and covers 9 imaging modalities and 22 anatomical regions. The authors divided the data into 49 modality–anatomy subsets, split reports into sentences, and removed duplicate sentences. For each subset they sampled 25 reference sentences and drew 1,000 candidate reports; BLEU, ROUGE, BERTScore, CIDEr, and RaTEScore were used to retrieve similar candidates, and the union of their retrieved pairs was sent for annotation. Two radiologists with more than five years of clinical practice rated each pair by counting candidate errors in six categories: false finding, omitted finding, wrong location or position, wrong severity, an unsupported comparison in the impression, and omission of a comparison describing change from a prior study.

    The rating divides the total error count by the number of potential errors, where potential errors are counted from both correct and incorrect findings in the reference sentence. This normalizes error counts for differences in sentence complexity. For metric parameter search, the annotated pairs were split 8:2 into training and test sets.

  6. Knowl 6 — RaTE-Eval adds paragraph ratings and controlled synonym–opposite pairs

    experimental setup

    RaTE-Eval also evaluates full paragraphs and controlled rewrites. Its paragraph-level component contains 1,856 MIMIC-IV reference–candidate report pairs. Radiologists assign a five-point score based on clinical judgment rather than attempting exhaustive error counts: 5 indicates that most diagnoses are correct and descriptions largely agree; 4 indicates about 75% of diagnoses are correct; 3 indicates about 50%; 2 indicates about 25%; 1 indicates an incorrect diagnosis, possibly with some matching negative descriptions; and 0 indicates no overlap in described information. The report sampling uses modality–anatomy subsets and the candidate-selection process used for sentence-level evaluation; the pairs are split 8:2 for parameter search.

    For a separate test of synonym and negation sensitivity, Mixtral 8x7B rewrote 847 MIMIC-IV reports under two instructions: preserve meaning while potentially replacing entities with synonyms, or rewrite the report to express the opposite meaning. This creates original, meaning-preserving synonym, and opposite-meaning versions. The intended comparison is whether a metric scores the synonym rewrite above the opposite-meaning rewrite.

  7. Knowl 7 — RaTEScore has the strongest reported alignment across the human-rated evaluations

    empirical result

    RaTEScore outperformed the reported baselines on ReXVal and on RaTE-Eval's sentence- and paragraph-level human ratings. ReXVal agreement is Kendall's τ\tau with radiologist error counts; sentence-level RaTE-Eval agreement is reported as Pearson correlation with normalized human ratings; paragraph-level results include Pearson, Kendall, and Spearman correlations. The table also gives accuracy on the synthetic test, defined as correctly scoring the synonym-preserving rewrite above its opposite-meaning counterpart.

    Metric ReXVal Kendall τ\tau Sentence Pearson Paragraph Pearson Paragraph Kendall τ\tau Paragraph Spearman τ\tau Synthetic accuracy
    RaTEScore 0.527 0.54 0.653 0.462 0.608 0.670
    RadGraph F1 0.515* 0.44 0.624 0.439 0.582 0.463
    BERTScore 0.511* 0.40 0.599 0.413 0.555 0.140
    CheXbert 0.499* 0.25 0.496 0.294 0.403 0.666
    BLEU 0.462* 0.27 0.409 0.289 0.404 0.119
    ROUGE-L – 0.34 0.572 0.396 0.567 0.117
    SPICE – 0.40 0.623 0.453 0.605 0.140
    METEOR – 0.39 0.599 0.422 0.567 0.168
    CIDEr – 0.25 – – – –

    On ReXVal, RaTEScore's Kendall correlation was 0.527, above the reported RadGraph F1, BERTScore, CheXbert, and BLEU values. On RaTE-Eval, it achieved the highest reported sentence correlation and the highest paragraph correlation in each of the three reported measures. Its synthetic accuracy of 0.670 was also higher than the other listed metrics. An asterisk marks ReXVal baseline results reported by the prior ReXVal study.

  8. Knowl 8 — Combining MIMIC-IV and Radiopaedia improves NER test performance

    empirical result

    The authors ablated both NER training architecture and data source on the RaTE-NER test set. The reported precision, recall, F1, and accuracy values are:

    Training scheme Initialization or training data Precision Recall F1 Accuracy
    IOB DeBERTa-v3 0.567 0.575 0.571 0.754
    IOB Medical-NER 0.559 0.572 0.565 0.759
    Span BioMedBERT 0.556 0.676 0.610 0.730
    Span SapBERT 0.560 0.658 0.605 0.731
    Span BlueBERT 0.554 0.657 0.601 0.726
    Span MedCPT-Q-Enc. 0.470 0.682 0.556 0.678
    Span BioLORD-2023-C 0.555 0.664 0.605 0.727
    Radiopaedia only – 0.525 0.558 0.541 0.727
    MIMIC-IV only – 0.515 0.550 0.531 0.744
    Radiopaedia + MIMIC-IV – 0.567 0.575 0.571 0.754

    The combined training data achieves higher F1 and accuracy than either source alone. DeBERTa-v3 with IOB labeling is the NER configuration used for the final metric, although the table shows that some span-based configurations have higher recall or F1. Medical-NER has the highest listed accuracy, 0.759, while the selected DeBERTa-v3 IOB model has F1 0.571.

  9. Knowl 9 — BioLORD performs best among the tested entity encoders for sentence correlation

    empirical result

    The entity-encoding ablation compares off-the-shelf encoders on sentence-level correlation with human ratings in RaTE-Eval. The reported values are BioLORD-2023-C, 0.540; MedCPT, 0.498; RadBERT, 0.519; CXR-BERT, 0.368; and BioViL-T, 0.465. BioLORD-2023-C has the highest score among the tested encoders and is the model selected for RaTEScore's entity-name embeddings. The authors attribute its suitability to training that targets medical entity normalization, which matches the encoder's role in mapping synonymous entity names to similar vectors.

  10. Knowl 10 — The evaluation scope remains limited to radiology and uses an untuned encoder

    limitation

    The authors identify two limitations. First, they selected an existing synonym-disambiguation encoder after comparing available models but did not fine-tune it specifically for the report-evaluation task. Second, although RaTEScore covers multiple radiology modalities and body regions, it is still designed for radiology reports; its applicability to other medical text settings and tasks such as medical question answering or summarization remains unestablished.

Coverage note — The per-entity-type train/dev/test frequency breakdown is omitted because the aggregate dataset counts and five-type label definitions capture the dataset's role without reproducing a long distribution table.

References

  1. 1.ICD-10-CM. https://www.icd10data.com/ICD10CM/Codes. Accessed: Dec.2023.
  2. 2.Radiopaedia.org. https://radiopaedia.org. Accessed: May 2023.
  3. 3.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35:23716–23736.
  4. 4.Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016. Spice: Semantic propositional image caption evaluation. In Proceedings of European Conference on Computer Vision (ECCV), pages 382–398.
  5. 5.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403.
  6. 6.Satanjeev Banerjee and Alon Lavie. 2005. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the Acl Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72.
  7. 7.Shruthi Bannur, Stephanie Hyland, Qianchu Liu, Fernando Perez-Garcia, Maximilian Ilse, Daniel C Castro, Benedikt Boecking, Harshita Sharma, Kenza Bouzid, Anja Thieme, et al. 2023. Learning to exploit temporal structure for biomedical vision-language processing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15016–15027.
  8. 8.James Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl. 2011. Algorithms for hyper-parameter optimization. Advances in Neural Information Processing Systems, 24.
  9. 9.Olivier Bodenreider. 2004. The unified medical language system (umls): integrating biomedical terminology. Nucleic Acids Research, 32(suppl_1):D267–D270.
  10. 10.Benedikt Boecking, Naoto Usuyama, Shruthi Bannur, Daniel C Castro, Anton Schwaighofer, Stephanie Hyland, Maria Wetscherek, Tristan Naumann, Aditya Nori, Javier Alvarez-Valle, et al. 2022. Making the most of text semantics to improve biomedical vision–language processing. In Proceedings of European Conference on Computer Vision (ECCV), pages 1–21. Springer.
  11. 11.Kathi Canese and Sarah Weis. 2013. Pubmed: the bibliographic database. The NCBI Handbook, 2(1).
  12. 12.Souradip Chakraborty, Ekaba Bisong, Shweta Bhatt, Thomas Wagner, Riley Elliott, and Francesco Mosconi. 2020. Biomedbert: A pre-trained biomedical language model for qa and ir. In Proceedings of the 28th international conference on computational linguistics, pages 669–679.
  13. 13.Pierre Chambon, Tessa S Cook, and Curtis P Langlotz. 2023. Improved fine-tuning of in-domain transformer model for inferring covid-19 presence in multi-institutional radiology reports. Journal of Digital Imaging, 36(1):164–177.
  14. 14.Peng Chen, Jian Wang, Hongfei Lin, Di Zhao, and Zhihao Yang. 2023. Few-shot biomedical named entity recognition via knowledge-guided instance generation and prompt contrastive learning. Bioinformatics, 39(8):btad496.
  15. 15.Clinical-AI-Apollo. 2023. Clinical-AI-Apollo Medical-NER. HuggingFace.
  16. 16.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186.
  17. 17.Kevin Donnelly et al. 2006. Snomed-ct: The advanced terminology and coding system for ehealth. Studies in Health Technology and Informatics, 121:279.
  18. 18.Hannah Eyre, Alec B Chapman, Kelly S Peterson, Jianlin Shi, Patrick R Alba, Makoto M Jones, Tamara L Box, Scott L DuVall, and Olga V Patterson. 2021. Launching into clinical space with medspacy: a new clinical text processing toolkit in python. In AMIA Annual Symposium Proceedings, volume 2021, page 438.
  19. 19.Christiane Fellbaum. 2010. Wordnet. In Theory and Applications of Ontology: Computer Applications, pages 231–243.
  20. 20.Shlomit Goldberg-Stein, L Alexandre Frigini, Scott Long, Zeyad Metwalli, Xuan V Nguyen, Mark Parker, and Hani Abujudeh. 2017. Acr radpeer committee white paper with 2016 updates: revised scoring system, new classifications, self-review, and subspecialized reports. Journal of the American College of Radiology, 14(8):1080–1086.
  21. 21.Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2022. Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing. In The Eleventh International Conference on Learning Representations.
  22. 22.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020. Deberta: Decoding-enhanced bert with disentangled attention. In International Conference on Learning Representations.
  23. 23.Saahil Jain, Ashwin Agrawal, Adriel Saporta, Steven Truong, Tan Bui, Pierre Chambon, Yuhao Zhang, Matthew P Lungren, Andrew Y Ng, Curtis Langlotz, et al. Radgraph: Extracting clinical entities and relations from radiology reports. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1).
  24. 24.Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088.
  25. 25.Qiao Jin, Won Kim, Qingyu Chen, Donald C Comeau, Lana Yeganova, W John Wilbur, and Zhiyong Lu. 2023. Medcpt: Contrastive pre-trained transformers with large-scale pubmed search logs for zero-shot biomedical information retrieval. Bioinformatics, 39(11):btad651.
  26. 26.Alistair Johnson, Lucas Bulgarelli, Tom Pollard, Steven Horng, Leo Anthony Celi, and Roger Mark. 2020. Mimic-iv. PhysioNet. Available online at: https://physionet. org/content/mimiciv/1.0/(accessed August 23, 2021), pages 49–55.
  27. 27.Alistair EW Johnson, Tom J Pollard, Seth J Berkowitz, Nathaniel R Greenbaum, Matthew P Lungren, Chih-ying Deng, Roger G Mark, and Steven Horng. 2019. Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports. Scientific Data, 6(1):317.
  28. 28.Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. 2016. Mimic-iii, a freely accessible critical care database. Scientific Data, 3(1):1–9.
  29. 29.Vipina K Keloth, Yan Hu, Qianqian Xie, Xueqing Peng, Yan Wang, Andrew Zheng, Melih Selek, Kalpana Raja, Chih Hsuan Wei, Qiao Jin, et al. 2024. Advancing entity recognition in biomedicine via instruction tuning of large language models. Bioinformatics, 40(4):btae163.
  30. 30.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International Conference on Machine Learning, pages 19730–19742.
  31. 31.Mingchen Li and Rui Zhang. 2023. How far is language model from 100% few-shot named entity recognition in medical domain. arXiv preprint arXiv:2307.00186.
  32. 32.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81.
  33. 33.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: Nlg evaluation using gpt-4 with better human alignment. In Processing of the 2023 Conference on Empirical Methods in Natural Language (EMNLP).
  34. 34.Masoud Monajatipoor, Jiaxin Yang, Joel Stremmel, Melika Emami, Fazlolah Mohaghegh, Mozhdeh Rouhsedaghat, and Kai-Wei Chang. 2024. Llms in biomedicine: A study on clinical named entity recognition. arXiv preprint arXiv:2404.07376.
  35. 35.Michael Moor, Oishi Banerjee, Zahra Shakeri Hossein Abad, Harlan M Krumholz, Jure Leskovec, Eric J Topol, and Pranav Rajpurkar. 2023. Foundation models for generalist medical artificial intelligence. Nature, 616(7956):259–265.
  36. 36.OpenAI. Gpt-4v(ision) system card.
  37. 37.OpenAI. 2023. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774.
  38. 38.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  39. 39.Yifan Peng, Shankai Yan, and Zhiyong Lu. 2019. Transfer learning in biomedical natural language processing: An evaluation of bert and elmo on ten benchmarking datasets. In Proceedings of the 18th BioNLP Workshop and Shared Task, pages 58–65.
  40. 40.Pengcheng Qiu, Chaoyi Wu, Xiaoman Zhang, Weixiong Lin, Haicheng Wang, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2024. Towards building multilingual language model for medicine. arXiv preprint arXiv:2402.13963.
  41. 41.François Remy, Kris Demuynck, and Thomas Demeester. 2024. Biolord-2023: semantic textual representations fusing large language models and clinical knowledge graph insights. Journal of the American Medical Informatics Association, page ocae029.
  42. 42.Akshay Smit, Saahil Jain, Pranav Rajpurkar, Anuj Pareek, Andrew Y Ng, and Matthew Lungren. 2020. Combining automatic labelers and expert annotations for accurate radiology report labeling using bert. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1500–1519.
  43. 43.Luca Soldaini and Nazli Goharian. 2016. Quickumls: a fast, unsupervised approach for medical concept extraction. In MedIR Workshop, Sigir, pages 1–4.
  44. 44.Tao Tu, Shekoofeh Azizi, Danny Driess, Mike Schaekermann, Mohamed Amin, Pi-Chuan Chang, Andrew Carroll, Charles Lau, Ryutaro Tanno, Ira Ktena, et al. 2024. Towards generalist biomedical ai. NEJM AI, 1(3):AIoa2300138.
  45. 45.Zilong Wang, Xufang Luo, Xinyang Jiang, Dongsheng Li, and Lili Qiu. 2024. Llm-radjudge: Achieving radiologist-level evaluation for x-ray report generation. arXiv preprint arXiv:2404.00998.
  46. 46.Jerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu, Nathan Hu, Dustin Tran, Daiyi Peng, Ruibo Liu, Da Huang, Cosmo Du, et al. 2024. Long-form factuality in large language models. arXiv preprint arXiv:2403.18802.
  47. 47.Chaoyi Wu, Jiayu Lei, Qiaoyu Zheng, Weike Zhao, Weixiong Lin, Xiaoman Zhang, Xiao Zhou, Ziheng Zhao, Ya Zhang, Yanfeng Wang, et al. 2023a. Can gpt-4v (ision) serve medical applications? case studies on gpt-4v for multimodal medical diagnosis. arXiv preprint arXiv:2310.09909.
  48. 48.Chaoyi Wu, Weixiong Lin, Xiaoman Zhang, Ya Zhang, Weidi Xie, and Yanfeng Wang. 2024. Pmc-llama: toward building open-source language models for medicine. Journal of the American Medical Informatics Association, page ocae045.
  49. 49.Chaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023b. Towards generalist foundation model for radiology by leveraging web-scale 2d&3d medical data. arXiv preprint arXiv:2308.02463.
  50. 50.Wen-wai Yim, Yujuan Fu, Asma Ben Abacha, Neal Snider, Thomas Lin, and Meliha Yetisgen. 2023. Acibench: a novel ambient clinical intelligence dataset for benchmarking automatic visit note generation. Scientific Data, 10(1):586.
  51. 51.Feiyang Yu, Mark Endo, Rayan Krishnan, Ian Pan, Andy Tsai, Eduardo Pontes Reis, Eduardo Kaiser Ururahy Nunes Fonseca, Henrique Min Ho Lee, Zahra Shakeri Hossein Abad, Andrew Y Ng, et al. 2023a. Evaluating progress in automatic chest x-ray radiology report generation. Patterns, 4(9).
  52. 52.Feiyang Yu, Mark Endo, Rayan Krishnan, Ian Pan, Andy Tsai, Eduardo Pontes Reis, EKU Fonseca, Henrique Lee, Zahra Shakeri, Andrew Ng, et al. 2023b. Radiology report expert evaluation (rexval) dataset.
  53. 53.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.
  54. 54.Xiaoman Zhang, Chaoyi Wu, Ziheng Zhao, Weixiong Lin, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023. Pmc-vqa: Visual instruction tuning for medical visual question answering. arXiv preprint arXiv:2305.10415.
  55. 55.Ziheng Zhao, Yao Zhang, Chaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023. One model to rule them all: Towards universal segmentation for medical images with text prompts. arXiv preprint arXiv:2312.17183.
  56. 56.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2024. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36.
  57. 57.Qiaoyu Zheng, Weike Zhao, Chaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023. Large-scale long-tailed disease diagnosis on radiology images. arXiv preprint arXiv:2312.16151.
  58. 58.Zexuan Zhong and Danqi Chen. 2021. A frustratingly easy approach for entity and relation extraction. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 50–61.
  59. 59.Xiao Zhou, Xiaoman Zhang, Chaoyi Wu, Ya Zhang, Weidi Xie, and Yanfeng Wang. 2024. Knowledge-enhanced visual-language pretraining for computational pathology. arXiv preprint arXiv:2404.09942.

Citation

MLA
Zhao, W., et al. “RaTEScore: A Metric for Radiology Report Generation”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 15004–19, https://doi.org/10.18653/v1/2024.emnlp-main.836.
APA
Zhao, W., Wu, C., Zhang, X., Zhang, Y., Wang, Y., & Xie, W. (2024). RaTEScore: A Metric for Radiology Report Generation. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 15004–15019. https://doi.org/10.18653/v1/2024.emnlp-main.836
Chicago
Zhao, W., C. Wu, X. Zhang, Y. Zhang, Y. Wang, and W. Xie. 2024. “RaTEScore: A Metric for Radiology Report Generation”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 15004–19. https://doi.org/10.18653/v1/2024.emnlp-main.836.
Harvard
Zhao, W. et al. (2024) “RaTEScore: A Metric for Radiology Report Generation”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 15004–15019. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.836.
Vancouver
1. Zhao W, Wu C, Zhang X, Zhang Y, Wang Y, Xie W (2024) RaTEScore: A Metric for Radiology Report Generation. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 15004–15019

BibTeX

@inproceedings{zhao-etal-2024-ratescore,
    title = "{R}a{TES}core: A Metric for Radiology Report Generation",
    author = "Zhao, Weike  and
      Wu, Chaoyi  and
      Zhang, Xiaoman  and
      Zhang, Ya  and
      Wang, Yanfeng  and
      Xie, Weidi",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.836/",
    doi = "10.18653/v1/2024.emnlp-main.836",
    pages = "15004--15019"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/