Built independently by an author, for readers. Read the story and support ChapterPal

keyword

RaTE-Eval benchmark

The RaTE-Eval benchmark is an evaluation framework and dataset designed to assess how closely automated evaluation metrics align with expert radiologist judgments when scoring artificial intelligence-generated radiology reports. Developed to improve the measurement of clinical accuracy in medical text generation across diverse anatomical regions and imaging modalities, the benchmark comprises multiple structured evaluation tasks, including sentence-level human ratings, paragraph-level human ratings, and assessments on synthetic reports with systematically introduced errors. By comparing the outputs of natural language generation metrics against verified expert annotations, it provides a standardized testbed to determine whether automated metrics can reliably capture medical entity accuracy, clinical synonyms, and negation in radiological descriptions.

1 item