keyword
robust automatic VQA metrics
Robust automatic VQA metrics are automated evaluation methods designed to assess the correctness and quality of answers generated by visual question answering models by reliably reflecting human judgment rather than relying solely on strict lexical matching. Unlike traditional accuracy measures that penalize candidate responses for minor phrasing differences, synonyms, or open-ended formatting variations, robust metrics evaluate semantic equivalence and contextual validity against ground-truth references. These evaluation frameworks often employ semantic similarity measures, learned scoring functions, or instruction-tuned language models to handle open-ended text and out-of-distribution predictions, providing a resilient and faithful proxy for human evaluation across diverse multimodal tasks.
1 item

