FRABench and UFEval: Unified Fine-grained Evaluation with Task and Aspect Generalization
ICLROral2026
TL;DR
Evaluating open-ended outputs of Multimodal Large Language Models has become a bottleneck as model capabilities, task diversity, and modality rapidly expand. Existing ``MLLM-as-a-Judge'' evaluators, though promising, remain constrained to specific tasks and aspects (i.e., specific evaluation criteria such as fluency for text and image quality fo...
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model generalization language model evaluation multimodal llm