FRABench and UFEval: Unified Fine-grained Evaluation with Task and Aspect Generalization

ICLROral2026

Authors
Shibo Hong, Jiahao Ying, Haiyuan Liang, Mengdi Zhang, Jun Kuang, Jiazheng Zhang, Yixin Cao
Affiliation
Fudan University
Venue
ICLR 2026
Track
Oral

TL;DR

Evaluating open-ended outputs of Multimodal Large Language Models has become a bottleneck as model capabilities, task diversity, and modality rapidly expand. Existing ``MLLM-as-a-Judge'' evaluators, though promising, remain constrained to specific tasks and aspects (i.e., specific evaluation criteria such as fluency for text and image quality fo...

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model generalization language model evaluation multimodal llm

← All ICLR 2026 Oral papers · Browse the whole archive