CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering

ICLROral2026

Authors
Yahan Li, Jifan Yao, John Bosco S. Bunyi, Adam C Frank, Angel Hsing-Chi Hwang, Ruishan Liu
Venue
ICLR 2026
Track
Oral

TL;DR

Medical question answering (QA) benchmarks often focus on multiple-choice or fact-based tasks, leaving open-ended answers to real patient questions underexplored. This gap is particularly critical in mental health, where patient questions often mix symptoms, treatment concerns, and emotional needs, requiring answers that balance clinical caution...

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model adversarial evaluation benchmark

← All ICLR 2026 Oral papers · Browse the whole archive