KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
NeurIPSSpotlight2025
TL;DR
Recent advancements in large language models (LLMs) underscore the need for more comprehensive evaluation methods to accurately assess their reasoning capabilities…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model evaluation reasoning llm
← All NeurIPS 2025 Spotlight papers · Browse the whole archive