ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data

ICLROral2026

Authors
Zhaoyang Liu, JingJing Xie, Zichen Ding, Zehao Li, Bowen Yang, Zhenyu Wu, Xuehui Wang, Qiushi Sun, Shi Liu, Weiyun Wang, Shenglong Ye, Qingyun Li, Zeyue Tian, Gen Luo, Xiangyu Yue, Biqing Qi, Kai Chen, Bowen Zhou, Yu Qiao, Qifeng Chen, Wenhai Wang
Affiliation
Hong Kong University of Science and Technology
Venue
ICLR 2026
Track
Oral

TL;DR

Vision-Language Models (VLMs) have enabled computer use agents (CUAs) that operate GUIs autonomously, showing great potential, yet progress is limited by the lack of large-scale, open-source computer use data and foundation models. In this work, we introduce ScaleCUA, a step toward scaling open-source CUAs.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

vision-language language model agent

← All ICLR 2026 Oral papers · Browse the whole archive