OpenApps: Simulating Environment Variations to Measure UI Agent Reliability

ICLROral2026

Authors
Karen Ullrich, Jingtong Su, Claudia Shi, Arjun Subramonian, Amir Bar, Ivan Evtimov, Nikolaos Tsilivis, Randall Balestriero, Julia Kempe, Mark Ibrahim
Affiliation
Meta AI
Venue
ICLR 2026
Track
Oral

TL;DR

We introduce a new environment, OpenApps, for generating thousands of versions of apps to test UI agent reliability.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

agent

← All ICLR 2026 Oral papers · Browse the whole archive