EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning

ICLROral2026

Authors
Xuan Ju, Tianyu Wang, Yuqian Zhou, He Zhang, Qing Liu, Nanxuan Zhao, Zhifei Zhang, Yijun Li, Yuanhao Cai, Shaoteng Liu, Daniil Pakhomov, Zhe Lin, Soo Ye Kim, Qiang Xu
Affiliation
Chinese University of Hong Kong
Venue
ICLR 2026
Track
Oral

TL;DR

Recent advances in foundation models highlight a clear trend toward unification and scaling, showing emergent capabilities across diverse domains. While image generation and editing have rapidly transitioned from task-specific to unified frameworks, video generation and editing remain fragmented due to architectural limitations and data scarcity.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

in-context learning image generation video rag

← All ICLR 2026 Oral papers · Browse the whole archive