EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
ICLROral2026
TL;DR
Recent advances in foundation models highlight a clear trend toward unification and scaling, showing emergent capabilities across diverse domains. While image generation and editing have rapidly transitioned from task-specific to unified frameworks, video generation and editing remain fragmented due to architectural limitations and data scarcity.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
in-context learning image generation video rag