Instilling an Active Mind in Avatars via Cognitive Simulation
ICLROral2026
TL;DR
This paper introduces a novel framework that uses a Large Language Model (LLM) for semantic guidance and a Multimodal Diffusion Transformer (DiT) for fusion to generate expressive, context-aware video avatars, demonstrating competitive performance…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model transformer multimodal diffusion video llm