Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator

ICLROral2026

Authors
Hyojun Go, Dominik Narnhofer, Goutam Bhat, Prune Truong, Federico Tombari, Konrad Schindler
Affiliation
ETHZ - ETH Zurich
Venue
ICLR 2026
Track
Oral

TL;DR

Text-to-3D scene generative modelling by unifying a video generative model with a foundational 3D model via model stitching and alignment.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

generative model alignment video 3d

← All ICLR 2026 Oral papers · Browse the whole archive