MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Models for Embodied Task Planning

ICLROral2026

Authors
Yuanchen Ju, Yongyuan Liang, Yen-Jen Wang, Gireesh Nandiraju, Yuanliang Ju, Seungjae Lee, Qiao Gu, Elvis Hsieh, Furong Huang, Koushil Sreenath
Venue
ICLR 2026
Track
Oral

TL;DR

We present MomaGraph, a unified scene representation for task-oriented understanding, along with a dataset and benchmark built upon it, and MomaGraph-R1, a 7B model that constructs MomaGraph representations and generates task plans.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

vision-language language model benchmark planning dataset graph

← All ICLR 2026 Oral papers · Browse the whole archive