MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Models for Embodied Task Planning
ICLROral2026
TL;DR
We present MomaGraph, a unified scene representation for task-oriented understanding, along with a dataset and benchmark built upon it, and MomaGraph-R1, a 7B model that constructs MomaGraph representations and generates task plans.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
vision-language language model benchmark planning dataset graph