WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning
Paper • 2609.34826 • Published • 12
None defined yet.
LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation
WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning