Building general visual intelligence across perception, reasoning, synthesis, and physical action.
Modern Vision-Language Models (VLMs) have demonstrated astonishing semantic breadth, but often lack precise 3D spatial grounding and physical depth understanding. Our research pioneers methods to integrate camera geometries directly into the language token space, enabling true spatial reasoning, contextual scene explanation, and fine-grained visual question answering.
Generative visual AI is rapidly transforming creative workflows and simulated worlds. We design structured diffusion and autoregressive token generation frameworks that obey 3D scene physics, multi-view consistency, and strict user constraints.
To bridge digital and physical worlds, accurate and real-time 3D representation is indispensable. Our work pushes the envelope of 3D Gaussian Splatting (3D-GS) to handle dynamic scenes, non-rigid deformations, specular reflections, and streaming video streams.
True visual intelligence must be tested through physical action. We create vision-language-action (VLA) foundation policies for robotic arms and mobile manipulators. By combining visual demonstration datasets with closed-loop predictive modeling, our robots execute complex manipulation tasks in dynamic unstructured environments.