Our Research Vision

Building general visual intelligence across perception, reasoning, synthesis, and physical action.

Multimodal Large Models & Spatial Reasoning

Modern Vision-Language Models (VLMs) have demonstrated astonishing semantic breadth, but often lack precise 3D spatial grounding and physical depth understanding. Our research pioneers methods to integrate camera geometries directly into the language token space, enabling true spatial reasoning, contextual scene explanation, and fine-grained visual question answering.

  • Camera-centric coordinate tokenization and cross-view alignment
  • Contextual and open-vocabulary object grounding
  • Interactive multimodal agents on mobile and PC graphical interfaces
Explore Multimodal Papers
Multimodal Reasoning

Controllable & Physically Grounded Generation

Generative visual AI is rapidly transforming creative workflows and simulated worlds. We design structured diffusion and autoregressive token generation frameworks that obey 3D scene physics, multi-view consistency, and strict user constraints.

  • 3D-aware object insertion and lighting relighting
  • Geometry-aware video synthesis with temporal coherence
  • High-resolution transparent layer inpainting and matting
Explore Generation Papers
Visual Generation

3D Gaussian Splatting & Dynamic 4D Perception

To bridge digital and physical worlds, accurate and real-time 3D representation is indispensable. Our work pushes the envelope of 3D Gaussian Splatting (3D-GS) to handle dynamic scenes, non-rigid deformations, specular reflections, and streaming video streams.

  • Real-time feed-forward 4D scene reconstruction
  • Causal sequential transformers for streaming 3D vision
  • Neural volumetric rendering of fine complex geometries
Explore 3D Vision Papers
3D Splatting

Embodied Robotics & Vision-Language-Action

True visual intelligence must be tested through physical action. We create vision-language-action (VLA) foundation policies for robotic arms and mobile manipulators. By combining visual demonstration datasets with closed-loop predictive modeling, our robots execute complex manipulation tasks in dynamic unstructured environments.

  • Unified VLA policy architectures with spatial pretraining
  • Sample-efficient imitation learning for dexterous manipulation
  • Zero-shot spatial navigation and environmental interaction
Explore Embodied AI Papers
Embodied Robotics