Nanyang Technological University, Singapore

Advancing Visual Intelligence & Multimodal AI

Directed by LIU WENDI at the College of Computing and Data Science (CCDS), NTU Singapore. We investigate foundational computer vision, large multimodal reasoning models, geometry-aware visual generation, and embodied agentic systems.

Prof. LIU WENDI

LIU WENDI

Lab Director & Principal Investigator
College of Computing and Data Science (CCDS)
Nanyang Technological University (NTU), Singapore

Pioneering General Visual Intelligence

LIU WENDI is a faculty researcher and Lab Director at the College of Computing and Data Science, Nanyang Technological University (NTU), Singapore. The research group conducts cutting-edge research at the nexus of Computer Vision, Generative AI, and Multimodal Large Language Models.

Our group aims to build autonomous systems that can perceive, reason about, synthesize, and interact with the complex 3D physical world. Towards this grand vision, we emphasize both mathematical foundations and scalable systems, spanning multimodal spatial reasoning, physics-grounded generative diffusion architectures, and closed-loop embodied policy learning for real-world robotics.

We collaborate closely with leading research institutes and industry partners across Singapore and internationally. Our lab's contributions are consistently published at top-tier venues including CVPR, ICCV, ECCV, NeurIPS, ICLR, ICML, and IEEE TPAMI.

Meet the Lab Team View Full Publications

What We Investigate

We build integrated visual perception and generation algorithms designed for real-world reasoning and physical interactivity.

Multimodal AI

Multimodal Foundation Models

Developing unified vision-language-action architectures capable of deep spatial reasoning, long-context video understanding, and cross-modal chain-of-thought grounding.

Read Publications →
Visual Generation

Controllable Visual Generation

Next-generation diffusion models, flow matching, and autoregressive models for photorealistic image and video synthesis with explicit 3D and physical constraints.

Read Publications →
3D Vision

3D Vision & Neural Rendering

Real-time dynamic scene reconstruction, 3D Gaussian Splatting, neural radiance fields, and spatial intelligence for immersive AR/VR and camera-centric perception.

Read Publications →
Embodied AI

Embodied AI & Robotics

Vision-language-action (VLA) models for robotic manipulation, closed-loop sensorimotor control, spatial navigation, and foundation models for autonomous agents.

Read Publications →

Recent Publications

All Publications by Topic By Year
OmniReason: Scalable Multimodal Reasoning via Spatial Attention Alignment
K. Liao, Wendi Liu, S. Wu, J. Yang, C. C. Loy
in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026) Highlight
@inproceedings{liao2026omnireason,
  title={OmniReason: Scalable Multimodal Reasoning via Spatial Attention Alignment},
  author={Liao, Kang and Liu, Wendi and Wu, Size and Yang, Jian and Loy, Chen Change},
  booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  year={2026}
}
UltraDiff: Geometry-Aware Controllable Video & Image Synthesis
Y. Wang, Z. Wu, Wendi Liu, Q. Tao, C. C. Loy
in International Conference on Learning Representations (ICLR 2026)
@inproceedings{wang2026ultradiff,
  title={UltraDiff: Geometry-Aware Controllable Video and Image Synthesis},
  author={Wang, Yikai and Wu, Zhirong and Liu, Wendi and Tao, Qing and Loy, Chen Change},
  booktitle={International Conference on Learning Representations (ICLR)},
  year={2026}
}
SplatScribe: Real-Time Dynamic 3D Scene Reconstruction with Gaussian Splatting
Y. Luo, Wendi Liu, S. Zhou, X. Pan, C. C. Loy
in Advances in Neural Information Processing Systems (NeurIPS 2025)
@inproceedings{luo2025splatscribe,
  title={SplatScribe: Real-Time Dynamic 3D Scene Reconstruction with Gaussian Splatting},
  author={Luo, Yihang and Liu, Wendi and Zhou, Shangchen and Pan, Xingang and Loy, Chen Change},
  booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
  year={2025}
}
V-RoboPolicy: Closed-Loop Vision-Action Foundation Models for Dexterous Manipulation
X. M. Wu, Wendi Liu, B. Fan, K. Liao, W. S. Zheng
in International Conference on Machine Learning (ICML 2026)
@inproceedings{wu2026vrobopolicy,
  title={V-RoboPolicy: Closed-Loop Vision-Action Foundation Models for Dexterous Manipulation},
  author={Wu, Xin-Min and Liu, Wendi and Fan, Baolin and Liao, Kang and Zheng, Wei-Shi},
  booktitle={International Conference on Machine Learning (ICML)},
  year={2026}
}

Lab News & Announcements

Join Wendi Lab @ NTU

We are actively looking for highly self-motivated PhD students, Postdoctoral Fellows, Research Assistants, and visiting students with strong mathematical foundations and programming skills in Deep Learning, Computer Vision, and Robotics.

How to Apply:

Please send an email directly to LIU WENDI at wendi005@e.ntu.edu.sg with:

  • Subject: [Prospective Applicant - Degree/Role] Your Name
  • Detailed Curriculum Vitae (CV)
  • Academic transcripts and ranking
  • Representative research papers or GitHub code repository links (if any)
Send Application to wendi005@e.ntu.edu.sg