Curated papers by research thrust • Directed by LIU WENDI (wendi005@e.ntu.edu.sg), CCDS, NTU Singapore
Multimodal AI
We develop general-purpose multimodal foundation models that unify vision, language, and spatial-geometric reasoning. Our research explores cross-modal chain-of-thought grounding, camera-centric tokenization, contextual reasoning, and visual agents for complex personal and physical tasks.
Multimodal Reasoning and Foundation Models
OmniReason: Scalable Multimodal Reasoning via Spatial Attention Alignment
K. Liao, Wendi Liu, S. Wu, J. Yang, C. C. Loy
in IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026 (CVPR) Highlight
@inproceedings{liao2026omnireason,
title={OmniReason: Scalable Multimodal Reasoning via Spatial Attention Alignment},
author={Liao, Kang and Liu, Wendi and Wu, Size and Yang, Jian and Loy, Chen Change},
booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year={2026}
}
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
K. Liao, S. Wu, Z. Wu, Wendi Liu, L. Jin, C. Wang, Y. Wang, F. Wang, W. Li, C. C. Loy
in International Conference on Learning Representations, 2026 (ICLR)
@inproceedings{liao2026thinking,
title={Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation},
author={Liao, Kang and Wu, Size and Wu, Zhirong and Liu, Wendi and Jin, Lin and Wang, Chao and Wang, Yikai and Wang, Fan and Li, Wei and Loy, Chen Change},
booktitle={International Conference on Learning Representations (ICLR)},
year={2026}
}
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
S. Wu, W. Zhang, L. Xu, S. Jin, Z. Wu, Q. Tao, Wendi Liu, W. Li, C. C. Loy
in IEEE/CVF International Conference on Computer Vision, 2025 (ICCV)
@inproceedings{wu2025harmonizing,
title={Harmonizing Visual Representations for Unified Multimodal Understanding and Generation},
author={Wu, Size and Zhang, Wen and Xu, Lin and Jin, Sheng and Wu, Zhirong and Tao, Qing and Liu, Wendi and Li, Wei and Loy, Chen Change},
booktitle={IEEE/CVF International Conference on Computer Vision (ICCV)},
year={2025}
}
F-LMM: Grounding Frozen Large Multimodal Models for Open-Vocabulary Reasoning
S. Wu, S. Jin, W. Zhang, L. Xu, Wendi Liu, W. Li, C. C. Loy
in IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025 (CVPR)
@inproceedings{wu2025flmm,
title={F-LMM: Grounding Frozen Large Multimodal Models},
author={Wu, Size and Jin, Sheng and Zhang, Wen and Xu, Lin and Liu, Wendi and Li, Wei and Loy, Chen Change},
booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year={2025}
}
Contextual Visual Reasoning with Multimodal Foundation Models
Y. Zang, Wendi Liu, W. Li, J. Han, K. Zhou, C. C. Loy
International Journal of Computer Vision, 2025 (IJCV)
@article{zang2025contextual,
title={Contextual Visual Reasoning with Multimodal Foundation Models},
author={Zang, Yuhang and Liu, Wendi and Li, Wei and Han, Jiaming and Zhou, Kaiyang and Loy, Chen Change},
journal={International Journal of Computer Vision (IJCV)},
year={2025}
}
Visual Generation
We explore generative models that are photorealistic, controllable, and grounded in physical and geometric structure. Our work covers structured diffusion models, continuous flow matching, transparent asset synthesis, and high-fidelity video motion dynamics.
Diffusion Models and Controllable Synthesis
UltraDiff: Geometry-Aware Controllable Video & Image Synthesis
Y. Wang, Z. Wu, Wendi Liu, Q. Tao, C. C. Loy
in International Conference on Learning Representations, 2026 (ICLR)
@inproceedings{wang2026ultradiff,
title={UltraDiff: Geometry-Aware Controllable Video and Image Synthesis},
author={Wang, Yikai and Wu, Zhirong and Liu, Wendi and Tao, Qing and Loy, Chen Change},
booktitle={International Conference on Learning Representations (ICLR)},
year={2026}
}
Direct 3D-Aware Object Insertion via Decomposed Visual Proxies
J. Gong, Y. Wang, Wendi Liu, Y. Lan, Z. Ouyang, R. Zhao, M.-M. Cheng, Q. Hou, C. C. Loy
in International Conference on Machine Learning, 2026 (ICML)
@inproceedings{gong2026direct,
title={Direct 3D-Aware Object Insertion via Decomposed Visual Proxies},
author={Gong, Jiahui and Wang, Yikai and Liu, Wendi and Lan, Yushi and Ouyang, Zhen and Zhao, Rui and Cheng, Ming-Ming and Hou, Qibin and Loy, Chen Change},
booktitle={International Conference on Machine Learning (ICML)},
year={2026}
}
SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style Transfer
Z. Gong, Z. Wu, Q. Tao, Wendi Liu, Q. Li, C. C. Loy
in IEEE/CVF International Conference on Computer Vision, 2025 (ICCV)
@inproceedings{gong2025salut,
title={SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style Transfer},
author={Gong, Zheng and Wu, Zhirong and Tao, Qing and Liu, Wendi and Li, Qi and Loy, Chen Change},
booktitle={IEEE/CVF International Conference on Computer Vision (ICCV)},
year={2025}
}
Trans-Adapter: A Plug-and-Play Framework for Transparent Image Synthesis and Inpainting
Y. Dai, H. Li, Wendi Liu, S. Zhou, C. C. Loy
in IEEE/CVF International Conference on Computer Vision, 2025 (ICCV)
@inproceedings{dai2025transadapter,
title={Trans-Adapter: A Plug-and-Play Framework for Transparent Image Synthesis and Inpainting},
author={Dai, Yunkai and Li, Hao and Liu, Wendi and Zhou, Shangchen and Loy, Chen Change},
booktitle={IEEE/CVF International Conference on Computer Vision (ICCV)},
year={2025}
}
3D Vision & Neural Rendering
Spatial intelligence is paramount for the next era of computing. We design novel 3D representations, dynamic Gaussian Splatting, multi-view sequential causal transformers, and neural implicit fields to model complex real-world scenes with high visual fidelity and real-time efficiency.
Neural Gaussian Splatting and 4D Scene Reconstruction
SplatScribe: Real-Time Dynamic 3D Scene Reconstruction with Gaussian Splatting
Y. Luo, Wendi Liu, S. Zhou, X. Pan, C. C. Loy
in Advances in Neural Information Processing Systems, 2025 (NeurIPS)
@inproceedings{luo2025splatscribe,
title={SplatScribe: Real-Time Dynamic 3D Scene Reconstruction with Gaussian Splatting},
author={Luo, Yihang and Liu, Wendi and Zhou, Shangchen and Pan, Xingang and Loy, Chen Change},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
year={2025}
}
STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
Y. Lan, Y. Luo, F. Hong, S. Zhou, Wendi Liu, H. Chen, Z. Lyu, S. Yang, B. Dai, C. C. Loy, X. Pan
in International Conference on Learning Representations, 2026 (ICLR)
@inproceedings{lan2026stream3r,
title={STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer},
author={Lan, Yushi and Luo, Yihang and Hong, Fangzhou and Zhou, Shangchen and Liu, Wendi and Chen, Hao and Lyu, Zhaoyang and Yang, Shuai and Dai, Bo and Loy, Chen Change and Pan, Xingang},
booktitle={International Conference on Learning Representations (ICLR)},
year={2026}
}
4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere
Y. Luo, S. Zhou, Wendi Liu, Y. Lan, X. Pan, C. C. Loy
in International Conference on Machine Learning, 2026 (ICML)
@inproceedings{luo20264rc,
title={4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere},
author={Luo, Yihang and Zhou, Shangchen and Liu, Wendi and Lan, Yushi and Pan, Xingang and Loy, Chen Change},
booktitle={International Conference on Machine Learning (ICML)},
year={2026}
}
Neural Volumetric Gaussian Rendering for Complex Geometry and Transparency
Z. Ouyang, Wendi Liu, J. Gong, R. Zhao, C. C. Loy
in European Conference on Computer Vision, 2026 (ECCV)
@inproceedings{ouyang2026neural,
title={Neural Volumetric Gaussian Rendering for Complex Geometry and Transparency},
author={Ouyang, Zhen and Liu, Wendi and Gong, Jiahui and Zhao, Rui and Loy, Chen Change},
booktitle={European Conference on Computer Vision (ECCV)},
year={2026}
}
Embodied Intelligence & Robotics
Connecting visual perception with physical interaction. We research vision-language-action (VLA) foundation policies, imitation and reinforcement learning from visual demonstrations, and robust spatial navigation for robotic manipulators and mobile agents.
Vision-Language-Action Models
VLANeXt: Recipes for Building Strong Vision-Language-Action Models
X. M. Wu, B. Fan, K. Liao, Wendi Liu, J. J. Jiang, R. Yang, Y. Luo, Z. Wu, W. S. Zheng, C. C. Loy
in International Conference on Machine Learning, 2026 (ICML)
@inproceedings{wu2026vlanext,
title={VLANeXt: Recipes for Building Strong VLA Models},
author={Wu, Xin-Min and Fan, Baolin and Liao, Kang and Liu, Wendi and Jiang, J. J. and Yang, R. and Luo, Yihang and Wu, Zhirong and Zheng, Wei-Shi and Loy, Chen Change},
booktitle={International Conference on Machine Learning (ICML)},
year={2026}
}
V-RoboPolicy: Closed-Loop Vision-Action Foundation Models for Dexterous Manipulation
X. M. Wu, Wendi Liu, B. Fan, K. Liao, W. S. Zheng
in IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026 (CVPR)
@inproceedings{wu2026vrobopolicy,
title={V-RoboPolicy: Closed-Loop Vision-Action Foundation Models for Dexterous Manipulation},
author={Wu, Xin-Min and Liu, Wendi and Fan, Baolin and Liao, Kang and Zheng, Wei-Shi},
booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year={2026}
}
HippoCamp: Benchmarking Contextual Agents on Personal Computing Interfaces
Z. Yang, S. Tian, K. Hu, Wendi Liu, S. Liu, H.-N. Nguyen, Y. Zhang, Z. Guo, M. Yu, Z. Zhang, J. Yang, C. C. Loy, Z. Liu
in European Conference on Computer Vision, 2026 (ECCV)
@inproceedings{yang2026hippocamp,
title={HippoCamp: Benchmarking Contextual Agents on Personal Computers},
author={Yang, Z. and Tian, S. and Hu, K. and Liu, Wendi and Liu, S. and Nguyen, H.-N. and Zhang, Y. and Guo, Z. and Yu, M. and Zhang, Z. and Yang, J. and Loy, Chen Change and Liu, Ziwei},
booktitle={European Conference on Computer Vision (ECCV)},
year={2026}
}
Representation Learning
Foundational representations form the backbone of visual intelligence. We study self-supervised learning, geometry-aware feature pre-training, parameter-efficient adaptation, and multi-modal alignment across diverse modalities.
Self-Supervised and Foundation Pre-training
Geometry-Preserving Self-Supervised Pretraining for Vision Transformers
Wendi Liu, S. Wu, Z. Wu, C. C. Loy
in IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025 (TPAMI)
@article{liu2025geometry,
title={Geometry-Preserving Self-Supervised Pretraining for Vision Transformers},
author={Liu, Wendi and Wu, Size and Wu, Zhirong and Loy, Chen Change},
journal={IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)},
year={2025}
}
Scalable Masked Autoencoding with Dynamic Spatial Tokens
Wendi Liu, Y. Wang, W. Li, C. C. Loy
in Advances in Neural Information Processing Systems, 2024 (NeurIPS)
@inproceedings{liu2024scalable,
title={Scalable Masked Autoencoding with Dynamic Spatial Tokens},
author={Liu, Wendi and Wang, Yikai and Li, Wei and Loy, Chen Change},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
year={2024}
}
Cross-View Invariant Representations for Embodied Agent Perception
K. Liao, Wendi Liu, J. Yang, C. C. Loy
in IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024 (CVPR)
@inproceedings{liao2024crossview,
title={Cross-View Invariant Representations for Embodied Agent Perception},
author={Liao, Kang and Liu, Wendi and Yang, Jian and Loy, Chen Change},
booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year={2024}
}