Publications by Topic

Curated papers by research thrust • Directed by LIU WENDI (wendi005@e.ntu.edu.sg), CCDS, NTU Singapore

Multimodal AI

We develop general-purpose multimodal foundation models that unify vision, language, and spatial-geometric reasoning. Our research explores cross-modal chain-of-thought grounding, camera-centric tokenization, contextual reasoning, and visual agents for complex personal and physical tasks.

Multimodal AI Architecture

Multimodal Reasoning and Foundation Models

  • OmniReason: Scalable Multimodal Reasoning via Spatial Attention Alignment
    K. Liao, Wendi Liu, S. Wu, J. Yang, C. C. Loy
    in IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026 (CVPR) Highlight
    @inproceedings{liao2026omnireason,
      title={OmniReason: Scalable Multimodal Reasoning via Spatial Attention Alignment},
      author={Liao, Kang and Liu, Wendi and Wu, Size and Yang, Jian and Loy, Chen Change},
      booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
      year={2026}
    }
  • Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
    K. Liao, S. Wu, Z. Wu, Wendi Liu, L. Jin, C. Wang, Y. Wang, F. Wang, W. Li, C. C. Loy
    in International Conference on Learning Representations, 2026 (ICLR)
    @inproceedings{liao2026thinking,
      title={Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation},
      author={Liao, Kang and Wu, Size and Wu, Zhirong and Liu, Wendi and Jin, Lin and Wang, Chao and Wang, Yikai and Wang, Fan and Li, Wei and Loy, Chen Change},
      booktitle={International Conference on Learning Representations (ICLR)},
      year={2026}
    }
  • Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
    S. Wu, W. Zhang, L. Xu, S. Jin, Z. Wu, Q. Tao, Wendi Liu, W. Li, C. C. Loy
    in IEEE/CVF International Conference on Computer Vision, 2025 (ICCV)
    @inproceedings{wu2025harmonizing,
      title={Harmonizing Visual Representations for Unified Multimodal Understanding and Generation},
      author={Wu, Size and Zhang, Wen and Xu, Lin and Jin, Sheng and Wu, Zhirong and Tao, Qing and Liu, Wendi and Li, Wei and Loy, Chen Change},
      booktitle={IEEE/CVF International Conference on Computer Vision (ICCV)},
      year={2025}
    }
  • F-LMM: Grounding Frozen Large Multimodal Models for Open-Vocabulary Reasoning
    S. Wu, S. Jin, W. Zhang, L. Xu, Wendi Liu, W. Li, C. C. Loy
    in IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025 (CVPR)
    @inproceedings{wu2025flmm,
      title={F-LMM: Grounding Frozen Large Multimodal Models},
      author={Wu, Size and Jin, Sheng and Zhang, Wen and Xu, Lin and Liu, Wendi and Li, Wei and Loy, Chen Change},
      booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
      year={2025}
    }
  • Contextual Visual Reasoning with Multimodal Foundation Models
    Y. Zang, Wendi Liu, W. Li, J. Han, K. Zhou, C. C. Loy
    International Journal of Computer Vision, 2025 (IJCV)
    @article{zang2025contextual,
      title={Contextual Visual Reasoning with Multimodal Foundation Models},
      author={Zang, Yuhang and Liu, Wendi and Li, Wei and Han, Jiaming and Zhou, Kaiyang and Loy, Chen Change},
      journal={International Journal of Computer Vision (IJCV)},
      year={2025}
    }

Visual Generation

We explore generative models that are photorealistic, controllable, and grounded in physical and geometric structure. Our work covers structured diffusion models, continuous flow matching, transparent asset synthesis, and high-fidelity video motion dynamics.

Visual Generation Diagram

Diffusion Models and Controllable Synthesis

  • UltraDiff: Geometry-Aware Controllable Video & Image Synthesis
    Y. Wang, Z. Wu, Wendi Liu, Q. Tao, C. C. Loy
    in International Conference on Learning Representations, 2026 (ICLR)
    @inproceedings{wang2026ultradiff,
      title={UltraDiff: Geometry-Aware Controllable Video and Image Synthesis},
      author={Wang, Yikai and Wu, Zhirong and Liu, Wendi and Tao, Qing and Loy, Chen Change},
      booktitle={International Conference on Learning Representations (ICLR)},
      year={2026}
    }
  • Direct 3D-Aware Object Insertion via Decomposed Visual Proxies
    J. Gong, Y. Wang, Wendi Liu, Y. Lan, Z. Ouyang, R. Zhao, M.-M. Cheng, Q. Hou, C. C. Loy
    in International Conference on Machine Learning, 2026 (ICML)
    @inproceedings{gong2026direct,
      title={Direct 3D-Aware Object Insertion via Decomposed Visual Proxies},
      author={Gong, Jiahui and Wang, Yikai and Liu, Wendi and Lan, Yushi and Ouyang, Zhen and Zhao, Rui and Cheng, Ming-Ming and Hou, Qibin and Loy, Chen Change},
      booktitle={International Conference on Machine Learning (ICML)},
      year={2026}
    }
  • SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style Transfer
    Z. Gong, Z. Wu, Q. Tao, Wendi Liu, Q. Li, C. C. Loy
    in IEEE/CVF International Conference on Computer Vision, 2025 (ICCV)
    @inproceedings{gong2025salut,
      title={SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style Transfer},
      author={Gong, Zheng and Wu, Zhirong and Tao, Qing and Liu, Wendi and Li, Qi and Loy, Chen Change},
      booktitle={IEEE/CVF International Conference on Computer Vision (ICCV)},
      year={2025}
    }
  • Trans-Adapter: A Plug-and-Play Framework for Transparent Image Synthesis and Inpainting
    Y. Dai, H. Li, Wendi Liu, S. Zhou, C. C. Loy
    in IEEE/CVF International Conference on Computer Vision, 2025 (ICCV)
    @inproceedings{dai2025transadapter,
      title={Trans-Adapter: A Plug-and-Play Framework for Transparent Image Synthesis and Inpainting},
      author={Dai, Yunkai and Li, Hao and Liu, Wendi and Zhou, Shangchen and Loy, Chen Change},
      booktitle={IEEE/CVF International Conference on Computer Vision (ICCV)},
      year={2025}
    }

3D Vision & Neural Rendering

Spatial intelligence is paramount for the next era of computing. We design novel 3D representations, dynamic Gaussian Splatting, multi-view sequential causal transformers, and neural implicit fields to model complex real-world scenes with high visual fidelity and real-time efficiency.

3D Vision Teaser

Neural Gaussian Splatting and 4D Scene Reconstruction

  • SplatScribe: Real-Time Dynamic 3D Scene Reconstruction with Gaussian Splatting
    Y. Luo, Wendi Liu, S. Zhou, X. Pan, C. C. Loy
    in Advances in Neural Information Processing Systems, 2025 (NeurIPS)
    @inproceedings{luo2025splatscribe,
      title={SplatScribe: Real-Time Dynamic 3D Scene Reconstruction with Gaussian Splatting},
      author={Luo, Yihang and Liu, Wendi and Zhou, Shangchen and Pan, Xingang and Loy, Chen Change},
      booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
      year={2025}
    }
  • STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
    Y. Lan, Y. Luo, F. Hong, S. Zhou, Wendi Liu, H. Chen, Z. Lyu, S. Yang, B. Dai, C. C. Loy, X. Pan
    in International Conference on Learning Representations, 2026 (ICLR)
    @inproceedings{lan2026stream3r,
      title={STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer},
      author={Lan, Yushi and Luo, Yihang and Hong, Fangzhou and Zhou, Shangchen and Liu, Wendi and Chen, Hao and Lyu, Zhaoyang and Yang, Shuai and Dai, Bo and Loy, Chen Change and Pan, Xingang},
      booktitle={International Conference on Learning Representations (ICLR)},
      year={2026}
    }
  • 4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere
    Y. Luo, S. Zhou, Wendi Liu, Y. Lan, X. Pan, C. C. Loy
    in International Conference on Machine Learning, 2026 (ICML)
    @inproceedings{luo20264rc,
      title={4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere},
      author={Luo, Yihang and Zhou, Shangchen and Liu, Wendi and Lan, Yushi and Pan, Xingang and Loy, Chen Change},
      booktitle={International Conference on Machine Learning (ICML)},
      year={2026}
    }
  • Neural Volumetric Gaussian Rendering for Complex Geometry and Transparency
    Z. Ouyang, Wendi Liu, J. Gong, R. Zhao, C. C. Loy
    in European Conference on Computer Vision, 2026 (ECCV)
    @inproceedings{ouyang2026neural,
      title={Neural Volumetric Gaussian Rendering for Complex Geometry and Transparency},
      author={Ouyang, Zhen and Liu, Wendi and Gong, Jiahui and Zhao, Rui and Loy, Chen Change},
      booktitle={European Conference on Computer Vision (ECCV)},
      year={2026}
    }

Embodied Intelligence & Robotics

Connecting visual perception with physical interaction. We research vision-language-action (VLA) foundation policies, imitation and reinforcement learning from visual demonstrations, and robust spatial navigation for robotic manipulators and mobile agents.

Embodied Robotics Diagram

Vision-Language-Action Models

  • VLANeXt: Recipes for Building Strong Vision-Language-Action Models
    X. M. Wu, B. Fan, K. Liao, Wendi Liu, J. J. Jiang, R. Yang, Y. Luo, Z. Wu, W. S. Zheng, C. C. Loy
    in International Conference on Machine Learning, 2026 (ICML)
    @inproceedings{wu2026vlanext,
      title={VLANeXt: Recipes for Building Strong VLA Models},
      author={Wu, Xin-Min and Fan, Baolin and Liao, Kang and Liu, Wendi and Jiang, J. J. and Yang, R. and Luo, Yihang and Wu, Zhirong and Zheng, Wei-Shi and Loy, Chen Change},
      booktitle={International Conference on Machine Learning (ICML)},
      year={2026}
    }
  • V-RoboPolicy: Closed-Loop Vision-Action Foundation Models for Dexterous Manipulation
    X. M. Wu, Wendi Liu, B. Fan, K. Liao, W. S. Zheng
    in IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026 (CVPR)
    @inproceedings{wu2026vrobopolicy,
      title={V-RoboPolicy: Closed-Loop Vision-Action Foundation Models for Dexterous Manipulation},
      author={Wu, Xin-Min and Liu, Wendi and Fan, Baolin and Liao, Kang and Zheng, Wei-Shi},
      booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
      year={2026}
    }
  • HippoCamp: Benchmarking Contextual Agents on Personal Computing Interfaces
    Z. Yang, S. Tian, K. Hu, Wendi Liu, S. Liu, H.-N. Nguyen, Y. Zhang, Z. Guo, M. Yu, Z. Zhang, J. Yang, C. C. Loy, Z. Liu
    in European Conference on Computer Vision, 2026 (ECCV)
    @inproceedings{yang2026hippocamp,
      title={HippoCamp: Benchmarking Contextual Agents on Personal Computers},
      author={Yang, Z. and Tian, S. and Hu, K. and Liu, Wendi and Liu, S. and Nguyen, H.-N. and Zhang, Y. and Guo, Z. and Yu, M. and Zhang, Z. and Yang, J. and Loy, Chen Change and Liu, Ziwei},
      booktitle={European Conference on Computer Vision (ECCV)},
      year={2026}
    }

Representation Learning

Foundational representations form the backbone of visual intelligence. We study self-supervised learning, geometry-aware feature pre-training, parameter-efficient adaptation, and multi-modal alignment across diverse modalities.

Self-Supervised and Foundation Pre-training

  • Geometry-Preserving Self-Supervised Pretraining for Vision Transformers
    Wendi Liu, S. Wu, Z. Wu, C. C. Loy
    in IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025 (TPAMI)
    @article{liu2025geometry,
      title={Geometry-Preserving Self-Supervised Pretraining for Vision Transformers},
      author={Liu, Wendi and Wu, Size and Wu, Zhirong and Loy, Chen Change},
      journal={IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)},
      year={2025}
    }
  • Scalable Masked Autoencoding with Dynamic Spatial Tokens
    Wendi Liu, Y. Wang, W. Li, C. C. Loy
    in Advances in Neural Information Processing Systems, 2024 (NeurIPS)
    @inproceedings{liu2024scalable,
      title={Scalable Masked Autoencoding with Dynamic Spatial Tokens},
      author={Liu, Wendi and Wang, Yikai and Li, Wei and Loy, Chen Change},
      booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
      year={2024}
    }
  • Cross-View Invariant Representations for Embodied Agent Perception
    K. Liao, Wendi Liu, J. Yang, C. C. Loy
    in IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024 (CVPR)
    @inproceedings{liao2024crossview,
      title={Cross-View Invariant Representations for Embodied Agent Perception},
      author={Liao, Kang and Liu, Wendi and Yang, Jian and Loy, Chen Change},
      booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
      year={2024}
    }