Publications by Year

Chronological record of peer-reviewed articles • PI: LIU WENDI (wendi005@e.ntu.edu.sg)

2026

  • OmniReason: Scalable Multimodal Reasoning via Spatial Attention Alignment
    K. Liao, Wendi Liu, S. Wu, J. Yang, C. C. Loy
    in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026) Highlight
  • Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
    K. Liao, S. Wu, Z. Wu, Wendi Liu, L. Jin, C. Wang, Y. Wang, F. Wang, W. Li, C. C. Loy
    in International Conference on Learning Representations (ICLR 2026)
  • UltraDiff: Geometry-Aware Controllable Video & Image Synthesis
    Y. Wang, Z. Wu, Wendi Liu, Q. Tao, C. C. Loy
    in International Conference on Learning Representations (ICLR 2026)
  • STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
    Y. Lan, Y. Luo, F. Hong, S. Zhou, Wendi Liu, H. Chen, Z. Lyu, S. Yang, B. Dai, C. C. Loy, X. Pan
    in International Conference on Learning Representations (ICLR 2026)
  • 4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere
    Y. Luo, S. Zhou, Wendi Liu, Y. Lan, X. Pan, C. C. Loy
    in International Conference on Machine Learning (ICML 2026)
  • VLANeXt: Recipes for Building Strong Vision-Language-Action Models
    X. M. Wu, B. Fan, K. Liao, Wendi Liu, J. J. Jiang, R. Yang, Y. Luo, Z. Wu, W. S. Zheng, C. C. Loy
    in International Conference on Machine Learning (ICML 2026)
  • HippoCamp: Benchmarking Contextual Agents on Personal Computers
    Z. Yang, S. Tian, K. Hu, Wendi Liu, S. Liu, H.-N. Nguyen, Y. Zhang, Z. Guo, M. Yu, Z. Zhang, J. Yang, C. C. Loy, Z. Liu
    in European Conference on Computer Vision (ECCV 2026)

2025

  • SplatScribe: Real-Time Dynamic 3D Scene Reconstruction with Gaussian Splatting
    Y. Luo, Wendi Liu, S. Zhou, X. Pan, C. C. Loy
    in Advances in Neural Information Processing Systems (NeurIPS 2025)
  • Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
    S. Wu, W. Zhang, L. Xu, S. Jin, Z. Wu, Q. Tao, Wendi Liu, W. Li, C. C. Loy
    in IEEE/CVF International Conference on Computer Vision (ICCV 2025)
  • F-LMM: Grounding Frozen Large Multimodal Models for Open-Vocabulary Reasoning
    S. Wu, S. Jin, W. Zhang, L. Xu, Wendi Liu, W. Li, C. C. Loy
    in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2025)
  • Geometry-Preserving Self-Supervised Pretraining for Vision Transformers
    Wendi Liu, S. Wu, Z. Wu, C. C. Loy
    IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI 2025)
  • SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style Transfer
    Z. Gong, Z. Wu, Q. Tao, Wendi Liu, Q. Li, C. C. Loy
    in IEEE/CVF International Conference on Computer Vision (ICCV 2025)
  • Trans-Adapter: A Plug-and-Play Framework for Transparent Image Synthesis and Inpainting
    Y. Dai, H. Li, Wendi Liu, S. Zhou, C. C. Loy
    in IEEE/CVF International Conference on Computer Vision (ICCV 2025)

2024

  • Scalable Masked Autoencoding with Dynamic Spatial Tokens
    Wendi Liu, Y. Wang, W. Li, C. C. Loy
    in Advances in Neural Information Processing Systems (NeurIPS 2024)
  • Cross-View Invariant Representations for Embodied Agent Perception
    K. Liao, Wendi Liu, J. Yang, C. C. Loy
    in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024)
  • Contextual Visual Reasoning with Multimodal Foundation Models
    Y. Zang, Wendi Liu, W. Li, J. Han, K. Zhou, C. C. Loy
    International Journal of Computer Vision (IJCV 2024)