Computer Vision · Embodied AI

About Me

I am an Associate Professor in the College of Artificial Intelligence at Nanjing University of Aeronautics and Astronautics, working closely with Prof. Jie Qin.

In August 2025 and August 2026, I was a visiting researcher in the Department of Computing at The Hong Kong Polytechnic University, where I collaborated closely with my Ph.D. advisor, Prof. Lei Zhang (IEEE Fellow). Previously, I completed my Ph.D. in the College of Computer Science and Technology at Zhejiang University, supervised by Prof. Jianke Zhu and Prof. Lei Zhang, in June 2024.

I serve as an Area Chair for top-tier AI conferences, including ICLR 2026 and 2027, NeurIPS 2026, and AAAI 2027.

Research

Vision and Language

Scene understanding, video understanding, fine-grained visual understanding, and efficient architectures.

Embodied Intelligence

Embodied understanding, reasoning, planning, and action for intelligent agents in complex environments.

Machine Learning

Multimodal learning, robust learning, and foundation models for vision, agent, and robotics.

Open positions. I am looking for self-motivated Master's students, research interns/assistants, and co-supervised Ph.D. students. Please email me if you are interested.

News

  • [2026.09]: AgentVLN was featured by 机器之心 (Synced) and CVer.
  • [2026.08]: Invited to serve as Area Chair for ICLR 2027.
  • [2026.07]: Invited to serve as Senior Program Committee for AAAI 2027.
  • [2026.05]: Our AgentVLN is accepted by ECCV 2026 (An Agentic Embodied Navigation Framework).
  • [2026.05]: Two papers are accepted by ICML 2026 (IDEAL-VLN is an embodied navigation approach with real-world deployment).
  • [2026.04]: We released a Survey of Object-level LMM. [GitHub Repo]
  • [2026.03]: Invited to serve as Area Chair for NeurIPS 2026.
  • [2026.02]: One paper is accepted by CVPR 2026 (embodied navigation approach DecoVLN with real-world deployment).
  • [2026.02]: Awarded the Outstanding Doctoral Dissertation Award of Zhejiang Province.
  • [2026.01]: Two papers are accepted by ICLR 2026.
  • [2025.12]: We released a Survey forging Spatial Intelligence for Autonomous Systems.
  • [2025.11]: One paper about Object-level Generation on Camouflage Images is accepted by AAAI 2026.
  • [2025.11]: Our PixelRefer is reported by PaperWeekly and 机器之心, respectively.
  • [2025.10]: We released PixelRefer, a new unified pixel-level MLLM framework for fine-grained regional understanding.
  • [2025.10]: Shared a talk@PRCV2025.[Slides]
  • [2025.9]: Two papers are accepted by NeurIPS 2025.
  • [2025.8]: Be funded by NSFC 🎉.
  • [2025.8]: Invited to serve as Area Chair for ICLR 2026.
  • [2025.8]: Visited The Hong Kong Polytechnic University, where I enjoyed the visit and shared a talk.[Slides]
  • [2025.6]: We released the EOC-Bench, an object-centric embodied cognition benchmark in dynamic egocentric scenarios.
  • [2025.5]: One paper is accepted by IJCV (TokenPacker, 57 citations at the time of acceptance).
  • [2025.4]: Our VideoRefer and VideoRefer-Bench have been discussed and adopted by NVIDIA & UC Berkely in their DAM work.
  • [2025.2]: Five papers are accepted by CVPR 2025 (One Highlight).
  • [2025.2]: We released the VideoRefer-700K dataset on HuggingFace. Please see the VideoRefer Suite for the details.
  • [2024.12]: Awarded Outstanding Doctoral Dissertation Award of ZJU (浙江大学优秀博士学位论文).
  • [2024.6]: Obtained my Ph.D. degree from ZJU.

Publications

*: equal contribution, #: corresponding author, +: project leader

Preprints

EmbodiedSkills framework overview
EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents
Wei Wang, Wenqiao Zhang, Yutong Lin, Yuqian Yuan, Tianwei Lin, Jinhao Mao, Zhenxuan Fan, Mingjian Gao, Yang Dai, Wentong Li, Zheqi Lv, Zheng Dong, Yingjie Niu, Jiaqi Zhu, Jun Xiao, Chao Li, Yueting Zhuang
arXiv:2609.01281

PaperCode GitHub stars

photo
InstructSAM: Segment Any Instance with Any Instructions
Yuqian Yuan*, Wentong Li*, Zhaocheng Li*, Yutong Lin*, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang, Wenqiao Zhang
arXiv:2605.26102
Project Leader
photo
Weighted Reverse Convolution for Feature Upsampling
Wentong Li*, Zhiyuan Qi*, Zichen Zhao, Kai Zhang, Lei Zhang
arXiv:2605.17472

PaperCode

photo
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
Yuqian Yuan, Wenqiao Zhang, Juekai Lin, Yu Zhong, Mingjian Gao, Binhe Yu, Yunqi Cao, Wentong Li, Yueting Zhuang, Beng Chin Ooi
arXiv:2604.11789
photo
PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
Yuqian Yuan, Wenqiao Zhang, Xin Li, Shihao Wang, Kehan Li, Wentong Li#, Jun Xiao, Lei Zhang, Beng Chin Ooi
arxiv: 2510.23603

Selected Publications

photo
AgentVLN: Towards Agentic Vision-and-Language Navigation
Zihao Xin, Wentong Li#, Yixuan Jiang, Ziyuan Huang, Bin Wang, Piji Li, Jianke Zhu, Jie Qin, Shengjun Huang
ECCV 2026
photo
Instruction Decomposition and Action Alignment for Vision-and-Language Navigation
Zihao Xin, Wentong Li#, Yixuan Jiang, Bin Wang, Piji Li, Jianke Zhu, Jie Qin, Shengjun Huang
ICML 2026

PaperCode

photo
DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigation
Zihao Xin, Wentong Li+, Yixuan Jiang, Bin Wang, Runmin Cong, Jie Qin, Shengjun Huang
CVPR, 2026
photo
VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration
Hanxun Yu*, Wentong Li*, Xuan Qu*, Song Wang, Junbo Chen, Jianke Zhu
ICLR, 2026

PaperCode

photo
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
Yuqian Yuan*, Ronghao Dang*, Long Li*, Wentong Li*, Diao Jiao, Xin Li, Deli Zhao, Fan Wang, Wenqiao Zhang, Jun Xiao, Yueting Zhuang
NeurIPS (DB Track), 2025
photo
TokenPacker: Efficient Visual Projector for Multimodal LLM
Wentong Li*, Yuqian Yuan*, Jian Liu, Dongqi Tang, Song Wang, Jie Qin, Jianke Zhu, Lei Zhang
IJCV, 2025
photo
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
Yuqian Yuan, Hang Zhang, Wentong Li, Zesen Cheng, Boqiang Zhang, Long Li, Xin Li, Deli Zhao, Wenqiao Zhang, Yueting Zhuang, Jianke Zhu, Lidong Bing
CVPR, 2025
photo
Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning
Hanxun Yu*, Wentong Li*, Song Wang, Junbo Chen, Jianke Zhu
CVPR, 2025 (Highlight, 2.9%)

PaperCode

photo
Osprey: Pixel Understanding with Visual Instruction Tuning
Yuqian Yuan*, Wentong Li*+, Jian Liu, Dongqi Tang, Xinjie Luo, Chi Qin, Lei Zhang, Jianke Zhu
CVPR, 2024 (Project Leader)
photo
Box2Mask: Box-supervised Instance Segmentation via Level-set Evolution
Wentong Li, Wenyu Liu, Jianke Zhu, Miaomiao Cui, Risheng Yu, Xiansheng Hua, Lei Zhang
T-PAMI, 2024

Research Experience

Honors

  • Outstanding Doctoral Dissertation Award of Zhejiang Province, 2026
  • Outstanding Doctoral Dissertation Award of Zhejiang University, 2024
  • Excellent Doctoral Graduates of Zhejiang Province, China (Top 1%), 2024
  • Excellent Doctoral Graduates of Zhejiang University, 2024
  • Tencent Scholarship, 2023
  • Five-A Postgraduate Student, 2023
  • Outstanding Postgraduate Student, 2020-2023
  • Longhu Scholarship, 2022
  • First-class Academic Scholarship, 2018-2023
  • National Scholarship, 2016

Academic Service

  • Area Chair:
    ICLR2026&2027, NeurIPS2026, AAAI2027
  • Conference Reviewer:
    AAAI2025, ICLR2025, CVPR2025, ICML2025, ICCV2025, NeurIPS2025, ACM MM2025
    CVPR2024, ICLR2024, ICML2024, ECCV2024, ACM MM2024, NeurIPS2024
    CVPR2023, ICCV2023, NeurIPS2023, ACM MM2023
  • Journal Reviewer:
    Transactions on Pattern Analysis and Machine Intelligence (TPAMI)
    International Journal of Computer Vision (IJCV)
    Transactions on Image Processing (TIP)
    Transactions on Circuits and Systems for Video Technology (TCSVT)
    Transactions on Multimedia (TMM)
    Transactions on Geoscience and Remote Sensing (TGRS)
    Pattern Recognition (PR)
    ACM Computing Surveys
    ISPRS Journal of Photogrammetry and Remote Sensing (P&RS)
    Neurcomputing

Selected Talks

  • Efficient Visual Understanding and Interaction with VLMs, PolyU HongKong, slides, 2025/08.
  • Fine-grained Image Understanding with VLMs, ECNU, Visual Perception+X(VPX) Group, 2024/09.
  • Osprey:Pixel Understanding with Visual Instruction Tuning, Video, slides, AI TIME, 2024/01.
  • Point-supervised Image Segmentation, AntGroup, Machine Intelligence Group, 2023/09.

Teaching

  • Intro. to AI: A Foundational Course, NUAA, Fall 2025.
  • Foundations and Frontiers of Multimodal Large Models, NUAA, Spring 2025.
  • Image Processing and Analysis, Police Brain of Zhejiang Province, Teaching Assistant, Fall 2022.
  • FDS2021: Foundation of Data Structure, Zhejiang University, Teaching Assistant, Fall 2021.