Vision and Language
Scene understanding, video understanding, fine-grained visual understanding, and efficient architectures.
Representative work
Computer Vision · Embodied AI
I am an Associate Professor in the College of Artificial Intelligence at Nanjing University of Aeronautics and Astronautics, working closely with Prof. Jie Qin.
In August 2025 and August 2026, I was a visiting researcher in the Department of Computing at The Hong Kong Polytechnic University, where I collaborated closely with my Ph.D. advisor, Prof. Lei Zhang (IEEE Fellow). Previously, I completed my Ph.D. in the College of Computer Science and Technology at Zhejiang University, supervised by Prof. Jianke Zhu and Prof. Lei Zhang, in June 2024.
I serve as an Area Chair for top-tier AI conferences, including CVPR 2027, ICLR 2026 and 2027, NeurIPS 2026, and AAAI 2027.
Scene understanding, video understanding, fine-grained visual understanding, and efficient architectures.
Representative work
Embodied understanding, reasoning, planning, and action for intelligent agents in complex environments.
Representative work
Multimodal learning, post-training (RL, OPD), optimization theory for multimodal foundation models in vision, agent and robotics.
Representative work
Open positions. I am looking for self-motivated Master's students, research interns/assistants, and co-supervised Ph.D. students. Please email me if you are interested.
*: equal contribution, #: corresponding author, +: project leader
Project Page |
Paper |
Code |
HuggingFace |
PaperWeekly |
机器之心
Project Page |
Paper |
Code |
机器之心 |
CVer
Project Page |
Paper |
Code
Paper |
Project Page |
Code |
HuggingFace |
LeaderBoard |
中文解读
Paper |
ArXiv |
Code
|
HuggingFace Model |
中文解读 |
Daily Papers
Paper |
Code
|
HuggingFace Model |
Dataset |
VideoRefer-Bench |
中文解读 |
视频解读
Paper |
Code
|
Online Demo |
Video Demo |
中文解读 |
视频解读