Vision and Language
Scene understanding, video understanding, fine-grained visual understanding, and efficient architectures.
Computer Vision · Embodied AI
I am an Associate Professor in the College of Artificial Intelligence at Nanjing University of Aeronautics and Astronautics, working closely with Prof. Jie Qin.
In August 2025 and August 2026, I was a visiting researcher in the Department of Computing at The Hong Kong Polytechnic University, where I collaborated closely with my Ph.D. advisor, Prof. Lei Zhang (IEEE Fellow). Previously, I completed my Ph.D. in the College of Computer Science and Technology at Zhejiang University, supervised by Prof. Jianke Zhu and Prof. Lei Zhang, in June 2024.
I serve as an Area Chair for top-tier AI conferences, including ICLR 2026 and 2027, NeurIPS 2026, CVPR 2027, and AAAI 2027.
Scene understanding, video understanding, fine-grained visual understanding, and efficient architectures.
Embodied understanding, reasoning, planning, and action for intelligent agents in complex environments.
Multimodal learning, robust learning, and foundation models for vision, agent, and robotics.
Open positions. I am looking for self-motivated Master's students, research interns/assistants, and co-supervised Ph.D. students. Please email me if you are interested.
*: equal contribution, #: corresponding author, +: project leader
Project Page |
Paper |
Code |
HuggingFace |
PaperWeekly |
机器之心
Project Page |
Paper |
Code |
机器之心 |
CVer
Project Page |
Paper |
Code
Paper |
Project Page |
Code |
HuggingFace |
LeaderBoard |
中文解读
Paper |
ArXiv |
Code
|
HuggingFace Model |
中文解读 |
Daily Papers
Paper |
Code
|
HuggingFace Model |
Dataset |
VideoRefer-Bench |
中文解读 |
视频解读
Paper |
Code
|
Online Demo |
Video Demo |
中文解读 |
视频解读