About

Hi there! I'm currently an undergraduate student in the Department of Computer Science and Technology at Harbin Institute of Technology (Shenzhen), supervised by Prof. Rui Shao. I'm pursuing a B.Eng. in Computer Science and Technology from September 2023 to June 2027.

My current research interests lie in vision-language-action models, embodied intelligence, robotic manipulation, world model, and efficient model adaptation for real-world decision making.

News

  1. Academic homepage moved to minijie3.github.io, with the blog hosted at /blog.
  2. Our work on hierarchical interaction refinement, H-GAR, has been accepted to AAAI 2026 (Oral).
  3. Our work on cognition-aligned routing, CogVLA, has been accepted to NeurIPS 2025.

Research Interests

  • Vision-Language-Action Models
  • Embodied AI
  • Robotic Manipulation
  • World Model
  • Efficient Multimodal Learning

Publications

* Equal contribution. † Corresponding author.

DeltaVLA paper preview

ΔVLA: Prior-Guided Vision-Language-Action Models via World Knowledge Variation

Yijie Zhu, Jie He, Rui Shao†, Kaishen Yuan, Tao Tan, Xiaochen Yuan, Zitong Yu†

arXiv preprint arXiv:2603.08361, 2026.

ΔVLA studies how world knowledge variation can guide vision-language-action models. The work uses prior knowledge to improve action reasoning and adaptation, aiming to make embodied policies more reliable under changing task and scene conditions.

H-GAR paper preview

H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic Manipulation

Yijie Zhu, Rui Shao†, Ziyang Liu, Jie He, Jizhihui Liu, Jiuru Wang, Zitong Yu†

AAAI 2026 (Oral).

H-GAR introduces a hierarchical interaction framework for robotic manipulation. It refines observations and actions according to task goals, improving the robot's ability to reason over long-horizon interactions and execute manipulation steps more precisely.

CogVLA paper preview

CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing and Sparsification

Wei Li, Renshan Zhang, Rui Shao†, Jie He, Liqiang Nie

NeurIPS 2025.

CogVLA aligns vision-language-action models with cognitive execution patterns through instruction-driven routing and sparsification. It targets efficient, task-aware reasoning so the model can activate the most relevant pathways for embodied decision making.

Education

Harbin Institute of Technology

2023.09 - now, B.Eng in Computer Science and Technology, Harbin Institute of Technology (Shenzhen).