I am a PhD student at Fudan University and Shanghai Innovation Institute. Prior to that, I received my Bachelor’s degree from Sun Yat-sen University. My research interests include robotics, particularly tactile sensing and perception.

🔥 News

  • 2026.07:  🎉🎉 DriveWeaver has been accepted by ECCV 2026.
  • 2026.05:  🎉🎉 MoLA has been accepted by ICML 2026.
  • 2025.06:  🎉🎉 BézierGS has been accepted by ICCV 2025.

📝 Publications

Tech Report 2026
N0-TWAM teaser

$\mathcal{N}_0$-TWAM: Scaling Tactile-Native World Action Model for Contact-Rich Manipulation

NeoteAI Team & Fudan TEAI Team

Project Code

  • TL;DR: We introduce $\mathcal{N}_0$-TWAM, a tactile-native world-action model that jointly predicts future vision, touch, and actions, enabling stronger contact-rich manipulation through large-scale visuo-tactile pretraining.
Tech Report 2026
N0-Foundation teaser

$\mathcal{N}_0$-Foundation: Towards the Age of Tactile Intelligence

NeoteAI Team & Fudan TEAI Team

Project Dataset

  • TL;DR: We present $\mathcal{N}_0$-Foundation, a tactile-centric foundation that unifies scalable sensing hardware, large-scale multimodal data, transferable tactile representations, and standardized real-world and simulated evaluation.
ECCV 2026
DriveWeaver pipeline

DriveWeaver: Point-Conditioned Video Inpainting for Controllable Vehicle Insertion in Autonomous Driving Simulation

Junzhe Jiang, Zipei Ma, Zijie Pan, Li Zhang†

Code

  • TL;DR: We propose DriveWeaver, a point-cloud-conditioned video inpainting framework that inserts controllable, temporally consistent vehicles and extracts their 3D Gaussian representations for real-time autonomous driving simulation.
arXiv 2026
Metis framework

Metis: A Generalizable and Efficient World-Action Model for Autonomous Driving and Urban Navigation

Jingyu Li*, Zhe Liu*, Dongnan Hu, Junjie Wu, Zipei Ma, Wenxiao Wu, Chao Han, Zhihui Hao, Zhikang Liu, Kun Zhan, Jiankang Deng, Xiatian Zhu, Li Zhang†

Code

  • TL;DR: We introduce Metis, an efficient world-action model that decouples video generation and action prediction through specialized transformer experts and asymmetric attention, retaining world-model supervision during training while bypassing video generation at inference.
ICML 2026
MoLA pipeline

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation

Yajie Li*, Bozhou Zhang*, Chun Gu, Zipei Ma, Jiahui Zhang, Jiankang Deng, Xiatian Zhu, Li Zhang

Project Code

  • TL;DR: We introduce MoLA, which transforms imagined future videos into executable representations through a mixture of pretrained inverse dynamics models, achieving consistent gains in simulation and real-world robot manipulation.
arXiv 2025
ProphRL thumbnail

Reinforcing Action Policies by Prophesying

Jiahui Zhang*, Ze Huang*, Chun Gu, Zipei Ma, Li Zhang†

Project Code

  • TL;DR: We introduce ProphRL, which learns a reusable action-to-video world model and a flow-aware RL post-training strategy to improve VLA policies efficiently, delivering strong gains in both simulation and real-robot tasks.
ICCV 2025
BézierGS pipeline

BézierGS: Dynamic Urban Scene Reconstruction with Bézier Curve Gaussian Splatting

Zipei Ma*, Junzhe Jiang*, Yurui Chen, Li Zhang†

Code

  • TL;DR: We propose a method that models object motion using Bézier curves, achieving accurate reconstruction and effective foreground-background separation.

📖 Educations

  • 2025.06 - Present, Fudan University & SII
  • 2021.09 - 2025.06, Sun Yat-sen University

🤣 Hobbies

  • 🏸 Badminton, ⚽️ Football(Visca el Barça! 💙❤️)
  • 🎮 Games(Soulslike, Multiplayer…)
  • 🎬 Movies, 🎵 Music