TopPaper
01 August 2026
Today's Paper

Contrastive Latent Action Learning from Human Videos for Robotic Manipulation

Weisheng Dai, Kai Lan, Jianyi Zhou, Xiu Su, Junwen Tong, Weili Guan, Bo Zhao, Shuo Yang • arXiv (cs.RO)

ConLA is an unsupervised pretraining framework that extracts semantically consistent latent actions from unlabelled human demonstration videos without relying on explicit action tags. By utilizing contrastive disentanglement with action category and temporal priors, it isolates pure motion dynamics from complex visual background noise. For the first time, pretraining on human videos with ConLA surpasses the performance achieved by pretraining directly on real robot teleoperation trajectories.

View Full Abstract DOI: 10.48550/arXiv.2602.00557 Share

Keywords

Embodied AI Vision-Language-Action Latent Action Quantization Contrastive Learning Robotic Manipulation