Category
Robotics & Embodied AI
1 Paper
August 01, 2026
Robotics & Embodied AI
Contrastive Latent Action Learning from Human Videos for Robotic Manipulation
ConLA is an unsupervised pretraining framework that extracts semantically consistent latent actions from unlabelled human demonstration videos without relying on explicit action tags. By utilizing contrastive disentanglement with action category and temporal priors, it isolates pure motion dynamics from complex …
Weisheng Dai, Kai …
2 views
0 downloads