dorsal/arxiv
View SchemaCEI: A Unified Interface for Cross-Embodiment Visuomotor Policy Learning in 3D Space
| Authors | Tong Wu, Shoujie Li, Junhao Gong, Changqing Guo, Xingting Li, Shilong Mu, Wenbo Ding |
|---|---|
| Categories | |
| ArXiv ID | 2601.09163vv1 |
| URL | https://arxiv.org/abs/2601.09163 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Robotic foundation models trained on large-scale manipulation datasets have shown promise in learning generalist policies, but they often overfit to specific viewpoints, robot arms, and especially parallel-jaw grippers due to dataset biases. To address this limitation, we propose Cross-Embodiment Interface (\CEI), a framework for cross-embodiment learning that enables the transfer of demonstrations across different robot arm and end-effector morphologies. \CEI introduces the concept of \textit{functional similarity}, which is quantified using Directional Chamfer Distance. Then it aligns robot trajectories through gradient-based optimization, followed by synthesizing observations and actions for unseen robot arms and end-effectors. In experiments, \CEI transfers data and policies from a Franka Panda robot to \textbf{16} different embodiments across \textbf{3} tasks in simulation, and supports bidirectional transfer between a UR5+AG95 gripper robot and a UR5+Xhand robot across \textbf{6} real-world tasks, achieving an average transfer ratio of 82.4\%. Finally, we demonstrate that \CEI can also be extended with spatial generalization and multimodal motion generation capabilities using our proposed techniques. Project website: https://cross-embodiment-interface.github.io/
{
"annotation_id": "19d1fa34-21c3-4176-ae5f-91b8773ee145",
"date_created": "2026-02-17T05:53:19.855000Z",
"date_modified": "2026-02-17T05:53:19.855000Z",
"file_hash": "faa8c5e14e005893c988eb9ea1720fe61c4e0f3c46e4701c5fb9988d8c5962be",
"private": false,
"record": {
"abstract": "Robotic foundation models trained on large-scale manipulation datasets have shown promise in learning generalist policies, but they often overfit to specific viewpoints, robot arms, and especially parallel-jaw grippers due to dataset biases. To address this limitation, we propose Cross-Embodiment Interface (\\CEI), a framework for cross-embodiment learning that enables the transfer of demonstrations across different robot arm and end-effector morphologies. \\CEI introduces the concept of \\textit{functional similarity}, which is quantified using Directional Chamfer Distance. Then it aligns robot trajectories through gradient-based optimization, followed by synthesizing observations and actions for unseen robot arms and end-effectors. In experiments, \\CEI transfers data and policies from a Franka Panda robot to \\textbf{16} different embodiments across \\textbf{3} tasks in simulation, and supports bidirectional transfer between a UR5+AG95 gripper robot and a UR5+Xhand robot across \\textbf{6} real-world tasks, achieving an average transfer ratio of 82.4\\%. Finally, we demonstrate that \\CEI can also be extended with spatial generalization and multimodal motion generation capabilities using our proposed techniques. Project website: https://cross-embodiment-interface.github.io/",
"arxiv_id": "2601.09163",
"authors": [
"Tong Wu",
"Shoujie Li",
"Junhao Gong",
"Changqing Guo",
"Xingting Li",
"Shilong Mu",
"Wenbo Ding"
],
"categories": [
"cs.RO"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "CEI: A Unified Interface for Cross-Embodiment Visuomotor Policy Learning in 3D Space",
"url": "https://arxiv.org/abs/2601.09163",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "bd890ee8-c831-473b-ace5-2a3cfb5b1c61",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}