dorsal/arxiv
View SchemaRotate Your Character: Revisiting Video Diffusion Models for High-Quality 3D Character Generation
| Authors | Jin Wang, Jianxiang Lu, Comi Chen, Guangzheng Xu, Haoyu Yang, Peng Chen, Na Zhang, Yifan Xu, Longhuang Wu, Shuai Shao, Qinglin Lu, Ping Luo |
|---|---|
| Categories | |
| ArXiv ID | 2601.05722vv1 |
| URL | https://arxiv.org/abs/2601.05722 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Generating high-quality 3D characters from single images remains a significant challenge in digital content creation, particularly due to complex body poses and self-occlusion. In this paper, we present RCM (Rotate your Character Model), an advanced image-to-video diffusion framework tailored for high-quality novel view synthesis (NVS) and 3D character generation. Compared to existing diffusion-based approaches, RCM offers several key advantages: (1) transferring characters with any complex poses into a canonical pose, enabling consistent novel view synthesis across the entire viewing orbit, (2) high-resolution orbital video generation at 1024x1024 resolution, (3) controllable observation positions given different initial camera poses, and (4) multi-view conditioning supporting up to 4 input images, accommodating diverse user scenarios. Extensive experiments demonstrate that RCM outperforms state-of-the-art methods in both novel view synthesis and 3D generation quality.
{
"annotation_id": "fb82b88d-dcf6-4e96-81af-825ff276b021",
"date_created": "2026-02-17T05:53:04.667000Z",
"date_modified": "2026-02-17T05:53:04.667000Z",
"file_hash": "4662634ecd5ec0c0cd5cecbe230a48cae7658abf016d6c1d9c46382c0b0ed2a0",
"private": false,
"record": {
"abstract": "Generating high-quality 3D characters from single images remains a significant challenge in digital content creation, particularly due to complex body poses and self-occlusion. In this paper, we present RCM (Rotate your Character Model), an advanced image-to-video diffusion framework tailored for high-quality novel view synthesis (NVS) and 3D character generation. Compared to existing diffusion-based approaches, RCM offers several key advantages: (1) transferring characters with any complex poses into a canonical pose, enabling consistent novel view synthesis across the entire viewing orbit, (2) high-resolution orbital video generation at 1024x1024 resolution, (3) controllable observation positions given different initial camera poses, and (4) multi-view conditioning supporting up to 4 input images, accommodating diverse user scenarios. Extensive experiments demonstrate that RCM outperforms state-of-the-art methods in both novel view synthesis and 3D generation quality.",
"arxiv_id": "2601.05722",
"authors": [
"Jin Wang",
"Jianxiang Lu",
"Comi Chen",
"Guangzheng Xu",
"Haoyu Yang",
"Peng Chen",
"Na Zhang",
"Yifan Xu",
"Longhuang Wu",
"Shuai Shao",
"Qinglin Lu",
"Ping Luo"
],
"categories": [
"cs.CV"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Rotate Your Character: Revisiting Video Diffusion Models for High-Quality 3D Character Generation",
"url": "https://arxiv.org/abs/2601.05722",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "cdd0c878-6fc6-4ff2-863b-7bca5d22abcc",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}