dorsal/arxiv
View SchemaInference-time Physics Alignment of Video Generative Models with Latent World Models
| Authors | Jianhao Yuan, Xiaofeng Zhang, Felix Friedrich, Nicolas Beltran-Velez, Melissa Hall, Reyhane Askari-Hemmat, Xiaochuang Han, Nicolas Ballas, Michal Drozdzal, Adriana Romero-Soriano |
|---|---|
| Categories | |
| ArXiv ID | 2601.10553vv1 |
| URL | https://arxiv.org/abs/2601.10553 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
State-of-the-art video generative models produce promising visual content yet often violate basic physics principles, limiting their utility. While some attribute this deficiency to insufficient physics understanding from pre-training, we find that the shortfall in physics plausibility also stems from suboptimal inference strategies. We therefore introduce WMReward and treat improving physics plausibility of video generation as an inference-time alignment problem. In particular, we leverage the strong physics prior of a latent world model (here, VJEPA-2) as a reward to search and steer multiple candidate denoising trajectories, enabling scaling test-time compute for better generation performance. Empirically, our approach substantially improves physics plausibility across image-conditioned, multiframe-conditioned, and text-conditioned generation settings, with validation from human preference study. Notably, in the ICCV 2025 Perception Test PhysicsIQ Challenge, we achieve a final score of 62.64%, winning first place and outperforming the previous state of the art by 7.42%. Our work demonstrates the viability of using latent world models to improve physics plausibility of video generation, beyond this specific instantiation or parameterization.
{
"annotation_id": "6a38f7fb-e5b6-49dd-ac68-e5c889b07fd6",
"date_created": "2026-02-17T05:53:24.060000Z",
"date_modified": "2026-02-17T05:53:24.060000Z",
"file_hash": "13cbb4556138fc0b7140278a4278b6d4ac5849cc406ad33c17cb071f35f43174",
"private": false,
"record": {
"abstract": "State-of-the-art video generative models produce promising visual content yet often violate basic physics principles, limiting their utility. While some attribute this deficiency to insufficient physics understanding from pre-training, we find that the shortfall in physics plausibility also stems from suboptimal inference strategies. We therefore introduce WMReward and treat improving physics plausibility of video generation as an inference-time alignment problem. In particular, we leverage the strong physics prior of a latent world model (here, VJEPA-2) as a reward to search and steer multiple candidate denoising trajectories, enabling scaling test-time compute for better generation performance. Empirically, our approach substantially improves physics plausibility across image-conditioned, multiframe-conditioned, and text-conditioned generation settings, with validation from human preference study. Notably, in the ICCV 2025 Perception Test PhysicsIQ Challenge, we achieve a final score of 62.64%, winning first place and outperforming the previous state of the art by 7.42%. Our work demonstrates the viability of using latent world models to improve physics plausibility of video generation, beyond this specific instantiation or parameterization.",
"arxiv_id": "2601.10553",
"authors": [
"Jianhao Yuan",
"Xiaofeng Zhang",
"Felix Friedrich",
"Nicolas Beltran-Velez",
"Melissa Hall",
"Reyhane Askari-Hemmat",
"Xiaochuang Han",
"Nicolas Ballas",
"Michal Drozdzal",
"Adriana Romero-Soriano"
],
"categories": [
"cs.CV"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Inference-time Physics Alignment of Video Generative Models with Latent World Models",
"url": "https://arxiv.org/abs/2601.10553",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "d1573511-7b21-45cd-acfe-9352595ad104",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}