DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
Published in arXiv preprint arXiv:2603.10448, 2026
Recommended citation: Ma, Teli and Zheng, Jia and Wang, Zifan and Jiang, Chunli and Cui, Andy and Liang, Junwei and Yang, Shuo. "DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control." arXiv preprint arXiv:2603.10448, 2026. https://arxiv.org/abs/2603.10448
DiT4DiT is an end-to-end video-action model that couples video and action Diffusion Transformers for generalizable robot control across simulation and real-world humanoid manipulation tasks.
