Despite recent progress, direct text-driven 4D object generation remains challenging yet highly desirable. In this paper, we introduce DiTex4D, a native text-to-4D generation framework that enables both text-driven 4D generation from scratch and 3D animation from static mesh.
Built upon large-scale pre-trained 3D generation models, our framework DiTex4D avoids intermediate text-to-video pipelines and costly per-object optimization. Specifically, (i) we achieve 4D spatiotemporal consistency via inflating 3D attention with mixed-4D RoPE and tailored correlated noise injection strategy. (ii) To enable 3D animation, we introduce a mask-based diffusion model conditioned on multi-view global context to maintain strict consistency with the initial frame. We further fine-tune the framework for 4D interpolation to synthesize high-frame-rate sequences with smoother motion.
Extensive experiments demonstrate that DiTex4D can achieve higher-quality, semantically align-ed, and spatiotemporally coherent 4D object generation, surpassing most existing state-of-the-art text-to-4D generation methods.
Prompt: “A man in red clothes is running.”
Prompt: “Banana is dancing.”
Prompt: “A cartoon character is raising his hands.”
Prompt: “The whale is swimming.”
Prompt: “A man wearing a blue coat is walking.”
Prompt: “A boy wearing glasses is jogging.”
Prompt: “The dragon is soaring with wings beating.”
Prompt: “Robot is running.”
Prompt: “The 3D airplane is rotating.”
Prompt: “A man is squatting.”
Prompt: “The 3D rabbit is walking forward.”
Prompt: “The spider is crawling.”
@article{chen2026ditex4d,
author = {Chen, Xiaozhe and Rong, Mengqi and Liu, Jian and Shen, Shuhan},
title = {DiTex4D: Direct Text-Driven 4D Generation with Structured Latent Diffusion},
journal = {ECCV},
year = {2026},
}