DualMoE-GAN technical report / preprint. Text-to-video diffusion models achieve strong text alignment and coarse motion but struggle with photorealistic detail, temporal sharpness, and flicker-free fine structure. A likely cause is objective interference: a single denoiser must simultaneously support stable diffusion learning across all noise levels while recovering high-frequency structure favored by adversarial objectives. DualMoE-GAN is a video generation framework combining two coordinated experts within a shared diffusion-transformer backbone. Expert-G is trained with diffusion supervision alone to preserve semantic coverage and denoising stability. Expert-R is trained with both diffusion and adversarial supervision from a text-conditioned temporal discriminator to emphasize realism and motion sharpness. A learnable spatial-temporal-timestep router assigns token-wise expert mixtures for each latent region and denoising step. To improve optimization stability, progressive adversarial injection gradually increases adversarial influence and initially restricts it to low-noise denoising stages. The model is formulated as a router-mediated minimax optimization over two generator experts and one discriminator. Theoretical analysis includes a local Nash equilibrium existence proof for the three-player game, Lyapunov stability analysis for the progressive adversarial injection schedule, and an FID improvement bound showing the adversarial expert provides a non-negative distributional correction. Empirically, DualMoE-GAN improves VBench from 80.1 to 82.7, reduces UCF-101 FVD from 411.0 to 356.8 relative to the strongest open baseline in the comparison set, and attains a 58.4% human win rate against a matched dense model, while incurring 1.49 training FLOPs and 1.08 inference FLOPs under hard top-1 routing. Existing OSF archival DOI: 10.17605/OSF.IO/KESWT; Existing OSF archival page: https://osf.io/keswt/. Files include the technical report PDF and the LaTeX source tarball when available.
Haopeng Jin (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: