Full text not available for this paper

Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

Summary (Overview)

  • Kandinsky 6.0 Video is a family of foundation diffusion models for synchronized text-to-audio-video (T2AV) and image-to-audio-video (I2AV) generation, released in two sizes: Lite (3B parameters) and Pro (29B parameters).
  • Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, and feature a built-in super-resolution model that upscales output to Full-HD (1920×1080).
  • The architecture uses a dual-stream CrossDiT design connecting a pretrained video stream (from Kandinsky 5.0) with a newly trained audio stream via bidirectional cross-attention for temporal and semantic alignment.
  • A comprehensive training pipeline includes continuous pretraining, supervised fine-tuning with model soup merging, OmniNFT-adapted reinforcement learning, and two-stage distillation (π-Flow + adversarial refinement) to 10 function evaluations.
  • In human evaluation, Kandinsky 6.0 Video Pro outperforms its predecessor Kandinsky 5.0 Video Pro and remains competitive with leading proprietary systems (Kling 2.6, Veo 3.1 Fast), particularly in speech quality. All code, weights, and diffusers integration are released under the MIT license.

Introduction and Theoretical Foundation

Background

Diffusion models [1, 2] and flow matching methods [3] have proven effective for image synthesis [4, 5]. Extending them to video generation has required specialized architectures based on diffusion transformers (DiT) [6] with efficient spatio-temporal attention mechanisms. The field includes closed-source systems (Sora [7], Veo [8]) and open-source models (HunyuanVideo [9], Wan [10], Kandinsky 5.0 [11]).

The Challenge of Joint Audio-Video Generation

Generating video with synchronized audio is substantially more challenging than generating either modality alone, as it requires semantic, temporal, and emotional consistency across modalities. Recent works targeting joint generation include ALIVE [14], SyncFlow [15], JavisDiT [16], 3MDiT [17], LTX-2.5 [18], MiniMax-H3 [19], and NVIDIA Cosmos 3 [20].

Key Contributions

The paper's contributions include:

  1. Two model scales (Lite 3B, Pro 29B) with T2AV and I2AV modes
  2. A continuous pretraining strategy for audio-video fusion
  3. A multi-stage post-training pipeline (SFT + RL + distillation)
  4. Comprehensive evaluation on VABench and human side-by-side comparisons
  5. Open-source release under MIT license

Methodology

Architecture: Dual-Stream CrossDiT

The core model uses a dual-stream CrossDiT architecture with asymmetric DiT backbones for video and audio, connected by blockwise bidirectional cross-attention:

  • Video stream: Inherited from Kandinsky 5.0 (19B parameters in Pro)
  • Audio stream: Trained from scratch (5B parameters in Pro)
  • Cross-attention layers: 5B parameters in Pro

Positional encoding uses Rotary Position Embedding (RoPE) [50]:

  • Video tokens indexed by frame number, normalized by frame rate:
RoPE(Nframefps/24)\text{RoPE}\left(\frac{N_{\text{frame}}}{\text{fps}/24}\right)
  • Audio tokens indexed by temporal position

Key Architectural Components

ParameterLitePro
Video model dim17924096
Audio model dim8962048
Video blocks3260
Audio blocks3260
Total parameters3B29B

Text encoding: Qwen2.5-VL-7B-Instruct + CLIP ViT-L/14 Video VAE: Hunyuan VAE Audio VAE: MMAudio autoencoder (44 kHz)

Super-Resolution Model

The SR pipeline consists of:

  1. Latent Upscaler (LU): Convolutional network that deterministically upsamples LQ latents by ×2 or ×4
  2. SR-DiT: Flow-matching diffusion transformer (1.41B parameters) that removes degradations

The SR-DiT operates in the latent space of K-VAE (64 latent channels, 16× spatial, 4× temporal compression). For a clip of T×H×WT \times H \times W pixels, the latent is T′×H16×W16×64T' \times \frac{H}{16} \times \frac{W}{16} \times 64, where T′=(T−1)/4+1T' = (T-1)/4 + 1.

Training Pipeline

1. Continuous Pre-training

All stages use flow matching with MSE loss:

  • T2A (audio-only): Trained from scratch; Pro: 30k steps on audio captions + 4k steps on AV captions
  • T2V (video-only): Initialized from Kandinsky 5.0 checkpoint; Pro: 14k steps on T2AV captions
  • Joint T2AV: Both streams trained together with independent diffusion time sampling
  • I2AV mode: 25% of samples use image conditioning with Token Role Embeddings (TRE)

2. Supervised Fine-Tuning

Two-stage process across 11 domains:

  • Stage 1 (visual quality): Video loss only, audio stream frozen
  • Stage 2 (sync/lip-sync): Weighted loss: video (0.85) + audio (0.15) + face loss (1.0)

Domain-specific models merged via model soup (uniform weight averaging).

3. Reinforcement Learning

Adapted from OmniNFT [24] with:

  • 8 reward objectives in three branches: video (HPSv3, VideoAlign), audio (AudioBox, CLAP, Whisper-ASR), sync (AV-desync, AV-align, LatentSync-gated)
  • GRPO groups of K=10 rollouts per prompt
  • LoRA fine-tuning (r=16, α=32) on 0.69% of Pro parameters
  • Layer-wise gradient surgery to prevent video gradients disrupting audio representations

The policy loss uses a DPO-like formulation:

pos_pred=bmix⋅x0cur+(1−bmix)⋅x0old\mathrm{pos\_pred} = b_{\text{mix}} \cdot x^{\text{cur}}_0 + (1 - b_{\text{mix}}) \cdot x^{\text{old}}_0 neg_pred=(1+bmix)⋅x0old−bmix⋅x0cur\mathrm{neg\_pred} = (1 + b_{\text{mix}}) \cdot x^{\text{old}}_0 - b_{\text{mix}} \cdot x^{\text{cur}}_0

With bmix=1b_{\text{mix}} = 1: pos_pred = x0curx^{\text{cur}}_0 and neg_pred = 2⋅x0old−x0cur2 \cdot x^{\text{old}}_0 - x^{\text{cur}}_0.

4. Distillation

Two-stage approach:

  1. π-Flow trajectory distillation: On-policy imitation learning to 10 NFE
  2. Sim-LADD adversarial refinement: Uses student's own backward rollouts with DRaFT-5 (gradients through final 5 solver steps)

Super-Resolution Training

The SR-DiT training objective combines flow matching with pixel-space losses:

L=Lflow+∑iλiLiL = L_{\text{flow}} + \sum_i \lambda_i L_i

Where:

  • Lflow=MSE(vθ(xt,t,a,m),z~LQ−zHQ)L_{\text{flow}} = \text{MSE}(v_\theta(x_t, t, a, m), \tilde{z}_{\text{LQ}} - z_{\text{HQ}})
  • Pixel L1 (weight 0.6), LPIPS (0.35), HFP (1.0), MFS (1.0)

The LQ latent is noised as: z~LQ=1−s2 zLQ+sε\tilde{z}_{\text{LQ}} = \sqrt{1-s^2}\, z_{\text{LQ}} + s\varepsilon, with s=0.7s = 0.7.

Empirical Validation / Results

VABench Results

Kandinsky 6.0 Video Pro leads on speech quality, audio aesthetics, text-video and audio-video alignment, lip-sync accuracy, desynchronization, and visual realism.

MetricLTX 2.5K6 LiteK6 Pro
speech_qn1.4271.4771.496
audio_aes3.3073.3233.345
tv_align0.1850.2270.229
ta_align0.3070.3740.367
av_align0.2580.2450.268
lipsync1.5671.7612.072
desync ↓0.5210.4740.396
visual_realism4.3664.4594.485
audio_realism3.6633.9003.893
WER ↓0.0990.1030.117

RL Impact on Speech Quality

RL post-training substantially improves speech intelligibility:

  • Pro: WER decreases from 0.235 → 0.124 (47% relative reduction, p < 10⁻³)
  • Lite: WER decreases from 0.179 → 0.127 (29% reduction, p < 10⁻³)
  • Confirmed on public Harvard sentences: Pro WER drops from 0.251 → 0.180 (p = 0.005)

Human Evaluation Summary

vs Kandinsky 5.0 Video Pro: Clear and consistent advantage across all criteria in both T2V and I2V modes.

vs Kling 2.6: Mixed results — Kling leads on visual aesthetics and audio quality; Kandinsky leads on audio-prompt alignment and camera control.

vs LTX 2.5: Kandinsky preferred on visual quality and speech quality (statistically significant); LTX retains small non-significant edge on general audio fidelity.

vs Veo 3.1 Fast: Clear split — Kandinsky produces cleaner visuals and better preserves input images; Veo better follows prompts and delivers stronger audio.

vs MiniMax H3: H3 preferred on most criteria; Kandinsky competitive on speech quality, camera motion, and AV synchronization.

vs Seedance 2.0: Seedance stronger on visual criteria; Kandinsky leads on speech quality and competitive on synchronization.

Super-Resolution Validation

SR-DiT metrics (per scale, higher is better for MUSIQ/DOVER/CLIP-IQA):

ScaleStageMUSIQ ↑DOVER ↑CLIP-IQA ↑
×4Stage 148.920.54670.4426
×4Stage 250.380.55890.4569
×4Stage 350.410.55530.4587
×2Stage 365.810.73480.5906

Latent Upscaler PSNR (×4 route): Achieves higher PSNR than pixel-space bilinear upsampling in all groups (+1.30 dB mean improvement).

Inference Performance

Generation times (seconds) for 5-second clips on consumer GPUs:

GPULite SDPro SDPro Full HD
RTX 40904379361247
RTX 5090309754854
H100239356402

Block offloading enables deployment on 16-32 GB consumer GPUs with peak memory as low as 7.9 GiB at SD resolution.

Theoretical and Practical Implications

Architectural Innovation

The dual-stream CrossDiT architecture with bidirectional cross-attention provides a principled approach to modality fusion. The continuous pretraining strategy — training the audio stream separately before joint training — preserves unimodal fidelity while enabling cross-modal alignment.

Practical Deployment

The paper demonstrates that large foundation models (29B parameters) can be made accessible on consumer hardware through:

  • Block offloading (peak memory reduction of ~3.4×)
  • VRAM budget presets for 32/24/16 GB GPUs
  • MagCache step skipping and regional compilation
  • Attention backend selection (FlashAttention 3 / SageAttention2++)

Training Methodology Insights

The RL post-training shows remarkable effectiveness for audio-video alignment, with the layer-wise gradient surgery and region-wise loss reweighting proving critical for preventing modality interference. The two-stage distillation (π-Flow + Sim-LADD) maintains quality at 10 NFE while preserving teacher semantics and diversity.

Conclusion

Kandinsky 6.0 Video represents a significant open-source contribution to synchronized audio-video generation. Key takeaways:

  1. Dual-stream CrossDiT with bidirectional cross-attention effectively unifies video and audio generation
  2. Continuous pretraining from separate to joint training enables efficient modality fusion
  3. Multi-stage post-training (SFT + RL + distillation) delivers substantial quality improvements, particularly in speech intelligibility (47% WER reduction)
  4. The model is competitive with proprietary systems in speech quality and synchronization, though gaps remain in visual quality and overall audio quality

Limitations: 5-second clip maximum, SD base resolution (Full HD via SR), and remaining gaps to strongest proprietary systems.

Future directions: Longer clips, higher native resolution, and additional reference-based conditioning.

The release of code, weights, and diffusers integration under the MIT license provides a strong foundation for further research in multimodal generation.

Related papers