# Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

> Kandinsky 6.0 Video releases open-source 3B and 29B diffusion models that generate synchronized 5-second video with 44 kHz audio, matching proprietary systems in speech quality.

- **Source:** [arXiv](https://arxiv.org/abs/2610.05608)
- **Published:** 2026-10-07
- **Permalink:** https://picx.dev/p/RnqLWZ
- **Whiteboard:** https://picx.dev/p/RnqLWZ/image

## Summary

# Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

## Summary (Overview)

- **Kandinsky 6.0 Video** is a family of foundation diffusion models for synchronized text-to-audio-video (T2AV) and image-to-audio-video (I2AV) generation, released in two sizes: **Lite (3B parameters)** and **Pro (29B parameters)**.
- Both models generate **5-second video clips with synchronized 44 kHz audio**, including lip-sync, and feature a built-in super-resolution model that upscales output to **Full-HD (1920×1080)**.
- The architecture uses a **dual-stream CrossDiT** design connecting a pretrained video stream (from Kandinsky 5.0) with a newly trained audio stream via **bidirectional cross-attention** for temporal and semantic alignment.
- A comprehensive training pipeline includes continuous pretraining, supervised fine-tuning with model soup merging, **OmniNFT-adapted reinforcement learning**, and two-stage distillation (π-Flow + adversarial refinement) to 10 function evaluations.
- In human evaluation, Kandinsky 6.0 Video Pro outperforms its predecessor Kandinsky 5.0 Video Pro and remains competitive with leading proprietary systems (Kling 2.6, Veo 3.1 Fast), particularly in speech quality. All code, weights, and diffusers integration are released under the **MIT license**.

## Introduction and Theoretical Foundation

### Background

Diffusion models [1, 2] and flow matching methods [3] have proven effective for image synthesis [4, 5]. Extending them to video generation has required specialized architectures based on diffusion transformers (DiT) [6] with efficient spatio-temporal attention mechanisms. The field includes closed-source systems (Sora [7], Veo [8]) and open-source models (HunyuanVideo [9], Wan [10], Kandinsky 5.0 [11]).

### The Challenge of Joint Audio-Video Generation

Generating video with synchronized audio is substantially more challenging than generating either modality alone, as it requires **semantic, temporal, and emotional consistency across modalities**. Recent works targeting joint generation include ALIVE [14], SyncFlow [15], JavisDiT [16], 3MDiT [17], LTX-2.5 [18], MiniMax-H3 [19], and NVIDIA Cosmos 3 [20].

### Key Contributions

The paper's contributions include:
1. Two model scales (Lite 3B, Pro 29B) with T2AV and I2AV modes
2. A continuous pretraining strategy for audio-video fusion
3. A multi-stage post-training pipeline (SFT + RL + distillation)
4. Comprehensive evaluation on VABench and human side-by-side comparisons
5. Open-source release under MIT license

## Methodology

### Architecture: Dual-Stream CrossDiT

The core model uses a **dual-stream CrossDiT architecture** with asymmetric DiT backbones for video and audio, connected by blockwise bidirectional cross-attention:

- **Video stream**: Inherited from Kandinsky 5.0 (19B parameters in Pro)
- **Audio stream**: Trained from scratch (5B parameters in Pro)
- **Cross-attention layers**: 5B parameters in Pro

**Positional encoding** uses Rotary Position Embedding (RoPE) [50]:
- Video tokens indexed by frame number, normalized by frame rate:
$$\text{RoPE}\left(\frac{N_{\text{frame}}}{\text{fps}/24}\right)$$
- Audio tokens indexed by temporal position

### Key Architectural Components

| Parameter | Lite | Pro |
|-----------|------|-----|
| Video model dim | 1792 | 4096 |
| Audio model dim | 896 | 2048 |
| Video blocks | 32 | 60 |
| Audio blocks | 32 | 60 |
| Total parameters | 3B | 29B |

**Text encoding**: Qwen2.5-VL-7B-Instruct + CLIP ViT-L/14
**Video VAE**: Hunyuan VAE
**Audio VAE**: MMAudio autoencoder (44 kHz)

### Super-Resolution Model

The SR pipeline consists of:
1. **Latent Upscaler (LU)**: Convolutional network that deterministically upsamples LQ latents by ×2 or ×4
2. **SR-DiT**: Flow-matching diffusion transformer (1.41B parameters) that removes degradations

The SR-DiT operates in the latent space of K-VAE (64 latent channels, 16× spatial, 4× temporal compression). For a clip of $T \times H \times W$ pixels, the latent is $T' \times \frac{H}{16} \times \frac{W}{16} \times 64$, where $T' = (T-1)/4 + 1$.

### Training Pipeline

#### 1. Continuous Pre-training

All stages use flow matching with MSE loss:

- **T2A (audio-only)**: Trained from scratch; Pro: 30k steps on audio captions + 4k steps on AV captions
- **T2V (video-only)**: Initialized from Kandinsky 5.0 checkpoint; Pro: 14k steps on T2AV captions
- **Joint T2AV**: Both streams trained together with independent diffusion time sampling
- **I2AV mode**: 25% of samples use image conditioning with Token Role Embeddings (TRE)

#### 2. Supervised Fine-Tuning

Two-stage process across 11 domains:
- **Stage 1 (visual quality)**: Video loss only, audio stream frozen
- **Stage 2 (sync/lip-sync)**: Weighted loss: video (0.85) + audio (0.15) + face loss (1.0)

Domain-specific models merged via **model soup** (uniform weight averaging).

#### 3. Reinforcement Learning

Adapted from OmniNFT [24] with:
- **8 reward objectives** in three branches: video (HPSv3, VideoAlign), audio (AudioBox, CLAP, Whisper-ASR), sync (AV-desync, AV-align, LatentSync-gated)
- **GRPO groups** of K=10 rollouts per prompt
- **LoRA fine-tuning** (r=16, α=32) on 0.69% of Pro parameters
- **Layer-wise gradient surgery** to prevent video gradients disrupting audio representations

The policy loss uses a DPO-like formulation:
$$\text{pos\_pred} = b_{\text{mix}} \cdot x^{\text{cur}}_0 + (1 - b_{\text{mix}}) \cdot x^{\text{old}}_0$$
$$\text{neg\_pred} = (1 + b_{\text{mix}}) \cdot x^{\text{old}}_0 - b_{\text{mix}} \cdot x^{\text{cur}}_0$$

With $b_{\text{mix}} = 1$: pos_pred = $x^{\text{cur}}_0$ and neg_pred = $2 \cdot x^{\text{old}}_0 - x^{\text{cur}}_0$.

#### 4. Distillation

Two-stage approach:
1. **π-Flow trajectory distillation**: On-policy imitation learning to 10 NFE
2. **Sim-LADD adversarial refinement**: Uses student's own backward rollouts with DRaFT-5 (gradients through final 5 solver steps)

### Super-Resolution Training

The SR-DiT training objective combines flow matching with pixel-space losses:

$$L = L_{\text{flow}} + \sum_i \lambda_i L_i$$

Where:
- $L_{\text{flow}} = \text{MSE}(v_\theta(x_t, t, a, m), \tilde{z}_{\text{LQ}} - z_{\text{HQ}})$
- Pixel L1 (weight 0.6), LPIPS (0.35), HFP (1.0), MFS (1.0)

The LQ latent is noised as: $\tilde{z}_{\text{LQ}} = \sqrt{1-s^2}\, z_{\text{LQ}} + s\varepsilon$, with $s = 0.7$.

## Empirical Validation / Results

### VABench Results

Kandinsky 6.0 Video Pro leads on speech quality, audio aesthetics, text-video and audio-video alignment, lip-sync accuracy, desynchronization, and visual realism.

| Metric | LTX 2.5 | K6 Lite | K6 Pro |
|--------|---------|---------|--------|
| speech_qn | 1.427 | 1.477 | **1.496** |
| audio_aes | 3.307 | 3.323 | **3.345** |
| tv_align | 0.185 | 0.227 | **0.229** |
| ta_align | 0.307 | **0.374** | 0.367 |
| av_align | 0.258 | 0.245 | **0.268** |
| lipsync | 1.567 | 1.761 | **2.072** |
| desync ↓ | 0.521 | 0.474 | **0.396** |
| visual_realism | 4.366 | 4.459 | **4.485** |
| audio_realism | 3.663 | **3.900** | 3.893 |
| WER ↓ | **0.099** | 0.103 | 0.117 |

### RL Impact on Speech Quality

RL post-training substantially improves speech intelligibility:
- **Pro**: WER decreases from 0.235 → 0.124 (**47% relative reduction**, p < 10⁻³)
- **Lite**: WER decreases from 0.179 → 0.127 (**29% reduction**, p < 10⁻³)
- Confirmed on public Harvard sentences: Pro WER drops from 0.251 → 0.180 (p = 0.005)

### Human Evaluation Summary

**vs Kandinsky 5.0 Video Pro**: Clear and consistent advantage across all criteria in both T2V and I2V modes.

**vs Kling 2.6**: Mixed results — Kling leads on visual aesthetics and audio quality; Kandinsky leads on audio-prompt alignment and camera control.

**vs LTX 2.5**: Kandinsky preferred on visual quality and speech quality (statistically significant); LTX retains small non-significant edge on general audio fidelity.

**vs Veo 3.1 Fast**: Clear split — Kandinsky produces cleaner visuals and better preserves input images; Veo better follows prompts and delivers stronger audio.

**vs MiniMax H3**: H3 preferred on most criteria; Kandinsky competitive on speech quality, camera motion, and AV synchronization.

**vs Seedance 2.0**: Seedance stronger on visual criteria; Kandinsky leads on speech quality and competitive on synchronization.

### Super-Resolution Validation

**SR-DiT metrics** (per scale, higher is better for MUSIQ/DOVER/CLIP-IQA):

| Scale | Stage | MUSIQ ↑ | DOVER ↑ | CLIP-IQA ↑ |
|-------|-------|---------|---------|------------|
| ×4 | Stage 1 | 48.92 | 0.5467 | 0.4426 |
| ×4 | Stage 2 | 50.38 | 0.5589 | 0.4569 |
| ×4 | Stage 3 | 50.41 | 0.5553 | 0.4587 |
| ×2 | Stage 3 | 65.81 | 0.7348 | 0.5906 |

**Latent Upscaler PSNR** (×4 route): Achieves higher PSNR than pixel-space bilinear upsampling in all groups (+1.30 dB mean improvement).

### Inference Performance

**Generation times (seconds)** for 5-second clips on consumer GPUs:

| GPU | Lite SD | Pro SD | Pro Full HD |
|-----|---------|--------|-------------|
| RTX 4090 | 437 | 936 | 1247 |
| RTX 5090 | 309 | 754 | 854 |
| H100 | 239 | 356 | 402 |

Block offloading enables deployment on 16-32 GB consumer GPUs with peak memory as low as 7.9 GiB at SD resolution.

## Theoretical and Practical Implications

### Architectural Innovation

The dual-stream CrossDiT architecture with bidirectional cross-attention provides a principled approach to modality fusion. The continuous pretraining strategy — training the audio stream separately before joint training — preserves unimodal fidelity while enabling cross-modal alignment.

### Practical Deployment

The paper demonstrates that large foundation models (29B parameters) can be made accessible on consumer hardware through:
- Block offloading (peak memory reduction of ~3.4×)
- VRAM budget presets for 32/24/16 GB GPUs
- MagCache step skipping and regional compilation
- Attention backend selection (FlashAttention 3 / SageAttention2++)

### Training Methodology Insights

The RL post-training shows remarkable effectiveness for audio-video alignment, with the layer-wise gradient surgery and region-wise loss reweighting proving critical for preventing modality interference. The two-stage distillation (π-Flow + Sim-LADD) maintains quality at 10 NFE while preserving teacher semantics and diversity.

## Conclusion

Kandinsky 6.0 Video represents a significant open-source contribution to synchronized audio-video generation. Key takeaways:

1. **Dual-stream CrossDiT** with bidirectional cross-attention effectively unifies video and audio generation
2. **Continuous pretraining** from separate to joint training enables efficient modality fusion
3. **Multi-stage post-training** (SFT + RL + distillation) delivers substantial quality improvements, particularly in speech intelligibility (47% WER reduction)
4. The model is **competitive with proprietary systems** in speech quality and synchronization, though gaps remain in visual quality and overall audio quality

**Limitations**: 5-second clip maximum, SD base resolution (Full HD via SR), and remaining gaps to strongest proprietary systems.

**Future directions**: Longer clips, higher native resolution, and additional reference-based conditioning.

The release of code, weights, and diffusers integration under the MIT license provides a strong foundation for further research in multimodal generation.

---

_Markdown view of https://picx.dev/p/RnqLWZ, served by PicX — AI-generated visual whiteboard summaries of research papers._
