# MaLiang-Harness: A Programmable Path to Image and Video Generation

> MaLiang-Harness shows that MLLM-driven programmatic generation achieves 100% success on image and video tasks, but visual quality requires revision-aware verification beyond program correctness alone.

- **Source:** [arXiv](https://arxiv.org/abs/2609.34309)
- **Published:** 2026-10-01
- **Permalink:** https://picx.dev/p/3vDx97
- **Whiteboard:** https://picx.dev/p/3vDx97/image

## Summary

## Summary (Overview)

- **MaLiang-Harness** is a unified framework for programmable image and video generation, where multimodal large language models (MLLMs) translate creative prompts into executable visual programs rendered by backends like Canvas, SVG, Scene2d, or Three.js.
- The paper introduces the **Program-to-Visual (P2V) gap**: a program can execute correctly yet produce visuals that violate user intent (wrong composition, appearance, or motion).
- Three core mechanisms—**Persistent Executable Generation (PEG) State**, **Traceable Generation Process (TGP)**, and **Revision-aware Editing and Verification (REV)**—coordinate planning, execution, and visual feedback across rendering backends.
- Evaluations on **MaLiang-IBench** (50 image tasks, 11 models) and **MaLiang-VBench** (13 video tasks, 4 models) show **GPT-6-Astra** achieves 100% generation success on both, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds.
- General MLLM capability scores (e.g., AA Intelligence Index) correlate only weakly with visual generation performance; similarly scored models can differ dramatically in visual outcomes (e.g., GPT-5.6-Luna vs. GPT-6-Luna: 44% vs. 88% quality pass rates despite identical index scores).

---

## Introduction and Theoretical Foundation

### Background and Motivation

Traditional image/video generation follows a **direct visual synthesis** paradigm (GANs, diffusion models, flow matching) that learns data distributions and directly synthesizes pixels. This approach has achieved remarkable success but remains **implicit**—users cannot inspect or control the construction process.

The rapid improvement of MLLMs in semantic understanding, reasoning, and code generation makes a **programmable path** viable: the model expresses creative intent as an executable visual program, which a renderer converts into an image or video. This enables:
- Explicit control over spatial composition and temporal dynamics
- Targeted editing of generated artwork
- Inspectable construction processes

### The Program-to-Visual (P2V) Gap

A program can execute correctly while producing visuals that fail user intent. The paper defines this as:

> **P2V Gap**: The discrepancy between program-level correctness and visual requirement satisfaction.

Bridging this gap requires the model to:
1. Reason about the **visual consequences** of its code
2. Account for **rendering backend** choices in initial planning
3. Use **rendered output as evidence** for revising both program and plan
4. **Preserve artwork across iterations** while addressing unmet requirements
5. **Re-verify** after each revision, since changes can affect previously satisfied requirements

### Theoretical Foundation

The framework is built on the insight that code directly specifies **how** visual content is constructed—from spatial composition to motion over time. Images and videos share the same creation framework: at a fixed program revision, an image is rendered at a specified content time $t$, while a video samples temporal behavior encoded by the program.

The paper draws on prior work in:
- **Executable visual representations** (VISPROG, Design2Code, BlenderAlchemy)
- **Agent harnesses** (Code as Agent Harness, Show-Harness, OmniHarness)

---

## Methodology

### Unified Programmatic Visual Generation Interface

Given a prompt $p$ and output specification $\omega$, MaLiang-Harness plans appearance, spatial composition, and temporal dynamics, selects a backend $b$, and generates drawing/animation code. The backend renders the representation into an image or video that the MLLM evaluates against task requirements.

### Persistent Executable Generation (PEG) State

At revision $k$, the state is defined as:

$$
S_{k} = \left(P_{k}, A_{k}, Z_{k}, C_{k}, k\right),\tag{1}
$$

where:
- $P_k$: visual program with backend identifier
- $A_k$: associated assets (retained with content hashes)
- $Z_k$: spatial composition and temporal dynamics (via scene attributes or program-embedded definitions)
- $C_k$: generation context (prompt $p$, output spec $\omega$, user requirements, current plan)

State transitions and rendering:

$$
S_{k+1} = \mathcal{E}(S_{k}, a_{k}), \quad I_{k}(t) = \mathcal{R}_{b}(S_{k}, t; \omega),\tag{2}
$$

where $\mathcal{E}$ commits an edit $a_k$ (validating inputs and checking source revision currency), and $\mathcal{R}_b$ evaluates the program at content time $t$.

### Traceable Generation Process (TGP)

For the $j$-th recorded operation:

$$
\tau_{j} = \big(o_{j}, x_{j}, y_{j}, k_{j}^{-}, k_{j}^{+}\big),\tag{3}
$$

where $o_j$ is the executed operation, $x_j$ and $y_j$ are inputs and results (including errors), and $k_j^-$, $k_j^+$ identify source and resulting PEG revisions. The operation index $j$ is distinct from revision index $k$—operations that don't commit state updates (rendering, inspection) retain the same revision.

### Revision-aware Editing and Verification (REV)

For each visual requirement $h_i$ in $C_k$, REV maintains a review $q_i = (k_i, E_i, v_i)$ where $E_i$ contains requirement-specific visual evidence from revision $k_i$ and $v_i \in \{\text{pass}, \text{fail}, \text{uncertain}\}$.

A review applies to the current state only when $k_i = k$. After each commit, the current revision must be re-reviewed. The delivery condition is:

$$
\begin{array}{c} \operatorname{Ready}(S_{k}) = \operatorname{ExportOK}(S_{k}) \wedge \operatorname{CheckpointOK}(k) \\ \wedge \bigwedge_{h_{i} \in \mathcal{H}_{k}} [k_{i} = k \wedge E_{i} \neq \varnothing \wedge v_{i} = \text{pass}], \end{array}\tag{4}
$$

where $\mathcal{H}_k$ contains mandatory requirements, `ExportOK` verifies source files/assets and output conformance to $\omega$, and `CheckpointOK` holds when all checkpoints pass for the current revision.

---

## Empirical Validation / Results

### MaLiang-IBench (Image Tasks)

**Table 2: Generation Success and Computational Cost**

| Model | Success ↑(%) | Failures ↓ | Token Limit ↓ | Time/Qualified ↓(min/img) | Calls | Output Tokens (k/case) |
|---|---|---|---|---|---|---|
| DeepSeek-V4.1-Flash | 12.0 | 44 | 5 | 27.15 | 2.48 | 24.2 |
| DeepSeek-V4-Pro | 18.0 | 41 | 3 | 48.12 | 4.50 | 39.6 |
| Kimi-K2.6 | 34.0 | 33 | 1 | 62.69 | 9.14 | 48.2 |
| Kimi-K2.7-Code | 40.0 | 30 | 1 | 29.84 | 12.94 | 59.0 |
| Kimi-K3 | 18.0 | 41 | 0 | 34.88 | 2.82 | 17.8 |
| GPT-5.6-Luna | 92.0 | 4 | 0 | 4.49 | 11.10 | 9.3 |
| GPT-5.6-Terra | 92.0 | 4 | 0 | 4.32 | 8.32 | 8.0 |
| GPT-5.6-Sol | 96.0 | 2 | 1 | 3.28 | 6.98 | 13.1 |
| GPT-6-Luna | 96.0 | 2 | 0 | 4.92 | 14.70 | 16.4 |
| GPT-6-Sol | 96.0 | 2 | 0 | 3.52 | 7.42 | 10.2 |
| **GPT-6-Astra** | **100.0** | **0** | **0** | **3.70** | **6.34** | **8.1** |

**Table 3: Visual Quality on MaLiang-IBench** (counts = samples scoring ≥4/5)

| Model | Successful ↑ | All Criteria ↑ | Alignment ↑ | Aesthetics ↑ | Composition ↑ |
|---|---|---|---|---|---|
| DeepSeek-V4.1-Flash | 6 | 6 | 6 | 6 | 6 |
| DeepSeek-V4-Pro | 9 | 8 | 8 | 8 | 9 |
| Kimi-K2.6 | 17 | 6 | 7 | 8 | 9 |
| Kimi-K2.7-Code | 20 | 12 | 13 | 13 | 13 |
| Kimi-K3 | 9 | 9 | 9 | 9 | 9 |
| GPT-5.6-Luna | 46 | 22 | 25 | 33 | 40 |
| GPT-5.6-Terra | 46 | 24 | 27 | 31 | 40 |
| GPT-5.6-Sol | 48 | 43 | 43 | 45 | 48 |
| GPT-6-Luna | 48 | 44 | 44 | 48 | 48 |
| GPT-6-Sol | 48 | 46 | 46 | 48 | 48 |
| **GPT-6-Astra** | **50** | **48** | **48** | **50** | **50** |

**Key findings:**
- GPT-5.6-Luna and GPT-5.6-Terra each generate 46/50 images successfully, but only 22 and 24 satisfy all quality criteria
- All successfully generated images from GPT-6 models meet aesthetics and composition thresholds; remaining failures concern prompt adherence
- DeepSeek and Kimi generate only 6–20 images, limiting comparability of their mean quality scores

### MaLiang-VBench (Video Tasks)

**Table 4: Generation Success and Cost on MaLiang-VBench**

| Model | Success ↑(%) | Failures ↓ | Token Limit ↓ | Time/Success ↓(min/video) |
|---|---|---|---|---|
| DeepSeek-V4.1-Flash (64K) | 0.0 | 13 | 5 | – |
| Kimi-K2.6 | 23.1 | 10 | 1 | 66.95 |
| GPT-5.6-Sol | 53.8 | 6 | 3 | 9.27 |
| **GPT-6-Astra** | **100.0** | **0** | **0** | **10.00** |

**Table 5: Visual Quality on MaLiang-VBench**

| Model | Successful ↑ | All Criteria ↑ | Alignment ↑ | Aesthetics ↑ | Composition ↑ | Motion ↑ |
|---|---|---|---|---|---|---|
| DeepSeek-V4.1-Flash (64K) | 0 | 0 | 0 | 0 | 0 | 0 |
| Kimi-K2.6 | 3 | 0 | 0 | 1 | 1 | 0 |
| GPT-5.6-Sol | 7 | 5 | 7 | 7 | 6 | 6 |
| **GPT-6-Astra** | **13** | **10** | **12** | **13** | **12** | **10** |

**Key findings:**
- **Motion coherence is the most restrictive criterion** for Astra (10/13), though all videos meet aesthetics
- Token-budget exhaustion explains only part of failures (5/13 DeepSeek, 1/10 Kimi, 3/6 GPT-5.6-Sol)
- Appearance alone does not capture temporal requirement satisfaction

### General Capability vs. Visual Performance

- Spearman rank correlation between AA Intelligence Index and drawing quality pass rate: $\rho = 0.65$
- **GPT-5.6-Luna and GPT-6-Luna share an index score of 37, yet quality pass rates are 44% and 88%**
- Kimi-K3 scores 44 on the index but satisfies all visual criteria on only 18% of tasks

### Qualitative Findings

- **Photorealism exploration**: Path-tracing backends enable material-dependent reflections, glass transmission, and contact shadows; however, prompted brushwork refinement alone cannot achieve photorealism—finer strokes lose existing detail without recovering coordinated shape, shading, and texture
- **Construction traceability**: The harness reveals how individual elements were constructed (e.g., a six-panel research poster where code controls layout and plots, while generated imagery handles photographic elements)

---

## Theoretical and Practical Implications

### Theoretical Contributions

1. **Formalization of the P2V gap**: Distinguishes program-level correctness from visual requirement satisfaction, providing a framework for evaluating executable visual generation
2. **Stateful generation paradigm**: PEG state provides a shared revision reference across planning, execution, and verification, unifying image and video generation under one protocol
3. **Revision-aware verification**: The Ready condition (Eq. 4) formalizes delivery criteria, tying visual verification to the specific revision being delivered

### Practical Implications

1. **Evaluation methodology**: Quality counts must be interpreted alongside success rates—mean scores over few successful samples can mislead (e.g., Kimi-K3's 4.56 alignment score based on only 9 images)
2. **Model selection**: General benchmark scores are insufficient predictors of visual generation ability; direct evaluation of program-to-visual translation is necessary
3. **Cost-quality trade-offs**: GPT-5.6-Sol produces 43 qualifying images at 3.28 min/image vs. Astra's 48 at 3.70 min/image—different cost-efficiency profiles
4. **Failure mode taxonomy**: Identifies reasoning stagnation, semantic livelock, and infinite agentic loops as distinct failure modes requiring different interventions

---

## Conclusion

MaLiang-Harness demonstrates a viable programmable path to image and video generation, organizing MLLM-driven visual creation into a persistent process of construction, inspection, and refinement. Key takeaways:

1. **Programmable generation works**: GPT-6-Astra achieves 100% generation success on both benchmarks, with high quality satisfaction rates
2. **Success ≠ satisfaction**: Generation success alone is insufficient; visual requirement satisfaction requires iterative refinement guided by rendered feedback
3. **General capability ≠ visual ability**: Similar general benchmark scores can correspond to substantially different visual outcomes
4. **Open challenges remain**:
   - **Photorealism**: Current Canvas-based backends produce stylized output; richer rendering backends (path tracing) show promise but need controlled comparisons
   - **Stalled refinement**: The harness separates termination from completion—token limits bound execution, but successful finalization requires meaningful progress
   - **Progress-aware strategies**: Future work should explore refinement strategies that detect when continued activity produces no meaningful visual improvement

The project is available at: https://github.com/gulucaptain/MaLiang-Harness

---

_Markdown view of https://picx.dev/p/3vDx97, served by PicX — AI-generated visual whiteboard summaries of research papers._
