Summary (Overview)

  • MaLiang-Harness is a unified framework for programmable image and video generation, where multimodal large language models (MLLMs) translate creative prompts into executable visual programs rendered by backends like Canvas, SVG, Scene2d, or Three.js.
  • The paper introduces the Program-to-Visual (P2V) gap: a program can execute correctly yet produce visuals that violate user intent (wrong composition, appearance, or motion).
  • Three core mechanisms—Persistent Executable Generation (PEG) State, Traceable Generation Process (TGP), and Revision-aware Editing and Verification (REV)—coordinate planning, execution, and visual feedback across rendering backends.
  • Evaluations on MaLiang-IBench (50 image tasks, 11 models) and MaLiang-VBench (13 video tasks, 4 models) show GPT-6-Astra achieves 100% generation success on both, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds.
  • General MLLM capability scores (e.g., AA Intelligence Index) correlate only weakly with visual generation performance; similarly scored models can differ dramatically in visual outcomes (e.g., GPT-5.6-Luna vs. GPT-6-Luna: 44% vs. 88% quality pass rates despite identical index scores).

Introduction and Theoretical Foundation

Background and Motivation

Traditional image/video generation follows a direct visual synthesis paradigm (GANs, diffusion models, flow matching) that learns data distributions and directly synthesizes pixels. This approach has achieved remarkable success but remains implicit—users cannot inspect or control the construction process.

The rapid improvement of MLLMs in semantic understanding, reasoning, and code generation makes a programmable path viable: the model expresses creative intent as an executable visual program, which a renderer converts into an image or video. This enables:

  • Explicit control over spatial composition and temporal dynamics
  • Targeted editing of generated artwork
  • Inspectable construction processes

The Program-to-Visual (P2V) Gap

A program can execute correctly while producing visuals that fail user intent. The paper defines this as:

P2V Gap: The discrepancy between program-level correctness and visual requirement satisfaction.

Bridging this gap requires the model to:

  1. Reason about the visual consequences of its code
  2. Account for rendering backend choices in initial planning
  3. Use rendered output as evidence for revising both program and plan
  4. Preserve artwork across iterations while addressing unmet requirements
  5. Re-verify after each revision, since changes can affect previously satisfied requirements

Theoretical Foundation

The framework is built on the insight that code directly specifies how visual content is constructed—from spatial composition to motion over time. Images and videos share the same creation framework: at a fixed program revision, an image is rendered at a specified content time tt, while a video samples temporal behavior encoded by the program.

The paper draws on prior work in:

  • Executable visual representations (VISPROG, Design2Code, BlenderAlchemy)
  • Agent harnesses (Code as Agent Harness, Show-Harness, OmniHarness)

Methodology

Unified Programmatic Visual Generation Interface

Given a prompt pp and output specification ω\omega, MaLiang-Harness plans appearance, spatial composition, and temporal dynamics, selects a backend bb, and generates drawing/animation code. The backend renders the representation into an image or video that the MLLM evaluates against task requirements.

Persistent Executable Generation (PEG) State

At revision kk, the state is defined as:

Sk=(Pk,Ak,Zk,Ck,k),(1)S_{k} = \left(P_{k}, A_{k}, Z_{k}, C_{k}, k\right),\tag{1}

where:

  • PkP_k: visual program with backend identifier
  • AkA_k: associated assets (retained with content hashes)
  • ZkZ_k: spatial composition and temporal dynamics (via scene attributes or program-embedded definitions)
  • CkC_k: generation context (prompt pp, output spec ω\omega, user requirements, current plan)

State transitions and rendering:

Sk+1=E(Sk,ak),Ik(t)=Rb(Sk,t;ω),(2)S_{k+1} = \mathcal{E}(S_{k}, a_{k}), \quad I_{k}(t) = \mathcal{R}_{b}(S_{k}, t; \omega),\tag{2}

where E\mathcal{E} commits an edit aka_k (validating inputs and checking source revision currency), and Rb\mathcal{R}_b evaluates the program at content time tt.

Traceable Generation Process (TGP)

For the jj-th recorded operation:

τj=(oj,xj,yj,kj−,kj+),(3)\tau_{j} = \big(o_{j}, x_{j}, y_{j}, k_{j}^{-}, k_{j}^{+}\big),\tag{3}

where ojo_j is the executed operation, xjx_j and yjy_j are inputs and results (including errors), and kj−k_j^-, kj+k_j^+ identify source and resulting PEG revisions. The operation index jj is distinct from revision index kk—operations that don't commit state updates (rendering, inspection) retain the same revision.

Revision-aware Editing and Verification (REV)

For each visual requirement hih_i in CkC_k, REV maintains a review qi=(ki,Ei,vi)q_i = (k_i, E_i, v_i) where EiE_i contains requirement-specific visual evidence from revision kik_i and vi∈{pass,fail,uncertain}v_i \in \{\text{pass}, \text{fail}, \text{uncertain}\}.

A review applies to the current state only when ki=kk_i = k. After each commit, the current revision must be re-reviewed. The delivery condition is:

Ready⁡(Sk)=ExportOK⁡(Sk)∧CheckpointOK⁡(k)∧⋀hi∈Hk[ki=k∧Ei≠∅∧vi=pass],(4)\begin{array}{c} \operatorname{Ready}(S_{k}) = \operatorname{ExportOK}(S_{k}) \wedge \operatorname{CheckpointOK}(k) \\ \wedge \bigwedge_{h_{i} \in \mathcal{H}_{k}} [k_{i} = k \wedge E_{i} \neq \varnothing \wedge v_{i} = \text{pass}], \end{array}\tag{4}

where Hk\mathcal{H}_k contains mandatory requirements, ExportOK verifies source files/assets and output conformance to ω\omega, and CheckpointOK holds when all checkpoints pass for the current revision.


Empirical Validation / Results

MaLiang-IBench (Image Tasks)

Table 2: Generation Success and Computational Cost

ModelSuccess ↑(%)Failures ↓Token Limit ↓Time/Qualified ↓(min/img)CallsOutput Tokens (k/case)
DeepSeek-V4.1-Flash12.044527.152.4824.2
DeepSeek-V4-Pro18.041348.124.5039.6
Kimi-K2.634.033162.699.1448.2
Kimi-K2.7-Code40.030129.8412.9459.0
Kimi-K318.041034.882.8217.8
GPT-5.6-Luna92.0404.4911.109.3
GPT-5.6-Terra92.0404.328.328.0
GPT-5.6-Sol96.0213.286.9813.1
GPT-6-Luna96.0204.9214.7016.4
GPT-6-Sol96.0203.527.4210.2
GPT-6-Astra100.0003.706.348.1

Table 3: Visual Quality on MaLiang-IBench (counts = samples scoring ≥4/5)

ModelSuccessful ↑All Criteria ↑Alignment ↑Aesthetics ↑Composition ↑
DeepSeek-V4.1-Flash66666
DeepSeek-V4-Pro98889
Kimi-K2.6176789
Kimi-K2.7-Code2012131313
Kimi-K399999
GPT-5.6-Luna4622253340
GPT-5.6-Terra4624273140
GPT-5.6-Sol4843434548
GPT-6-Luna4844444848
GPT-6-Sol4846464848
GPT-6-Astra5048485050

Key findings:

  • GPT-5.6-Luna and GPT-5.6-Terra each generate 46/50 images successfully, but only 22 and 24 satisfy all quality criteria
  • All successfully generated images from GPT-6 models meet aesthetics and composition thresholds; remaining failures concern prompt adherence
  • DeepSeek and Kimi generate only 6–20 images, limiting comparability of their mean quality scores

MaLiang-VBench (Video Tasks)

Table 4: Generation Success and Cost on MaLiang-VBench

ModelSuccess ↑(%)Failures ↓Token Limit ↓Time/Success ↓(min/video)
DeepSeek-V4.1-Flash (64K)0.0135–
Kimi-K2.623.110166.95
GPT-5.6-Sol53.8639.27
GPT-6-Astra100.00010.00

Table 5: Visual Quality on MaLiang-VBench

ModelSuccessful ↑All Criteria ↑Alignment ↑Aesthetics ↑Composition ↑Motion ↑
DeepSeek-V4.1-Flash (64K)000000
Kimi-K2.6300110
GPT-5.6-Sol757766
GPT-6-Astra131012131210

Key findings:

  • Motion coherence is the most restrictive criterion for Astra (10/13), though all videos meet aesthetics
  • Token-budget exhaustion explains only part of failures (5/13 DeepSeek, 1/10 Kimi, 3/6 GPT-5.6-Sol)
  • Appearance alone does not capture temporal requirement satisfaction

General Capability vs. Visual Performance

  • Spearman rank correlation between AA Intelligence Index and drawing quality pass rate: ρ=0.65\rho = 0.65
  • GPT-5.6-Luna and GPT-6-Luna share an index score of 37, yet quality pass rates are 44% and 88%
  • Kimi-K3 scores 44 on the index but satisfies all visual criteria on only 18% of tasks

Qualitative Findings

  • Photorealism exploration: Path-tracing backends enable material-dependent reflections, glass transmission, and contact shadows; however, prompted brushwork refinement alone cannot achieve photorealism—finer strokes lose existing detail without recovering coordinated shape, shading, and texture
  • Construction traceability: The harness reveals how individual elements were constructed (e.g., a six-panel research poster where code controls layout and plots, while generated imagery handles photographic elements)

Theoretical and Practical Implications

Theoretical Contributions

  1. Formalization of the P2V gap: Distinguishes program-level correctness from visual requirement satisfaction, providing a framework for evaluating executable visual generation
  2. Stateful generation paradigm: PEG state provides a shared revision reference across planning, execution, and verification, unifying image and video generation under one protocol
  3. Revision-aware verification: The Ready condition (Eq. 4) formalizes delivery criteria, tying visual verification to the specific revision being delivered

Practical Implications

  1. Evaluation methodology: Quality counts must be interpreted alongside success rates—mean scores over few successful samples can mislead (e.g., Kimi-K3's 4.56 alignment score based on only 9 images)
  2. Model selection: General benchmark scores are insufficient predictors of visual generation ability; direct evaluation of program-to-visual translation is necessary
  3. Cost-quality trade-offs: GPT-5.6-Sol produces 43 qualifying images at 3.28 min/image vs. Astra's 48 at 3.70 min/image—different cost-efficiency profiles
  4. Failure mode taxonomy: Identifies reasoning stagnation, semantic livelock, and infinite agentic loops as distinct failure modes requiring different interventions

Conclusion

MaLiang-Harness demonstrates a viable programmable path to image and video generation, organizing MLLM-driven visual creation into a persistent process of construction, inspection, and refinement. Key takeaways:

  1. Programmable generation works: GPT-6-Astra achieves 100% generation success on both benchmarks, with high quality satisfaction rates
  2. Success ≠ satisfaction: Generation success alone is insufficient; visual requirement satisfaction requires iterative refinement guided by rendered feedback
  3. General capability ≠ visual ability: Similar general benchmark scores can correspond to substantially different visual outcomes
  4. Open challenges remain:
    • Photorealism: Current Canvas-based backends produce stylized output; richer rendering backends (path tracing) show promise but need controlled comparisons
    • Stalled refinement: The harness separates termination from completion—token limits bound execution, but successful finalization requires meaningful progress
    • Progress-aware strategies: Future work should explore refinement strategies that detect when continued activity produces no meaningful visual improvement

The project is available at: https://github.com/gulucaptain/MaLiang-Harness

Related papers