Summary (Overview)
- MaLiang-Harness is a unified framework for programmable image and video generation, where multimodal large language models (MLLMs) translate creative prompts into executable visual programs rendered by backends like Canvas, SVG, Scene2d, or Three.js.
- The paper introduces the Program-to-Visual (P2V) gap: a program can execute correctly yet produce visuals that violate user intent (wrong composition, appearance, or motion).
- Three core mechanisms—Persistent Executable Generation (PEG) State, Traceable Generation Process (TGP), and Revision-aware Editing and Verification (REV)—coordinate planning, execution, and visual feedback across rendering backends.
- Evaluations on MaLiang-IBench (50 image tasks, 11 models) and MaLiang-VBench (13 video tasks, 4 models) show GPT-6-Astra achieves 100% generation success on both, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds.
- General MLLM capability scores (e.g., AA Intelligence Index) correlate only weakly with visual generation performance; similarly scored models can differ dramatically in visual outcomes (e.g., GPT-5.6-Luna vs. GPT-6-Luna: 44% vs. 88% quality pass rates despite identical index scores).
Introduction and Theoretical Foundation
Background and Motivation
Traditional image/video generation follows a direct visual synthesis paradigm (GANs, diffusion models, flow matching) that learns data distributions and directly synthesizes pixels. This approach has achieved remarkable success but remains implicit—users cannot inspect or control the construction process.
The rapid improvement of MLLMs in semantic understanding, reasoning, and code generation makes a programmable path viable: the model expresses creative intent as an executable visual program, which a renderer converts into an image or video. This enables:
- Explicit control over spatial composition and temporal dynamics
- Targeted editing of generated artwork
- Inspectable construction processes
The Program-to-Visual (P2V) Gap
A program can execute correctly while producing visuals that fail user intent. The paper defines this as:
P2V Gap: The discrepancy between program-level correctness and visual requirement satisfaction.
Bridging this gap requires the model to:
- Reason about the visual consequences of its code
- Account for rendering backend choices in initial planning
- Use rendered output as evidence for revising both program and plan
- Preserve artwork across iterations while addressing unmet requirements
- Re-verify after each revision, since changes can affect previously satisfied requirements
Theoretical Foundation
The framework is built on the insight that code directly specifies how visual content is constructed—from spatial composition to motion over time. Images and videos share the same creation framework: at a fixed program revision, an image is rendered at a specified content time , while a video samples temporal behavior encoded by the program.
The paper draws on prior work in:
- Executable visual representations (VISPROG, Design2Code, BlenderAlchemy)
- Agent harnesses (Code as Agent Harness, Show-Harness, OmniHarness)
Methodology
Unified Programmatic Visual Generation Interface
Given a prompt and output specification , MaLiang-Harness plans appearance, spatial composition, and temporal dynamics, selects a backend , and generates drawing/animation code. The backend renders the representation into an image or video that the MLLM evaluates against task requirements.
Persistent Executable Generation (PEG) State
At revision , the state is defined as:
where:
- : visual program with backend identifier
- : associated assets (retained with content hashes)
- : spatial composition and temporal dynamics (via scene attributes or program-embedded definitions)
- : generation context (prompt , output spec , user requirements, current plan)
State transitions and rendering:
where commits an edit (validating inputs and checking source revision currency), and evaluates the program at content time .
Traceable Generation Process (TGP)
For the -th recorded operation:
where is the executed operation, and are inputs and results (including errors), and , identify source and resulting PEG revisions. The operation index is distinct from revision index —operations that don't commit state updates (rendering, inspection) retain the same revision.
Revision-aware Editing and Verification (REV)
For each visual requirement in , REV maintains a review where contains requirement-specific visual evidence from revision and .
A review applies to the current state only when . After each commit, the current revision must be re-reviewed. The delivery condition is:
where contains mandatory requirements, ExportOK verifies source files/assets and output conformance to , and CheckpointOK holds when all checkpoints pass for the current revision.
Empirical Validation / Results
MaLiang-IBench (Image Tasks)
Table 2: Generation Success and Computational Cost
| Model | Success ↑(%) | Failures ↓ | Token Limit ↓ | Time/Qualified ↓(min/img) | Calls | Output Tokens (k/case) |
|---|---|---|---|---|---|---|
| DeepSeek-V4.1-Flash | 12.0 | 44 | 5 | 27.15 | 2.48 | 24.2 |
| DeepSeek-V4-Pro | 18.0 | 41 | 3 | 48.12 | 4.50 | 39.6 |
| Kimi-K2.6 | 34.0 | 33 | 1 | 62.69 | 9.14 | 48.2 |
| Kimi-K2.7-Code | 40.0 | 30 | 1 | 29.84 | 12.94 | 59.0 |
| Kimi-K3 | 18.0 | 41 | 0 | 34.88 | 2.82 | 17.8 |
| GPT-5.6-Luna | 92.0 | 4 | 0 | 4.49 | 11.10 | 9.3 |
| GPT-5.6-Terra | 92.0 | 4 | 0 | 4.32 | 8.32 | 8.0 |
| GPT-5.6-Sol | 96.0 | 2 | 1 | 3.28 | 6.98 | 13.1 |
| GPT-6-Luna | 96.0 | 2 | 0 | 4.92 | 14.70 | 16.4 |
| GPT-6-Sol | 96.0 | 2 | 0 | 3.52 | 7.42 | 10.2 |
| GPT-6-Astra | 100.0 | 0 | 0 | 3.70 | 6.34 | 8.1 |
Table 3: Visual Quality on MaLiang-IBench (counts = samples scoring ≥4/5)
| Model | Successful ↑ | All Criteria ↑ | Alignment ↑ | Aesthetics ↑ | Composition ↑ |
|---|---|---|---|---|---|
| DeepSeek-V4.1-Flash | 6 | 6 | 6 | 6 | 6 |
| DeepSeek-V4-Pro | 9 | 8 | 8 | 8 | 9 |
| Kimi-K2.6 | 17 | 6 | 7 | 8 | 9 |
| Kimi-K2.7-Code | 20 | 12 | 13 | 13 | 13 |
| Kimi-K3 | 9 | 9 | 9 | 9 | 9 |
| GPT-5.6-Luna | 46 | 22 | 25 | 33 | 40 |
| GPT-5.6-Terra | 46 | 24 | 27 | 31 | 40 |
| GPT-5.6-Sol | 48 | 43 | 43 | 45 | 48 |
| GPT-6-Luna | 48 | 44 | 44 | 48 | 48 |
| GPT-6-Sol | 48 | 46 | 46 | 48 | 48 |
| GPT-6-Astra | 50 | 48 | 48 | 50 | 50 |
Key findings:
- GPT-5.6-Luna and GPT-5.6-Terra each generate 46/50 images successfully, but only 22 and 24 satisfy all quality criteria
- All successfully generated images from GPT-6 models meet aesthetics and composition thresholds; remaining failures concern prompt adherence
- DeepSeek and Kimi generate only 6–20 images, limiting comparability of their mean quality scores
MaLiang-VBench (Video Tasks)
Table 4: Generation Success and Cost on MaLiang-VBench
| Model | Success ↑(%) | Failures ↓ | Token Limit ↓ | Time/Success ↓(min/video) |
|---|---|---|---|---|
| DeepSeek-V4.1-Flash (64K) | 0.0 | 13 | 5 | – |
| Kimi-K2.6 | 23.1 | 10 | 1 | 66.95 |
| GPT-5.6-Sol | 53.8 | 6 | 3 | 9.27 |
| GPT-6-Astra | 100.0 | 0 | 0 | 10.00 |
Table 5: Visual Quality on MaLiang-VBench
| Model | Successful ↑ | All Criteria ↑ | Alignment ↑ | Aesthetics ↑ | Composition ↑ | Motion ↑ |
|---|---|---|---|---|---|---|
| DeepSeek-V4.1-Flash (64K) | 0 | 0 | 0 | 0 | 0 | 0 |
| Kimi-K2.6 | 3 | 0 | 0 | 1 | 1 | 0 |
| GPT-5.6-Sol | 7 | 5 | 7 | 7 | 6 | 6 |
| GPT-6-Astra | 13 | 10 | 12 | 13 | 12 | 10 |
Key findings:
- Motion coherence is the most restrictive criterion for Astra (10/13), though all videos meet aesthetics
- Token-budget exhaustion explains only part of failures (5/13 DeepSeek, 1/10 Kimi, 3/6 GPT-5.6-Sol)
- Appearance alone does not capture temporal requirement satisfaction
General Capability vs. Visual Performance
- Spearman rank correlation between AA Intelligence Index and drawing quality pass rate:
- GPT-5.6-Luna and GPT-6-Luna share an index score of 37, yet quality pass rates are 44% and 88%
- Kimi-K3 scores 44 on the index but satisfies all visual criteria on only 18% of tasks
Qualitative Findings
- Photorealism exploration: Path-tracing backends enable material-dependent reflections, glass transmission, and contact shadows; however, prompted brushwork refinement alone cannot achieve photorealism—finer strokes lose existing detail without recovering coordinated shape, shading, and texture
- Construction traceability: The harness reveals how individual elements were constructed (e.g., a six-panel research poster where code controls layout and plots, while generated imagery handles photographic elements)
Theoretical and Practical Implications
Theoretical Contributions
- Formalization of the P2V gap: Distinguishes program-level correctness from visual requirement satisfaction, providing a framework for evaluating executable visual generation
- Stateful generation paradigm: PEG state provides a shared revision reference across planning, execution, and verification, unifying image and video generation under one protocol
- Revision-aware verification: The Ready condition (Eq. 4) formalizes delivery criteria, tying visual verification to the specific revision being delivered
Practical Implications
- Evaluation methodology: Quality counts must be interpreted alongside success rates—mean scores over few successful samples can mislead (e.g., Kimi-K3's 4.56 alignment score based on only 9 images)
- Model selection: General benchmark scores are insufficient predictors of visual generation ability; direct evaluation of program-to-visual translation is necessary
- Cost-quality trade-offs: GPT-5.6-Sol produces 43 qualifying images at 3.28 min/image vs. Astra's 48 at 3.70 min/image—different cost-efficiency profiles
- Failure mode taxonomy: Identifies reasoning stagnation, semantic livelock, and infinite agentic loops as distinct failure modes requiring different interventions
Conclusion
MaLiang-Harness demonstrates a viable programmable path to image and video generation, organizing MLLM-driven visual creation into a persistent process of construction, inspection, and refinement. Key takeaways:
- Programmable generation works: GPT-6-Astra achieves 100% generation success on both benchmarks, with high quality satisfaction rates
- Success ≠ satisfaction: Generation success alone is insufficient; visual requirement satisfaction requires iterative refinement guided by rendered feedback
- General capability ≠ visual ability: Similar general benchmark scores can correspond to substantially different visual outcomes
- Open challenges remain:
- Photorealism: Current Canvas-based backends produce stylized output; richer rendering backends (path tracing) show promise but need controlled comparisons
- Stalled refinement: The harness separates termination from completion—token limits bound execution, but successful finalization requires meaningful progress
- Progress-aware strategies: Future work should explore refinement strategies that detect when continued activity produces no meaningful visual improvement
The project is available at: https://github.com/gulucaptain/MaLiang-Harness
Related papers
- YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
YuE2 unifies symbolic and audio music generation in one model, achieving frontier song quality where explicit score planning improves perceived musicality over score-free generation.
- Omni-IO Skills: Harnessing Your Agent Omni-Native
Omni-IO Skills is a plug-and-play harness that makes general-purpose agents omni-native without retraining, achieving 100% input-support and large quality gains on UniM-90.
- What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling
Simple-WAM shows a single forward pass over noised future tokens, not iterative denoising, recovers nearly all of explicit world model generalization at latent-level efficiency.