MaLiang-Harness generates images and video from executable programs: GPT-6-Astra reaches 100% generation success on both benchmarks, with 96.0% of image and 76.9% of video tasks meeting all quality thresholds
Synopsis
The work introduces MaLiang-Harness, a framework that organizes MLLM-driven image and video generation as a persistent process of construction, inspection, and revision, in which a Persistent Executable Generation state, a Traceable Generation Process, and Revision-aware Editing and Verification share a common revision reference; it evaluates 11 and 4 closed-source MLLMs on MaLiang-IBench (50 text-to-image prompts) and MaLiang-VBench (13 text-to-video prompts), measuring GPT-6-Astra at 100% generation success on both benchmarks, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds.
Interpretation
The paper defines the Program-to-Visual (P2V) gap, the discrepancy in which a program executes correctly yet violates the requested composition, appearance, or motion, and frames visual program generation as a stateful process of construction, inspection, and revision. Prior work on executable visual representations (such as VISPROG, Design2Code, and BlenderAlchemy) established that programs can mediate visual understanding and construction; this work extends that premise into a unified image-and-video generation process organized around a persistent artwork state across several rendering backends. The definition is supported by both the framework design and the evaluation, which separates successful generation from satisfaction of visual requirements: GPT-5.6-Luna and GPT-5.6-Terra each successfully generate 46 of 50 images, but only 22 and 24 respectively satisfy all three quality criteria.
The framework comprises three mechanisms: Persistent Executable Generation (PEG) state preserves programs, assets, task requirements, the current plan, and a revision index; the Traceable Generation Process (TGP) links operations, execution results, and rendered evidence to specific revisions; and Revision-aware Editing and Verification (REV) supports restoring earlier content and re-checking the current revision before delivery. Unlike the direct visual synthesis paradigm (diffusion and flow matching), this path does not rely on a diffusion or flow-matching process; instead an MLLM plans and generates drawing and animation code that rendering backends such as Canvas, SVG, Scene2d, or Three.js turn into images or video. The mechanisms are given as formal definitions of state, operations, and reviews, and are illustrated through qualitative cases including a path-tracing still life, hyperrealism-inspired brushwork refinement, and the recorded construction of a six-panel research illustration.
On MaLiang-IBench, GPT-6-Astra reaches 100% generation success with 48 of 50 tasks satisfying all quality criteria; GPT-6-Sol and GPT-6-Luna reach 46 and 44, and GPT-5.6-Sol reaches 43. The results separate generation success from satisfaction of visual requirements and show that all successfully generated images from the three GPT-6 models meet the aesthetics and composition thresholds, with their remaining quality failures concentrated in prompt adherence. The evaluation covers 11 models and 50 text-to-image prompts, with quality scored 1-5 by GPT-6-Sol on prompt alignment, aesthetics, and composition; threshold counts use the full task set as denominator, and on cost GPT-5.6-Sol produces a qualifying image in 3.28 minutes versus 3.70 minutes for Astra.
On MaLiang-VBench, GPT-6-Astra reaches 100% generation success with 10 of 13 tasks satisfying all four thresholds; GPT-5.6-Sol succeeds on 7 and meets all thresholds on 5; Kimi-K2.6 completes 3 but none meets all criteria; DeepSeek-V4.1-Flash completes none. The video results identify motion coherence as the most restrictive dimension: all 13 Astra videos meet the aesthetics threshold but only 10 meet the motion-coherence threshold, and token-budget exhaustion explains only part of the failures (5 of DeepSeek's 13, 1 of Kimi's 10, and 3 of GPT-5.6-Sol's 6). Video evaluation was performed by a single reviewer using the original prompt and 12 temporally ordered sampled frames, covering 23 successful videos across four models; the review was not formally blinded, and motion coherence measures progression visible in the sampled frames without establishing frame-by-frame fluidity.
Perspective
The framework targets programmatic image and video creation that needs explicit spatial and temporal control, in settings where rendering backends such as Canvas, SVG, Scene2d, or Three.js can be connected and an MLLM can iteratively plan and revise code; for users who want to inspect construction history, compare revision-specific evidence, and retain earlier versions, PEG, TGP, and REV provide a shared revision reference. The evaluation conclusions apply to the 50 text-to-image prompts of MaLiang-IBench and the 13 text-to-video prompts of MaLiang-VBench, and to the 11 and 4 evaluated model configurations.
Cost columns use different aggregation rules across models (for example, Kimi's Time/Task is the median over successful tasks and its token statistics divide by the number of successful tasks), so efficiency comparisons describe the evaluated configurations rather than a uniform measure; mean image quality covers successful generations only, and the DeepSeek and Kimi means rest on a limited subset of 6-20 images, with Kimi-K3's alignment score of 4.56 based on just nine images; video review was performed by a single reviewer without formal blinding, and motion coherence does not establish frame-by-frame fluidity; the general-capability comparison uses public index scores, and the correspondence between the public DeepSeek-V4-Pro score and the evaluated checkpoint is unverified, so that analysis characterizes an association rather than isolating capability effects; the photorealism direction remains exploratory, and the paper notes that controlled comparisons are still needed to establish the contribution of richer rendering backends to perceptual realism, while refinement can stall and the mechanisms do not guarantee effective corrections.
