Letting models write code to draw has a pitfall everyone underestimates: the code runs, but the picture is wrong. Composition drifts, colors misfire, motion goes off track — the program is flawless at the syntax level and a mess at the visual level. A paper posted to arXiv on September 28 gives the phenomenon a name — the Program-to-Visual (P2V) gap — and ships a full construction-inspection-revision pipeline alongside it, reaching #2 on Hugging Face's daily paper ranking.

What the P2V gap is

Executable programs offer explicit control over how images and videos are constructed — that is their edge over prompts. But generating runnable code is only the beginning of visual creation: a program can execute without errors while violating the requested composition, appearance, or motion. The paper defines the discrepancy between "code correct" and "picture acceptable" as the P2V gap. The definition itself is useful: instead of vaguely saying a model "draws poorly", you get measurable slices — generation success, visual quality, and computational cost, tracked separately.

The fix: a draft-inspect-revise loop

MaLiang-Harness has multimodal LLMs write drawing and animation programs, renders them to PNG images or silent MP4 videos through backends such as Canvas, SVG, and Three.js, then inspects the output and revises the code in a loop. Three mechanisms hold it together: PEG persists programs and task context, snapshotting every committed revision; TGP ties each edit to its rendered evidence; REV restores from historical revisions and re-checks the current one before export. Built on Deep Agents, the project ships a CLI, a chat interface, and batch evaluation tools, with code open-sourced on GitHub.

Results: 100% can generate, 76.9% actually pass

The team evaluated 11 closed-source MLLMs on their MaLiang-IBench (50 image tasks) and 4 models on MaLiang-VBench (13 video tasks). GPT-6-Astra hits 100% generation success on both benchmarks, with 96.0% of image tasks and 76.9% of video tasks passing all quality thresholds. The stinging part is the other finding: general capability scores mismatch visual-generation performance — models with similar overall scores differ substantially in satisfying visual requirements. Generic leaderboards lose nearly all predictive power here.

So what

Two takeaways. First, programmable generation is a real alternative to diffusion: when you need precise control over layout, structure, and motion, writing code is inherently more controllable than sampling — provided you close the self-check-and-revise loop, which MaLiang engineers into place. Second, stop trusting general benchmarks for model selection; composite abilities like visual generation demand task-specific benches. The project is named after Ma Liang — in Chinese context, the boy with the magic brush. The magic was never the brush; it was the eye that looks, and revises.

References: arXiv paper, GitHub repo.