Reve Lab released its Reve 2.1 text-to-image model on July 9, completing a key iteration just one month after 2.0. The new model recaptures the global #2 spot on the Arena text-to-image leaderboard with 1306 Elo, ranks #8 on the image-editing sub-leaderboard, and continues to hold the title of "highest-resolution 4K independent model". Technically, Reve 2.1 continues its core bet — "images as code". The model first generates a hierarchical layout plan, then renders each region independently, so all elements are naturally addressable and can be individually redrawn. This generation lifts planning precision, prompt understanding, and foreign-language rendering together: 4K native 16MP output achieves SOTA controllability in dense scenes, small text, and multi-language text in the same frame. The most noteworthy counter-signal is the compute curve. The Reve team explicitly disclosed that 2.1's total training compute is less than one-tenth of that of the leading players (image-generation products from Microsoft, Google, Meta and other big labs), yet still hits Arena Top 2 in real-world testing. At a time when multinational AI labs are universally piling on compute, training diffusion Transformers on billion-scale image-text pairs, Reve uses "code-style representation + extreme engineering efficiency" to come out on top — showing that the density dividend of layout planning as an intermediate representation still has a lot of untapped room. This means image-generation competition in the second half of 2026 is splitting from the "model size race" into an independent track of "representational efficiency". For independent labs and small-to-medium teams, this may carry more methodological value than yet another large-model release: when representation itself becomes a programmable object, the diffusion model's "end-to-end pixel" assumption is no longer the only optimal answer, and design tools, Agent image generation, and batch asset pipelines will all see new engineering entry points.