[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-reve-2-0-layout-first-image-arena-2":3,"news-related-e6cc0fff-b4e1-425e-854d-b5f4fbd779c3":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"e6cc0fff-b4e1-425e-854d-b5f4fbd779c3","从 prompt 到像素之间插入代码：Reve2.0 用 layout-first架构把图像变成可编辑的结构化对象","6 月3 日，硅谷初创公司 Reve 发布第二代图像生成模型 **Reve2.0**，上线当日冲上 Image Arena文本到图像榜第二，仅次于 OpenAI gpt-image-2，把 Google Nano Banana2 与 Microsoft MAI-Image-2.5甩在身后。真正让技术圈侧目的不是 SOTA排名，而是它重新定义了图像生成：**先写代码，再渲染像素**。\n\n### 一、过去四年的「烟火困境」\n\nReve官方把此前图像生成模式形容为「fireworks phase」：把 prompt塞进黑盒，祈祷像素够好看；想编辑局部就重新生成、破坏整体构图。一锤子定音式工作流是图像工具多年没真正解决的事。\n\n### 二、Reve2.0 的核心：layout-first架构\n\nReve2.0 把规划与渲染解耦成两步：Planning 先生成结构化图像布局——每块区域有标签、坐标、语义关系，相当于图像的源代码；Rendering拿到布局后，用渲染器画成原生4K像素。\n\n这种 code-based image思路带来三件事：可寻址编辑（直接改布局某区域，不必重生整图）、agent-native（LLM 直接读改布局）、算力效率（转化为 next-token prediction思路，降低 diffusion反复去噪开销）。Reve1.0 已用数据结构代替 caption验证假设，Reve2.0 参数扩到3 倍并引入新规划架构，进入 SOTA梯队。\n\n### 三、范式变化在哪里？\n\nReve2.0推动图像模型的新范式：图像从黑盒输出变成结构化对象，可 review、可 diff、可合并；LLM 直接成为图像模型的指挥棒，不再只是写 prompt；迭代式创作成为默认工作流，设计师第一次可以只换局部而不破坏整体。\n\n代码作为图像的中间表示——与 GPT-Rosalind 把推理注入图像、FLUX.2引入协作分工、Ideogram4.0转向开源权重共同指向一个判断：**图像生成正从一次性放烟火转向可编程的工业化流程**。如果这条路线走通，下一波图像工具的形态，可能比想象的更接近 IDE。","https:\u002F\u002Freve.com\u002F","36f11d4d-7a06-4c5c-9206-da8ae76b5283",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":21,"name":22,"slug":22,"description":13,"color":13},"c883fd20-1d66-4fb7-9fc7-320fa7f87023","text-to-image",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"604c1593-c65b-42a9-9ba6-09780b02bdf1","en","Reve2.0 inserts code between prompt and pixels","On June 3, Silicon Valley startup Reve released the second-generation image generation model Reve2.0, which on launch day rocketed to #2 on the Image Arena text-to-image leaderboard — only behind OpenAI gpt-image-2, and ahead of Google Nano Banana2 and Microsoft MAI-Image-2.5. What really caught the tech world's eye wasn't the SOTA ranking, but the fact that it redefines image generation: **write the code first, then render the pixels**.\n\n### I. The \"fireworks dilemma\" of the past four years\n\nReve officially described the previous image generation mode as the \"fireworks phase\": stuff the prompt into a black box, pray the pixels look good enough; want to edit a part and you have to regenerate, breaking the overall composition. The one-shot \"make or break\" workflow is something image tools have never truly solved.\n\n### II. Reve2.0's core: layout-first architecture\n\nReve2.0 decouples planning and rendering into two steps. Planning first generates a structured image layout — each region has a label, coordinates, and semantic relationships, equivalent to the image's source code. Rendering then takes the layout and uses the renderer to draw it into native 4K pixels.\n\nThis code-based image approach brings three things: addressable editing (directly modify a region of the layout, without regenerating the whole image), agent-native (LLMs can directly read and modify the layout), and compute efficiency (translating into the next-token prediction mindset, reducing the repeated denoising overhead of diffusion). Reve1.0 used data structures to replace caption to verify the hypothesis, and Reve2.0 expands parameters by 3× and introduces a new planning architecture, entering the SOTA ranks.\n\n### III. Where does the paradigm shift lie?\n\nReve2.0 drives a new paradigm in image models: images shift from black-box output to structured objects, reviewable, diff-able, and mergeable; LLMs become the conductor's baton of image models, no longer just writing prompts; iterative creation becomes the default workflow, and designers can for the first time swap local parts without breaking the whole.\n\nCode as the image's intermediate representation — together with GPT-Rosalind injecting reasoning into images, FLUX.2 introducing collaborative division of labor, and Ideogram 4.0 turning to open-source weights — points to one judgment: **image generation is shifting from one-shot fireworks to a programmable industrial process.** If this path holds, the form of the next wave of image tools may be closer to an IDE than imagined.","reve-2-0-layout-first-image-arena-2","2026-06-09T14:30:00Z","2026-06-09T14:12:46.876130Z","2026-08-19T02:08:40.142862Z",true,"agent",106,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"747917d5-e65b-46dd-b0db-40dfa119cdd1","Reve 2.1 用 Layout-First 架构 + 4K 输出登顶 Arena #2：用不到头部 1\u002F10 算力做独立图像生成实验室","reve-2-1-layout-first","2026-07-16T02:14:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"18df2ab9-affd-4ba8-adaa-03cc7e3a2317","GPT Image 2 发布：OpenAI 首次将推理能力注入图像生成","gpt-image-2-openai-agentic-2k-reasoning","2026-04-28T10:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"b51af942-496b-4b58-95fb-37980d12a743","逆向工程发现:微软画图本地生成的 AI 图像,像素里埋着服务器下发的水印 GUID","mspaint-invisible-watermark-guid","2026-08-26T13:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"774de6ac-98e1-4343-a67a-bfdc72d377bb","INFORMS 实证:AI 广告真实投放胜过设计师,18 个月后仍领先","informs-ai-ads-beat-human-designers-18-months","2026-08-22T14:00:00+00:00"]