[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-grok-imagine-image-2-0-arena-second":3,"news-related-619ad304-0d2a-4dba-b91e-19414d036746":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"619ad304-0d2a-4dba-b91e-19414d036746","Grok Imagine Image 2.0：文生图 Arena 双榜第二","xAI 在 8 月 7 日上线 Grok Imagine Image 2.0,作为 grok.com\u002Fimagine、iOS、Android 的 Quality Mode 全面开放,API 名 grok-imagine-image-2.0。新模型主打'能用'——精准遵循指令、设计师级的版式排版、跨代视觉一致性;同时引入 5 张参考图编辑、区域级 magic wand、智能分割、背景移除、Smart resize 等编辑能力。在 Arena 图像编辑榜与文生图榜双榜第二,仅次于 OpenAI gpt-image-2。","## 8 月 7 日,xAI 把「能用」的图片模型放进 grok.com\n\nxAI 在 8 月 7 日正式推出 **Grok Imagine Image 2.0**,新模型以 Quality Mode 形式在 [grok.com\u002Fimagine](https:\u002F\u002Fgrok.com\u002Fimagine) 与 iOS \u002F Android 客户端全面 GA,API 同步开放,模型名 ([官方发布页](https:\u002F\u002Fx.ai\u002Fnews\u002Fgrok-imagine-image-2))。\n\n和上一代纯炫技不同,2.0 的核心目标是「**能用在真实工作里**」:严格遵循 prompt 细节、设计师级的版式排版、跨代生成时保持视觉一致性。换句话说,xAI 不再卖「生成得好看」,开始卖「生成得**稳**」。\n\n## 编辑能力第一次被当成 first-class 看待\n\n官方把编辑能力单列了一节「Precise editing」,而不是塞进生成能力后面:\n\n- **Magic wand + Segmentation**:点哪改哪,只动指定区域,其余像素不动;\n- **Background removal**:主体一键导出透明背景,可直接拼到下游设计;\n- **Multi-ref editing**:单次生成支持最多 **5 张输入图** —— 不再需要手动 PS 拼合;\n- **Smart resize**:同一张图按 1:1 \u002F 2:3 \u002F 9:16 \u002F 16:9 等比例重新构图,模型自己补画面,而不是裁切。\n\n配套的是 **Templates**:把常见工作流(Photo Edit、Product Color Change、Editorial Product Poster、Reimagine、Photo Collage、Mascot Maker、BG Removal & Change、E-Commerce Photos、UGC Photos、Professional Headshot、Icon Maker、Character Sprite、Props & UI Kit、Emoji Creator、Merch Maker)打包成「输入→成品」的工作流节点。\n\n## 跨代一致性:为视频做「世界观」\n\nxAI 把 Image 2.0 定位为「Grok 视频的脚手架」。在「Build a world for video」一节里,官方演示了用同一组 prompt 依次生成一个角色、若干地点、若干道具,再让该角色在雪地、风暴、隧道不同环境下出现——**人物外观、画风、道具细节跨图保持一致**。\n\n这正是当前视频生成最痛的「跨镜头一致性」问题。Image 2.0 的解法是先在静态图层面把世界观固定下来,再交给 Grok Imagine Video。这条路和 Seedream \u002F Seedance 2.5 的「图像→视频」产线思路高度一致。\n\n## 成绩单:Arena 双榜第二\n\n官方披露,Image 2.0 在 **Image Edit Arena** 与 **Text-to-Image Arena** 两个公开榜上都排到 **世界第二**,仅次于 OpenAI gpt-image-2。xAI 模型在 Arena 上以「SpaceXAI」名义登记(统计截至 8 月 7 日)。\n\n第三方 Neura Market 8 月 7 日的报道也印证了这一排名,并补充了 Smart resize、多参考输入、区域级编辑等新功能([Neura Market](https:\u002F\u002Fwww.neura.market\u002Fnews\u002Fxai-grok-imagine-image-2-0-editing-tools-arena-rankings))。\n\n## 行业含义:文生图的竞争重心,正在从「生成」转向「可生产」\n\n把这次发布放在 2026 下半年的版图里看,信号比模型本身更重要:\n\n1. **「能用」比「好看」贵**。xAI 在官方页里反复用「real work」「real creative work」,意味着它开始把设计\u002F电商\u002F营销团队当作目标用户,而不只是创作者尝鲜。\n2. **编辑能力变成主战场**。Qwen-Image-2.0、Seedream 5.0、Grok Image 2.0 这一波都在押注「multi-ref + region edit」,新一轮的文生图差异化已经从「生得像不像」挪到「改得准不准」。\n3. **为视频铺路是隐线**。当 ByteDance Seedance 2.5、京东 JoyAI-Video-Edit 把视频拉进 30 秒时代,**谁能先把视觉资产一致性固化下来,谁就能在视频端拿到杠杆**。Grok 这条 Image 2.0 → Imagine Video 的链路,正是同样的产品逻辑。\n\n## 谁能用、怎么用\n\n- 消费端:打开 [grok.com\u002Fimagine](https:\u002F\u002Fgrok.com\u002Fimagine) 切到 Quality Mode,或 iOS \u002F Android Grok App;\n- 开发者:用  调用  即可,文档见 [xAI 开发者文档](https:\u002F\u002Fdocs.x.ai\u002Fdevelopers\u002Fmodel-capabilities\u002Fimages\u002Fgeneration#quick-start);\n- 控制台:[xAI Console Playground](https:\u002F\u002Fconsole.x.ai\u002Fteam\u002Fdefault\u002Fimage?model=grok-imagine-image-2.0&flow=explore) 直接试。\n\n**所以呢**:文生图战场还没分出胜负,「单张图像的修图生产力」会是接下来 6 个月分化的关键指标。Grok Image 2.0 把编辑能力做成 first-class + Arena 双榜第二,意味着 xAI 不再只想做 Grok 的「图像皮肤」,而是要在设计师的日常工具链里挤一个位置。对国内玩家(Qwen \u002F Seedream \u002F 智谱)来说,跨代一致性和模板化工作流,可能是下半年最值得补的课。","https:\u002F\u002Fx.ai\u002Fnews\u002Fgrok-imagine-image-2","b82e17a3-1dbd-4b5d-88dc-9f518f917cc0",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",{"id":22,"name":23,"slug":23,"description":14,"color":14},"c883fd20-1d66-4fb7-9fc7-320fa7f87023","text-to-image",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"5908e484-8bdb-4dd7-86d3-f3f41d1a0324","en","Grok Imagine Image 2.0: #2 on both Arena boards","xAI shipped Grok Imagine Image 2.0 on August 7, 2026, rolling it out as the Quality Mode on grok.com\u002Fimagine, iOS, and Android, with the API as grok-imagine-image-2.0. The new model is built for real work: tight prompt fidelity, designer-grade typography and layout, and stable visual identity across generations. Editing is a first-class citizen — magic wand, segmentation, background removal, 5-reference editing, and Smart Resize. Image 2.0 ranks #2 worldwide on both the Image Edit Arena and the Text-to-Image Arena, behind only OpenAI gpt-image-2.","## August 7: xAI puts a production-grade image model on grok.com\n\nxAI officially launched **Grok Imagine Image 2.0** on August 7. The new model ships as the Quality Mode on [grok.com\u002Fimagine](https:\u002F\u002Fgrok.com\u002Fimagine) and the iOS \u002F Android Grok apps, with the API going live the same day under the model name  ([official post](https:\u002F\u002Fx.ai\u002Fnews\u002Fgrok-imagine-image-2)).\n\nUnlike the previous generation's pure showmanship, 2.0 is explicitly built to be **usable in real work**: tight prompt following, designer-grade layout and typography, and stable visual identity across generations. In other words, xAI is no longer selling \"looks good\" — it is selling \"looks the same, every time\".\n\n## Editing as a first-class capability\n\nThe official post gives editing its own section, instead of tucking it under generation:\n\n- **Magic wand + Segmentation**: edit exactly the region you point at, leave the rest untouched;\n- **Background removal**: export the subject with a transparent background, ready to drop into downstream designs;\n- **Multi-ref editing**: a single generation accepts up to **5 input images**, removing the need to manually composite in Photoshop;\n- **Smart Resize**: pick a target aspect ratio (1:1, 2:3, 9:16, 16:9, ...) and the model refills the frame rather than cropping.\n\nOn top of that, **Templates** package common workflows (Photo Edit, Product Color Change, Editorial Product Poster, Reimagine, Photo Collage, Mascot Maker, BG Removal & Change, E-Commerce Photos, UGC Photos, Professional Headshot, Icon Maker, Character Sprite, Props & UI Kit, Emoji Creator, Merch Maker) as ready-made input-to-output pipelines.\n\n## Cross-generation consistency: scaffolding the world for video\n\nxAI positions Image 2.0 as the scaffolding for Grok's video stack. In the \"Build a world for video\" section, the official demo shows the same character, multiple locations, and several props being generated from a single prompt set, and then the same character reappearing in snow, blizzard, and a glacier tunnel — with **consistent character, art style, and prop detail across all images**.\n\nThis directly attacks the most painful problem in current video generation: cross-shot consistency. Image 2.0's answer is to fix the world's visual assets at the static-image layer, then hand off to Grok Imagine Video. The same logic runs through ByteDance's Seedream 5.0 → Seedance 2.5 image-to-video production line.\n\n## The scoreboard: Arena #2 on both leaderboards\n\nxAI discloses that Image 2.0 ranks **#2 worldwide on both the Image Edit Arena and the Text-to-Image Arena**, behind only OpenAI gpt-image-2. xAI models are listed on Arena under the name \"SpaceXAI\" (figures as of August 7, 2026).\n\nThe third-party Neura Market report on August 7 corroborates the same ranking and adds detail on Smart Resize, multi-reference input, and region-level editing ([Neura Market](https:\u002F\u002Fwww.neura.market\u002Fnews\u002Fxai-grok-imagine-image-2-0-editing-tools-arena-rankings)).\n\n## What it means for the industry: from \"generate\" to \"produce\"\n\nRead against the second half of 2026, the signal is bigger than the model itself:\n\n1. **\"Usable\" is more expensive than \"pretty\"**. The official page keeps saying \"real work\" and \"real creative work\" — xAI is now targeting design \u002F e-commerce \u002F marketing teams, not just creative tinkerers.\n2. **Editing is the new battleground**. Qwen-Image-2.0, Seedream 5.0, and Grok Image 2.0 are all betting on \"multi-ref + region edit\". The next six months of differentiation will shift from \"does the image look right\" to \"can it be edited precisely\".\n3. **Scaffolding video is the hidden agenda**. As ByteDance Seedance 2.5 and JD JoyAI-Video-Edit push video into the 30-second era, whoever can lock in visual asset consistency first gets the leverage on the video side. Grok's Image 2.0 → Imagine Video chain is exactly the same product logic.\n\n## How to use it\n\n- **Consumers**: open [grok.com\u002Fimagine](https:\u002F\u002Fgrok.com\u002Fimagine) and switch to Quality Mode, or use the iOS \u002F Android Grok app;\n- **Developers**: call ; docs at [xAI developer docs](https:\u002F\u002Fdocs.x.ai\u002Fdevelopers\u002Fmodel-capabilities\u002Fimages\u002Fgeneration#quick-start);\n- **Playground**: [xAI Console](https:\u002F\u002Fconsole.x.ai\u002Fteam\u002Fdefault\u002Fimage?model=grok-imagine-image-2.0&flow=explore).\n\n**So what**: the text-to-image race is not settled yet, and \"how productively you can edit a single image\" is the metric that will separate the field over the next six months. With editing as a first-class capability and an Arena #2 on both leaderboards, xAI no longer wants to be the \"image skin\" of Grok — it wants a slot in the designer's daily toolchain. For domestic players (Qwen \u002F Seedream \u002F Zhipu), cross-generation consistency and templated workflows are probably the most urgent lessons to take into the second half of the year.","grok-imagine-image-2-0-arena-second","2026-08-13T02:00:00Z","2026-08-12T22:03:23.353273Z","2026-08-19T01:48:03.231362Z",true,"agent",96,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"671092ff-0ee4-4585-b703-af763a8afc60","微软MAI-Image-2.5闯入Arena图像编辑榜第二：局部编辑是图像模型的生产级分水岭","microsoft-mai-image-2-5-arena-edit-2","2026-06-06T04:01:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"b115486a-b837-4de1-9dac-d2237723ee85","宇树 UnifoLM-OminiA-0.3:G1 上跑通\"感知—行动\"端到端大模型","unitree-unifolm-ominia-0-3","2026-07-20T08:01:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"f1397080-206a-469f-846c-932a4b3ab8f9","京东开源 JoyAI-Image：统一多模态基础模型，把「理解-生成-编辑」拧成一个闭环","jd-joyai-image","2026-07-20T06:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"1ecbb79a-d843-43ad-b533-c01ae396275f","Qwen-Audio-3.0-Realtime：蒸馏拉满实时语音智商与延迟","qwen-audio-3-realtime","2026-07-15T10:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"19566223-1b02-4e48-8c44-518694edb049","Meta Muse Image 落地：Superintelligence Labs 把多模态推理与图生能力拧成一股","meta-muse-image","2026-07-07T20:01:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"ab4cf67e-175a-470b-9512-9d767be79fc6","Boogu-Image-0.1 开源家族：用比对手少一个数量级的数据，把\"理解+生成\"统一做到闭源水平","boogu-image-0-1-10b-unified-turbo","2026-06-26T14:00:00+00:00"]