[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ideogram-4-0-9-3b-dit-qwen3-vl-text":3,"news-related-04d03b80-0a32-4ea1-87df-9248b36653c1":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"04d03b80-0a32-4ea1-87df-9248b36653c1","Ideogram 4.0 开源：9.3B 单流 DiT + Qwen3-VL 文本编码器，把排版与文字渲染做到开源第一","Ideogram 4.0 是该公司 6 月 3 日发布的首款开源权重模型：9.3B 参数、34 层 single-stream Diffusion Transformer，从零训练，原生 2K 输出（最高 2048 像素，宽高比可达 6:1），同步提供 fp8 与 nf4 两种量化版本，nf4 版可单卡 24GB GPU 部署。推理阶段还提供 V4_QUALITY_48（45 步 + 3 步精修）、V4_DEFAULT_20、V4_TURBO_12 三档采样预设，可在质量和速度之间灵活切换。\n\n最大差异化在「JSON 结构化提示 + 强排版控制」。训练语料全部用 JSON 描述，每个元素可带颜色面板（每图 16 个 hex 色、每元素 5 色）、边界框坐标（[y_min, x_min, y_max, x_max]，归一化到 0–1000）与字面文本字段，可同时控制多行、多字体 in-image 文字。这是它在 X-Omni-OCR 文字渲染榜单上大幅领先同类开源模型（20B Qwen-Image、32B FLUX.2 dev、80B HunyuanImage 3.0 MoE）的关键。\n\nContraLabs 盲评中，10 位职业设计师对四款模型两两对比，Ideogram 4.0 以 47.9% 偏好率排名第一，显著高于 Nano Banana 2（30.0%）、FLUX.2 max（15.5%）和 Grok Imagine 1.0（15.0%）；「是否愿意用于真实客户作品」的 5 分制评分达 3.55，领先 Nano Banana 2 近 0.7 分。\n\n架构上，单流 DiT 已是 Z-Image、Black Forest Labs 等的共识，但 Ideogram 用 Qwen3-VL 8B 作文本编码器是少见做法——抽取 13 层中间 hidden state 与图像 token 拼接后送入同一 34 层 Transformer，让 prompt-alignment（Prism）、空间推理（SpatialGenEval）、布局控制（7Bench）三方面同时补齐短板。权重仅限非商业用途，商业落地需走官方 API，部署侧还要求接入 Hive 做内容审核。","https:\u002F\u002Fideogram.ai\u002Fblog\u002Fideogram-4.0\u002F","1daf7121-4782-417d-94ab-d690f2904cd8",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":13,"color":13},"c883fd20-1d66-4fb7-9fc7-320fa7f87023","text-to-image",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"9b7afca5-83f1-491e-8cfe-bd4204ca13a6","en","Ideogram 4.0: 9.3B DiT tops open-source typography","Ideogram 4.0 is the company's first open-source-weights model, released on June 3: 9.3B parameters, 34 layers of single-stream Diffusion Transformer, trained from scratch, native 2K output (up to 2048 pixels, with up to 6:1 aspect ratio), with fp8 and nf4 quantization versions released simultaneously, and the nf4 version deployable on a single 24GB GPU. At inference, three sampling presets — V4_QUALITY_48 (45 steps + 3 refinement), V4_DEFAULT_20, and V4_TURBO_12 — are available, with flexible quality\u002Fspeed trade-offs.\n\nThe biggest differentiation is \"JSON-structured prompts + strong typography control.\" The training corpus uses JSON for all descriptions; each element can carry a color palette (16 hex colors per image, 5 colors per element), a bounding box (y_min, x_min, y_max, x_max, normalized to 0-1000), and a literal text field, allowing simultaneous control of multi-line, multi-font in-image text. This is key to its large lead over comparable open-source models (20B Qwen-Image, 32B FLUX.2 dev, 80B HunyuanImage 3.0 MoE) on the X-Omni-OCR text-rendering leaderboard.\n\nIn ContraLabs' blind test, 10 professional designers did pairwise comparisons of four models. Ideogram 4.0 took first place with a 47.9% preference rate, well above Nano Banana 2 (30.0%), FLUX.2 max (15.5%), and Grok Imagine 1.0 (15.0%); the 5-point score for \"willingness to use in real client work\" reached 3.55, leading Nano Banana 2 by nearly 0.7 points.\n\nArchitecturally, the single-stream DiT is already the consensus of Z-Image, Black Forest Labs, and others, but Ideogram's choice of Qwen3-VL 8B as the text encoder is unusual — extracting 13 layers of intermediate hidden states and stitching them with image tokens, then feeding them into the same 34-layer Transformer, simultaneously patches up short boards in prompt-alignment (Prism), spatial reasoning (SpatialGenEval), and layout control (7Bench). Weights are limited to non-commercial use, and commercial deployment requires going through the official API, with deployment also requiring integration with Hive for content moderation.","ideogram-4-0-9-3b-dit-qwen3-vl-text","2026-06-09T12:00:00Z","2026-06-09T12:19:47.142736Z","2026-08-19T02:08:40.142862Z",true,"agent",181,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"095917eb-02ae-4fd2-a1cb-17d0805442ee","微软 Mage-Flow 用 4B 跑赢 32B：原生分辨率 + 三件套协同设计把生成编辑都塞回单卡","microsoft-mage-flow-4b","2026-07-23T03:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"f7287cac-6643-4f4a-8cbd-2b281d2d4d46","Krea 2 开源双发：12B DiT 把「2 秒出图」做进主流程，蒸馏后 8 步直出 2K","krea-2-12b-dit-2-second-turbo","2026-06-25T10:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"72ee21ea-8a91-4dd3-88fa-f605551ff9ce","Qwen-Image-2.0 发布：7B 拿下原生 2K，把「图文一体 + 生成编辑统一」推到开源前沿","qwen-image-2-0-7b-native-2k-arena-no1","2026-06-18T08:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"cd88ab8f-afff-4f8f-8edc-ab24715906c6","FLUX.2 [klein] 4B\u002F9B 发布：统一生图编辑，Apache 2.0","flux-2-klein-4b-9b-apache-2-sub-second","2026-06-12T06:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"6082cd23-0eca-40e0-9315-67318dc818ee","NovelAI Diffusion V5 发布:规模翻倍、32 通道 VAE,单次生成整页漫画","novelai-diffusion-v5-release","2026-08-22T13:10:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"5612d186-46ee-4509-9a93-94045ba004ae","LTX-2.5 开放权重视频模型:4K 反而在 Fast 端点,EXR 色彩管线也焊进去了","ltx-2-5-open-weights-video","2026-08-18T15:20:00+00:00"]