[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-novelai-diffusion-v5-release":3,"news-related-6082cd23-0eca-40e0-9315-67318dc818ee":35},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":21,"news_slug":28,"published_at":29,"created_at":30,"modified_at":31,"is_published":32,"publish_type":33,"image_url":14,"view_count":34},"6082cd23-0eca-40e0-9315-67318dc818ee","NovelAI Diffusion V5 发布:规模翻倍、32 通道 VAE,单次生成整页漫画","图像生成服务商 NovelAI 发布 Diffusion V5:规模超 V4.5 两倍,训练耗 26.8 万 B200 GPU 小时,VAE 从 16 通道升级到 32 通道。新模型强化自然语言与日语提示,测试中 22 角色同屏,单次生成多面板漫画页,原生支持 alpha 透明通道。","动漫图像生成的老牌玩家 NovelAI 拿出了新一代图像模型。8 月 21 日,NovelAI Diffusion V5 正式发布,同时推出 V5 Curated 与 V5 Full 两个版本。官方公告给出的核心数字很直接:基于自有架构,模型规模是 V4.5 的两倍以上,训练消耗 268,000 B200 GPU 小时([官方发布日志](https:\u002F\u002Fjournal.novelai.net\u002Fimage-generation-novelai-diffusion-v5-is-here-c2df7c6b8d2d\u002F))。\n\n## 架构升级的重点:一枚 32 通道 VAE\n\nV5 最值得注意的技术改动藏在压缩端:官方自训的 VAE 从 V4.5 的 16 通道扩展到 32 通道。VAE 负责在潜空间里表示图像,通道数翻倍直接决定了模型能\"看清\"多少细节。官方举例说,眼睛里的细线、珠宝、文字、配饰这类容易糊成一团的元素,现在渲染精度明显提高,\"融化感\"大幅减少。配合更大的模型体量,建筑、植被等复杂背景的连贯性也同步提升。\n\n## 提示词:继续向自然语言迁移\n\nV4.5 已经开始支持用自然语言代替标签写提示词,V5 把这条路走得更远。官方支持的语言是英语和日语两种,日语有专门的日文转写数据训练支撑;测试中发现中文、德语、西班牙语、葡萄牙语也能写提示词,但效果不稳定,因为这些语言不在训练重点里。\n\n## 一次生成整页漫画\n\n角色控制是这次更新的另一个重心。测试中最多 22 个角色同屏;Character Positioning 从固定小网格改成画布上的自由摆放,模型会贴近你指定的位置,减少角色之间特征\"串味\"。更长 prompt 的支持也随之开放。漫画能力是 V5 的差异化卖点:用自然语言描述版面布局,单次生成就能输出完整的多面板漫画页,而旧版的漫画数据只覆盖双格(2koma)级别。文字渲染支持英语、日语、中文等多种语言,前端现在会在你给文字加引号时自动生成 Text 块,不再需要手动准备。\n\n## 一些\"没做完\"的部分\n\nV5 不是全家桶:inpainting 目前只随 V5 Full 一起提供,curated 版还在训练,V5 Curated 暂时沿用 V4.5 的方案;Precise Reference、Vibe Transfer 也没有随发布上线。官方表示会像以往一样在发布后持续训练补齐。此外新自定义 VAE 带来了原生 alpha 透明通道:用 \"transparent background\"、\"has alpha\" 等标签就能生成透明背景或半透明效果,Enhance 功能新增 \"Max✨\" 档位,配套的新 upscaler 也可以单独使用。UI 同步整体翻新,订阅与用量规则也做了调整。\n\n## 值得关注的角度\n\n把 V5 放在整个图像生成市场里看,这是一次典型的\"垂直玩家加固护城河\"的更新。通用模型的自然语言理解在变强,但对二次元风格、多角色控制、漫画版面这类细分需求,通用模型很难像专门训练的模型这样下功夫;而 268,000 B200 GPU 小时的投入,也说明这个细分市场的算力门槛在抬高。对创作者来说,真正的问题不是\"哪个模型更强\",而是你的工作流落在哪个细分场景里——单次成页的漫画生成如果真的稳,它会改变的不只是出图速度,还有分镜创作的迭代方式。","https:\u002F\u002Fjournal.novelai.net\u002Fimage-generation-novelai-diffusion-v5-is-here-c2df7c6b8d2d\u002F","1f340b0b-504a-4d12-86d2-55ad51a5063a",[11,15,18],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"c883fd20-1d66-4fb7-9fc7-320fa7f87023","text-to-image",[22],{"id":23,"lang":24,"title":25,"summary":26,"content":27},"eae5e44c-795e-4f89-98c1-1ea7f070ec15","en","NovelAI Diffusion V5: Twice the Size, 32-Channel VAE, Comic Pages","NovelAI ships Diffusion V5 (Curated and Full): over twice the size of V4.5, trained on 268,000 B200 GPU-hours, with a custom 32-channel VAE. Stronger natural-language and Japanese prompting, up to 22 characters per scene, full multi-panel comic generation, and native alpha transparency.","NovelAI, the long-running anime-focused image generation service, has shipped its next-generation model. Announced on August 21, NovelAI Diffusion V5 arrives in two flavors, V5 Curated and V5 Full. The headline numbers from the official announcement are direct: built on the NovelAI in-house architecture, the model is more than twice the size of V4.5 and was trained in-house on 268,000 B200 GPU-hours ([official release post](https:\u002F\u002Fjournal.novelai.net\u002Fimage-generation-novelai-diffusion-v5-is-here-c2df7c6b8d2d\u002F)).\n\n## The Core Architectural Change: A 32-Channel VAE\n\nThe most technically interesting upgrade sits in the compression stage: the custom-trained VAE expands from 16 channels in V4.5 to 32 channels. The VAE represents images in latent space, and doubling the channel count directly determines how much detail the model can \"see\". The team cites fine lines in intricately drawn eyes, jewelry, text, and accessories as elements that now render with noticeably higher accuracy — less of the familiar \"melting goop\". Combined with the larger model size, complex environments featuring architecture and foliage also come through with more detail and coherence.\n\n## Prompting: Continuing the Shift to Natural Language\n\nV4.5 already let users write prompts in natural language instead of tags; V5 pushes that much further — the official framing is an unprecedented level of understanding. Officially supported languages are English and Japanese, with Japanese backed by training on Japanese text transcriptions. Testers also found Chinese, German, Spanish, and Portuguese prompts working, though results may vary since those were not a training focus.\n\n## Full Comic Pages in One Generation\n\nCharacter control is the other centerpiece. In testing, up to 22 distinct characters appeared on screen at once. Character Positioning moves from small fixed grids to free placement on the canvas, with the model following the positions you set and reducing feature bleed between characters. Longer prompts are supported as well. The comic capability is the V5 differentiator: describe the layout in natural language and a single generation produces a fully paneled page — the old training data only covered 2koma strips. Text rendering supports multiple languages including English, Japanese, and Chinese, and the frontend now auto-prepares the \"Text:\" block when you quote text.\n\n## What Is Not Finished Yet\n\nV5 does not ship complete: inpainting is available only with V5 Full at launch, while the curated inpainting model is still training — V5 Curated temporarily reuses the V4.5 inpainting function. Precise Reference and Vibe Transfer are also absent from launch day. The team says additional features will roll out after release, as usual. Meanwhile, the new custom VAE brings native alpha transparency: prompting \"transparent background\", \"has alpha\", or \"alpha transparency\" yields transparent backgrounds or translucent effects. The Enhance function gains a new \"Max✨\" tier backed by a new upscaler that also works standalone. The UI got a total overhaul alongside the release, and subscription usage limits were adjusted.\n\n## Why It Matters\n\nPlaced in the broader image-generation market, V5 reads as a vertical player reinforcing its moat. General-purpose models keep getting better at natural language, but anime styling, multi-character control, and comic layout are niche demands that general models rarely address with this depth — and the 268,000 B200 GPU-hour budget signals that the compute bar for this niche keeps rising. For creators, the real question is not which model is \"strongest\", but which workflow your work lives in. If single-shot full-page comic generation holds up in practice, it changes not just output speed but how storyboard iteration works.","novelai-diffusion-v5-release","2026-08-22T13:10:00Z","2026-08-22T13:07:28.385768Z","2026-08-22T13:07:28.385777Z",true,"agent",441,{"items":36},[37,42,47,52,57,62],{"id":38,"title":39,"news_slug":40,"published_at":41},"095917eb-02ae-4fd2-a1cb-17d0805442ee","微软 Mage-Flow 用 4B 跑赢 32B：原生分辨率 + 三件套协同设计把生成编辑都塞回单卡","microsoft-mage-flow-4b","2026-07-23T03:30:00+00:00",{"id":43,"title":44,"news_slug":45,"published_at":46},"22aad878-387f-4166-b68c-896b39a27de3","Sourceful Riverflow 2.5：把「评分函数」塞进图像生成，让「什么是好图」变成可编程的","sourceful-riverflow-2-5","2026-07-01T10:00:00+00:00",{"id":48,"title":49,"news_slug":50,"published_at":51},"f7287cac-6643-4f4a-8cbd-2b281d2d4d46","Krea 2 开源双发：12B DiT 把「2 秒出图」做进主流程，蒸馏后 8 步直出 2K","krea-2-12b-dit-2-second-turbo","2026-06-25T10:30:00+00:00",{"id":53,"title":54,"news_slug":55,"published_at":56},"72ee21ea-8a91-4dd3-88fa-f605551ff9ce","Qwen-Image-2.0 发布：7B 拿下原生 2K，把「图文一体 + 生成编辑统一」推到开源前沿","qwen-image-2-0-7b-native-2k-arena-no1","2026-06-18T08:00:00+00:00",{"id":58,"title":59,"news_slug":60,"published_at":61},"cd88ab8f-afff-4f8f-8edc-ab24715906c6","FLUX.2 [klein] 4B\u002F9B 发布：统一生图编辑，Apache 2.0","flux-2-klein-4b-9b-apache-2-sub-second","2026-06-12T06:30:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"04d03b80-0a32-4ea1-87df-9248b36653c1","Ideogram 4.0 开源：9.3B 单流 DiT + Qwen3-VL 文本编码器，把排版与文字渲染做到开源第一","ideogram-4-0-9-3b-dit-qwen3-vl-text","2026-06-09T12:00:00+00:00"]