NovelAI, the long-running anime-focused image generation service, has shipped its next-generation model. Announced on August 21, NovelAI Diffusion V5 arrives in two flavors, V5 Curated and V5 Full. The headline numbers from the official announcement are direct: built on the NovelAI in-house architecture, the model is more than twice the size of V4.5 and was trained in-house on 268,000 B200 GPU-hours (official release post).

The Core Architectural Change: A 32-Channel VAE

The most technically interesting upgrade sits in the compression stage: the custom-trained VAE expands from 16 channels in V4.5 to 32 channels. The VAE represents images in latent space, and doubling the channel count directly determines how much detail the model can "see". The team cites fine lines in intricately drawn eyes, jewelry, text, and accessories as elements that now render with noticeably higher accuracy — less of the familiar "melting goop". Combined with the larger model size, complex environments featuring architecture and foliage also come through with more detail and coherence.

Prompting: Continuing the Shift to Natural Language

V4.5 already let users write prompts in natural language instead of tags; V5 pushes that much further — the official framing is an unprecedented level of understanding. Officially supported languages are English and Japanese, with Japanese backed by training on Japanese text transcriptions. Testers also found Chinese, German, Spanish, and Portuguese prompts working, though results may vary since those were not a training focus.

Full Comic Pages in One Generation

Character control is the other centerpiece. In testing, up to 22 distinct characters appeared on screen at once. Character Positioning moves from small fixed grids to free placement on the canvas, with the model following the positions you set and reducing feature bleed between characters. Longer prompts are supported as well. The comic capability is the V5 differentiator: describe the layout in natural language and a single generation produces a fully paneled page — the old training data only covered 2koma strips. Text rendering supports multiple languages including English, Japanese, and Chinese, and the frontend now auto-prepares the "Text:" block when you quote text.

What Is Not Finished Yet

V5 does not ship complete: inpainting is available only with V5 Full at launch, while the curated inpainting model is still training — V5 Curated temporarily reuses the V4.5 inpainting function. Precise Reference and Vibe Transfer are also absent from launch day. The team says additional features will roll out after release, as usual. Meanwhile, the new custom VAE brings native alpha transparency: prompting "transparent background", "has alpha", or "alpha transparency" yields transparent backgrounds or translucent effects. The Enhance function gains a new "Max✨" tier backed by a new upscaler that also works standalone. The UI got a total overhaul alongside the release, and subscription usage limits were adjusted.

Why It Matters

Placed in the broader image-generation market, V5 reads as a vertical player reinforcing its moat. General-purpose models keep getting better at natural language, but anime styling, multi-character control, and comic layout are niche demands that general models rarely address with this depth — and the 268,000 B200 GPU-hour budget signals that the compute bar for this niche keeps rising. For creators, the real question is not which model is "strongest", but which workflow your work lives in. If single-shot full-page comic generation holds up in practice, it changes not just output speed but how storyboard iteration works.