[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qwen-image-2-1-open-source-7b":3,"topics-all":41,"news-related-88a287ca-c7c4-4f31-aff9-8a5116d02e5e":60},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"88a287ca-c7c4-4f31-aff9-8a5116d02e5e","Qwen-Image-2.1:7B 轻量模型把生成+编辑焊进一套权重","阿里千问 9 月 20 日开源 Qwen-Image-2.1 图像模型,视觉生成仅 7B 参数,文生图与图像编辑合并为同一套权重,原生支持透明图像生成与编辑,最多支持 10 张参考图输入,官方基准反超 Nano Banana 2.0 等闭源对手。","9 月 20 日,阿里千问开源了 Qwen-Image-2.1 —— 一款把「文生图」与「图像编辑」塞进同一套权重的开源扩散模型,视觉生成部分只有 7B 参数,GitHub、Hugging Face、ModelScope 三处权重同步放出,推荐 16GB 以上显存的 GPU 即可本地部署。\n\n## 一套权重覆盖「生成 + 编辑」两条管线\n\n最显眼的能力是原生透明图像( RGBA )生成与编辑。过去文生图模型只输出 RGB,要做电商贴图或海报设计必须走一遍抠图流程;Qwen-Image-2.1 直接把透明通道纳入训练目标,既能根据提示词生成带 alpha 通道的图层,也能在保留背景透明的前提下修改主体表情、替换图层内文字,透明能力继承自此前专用模型 Qwen-Image-Layered。同一个模型还把多图参考与局部编辑纳入输入端:最多 10 张参考图、圈选 \u002F 涂抹 \u002F 掩码三种显式控制方式,人像面部细节与商品文字纹理在多次编辑后保持一致。\n\n## 技术栈:Single-Stream DiT + 混合粒度注意力\n\n走的是 32 层 Single-Stream DiT 架构,文本与图像 Token 在同一 Transformer 流里统一处理;系统前缀与编辑指令用 Token 级因果掩码、图像块走 Chunk 级掩码,两套粒度并行;参考图与编辑指令视为静态上下文,首步预计算 Key-Value 后复用,官方称多图场景的推理效率因此显著提升。\n\n## 基准与生态站位\n\n官方评测里,Qwen-Image-2.1 在文生图 + 编辑合并基准上反超了 Google 的 Nano Banana 2.0 等闭源对手,被多家媒体列入近期开源榜首;权重全量开源、文生图与编辑二合一、原生 RGBA 输出,这些特性 Google 同类产品尚未提供。\n\n对设计、电商、内容创作这条工作流来说,Qwen-Image-2.1 的意义不是「更会画」,而是「免抠图 + 免切换模型」—— 一套权重解决生成与修改,加上原生透明通道,贴近生产管线的可用性。当开源生态把这种「生成 + 后期」合并能力做到闭源 API 同档,设计稿的边际成本会再被压一刀。","https:\u002F\u002Fqwen.ai\u002Fblog?id=qwen-image-2.1","c36a21ac-2a77-421b-9519-1e150695732a",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"e676a5cf-1f24-472f-a765-86fa21a1bc3c","ai-model",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",{"id":25,"name":26,"slug":26,"description":14,"color":14},"c883fd20-1d66-4fb7-9fc7-320fa7f87023","text-to-image",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"f626c331-a8e6-4ace-a7b4-da6f6b83e6df","en","Qwen-Image-2.1: 7B Model Bakes Generation and Editing Into One Set of Weights","Alibaba's Qwen team open-sourced Qwen-Image-2.1 on September 20. The vision generation module weighs in at just 7B parameters, folding text-to-image and image editing into the same weights. It natively supports RGBA transparent image generation and editing, takes up to 10 reference images as input, and its official benchmark edged out closed-source peers such as Google's Nano Banana 2.0.","On September 20, Alibaba's Qwen team open-sourced Qwen-Image-2.1 — a diffusion model that collapses text-to-image and image editing into a single set of weights. The vision generation component is only 7B parameters, with weights mirrored to GitHub, Hugging Face, and ModelScope. A consumer GPU with 16GB or more of VRAM is enough to run it locally.\n\n## One checkpoint covers both generation and editing\n\nThe headline capability is native transparent (RGBA) image generation and editing. Most text-to-image models only emit RGB, forcing designers and e-commerce operators to run an extra matting pass to drop a subject onto a poster or product page. Qwen-Image-2.1 bakes the alpha channel into the training objective: prompts can request a layer with transparency built in, and edits preserve the transparent background while swapping expressions, replacing in-layer text, and so on. The transparent-image capability is inherited from the team's earlier specialised Qwen-Image-Layered model.\n\nThe same checkpoint also folds multi-reference conditioning and local editing into its input stage. Up to 10 reference images can be supplied, and three explicit control modes — marquee selection, brush strokes, and explicit masks — drive local edits. Across multiple rounds, facial details and product textures (logos, fabric patterns) stay consistent, which is the property e-commerce and portrait workflows actually need.\n\n## Stack: Single-Stream DiT with mixed-granularity attention\n\nThe architecture is a 32-layer Single-Stream Diffusion Transformer. Text and image tokens flow through one unified Transformer stream. System prefixes and editing instructions are routed through a token-level causal mask, while image patches use a chunk-level mask — two granularities running in parallel. Reference images and editing instructions are treated as static context: their Key-Value pairs are precomputed on the first step and reused on every subsequent step, which the team credits for the visible efficiency gains in multi-reference scenarios.\n\n## Benchmarks and where it sits in the ecosystem\n\nOn the team's own benchmark combining text-to-image and editing, Qwen-Image-2.1 edged out closed-source rivals such as Google's Nano Banana 2.0 and was listed by multiple outlets as the current open-source leader. The combination of fully open weights, unified generation-plus-editing, and native RGBA output is something Google's comparable products still don't offer.\n\nFor design, e-commerce, and content creation workflows, Qwen-Image-2.1's significance is not that it draws better pictures — it is that it removes the matting step and the model-switching step. One checkpoint handles generation and editing, and the native alpha channel pushes the workflow closer to a production-ready pipeline. When the open-source stack matches closed APIs on this kind of combined generation-plus-post capability, the marginal cost of design assets gets cut another notch.","qwen-image-2-1-open-source-7b","2026-09-20T16:00:00Z","2026-09-25T07:04:46.109836Z","2026-09-25T07:04:46.109846Z",true,"agent",787,[42,51],{"slug":43,"tag_slug":43,"title_zh":44,"title_en":45,"intro_zh":46,"intro_en":47,"id":48,"is_active":38,"created_at":49,"modified_at":50},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":52,"tag_slug":52,"title_zh":53,"title_en":54,"intro_zh":55,"intro_en":56,"id":57,"is_active":38,"created_at":58,"modified_at":59},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":61},[62,67,72,77,82,87],{"id":63,"title":64,"news_slug":65,"published_at":66},"37eb94e4-2155-416a-a8d9-bc5075541a27","Qwen-Image-3.0 发布:把文生图从「好看」推向「好用」","qwen-image-3-0","2026-07-21T08:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"59d14de3-0e17-4202-8e7d-ad0bc51e3471","Qwen-Image-2.1-Turbo开源:8步去噪出图","qwen-image-2-1-turbo-8-step","2026-10-11T15:05:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"cb5ee922-ad27-4a5a-9b2b-8382903876df","Mozilla 把模型选择权交还给用户:Mistral Small 4 进 Firefox 默认菜单","mistral-small-4-firefox-smart-window-beta","2026-09-22T03:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"d056f67b-7e0d-4e44-8d39-e31ea50deeae","Bonsai 2 27B 三元压缩:Qwen3.8 压到 5.9 GB,benchmark 留存 98.2%","bonsai-2-27b-ternary-qwen3-8-compression","2026-09-17T15:47:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"c480d2d0-9156-4aa7-826f-fba2f252b6b7","蚂蚁开源 LLaDA-Image:6B 参数生成编辑一体,Turbo 版 4 步出图","ant-llada-image-open-generation-editing","2026-09-04T13:09:21+00:00",{"id":88,"title":89,"news_slug":90,"published_at":91},"d941056b-c2e7-42e5-965a-a982c20b1169","Qwen3.8-Flash-Next 架构细节:Gated Residual 多分支残差 + QSA micro-block 稀疏注意力","qwen3-8-flash-next-cost-efficiency-architecture","2026-09-02T02:00:00+00:00"]