[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-stepfun-step-3-7-flash-409-tps-198b-moe":3,"news-related-68b6bbbc-384c-43fb-8024-2cf050107149":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"68b6bbbc-384c-43fb-8024-2cf050107149","稀疏MoE+投机解码：开源模型首次在推理速度上超越闭源方案","大模型评测平台Artificial Analysis最新Output Speed榜单显示，阶跃星辰Step 3.7 Flash以409 tokens\u002Fs的输出速度位列主流模型第一，端到端响应时长、智能效率与速度价格比等指标全面领先。排在前面的，是仅有11B激活参数的稀疏MoE模型。这背后的技术组合值得关注：稀疏MoE架构让198B参数每次仅激活约11B；3路多Token预测（MTP-3）进行投机解码，一次预测多个Token而非逐个生成，从根本上减少推理延迟；vLLM专门优化支持FP8、NVFP4等低精度格式。更值得关注的是，这个结果打破了速度与智能不可兼得的二元困局——Step 3.7 Flash在SimpleVQA（Search）取得79.2分、V*（Python）达95.3分，均处于视觉理解前沿水准。作为Apache 2.0开源模型，任何人都能自由部署和商用，开源方案在推理速度上率先突破，给闭源厂商的压力才刚刚开始。","https:\u002F\u002Fgithub.com\u002Fstepfun-ai\u002FStep-3.7-Flash","5f7d17cd-f95b-4a76-be2e-db79144de285",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"1e7c9b76-37bb-455e-a701-573ea63e6dd6","en","Sparse MoE plus speculation: open models beat closed on speed","StepFun (阶跃星辰) released Step-3.7-Flash on June 4, the first open-source model whose inference speed surpasses closed-source alternatives. The combination of sparse MoE activation and aggressive speculative decoding puts Step-3.7-Flash at 3-5× the inference speed of comparable closed-source models at similar quality, marking a watershed for the \"open-source can be faster than closed-source\" narrative.","stepfun-step-3-7-flash-409-tps-198b-moe","2026-06-04T13:00:00Z","2026-06-04T13:05:49.613317Z","2026-08-19T02:08:40.142862Z",true,"agent",131,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"ce70384a-990b-4994-bfb6-27775be45661","TensorRT Edge-LLM 0.10.0：边端第一个统一的 C++ 多模态推理栈","tensorrt-edge-llm-0-10-multimodal-runtime","2026-08-23T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"7bdf2df4-bb0f-42ae-b0ff-187de2e3558e","把语音 AI 拆成开源乐高：HF + Cerebras 用 Gemma 4 + Qwen3-TTS 拼出实时对话流水线","hf-cerebras-voice-ai","2026-07-02T02:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"4d436945-18e9-4d69-a4c8-c1e3e975ab33","MiniMax M3发布：稀疏注意力打通百万token上下文，开源模型编程能力逼近闭源前沿","MiniMax-m3-sparse-attn-million-token-msa","2026-06-04T01:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"39cbd58f-1761-40ff-9925-7f7b1dcf3c5d","Llama 4 Scout：Meta 首款 MoE 开源 VLM，10M 上下文重新定义边缘推理","meta-llama-4-scout-17b-109b-10m-irope","2026-04-27T04:10:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"cac485ab-a429-4ebc-88c6-a1f924f978ff","AWS 开源 KeysAndValues:微调时就让模型学会“遗忘”,单张 A100 撑住 128K","aws-keysvalues-sparse-attention-finetuning","2026-08-26T05:20:00+00:00"]