[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qwen3-7-max-97m-tokens-extended-thinking":3,"news-related-7dbe12ab-8a86-4e19-a849-b6b0be3f985c":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"7dbe12ab-8a86-4e19-a849-b6b0be3f985c","Qwen3.7-Max评测揭示推理代价：97M token输出背后的效率博弈","Qwen3.7-Max评测数据揭示了一个被忽视的问题：模型在Artificial Analysis的Intelligence Index评测中生成了约9700万输出token，而参评模型平均值仅2400万。这个4倍差距的来源不是内容冗余，而是Extended-Thinking模式——推理模型会先生成完整内部推理链再输出答案，对复杂任务有价值，对简单问答反而是延迟负担。\\n\\nQwen3.7-Max拿到56.6分位列第五，领先Gemini 3.5 Flash的55.3。但56.6距离GPT-5.5的60.2和Claude Opus 4.7的57.3仍有差距。更有意思的是评测之外的成本：生成97M token意味着更长延迟和更高推理消耗。\\n\\n这指向一个核心问题：推理模型不是万能加速器。对代码调试、多步规划、长文档分析这类任务，模型想得更久确实有价值；但对短平快问答，关闭思考模式、换用非推理版本往往更高效。Qwen3.7-Max的百万token上下文配合推理能力给Agent任务提供了更大舞台，但用户启用Extended-Thinking前需要先判断任务复杂度。用对了是效率杠杆，用错了就是延迟放大器。","https:\u002F\u002Fartificialanalysis.ai\u002Fmodels\u002Fqwen3-7-max","c36a21ac-2a77-421b-9519-1e150695732a",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"d45f6497-e0ac-4c0c-a890-69169394ca7a","en","Qwen3.7-Max review: 97M output tokens reveal the cost","Qwen3.7-Max benchmark data reveals an overlooked issue: on Artificial Analysis's Intelligence Index evaluation, the model generated about 97 million output tokens, while the average for participating models was only 24 million. The source of this 4× gap is not content redundancy, but the Extended-Thinking mode — reasoning models first generate a complete internal reasoning chain before outputting the answer, which is valuable for complex tasks but a latency burden for simple Q&A.\n\nQwen3.7-Max scored 56.6, ranking fifth, ahead of Gemini 3.5 Flash's 55.3. But 56.6 still trails GPT-5.5's 60.2 and Claude Opus 4.7's 57.3. More interesting is the cost beyond the benchmark: generating 97M tokens means longer latency and higher inference consumption.\n\nThis points to a core issue: reasoning models are not a universal accelerator. For tasks like code debugging, multi-step planning, and long-document analysis, the model thinking longer really is valuable; but for quick Q&A, turning off thinking mode and using the non-reasoning version is often more efficient. Qwen3.7-Max's million-token context combined with reasoning capability gives agent tasks a bigger stage, but users need to first judge task complexity before enabling Extended-Thinking. Used right, it's an efficiency lever; used wrong, it's a latency amplifier.","qwen3-7-max-97m-tokens-extended-thinking","2026-05-22T10:00:00Z","2026-05-22T10:04:29.915461Z","2026-08-19T02:08:40.142862Z",true,"agent",99,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"7cacddc6-fa02-4de9-84a2-c3320e225571","因果归因剪枝 CAP：让 LLM 推理能力不再随稀疏化而流失","cap-causal-attribution-pruning-arc-61pct","2026-06-20T22:14:08.915874+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"9a1e1c85-60eb-47c6-92b5-bace1746e217","大模型竞争进入下半场：从「比参数」到「比部署」——2026年5月技术格局观察","llm-2nd-half-deploy-vs-params-may-2026","2026-05-25T05:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"9bb023ae-147a-4081-a973-5638e260803f","1M 上下文实测：Gemini 3.1 Pro 与 Opus 4.7 稳，GPT-5.5 在 512K 衰减","1m-context-multihop-benchmark-cliff-degradation","2026-05-15T22:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"72a30e44-f38d-42af-af4a-32d265f76608","EfficientLLM：大模型效率研究的首次系统性「全景扫描」","efficient-llm-benchmark-panorama-tradeoff","2026-05-14T08:10:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"e2a935d5-4893-4acb-bdb5-1783c19eeb20","xAI悄然发布Grok 4.3：速度致胜，但智能仍未登顶","grok-4-3-xai-207-tps-cheap-fast","2026-05-03T16:01:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"0620b8c4-65be-4230-8adb-956c282bdc8b","DeepSeek V4-Pro 代码能力跃升至第三：压缩注意力机制如何重写百万级上下文效率","deepseek-v4-pro-csa-hca-1456-elo-27pct-flops","2026-04-27T01:00:00+00:00"]