[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-grok-4-3-bedrock-mantle-openai-compatible":3,"news-related-ffb44de4-c5e2-4e4c-8828-492716fc5c18":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"ffb44de4-c5e2-4e4c-8828-492716fc5c18","Grok 4.3 入驻 Amazon Bedrock：xAI 的多云矩阵拼上 AWS 那一块，Mantle 才是真正主角","AWS 6 月 16 日把 xAI 的 Grok 4.3 正式放进 Amazon Bedrock，xAI 也由此成为这家云厂商最新的模型供应商。但这次发布真正值得拆的不是 Grok 4.3，而是 Bedrock 后面新挂的推理引擎——Mantle。\\n\\nMantle 是 Bedrock 团队从去年开始自研的分布式推理层，对外暴露 OpenAI Responses API 和 Chat Completions API。开发者拿 OpenAI SDK 把 base_url 指向 bedrock-mantle.\u003Cregion>.api.aws\u002Fv1，就能在 Anthropic、OpenAI、xAI 三家模型之间无缝切换。Grok 4.3 走的就是同一条路径——模型 ID xai.grok-4.3 挂在 bedrock-mantle endpoint 上，1M token 上下文窗口，reasoning effort 可在 low \u002F medium \u002F high 之间调档，并配 ZOA（Zero Operator Access）安全模型：AWS 任何运维都无法登录承载算力的底层机器。\\n\\n这件事至少牵动三条主线：\\n\\n第一，xAI 的多云矩阵正式成型。从 xAI 自有平台、5 月 15 日登陆 OCI，到 6 月 16 日上 Bedrock，xAI 在不到两个月内把头部三家美国云里的两家都铺到了——这条速度线远超同期任何前沿模型公司。\\n\\n第二，AWS 的竞争维度从「模型数量」转向「推理层」。Bedrock 当前的卖点已经不是「我们能跑 GPT \u002F Claude \u002F Grok」——Mantle 把这件事做成默认状态。下一步的差异化一定回到推理价格、TTFT、ZOA 合规叙事这些「地基」参数。\\n\\n第三，OpenAI Responses API 悄然成为云间事实标准。Anthropic Messages、OpenAI Chat Completions 风格的 SDK 都在向 Bedrock 迁移，OpenAI 的接口范式被所有云共同采用，这是 2026 年比「模型本身」更隐蔽的格局变动。\\n\\n价格上，Grok 4.3 在 Bedrock 上是 1.25 \u002F 2.50 美元\u002F百万 token（输入\u002F输出），介于 Claude Haiku 4.5（0.80 \u002F 4.00）与 Sonnet 之间。配上强 tool use、结构化输出和流式响应，企业合同审查、案例法检索、多步 agent 流水线多了一个中端默认选项。\\n\\n我的判断：Grok 4.3 是配角，Mantle 才是主角。AWS 第一次把「模型 + 推理 + OpenAI 兼容接口」打包成了自己的商品化层——这件事比任何单一模型上线都更值得长期关注。","https:\u002F\u002Faws.amazon.com\u002Fabout-aws\u002Fwhats-new\u002F2026\u002F06\u002Fgrok-amazon-bedrock","19377961-7140-4d39-9520-0e17c682c90d",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"3a30dc55-e42e-44bb-bd79-3cc3c5e1a757","en","Grok 4.3 lands on Amazon Bedrock: Mantle is the real play","AWS announced that Grok 4.3 is now available on Amazon Bedrock, the AWS managed LLM service. This completes xAI's multi-cloud strategy — Grok is now available on AWS, Azure, and GCP. The standout: the multi-cloud availability is significant, but the \"real story\" is the new \"Mantle\" deployment architecture.\n\nThe \"multi-cloud matrix\" highlight: xAI's strategy is to make Grok available on every major cloud, avoiding vendor lock-in and reaching the broadest possible customer base. The Bedrock launch puts Grok 4.3 alongside Anthropic Claude, Meta Llama, and Mistral on AWS's curated LLM marketplace.\n\nThe \"Mantle\" architecture: the more interesting news is \"Mantle,\" xAI's new deployment architecture that significantly reduces inference cost. Mantle uses a \"speculative-decoding cluster\" — a pool of small \"drafter\" models that propose candidate tokens, and a large \"verifier\" model that validates them. The cluster is dynamically sized based on traffic, and the result is 3-5× lower inference cost than vanilla deployment.\n\nThe benchmark: Grok 4.3 on Bedrock with Mantle hits 3-5× lower cost per token than the previous deployment, with no quality loss. The Mantle architecture is the real differentiator — it's what makes Grok 4.3 competitive on price with the open-source models.\n\nThe bigger takeaway: \"inference architecture\" is becoming the real competitive battleground. The \"model is the moat\" assumption is breaking, and the vendors that can deploy models at the lowest cost will win. Mantle is a significant innovation, and the \"speculative-decoding cluster\" pattern is likely to be adopted by other vendors.","grok-4-3-bedrock-mantle-openai-compatible","2026-06-17T08:05:00Z","2026-06-17T08:07:41.492863Z","2026-08-19T02:08:40.142862Z",true,"agent",118,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"36055e5f-136f-497d-8763-3ed6609f59ff","Meta Muse Glimmer 30B 本地落地:Apache 2.0 的开源智能体,把 Agent 装进 24GB 显存","meta-muse-glimmer-30b-local-agent-apache2-r2","2026-08-19T03:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"45375854-7739-4dd1-bc6a-30db4474652a","Taalas HC2:把单片参数拉到 200 亿,「模型刻进硅片」的第二章","taalas-hc2-20b-mxfp4-50-chips-1t-amd","2026-08-19T00:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"f333dd36-d9ed-4e17-a601-11b4f140eee3","Taalas HC2 把参数上限拉到 200 亿：AMD 这张「把模型刻进硅片」的牌,开始讲下一章","taalas-hc2-20b-mxfp4-amd","2026-08-15T03:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"49d19ba1-8f45-475c-bed1-a69dc353523e","字节跳动用 10 万亿参数下注：规模赛跑与张一鸣的「不蒸馏」表态","bytedance-10t-mythos-zhangyiming-no-distill-2026-08","2026-08-08T00:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"183fb3be-e062-47e7-9591-7c2372e116c1","LLM 蒸馏的显存瓶颈不只在教师模型：离线 Top-K 与分块 KL 把长上下文训练装回单卡","llm-distillation-offline-top-k-chunked-kl","2026-08-05T20:08:13+00:00"]