[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-kog-laneformer-2b":3,"news-related-93fb05a4-79cd-4abf-b263-c7d1910dbea7":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"93fb05a4-79cd-4abf-b263-c7d1910dbea7","Kog Laneformer 2B 开源:把推理引擎焊进 Transformer 架构,2B 模型单请求解码跑到 3000 tok\u002Fs","巴黎 AI 基础设施初创 Kog 在 Hugging Face 开源 Laneformer 2B(2.3B 参数代码模型),采用 Delayed Tensor Parallelism + 8 通道架构,把 Transformer 架构本身为推理引擎让路,在 8×MI300X 上跑到单请求 3000 tok\u002Fs、8×H200 上 2100 tok\u002Fs,HumanEval+ 45.1%、MBPP+ 51.6%,权重以 Apache 2.0 发布。","https:\u002F\u002Fhuggingface.co\u002Fblog\u002Fkogai\u002Fkog-laneformer-2b-the-latency-first-model","24d5c6c5-6573-4180-a1fd-f1459842d1af",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"11d54952-8fb4-42b5-b7fc-f4d1928f65b3","en","Kog Laneformer 2B: 3000 tok\u002Fs decode with a built-in engine","Parisian AI infrastructure startup Kog open-sources Laneformer 2B (2.3B-parameter code model) on Hugging Face, using Delayed Tensor Parallelism + 8-channel architecture, giving the Transformer architecture itself over to the inference engine, running 3000 tok\u002Fs single-request on 8×MI300X, 2100 tok\u002Fs on 8×H200, HumanEval+ 45.1%, MBPP+ 51.6%, weights released under Apache 2.0.","kog-laneformer-2b","2026-06-24T14:00:00Z","2026-07-03T16:07:21.600243Z","2026-08-19T02:08:40.142862Z",true,"agent",84,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"491f4904-c854-4925-b3e3-e34b8afd5e50","KDA+MLA 混合栈下沉到 1.3B 激活:Ling-3.0-tiny 把 MoE 端侧化,INT4 跑出 115 tok\u002Fs","ling-3-tiny-kda-mla-edge-deployment","2026-08-18T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"3d8b9b1a-e038-466f-9b6b-304f911e35a7","Kimi K3 开源三件套 MoonEP\u002FFlashKDA\u002FAgentEnv:Moonshot 把 2.8T MoE 训练栈完整交底","kimi-k3-moonep-flashkda-agentenv","2026-07-28T04:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"4d436945-18e9-4d69-a4c8-c1e3e975ab33","MiniMax M3发布：稀疏注意力打通百万token上下文，开源模型编程能力逼近闭源前沿","MiniMax-m3-sparse-attn-million-token-msa","2026-06-04T01:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"808c9e2b-8852-4996-9bbf-48668e259da6","开源与闭源AI模型的对决：2026年四月的技术格局","open-vs-closed-source-llm-divide-april-2026","2026-04-25T22:07:32+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"f6e4aab0-7693-4c2c-bb66-c1641fc2cc3e","Ox Alpha 谜底揭晓:智谱 GLM-5.3-Flash,MIT 开源 320B MoE","ox-alpha-glm-5-3-flash-reveal","2026-08-27T13:30:00+00:00"]