[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-thinking-machines-interaction-models-micro-turn":3,"topics-all":36,"news-related-fb260024-6d7c-465f-82f4-16d351cdfaa0":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"fb260024-6d7c-465f-82f4-16d351cdfaa0","Thinking Machines发布交互模型：让AI从\"问答\"走向\"协作\"","5月11日，AI创业公司Thinking Machines发布了交互模型（Interaction Models）研究预览版，这是AI领域一次值得关注的基础架构层面的尝试。该公司由前OpenAI CTO Mira Murati与联合创始人John Schulman共同创办，聚焦于将多模态实时交互能力直接融入模型架构本身，而非依赖外部套娃式的工程方案。\n\n当前主流模型的交互逻辑本质上是单线程的：必须等用户说完，模型才开始处理；模型生成期间，整个感知就冻结了。这种turn-based范式对于快速问答足够用，但一旦涉及需要持续协作的真实工作流程——比如调试代码、审视文档、共同分析问题——交互带宽就成了瓶颈。用户被迫凑合AI的节奏，而不是AI适应人的节奏。\n\nThinking Machines的解法是从头训练具备原生交互能力的模型。核心技术是时间对齐的micro-turn机制：模型将连续的音频、视频、文本输入切分为极细粒度的微轮次，而非传统模型的整句\u002F整段交替。这种设计让模型能感知用户正在说话（而不是说完）、能同时执行工具调用和内容生成、也能在生成过程中主动打断或补充。\n\n从benchmark结果看，交互模型在保持智能力的同时，响应延迟显著降低。更重要的是，它解锁了以往需要额外工程才能实现的能力：实时同声传译、口语打断修正、边说边搜索并把结果织入对话。模型不再只是回答问题，而是真的在协作。\n\n这背后的理念值得注意：用scaling来同时提升智能水平和交互质量，而不是分别优化。这与当前主流做法——在已有语言模型外面套一个实时交互层——有本质区别。如果这条路work，它意味着未来更强大的AI本身就会是更好的协作伙伴；反之，如果交互必须靠外挂，真协作的能力就会永远受限于接口工程。","https:\u002F\u002Fthinkingmachines.ai\u002Fblog\u002Finteraction-models\u002F","95239a8d-29f2-486d-84ca-28174cab2405",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"9d916545-b7f4-4b4f-bed8-2796776d86c4","en","Thinking Machines debuts interactive models beyond Q&A","On May 11, AI startup Thinking Machines released a research preview of Interaction Models — a noteworthy foundational architecture-level attempt in the AI field. The company was co-founded by former OpenAI CTO Mira Murati and co-founder John Schulman, focused on integrating multimodal real-time interaction capability directly into the model architecture itself, rather than relying on external matryoshka-style engineering solutions.\n\nThe interaction logic of current mainstream models is essentially single-threaded: it must wait for the user to finish speaking before the model begins processing; during the model's generation, the entire perception freezes. This turn-based paradigm is sufficient for quick Q&A, but once it involves real work flows requiring continuous collaboration — such as debugging code, reviewing documents, jointly analyzing problems — interaction bandwidth becomes a bottleneck. Users are forced to adapt to the AI's pace, rather than the AI adapting to the human's.\n\nThinking Machines' solution is to train from scratch a model with native interaction capability. The core technology is a time-aligned micro-turn mechanism: the model slices continuous audio, video, and text input into extremely fine-grained micro-turns, rather than the traditional model's whole-sentence\u002Fwhole-segment alternation. This design lets the model perceive the user speaking (rather than having finished), perform tool calls and content generation simultaneously, and also actively interrupt or supplement during generation.\n\nFrom the benchmark results, the interaction model maintains intelligence capability while significantly reducing response latency. More importantly, it unlocks capabilities that previously required additional engineering to achieve: real-time simultaneous interpretation, spoken interrupt correction, search-while-speaking with results woven into the conversation. The model is no longer just answering questions, but truly collaborating.\n\nThe underlying philosophy is worth attention: using scaling to simultaneously improve intelligence level and interaction quality, rather than optimizing them separately. This differs fundamentally from the current mainstream approach — wrapping a real-time interaction layer around an existing language model. If this path works, it means future more powerful AI will itself be a better collaboration partner; conversely, if interaction must rely on external add-ons, true collaboration capability will forever be limited by interface engineering.","thinking-machines-interaction-models-micro-turn","2026-05-12T10:00:00Z","2026-05-12T10:08:42.432368Z","2026-08-19T02:08:40.142862Z",true,"agent",145,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"b7eb05aa-bc46-4749-a57b-47fbd19644e3","企业AI架构新趋势：从单模型到多模型编排","multi-model-orchestration-enterprise-ai","2026-05-18T01:05:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"9f42d03a-ef19-4fa7-aaf9-f33c481f19fa","大模型正在重塑搜索引擎：Gemini与AI Overview的实测研究","gemini-ai-overview-search-sigir-2026-11500-queries","2026-05-01T22:05:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"12c67d52-17a2-4df5-8386-35d18ffd221a","JEPA-Anything:一套预测框架打通七个领域,湿实验也给了背书","jepa-anything-orthogonal-predictive-factorization","2026-09-19T23:10:37+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"ce70384a-990b-4994-bfb6-27775be45661","TensorRT Edge-LLM 0.10.0：边端第一个统一的 C++ 多模态推理栈","tensorrt-edge-llm-0-10-multimodal-runtime","2026-08-23T00:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"183fb3be-e062-47e7-9591-7c2372e116c1","LLM 蒸馏的显存瓶颈不只在教师模型：离线 Top-K 与分块 KL 把长上下文训练装回单卡","llm-distillation-offline-top-k-chunked-kl","2026-08-05T20:08:13+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"217f417d-1b9c-475b-99f4-e21e7c909711","MHAR 把 Transformer 残差流从「单车道」拆成 H 条独立路由:子空间第一次有权自己挑历史层","multi-head-attention-residuals-mhar","2026-08-01T07:30:00+00:00"]