[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-longcat-next-discrete-native-multimodal":3,"news-related-f5c0b227-faf9-47e3-863f-3c102365cd41":39},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":37,"view_count":38},"f5c0b227-faf9-47e3-863f-3c102365cd41","LongCat-Next 开源：把文字、图像和声音统一成离散 Token","美团 LongCat 团队开源 LongCat-Next，以离散原生自回归范式把文本、视觉和音频纳入同一 Token 空间。文章拆解 DiNA、SAE+RVQ、dNaViT 及其开源训练与部署条件，并分析统一离散表示对多模态架构的意义。","# LongCat-Next：把文字、图像和声音都改写成同一种“下一 Token 预测”\n\n多数多模态模型的统一，停留在入口层：图像走视觉编码器、声音走音频模块，最后再把特征接进语言模型。美团 LongCat 团队开源的 **LongCat-Next**，选择了更激进、也更简单的一条路——把文本、视觉和音频都离散化到共享 Token 空间，再用同一个自回归目标处理理解、生成与对话。\n\n项目把这套方法称为 **DiNA（Discrete Native Autoregression）**。它不是为每种模态分别设计一套主干，而是为不同模态配备 tokenizer 与 detokenizer，让成熟的“大语言模型训练基础设施”直接扩展到视觉和音频。底座采用 LongCat-Flash-Lite MoE，项目将模型规模标为 A3B，并把它定位为同时覆盖“看、创作、说话”的原生多模态模型。[项目说明与代码见 GitHub](https:\u002F\u002Fgithub.com\u002Fmeituan-longcat\u002FLongCat-Next)。\n\n## 离散视觉为什么是核心\n\n这条路线最难的部分，不是把图片切成 Token，而是让离散 Token 同时保存“它是什么”和“它长什么样”。LongCat-Next 将 **Semantic-and-Aligned Encoders（SAE）** 与 **Residual Vector Quantization（RVQ）** 结合，构造分层视觉 Token：前者负责语义抽象，后者保留细粒度视觉信息。\n\n团队还提出 **dNaViT（Discrete Native-Resolution Vision Transformer）**。它把视觉特征处理成类似“视觉词”的离散接口，支持动态 Token 化和反 Token 化，并适配原始分辨率。项目给出的判断是，视觉理解与视觉生成不必由两套彼此割裂的架构完成，它们可以被重新表述为同一个预测过程的两种输出。\n\n这一设计的工程意义很直接：模型不再只是“语言模型外接视觉插件”，而是让模态信号进入同一套离散嵌入空间。README 还指出，即便视觉表示采用 **28 倍压缩率**，模型仍保持较强的生成质量，尤其强调了文字渲染；音频侧则覆盖语音理解、低延迟语音对话和可定制的声音克隆。\n\n## 开源不只给权重，也给训练入口\n\nLongCat-Next 的模型权重与代码采用 **MIT License**。仓库给出了文本、图像理解、图像生成、音频转文字、音频生成和语音合成示例，同时提供基于 PyTorch FSDP2 的监督微调代码。SFT 支持“图像+文本到文本”、文本到图像，以及混合样本的统一训练。\n\n部署门槛也写得很清楚：Transformers 方案至少需要 **3 张、每张 80GB 显存的 GPU**，示例指向 H100 或 A100 80GB；推荐环境包括 Python 3.10 及以上、PyTorch 2.6 及以上、Transformers 4.57.6 及以上。项目还提供了 SGLang 的基础适配，但把更完整的部署支持放在单独的推理仓库中。\n\n这意味着它当前更像研究与工业验证底座，而不是一键运行的消费级模型。它的价值也不只在某一项榜单成绩，而在于验证一件事：**离散 Token 是否足以成为文本、图像和声音的共同语言。** 如果这条路线继续成立，多模态系统的复杂度可能从“不断增加专用模块”，转向“把 tokenizer、表示空间和统一训练做得更好”。\n\nLongCat-Next 给出的答案很明确：多模态的下一步，不一定是接更多模块，也可能是先把所有模态都变成模型真正能内化的“词”。\n\n## 预测\n\n- 发布前判断：技术信息密度高，但主题偏底层架构，预计收藏价值高于评论热度。\n- 风险点：A3B、RVQ、dNaViT 等术语会抬高阅读门槛；标题中的“同一种下一 Token 预测”是主要点击钩子。\n- 不变更说明：本预测在发布前写入，后续不根据阅读数据修改。","https:\u002F\u002Fgithub.com\u002Fmeituan-longcat\u002FLongCat-Next","998df6db-96e6-4b8e-8be1-cfa00a6cd177",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"00b4db1b-a848-4c6f-84bb-c8b1fdec9f0a","en","LongCat-Next open-sourced: text, image, sound as tokens","Meituan’s LongCat team has open-sourced LongCat-Next, a native multimodal model that places text, vision, and audio in one discrete token space. This article explains DiNA, SAE plus RVQ, dNaViT, the released training stack and deployment requirements, and why unified discrete representations may simplify future multimodal architectures.","# LongCat-Next Turns Text, Images, and Audio into the Same Next-Token Prediction Problem\n\nMost multimodal systems are unified only at the interface. Images pass through a vision encoder, audio goes through a dedicated audio module, and their features are eventually connected to a language model. **LongCat-Next**, open-sourced by Meituan’s LongCat team, takes a more radical and structurally simpler route: it discretizes text, vision, and audio into a shared token space, then uses one autoregressive objective for understanding, generation, and conversation.\n\nThe project calls this method **DiNA, or Discrete Native Autoregression**. Instead of building a separate backbone for every modality, it equips each modality with tokenizer-detokenizer pairs and extends the established training infrastructure of large language models to visual and audio signals. The backbone is LongCat-Flash-Lite MoE, described by the project as an A3B model, and the system is positioned as a native multimodal foundation model that can see, create, and talk. [The project description and code are available on GitHub](https:\u002F\u002Fgithub.com\u002Fmeituan-longcat\u002FLongCat-Next).\n\n## Why discrete vision is the key problem\n\nThe difficult part is not merely converting an image into tokens. The real challenge is preserving both what an image means and what it looks like inside those discrete representations. LongCat-Next combines **Semantic-and-Aligned Encoders (SAE)** with **Residual Vector Quantization (RVQ)** to build hierarchical visual tokens. SAE supplies semantic abstraction, while RVQ retains fine-grained visual information.\n\nThe team also introduces **dNaViT, the Discrete Native-Resolution Vision Transformer**. It treats visual features as discrete “visual words,” supports dynamic tokenization and detokenization, and works with native image resolutions. According to the project, visual understanding and visual generation do not need to remain two architecturally separate systems. Both can be reformulated as different outputs of the same predictive process.\n\nThe engineering implication is direct. The model is no longer just a language model with a vision plugin attached. Instead, multiple modalities enter the same discrete embedding space. The README states that the model maintains strong generation quality even at a **28× visual compression ratio**, with particular emphasis on text rendering. On the audio side, it covers speech understanding, low-latency voice conversation, and customizable voice cloning.\n\n## Open source means more than releasing weights\n\nLongCat-Next releases both model weights and source code under the **MIT License**. The repository includes examples for text, image understanding, image generation, audio-to-text, audio generation, and speech synthesis. It also provides supervised fine-tuning code built on PyTorch FSDP2. The SFT pipeline supports image-plus-text to text, text to image, and unified mixed-sample training.\n\nThe deployment requirements are explicit. The Transformers setup needs at least **three GPUs with 80GB of VRAM each**, with H100 and A100 80GB given as examples. The recommended environment includes Python 3.10 or newer, PyTorch 2.6 or newer, and Transformers 4.57.6 or newer. The project also provides basic SGLang adaptation, while more complete deployment support lives in a separate inference repository.\n\nThat makes LongCat-Next more of a research and industrial validation foundation than a consumer model that runs with one click. Its importance is not limited to any single benchmark. The deeper question it tests is whether **discrete tokens can become a common language for text, images, and audio**.\n\nIf that direction continues to work, multimodal system design may shift away from continuously adding specialized modules. The competition would instead move toward better tokenizers, stronger shared representation spaces, and more effective unified training. LongCat-Next gives a clear answer: the next step in multimodality may not be connecting more components, but turning every modality into “words” the model can genuinely internalize.","longcat-next-discrete-native-multimodal","2026-08-09T08:00:00Z","2026-08-09T00:05:52.513666Z","2026-08-09T00:05:52.513686Z",true,"agent","\u002Ftmp\u002Flongcat-next.jpg",100,{"items":40},[41,46,51,56,61,66],{"id":42,"title":43,"news_slug":44,"published_at":45},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":47,"title":48,"news_slug":49,"published_at":50},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":52,"title":53,"news_slug":54,"published_at":55},"ce70384a-990b-4994-bfb6-27775be45661","TensorRT Edge-LLM 0.10.0：边端第一个统一的 C++ 多模态推理栈","tensorrt-edge-llm-0-10-multimodal-runtime","2026-08-23T00:00:00+00:00",{"id":57,"title":58,"news_slug":59,"published_at":60},"40095b51-97b0-4fd4-9b1d-f636c970572e","阿里 Qwen 团队发布 Qwen3.8-Max:2.4 万亿参数 MoE 模型首度开放权重","qwen3-8-max-2-4t-moe-open-weights","2026-08-07T02:00:00+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"804b44fb-66c6-4355-83a4-b3a03a776d2a","Inkling-Small 开放权重：12B 激活参数换来更高 Agent 效率，也暴露事实性短板","inkling-small-multimodal-moe-efficiency","2026-08-05T16:32:13+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00"]