Sensetime's "Riri Xin" (SenseNova) brand released the next-generation base model SenseNova-U1 Pro, with the biggest innovation: a natively unified "understanding + generation + action" architecture. The model is scheduled for closed beta in July 2026.
The technical details: SenseNova-U1 Pro is a 200B-parameter MoE model with three coupled heads: (1) an understanding head (text + image + video input → semantic representation); (2) a generation head (semantic representation → text + image + video output); (3) an action head (semantic representation → tool call / robot action / UI interaction). The three heads share the same backbone but are trained with a multi-task loss that encourages them to share representations.
The unification value: traditional AI systems have separate models for understanding, generation, and action. SenseNova-U1 Pro's unified architecture means a single model can:
- Understand a user's request (e.g., "summarize this video")
- Generate a response (e.g., a text summary + a thumbnail image)
- Take an action (e.g., post the summary to social media, save the thumbnail to a folder)
The result is a "one-model-fits-all" AI agent that can handle complex, multi-step tasks without the coordination overhead of multi-model systems.
The bigger takeaway: the "unified base model" is becoming the new battleground. The traditional "one model, one task" approach is being challenged by "one model, many tasks" architectures. SenseNova-U1 Pro is Sensetime's bet that the future of AI is "one unified model" rather than "many specialized models."
For the industry, this signals that the next round of competition is in "model unification" — i.e., who can build the most general-purpose base model without sacrificing task-specific quality.