On July 20, Unitree released UnifoLM-OminiA-0.3, folding four capabilities — visual recognition, semantic understanding, fine-grained manipulation and device control — into one end-to-end large model, and running a full perception-to-action closed loop on its own humanoid robot G1. The demo videos show it putting a pillow on a sofa, recognizing a pill box's color and quantity and taking out a specific one, adjusting a hospital bed's height and responding in real time to a "stop" command. Spanning the four layers of "dialogue — recognition — planning — execution" — which used to require coordinated vision, voice, decision, and control models — is now done in a single pass. Technically, "all-modality" has finally moved from PowerPoint to a demo. The model supports joint input of vision, voice and action instructions, and directly outputs robot motion control commands, eliminating intermediate representation transitions in the traditional pipeline. In strongly interference-prone, cross-task environments like elderly care and the home, one fewer modality transition means one less place for error to accumulate, and anti-interference capability naturally improves. The industry implications are three layers: pulling "physical AI" from papers to a closed loop in hardware; embodied intelligence finally has a demo-evaluable baseline; this is a rare model+hardware tightly-coupled end-to-end solution domestically, not just APIs to wire up. Of course, it's still at demo level — long-tail scenarios and generalization to unseen objects all need large-scale deployment to verify. But in the second half of 2026, embodied-intelligence players have moved from "I can run" to "I can run stably", and Unitree's step pushes the bar up a notch.