Black Forest Labs pushed FLUX 3 to Early Access on July 23. This isn't a routine version bump — it's an attempt to squeeze image, video, audio, and even robot action prediction into a single flow-matching backbone. The technical foundation is BFL's self-developed Self-Flow — a method that aligns multimodal generation and understanding within the same architecture, with visible advantages over pure Flow Matching in per-modality generation error and action-task success rates. Video-side capabilities are opened first: native audio, up to 20 seconds, four input types (text / image / keyframe / reference video), cross-shot consistency, and multi-language dialogue all work; the initial version in BFL's self-evaluation beats Runway Gen-4.5 (77%) and Luma Ray 3.2 (93%), with win rates of 52%–60% against Kling v3 Pro, Gemini Omni Flash, and Seedance 2.0. Images and open-source weights enter the next release window, while the robot-action branch is a collaboration with mimic robotics and has been piloted in Audi's production environment. The more realistic judgment: FLUX 3 is still just a "staged checkpoint" on the multimodal-flow-model route — unifying perception, action, and language prediction is BFL's next stop.