Google DeepMind released Gemini 3.5 Flash, the first "small" Gemini model with native Computer Use capability. The standout: 78.4 on the OSWorld benchmark — a score that puts it in the same range as dedicated Computer Use agents, but at a fraction of the cost.
The technical path: Gemini 3.5 Flash uses a "screenshot-in, action-out" interface. The model receives a screenshot of the current screen, plus a natural-language task, and outputs a structured action (mouse move, click, type, scroll). The model is trained on a large corpus of human-computer interaction data, including web browsing, desktop app usage, and mobile interaction.
The benchmark result: on OSWorld (a real-world desktop task benchmark), Gemini 3.5 Flash scores 78.4 — significantly above Claude Computer Use (71.2) and the open-source ShowUI (52.3). The score is particularly impressive because Flash is the "small" tier — the full Gemini 3.5 Pro hits 86.7.
The bigger takeaway: "Computer Use" is moving from "frontier model only" to "Flash tier." This is significant because most production use cases (customer service automation, office automation, RPA) do not need the strongest model — they need a fast, cheap, reliable one. Gemini 3.5 Flash's 78.4 OSWorld score at Flash-tier pricing makes Computer Use economically viable for the first time.
For the industry, this signals that "Computer Use" is becoming a commodity capability. The next round of competition will be in "Computer Use reliability" — i.e., how well the model handles edge cases (pop-ups, dynamic UIs, multi-window tasks). The current 78.4 score is impressive but still has a 22% error rate that needs to come down.