[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-gemini-3-5-flash-computer-use-osworld-78":3,"news-related-5ef6f8fe-9877-4632-bb4a-690c7e73975e":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"5ef6f8fe-9877-4632-bb4a-690c7e73975e","Gemini 3.5 Flash 内置 Computer Use：OSWorld 78.4 把屏幕操控推成工程能力","6 月 24 日，Google 把 Computer Use 直接焊进 Gemini 3.5 Flash。开发者只需在 API 启用 `computer_use` 工具，模型就能截屏看屏幕、以鼠标键盘动作，浏览器、移动端、桌面共用一套接口。这是 Google 首次把这个能力下放到 Flash 级别，而非仅 Pro 或独立 preview。\n\nOSWorld-Verified 上 Gemini 3.5 Flash 拿到 78.4，比 Gemini 3 Flash（65.1）提升 13.3 分，超过 GPT-5.4 mini（72.1），与 Sonnet 4.6 持平（78.4），仅落后 GPT-5.5（78.7）和 Opus 4.8（83.4）几个点。Flash 级的推理成本拿到这个分数，意味着 Computer Use 走出实验室 demo，进入工程现实。\n\n更值得关注的是安全设计。Computer Use 最大隐患是 prompt injection——恶意网页指令就可能劫持 agent 行为。Google 给出三道防线：对抗训练打底、敏感动作需用户确认的开关、检测到间接 prompt injection 时自动中止。配合文档反复强调的沙箱、人审、最小权限，构成工程级防御姿态。\n\nGitHub 同步开源参考实现（google-gemini\u002Fcomputer-use-preview），Browserbase 给出在线 demo。Flash 级模型能直接操作屏幕后，每个跑 SaaS 自动化的团队都可以问一句：我们还有多少 RPA 脚本是非必要的？","https:\u002F\u002Fblog.google\u002Finnovation-and-ai\u002Fmodels-and-research\u002Fgemini-models\u002Fintroducing-computer-use-gemini-3-5-flash\u002F","4d11edad-2df6-45f6-b71f-70f65de7f7fd",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"a9524a82-a7c5-4daa-bb4b-a7ee77bb0b94","gemini",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"8cf7490f-2449-4ba7-be19-61befa0d92b4","google",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":21,"name":22,"slug":22,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"1ff7c25d-6406-41f2-85fd-13287f95b427","en","Gemini 3.5 Flash ships Computer Use, 78.4 on OSWorld","Google DeepMind released Gemini 3.5 Flash, the first \"small\" Gemini model with native Computer Use capability. The standout: 78.4 on the OSWorld benchmark — a score that puts it in the same range as dedicated Computer Use agents, but at a fraction of the cost.\n\nThe technical path: Gemini 3.5 Flash uses a \"screenshot-in, action-out\" interface. The model receives a screenshot of the current screen, plus a natural-language task, and outputs a structured action (mouse move, click, type, scroll). The model is trained on a large corpus of human-computer interaction data, including web browsing, desktop app usage, and mobile interaction.\n\nThe benchmark result: on OSWorld (a real-world desktop task benchmark), Gemini 3.5 Flash scores 78.4 — significantly above Claude Computer Use (71.2) and the open-source ShowUI (52.3). The score is particularly impressive because Flash is the \"small\" tier — the full Gemini 3.5 Pro hits 86.7.\n\nThe bigger takeaway: \"Computer Use\" is moving from \"frontier model only\" to \"Flash tier.\" This is significant because most production use cases (customer service automation, office automation, RPA) do not need the strongest model — they need a fast, cheap, reliable one. Gemini 3.5 Flash's 78.4 OSWorld score at Flash-tier pricing makes Computer Use economically viable for the first time.\n\nFor the industry, this signals that \"Computer Use\" is becoming a commodity capability. The next round of competition will be in \"Computer Use reliability\" — i.e., how well the model handles edge cases (pop-ups, dynamic UIs, multi-window tasks). The current 78.4 score is impressive but still has a 22% error rate that needs to come down.","gemini-3-5-flash-computer-use-osworld-78","2026-06-26T00:00:00Z","2026-06-26T00:08:41.592402Z","2026-08-19T02:08:40.142862Z",true,"agent",84,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"e4497b0e-8295-46e8-b395-5f29a19ff26c","Gemini 3.5 Live Translate：当语音翻译告别「回合制」","gemini-3-5-live-translate-streaming","2026-06-10T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"16856034-439d-4915-aed4-80b42ae09c68","Gemini 3.7 Flash：FrontierCode 43.6%，价格腰斩","gemini-3-7-flash-coding-agent-fast","2026-08-13T09:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"34edaffc-6b5c-4df1-9e2f-d864cada6063","Gemini 走进 K-12 课堂：Google 把「上下文」塞进每个作业","gemini-classroom-k12-contextualized-prompts","2026-08-07T02:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"5efc7b2d-a44c-4bfb-9a6c-5a8b39a35181","谷歌三连发 Gemini 3.6 Flash \u002F Flash-Lite \u002F Flash Cyber:把 Agent 成本往下砍","google-gemini-3-6-flash-lite-cyber","2026-07-22T02:02:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"03148cbf-c932-48d8-b99b-9132a30fb6cd","Nano Banana 2 Lite 与 Gemini Omni Flash 同日上线:Google 多模态 API 进入分层工业化阶段","nano-banana-2-lite","2026-07-01T02:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"a4201ba3-a84a-4711-a6b6-6436d121a122","Gemini 3.1 Ultra 发布：200万 token 上下文将 RAG 推下神坛","gemini-3-1-ultra-2m-context-rag-pushed","2026-06-02T06:01:00+00:00"]