
Google bakes computer control directly into Gemini 3.5 Flash, letting the model see and operate your screen
Quick Answer
Google has embedded 'Computer Use' functionality in Gemini 3.5 Flash, enabling it to autonomously control devices.
Quick Take
Scoring 78.4 on the OSWorld benchmark, it rivals GPT-5.5, allowing developers to create agents for software testing and office automation.
Key Points
- Gemini 3.5 Flash can operate computers, browsers, and mobile devices autonomously.
- Achieved a score of 78.4 on the OSWorld benchmark, comparable to GPT-5.5.
- Developers can utilize the Gemini API for software testing and office automation.
- This integration enhances user interaction with AI in everyday tasks.
- Potential applications include automated testing and streamlined office workflows.
📖 Reader Mode
~1 min readGoogle has integrated "Computer Use" directly into Gemini 3.5 Flash. The model can now see, understand, and interact with computers, browsers, and mobile devices on its own. Previously, this was only available as a separate Gemini 2.5 model. Combined with existing tools like function calls, Search, and Maps, developers can now build agents that work across browser, mobile, and desktop environments for tasks like software testing or office automation.
On the OSWorld benchmark, Gemini 3.5 Flash scores 78.4, beating Gemini 3 Flash (65.1) and GPT-5.4 mini (72.1). GPT-5.5 sits just ahead at 78.7, while Anthropic's Opus 4.8 leads at 83.4. Sonnet 4.6 also hits 78.4, and Gemini 3.1 Pro lands at 76.2.
To guard against prompt injection attacks, Google uses adversarial training and two optional enterprise safeguards. One requires user confirmation for sensitive or irreversible actions, while the other automatically stops tasks when it detects indirect prompt injections. Google also recommends sandboxing, human oversight, and strict access controls, with more details in its best practices documentation. The feature is available through the Gemini API and the Gemini Enterprise Agent Platform. A Browserbase demo and a GitHub reference implementation are also available.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

