After a stretch in which OpenAI and Anthropic set the pace at the top of the AI market, Google is back with Gemini 4 Argon. The company is positioning its new frontier model as a specialist in long, multi-step work such as software engineering, enterprise knowledge tasks and cybersecurity. The launch also comes after reports that the Gemini 3 generation fell short of internal expectations, particularly on coding, so Google has some ground to recover.
The most significant technical change is the output ceiling. Argon can produce up to 1 million tokens in a single response, up from 64,000. In practice, that means the model can carry a large code migration or an extended reasoning chain through to completion in one pass, rather than having to split the job into separate runs that each lose some context along the way. For agent-style workloads, that matters more than most benchmark gains.
Google’s own figures are strong. Argon scores 77.9% on DeepSWE v1.1, a long-horizon software engineering test, 51.3% on Zapier’s AutomationBench and 91.7% on the LVBench long-video benchmark. Google also says Argon leads the Vals Index of finance, legal, tax and coding tasks, and that it beats GPT-6 Astra and Claude Opus 5.5 on several of the comparisons it published. As with any launch, these are benchmarks the vendor chose to highlight. Independent testing over the coming weeks will show whether the lead holds up in everyday use.
Security is the central pitch. Google says Argon can find, confirm and patch vulnerabilities on its own, and it reports a 68% score on CWE-bench v1, tied for first place. The Wiz security team is already using the model through its Scan for Good program. Internally, Google says thousands of employees have used Argon for work ranging from quantum algorithm tuning to moving C and C++ codebases to Rust.
Those security capabilities also explain the cautious rollout. Argon is initially going only to vetted cybersecurity defenders through Google’s Fairwind Program while safety testing continues. This follows an industry pattern of giving cyber-capable models to defenders before the general public. Paid API customers and Google AI Ultra subscribers are next in line, but Google has not given a firm date for wider access.
The pricing is aggressive, but only for a while. The introductory rate is $2 per million input tokens and $10 per million output tokens, with a 95% discount on cached input. After that, prices double to $4 and $20, and Google hasn’t said when the switch happens. At the standard rate, a single response that uses the full million-token output would cost about $20. The long output window is a real capability, but teams will need to watch what it costs them.
