What happened?

On 30 September 2026 Google announced Gemini 4 Argon. Koray Kavukcuoglu wrote the post. Google posted it the same evening. It is the first Gemini above the Flash line in months, and a spokesperson told Reuters it is larger than the old Pro models. The Gemini 3.5 Pro that was supposed to land in June is not coming.

Argon is built for work that takes a while. A real code change. A finance or legal file. A security review that has to finish, not just sound finished. The part worth caring about is how long it can keep going. Google says the output limit is 1 million tokens, up from 64,000. That is output, not a new context trick. Context was already long.

You cannot call it yet. A set of trusted cyber defenders in the Fairwind program have it. Everyone else waits. Google says the next step is paid API customers and Google AI Ultra subscribers, and it did not give a date. Reuters checked: there is no public timetable. The model is also in the US government's voluntary pre-release review. It is not on the public Gemini model list, the pricing page, or the changelog.

The intro price is $2 per million input tokens and $10 per million output tokens. Cached input is 95% off that input price, which is $0.10 per million if you take the footnote literally. After the intro, the sticker becomes $4 and $20. Google did not say when the intro ends. Artificial Analysis was told the discount runs at least a month.

One wrinkle on the million tokens. Vals, which actually ran the model, lists a 1 million token context and a 262,144 output cap, and that is the cap they used. Artificial Analysis tested a separate switch, Long Decode Continuation, that pauses a long answer and resumes it. They did reach a million output tokens that way. So the million is real as a stitched run. It is not what the public eval harness got in one call. There is still no API spec that settles it.

Why this is interesting

  • It is good at the office loop - On the Vals Index, a mix of finance, coding, legal, and tax weighted by US GDP, Argon is first: 68.90% (plus or minus 0.97), $15.68 a test. Claude Sonnet 5.5 is next at 67.04% and $21.34. Claude Opus 5.5 is 66.97% and $32.14. Vals priced that run at the later sticker, $4 / $20, not the intro. On Zapier's AutomationBench it is also first: 51.29% at high effort, 50.08% at medium. Sonnet 5.5 is third at 44.75%. Zapier ranks it at list price, $1.70 a task, and notes the promo price is $0.85. Finance is the exception on that board. Sonnet 5.5 leads that slice at 50%. Argon at medium effort is just behind, at 49.17%.

  • The cheap sticker writes a lot - Artificial Analysis puts Argon level with GPT-6 Astra on their Intelligence Index. Both score 53. At the intro price that task costs $1.99, against $3.26 for Astra. The saving is the rate, not fewer tokens. Argon wrote about 62,000 output tokens a task. Astra wrote about 27,000. When the intro ends, the same task is $3.98, a bit above Astra. A low rate and a short bill are different things.

  • It would rather admit it does not know - On AA-Omniscience, Argon's hallucination rate is 15%, against 51% for Astra. Accuracy goes the other way: 50% for Argon, 63% for Astra. It guesses less. It also gets fewer answers right. Both belong in the same sentence.

  • Legal is a selected win, not the whole board - Google calls Argon leading on Harvey's Legal Agent Benchmark. The full Vals board says fifth: 19.58% (plus or minus 3.31). Muse Spark 1.2 is first at 25.42%, and three other Muse setups sit above Argon. Argon does beat the models Google chose to print next to it. It does not beat the board. On Finance Agent v2, which is a different test, it is first at 65.40%, ahead of Gemini 3.8 Flash at 61.44%.

  • Coding is mixed, and one big number is theirs - Vals has Argon second on Vibe Code Bench at 91.91%, just behind Sonnet 5.5 at 92.39%, and fifth on Terminal-Bench 4.0 at 57.58%. Reuters noted that in Google's own release Argon trailed on two of the four coding comparisons. DeepSWE v1.1 at 77.9% is Google's self-run with mini-swe-agent. It is not on Datacurve's public board. Treat 77.9% as their number.

  • Cyber defense is why access is gated - On CWE-bench v1, a private set of 120 audit-and-patch tasks, Argon ties Grok 4.7 and GPT-6 Astra at 68% on the first try. Give it four tries and it reaches 75%, behind Grok at 81%. The run cost about $6.63, against $2.75 for Grok and $0.79 for Opus 5.5. Collinear priced that Argon run at $2 input, $0.20 cached, and $10 output. That cached rate is double the 95% off in Google's footnote. Keep both. Google also says Argon found a critical exposure of personal data in hospital software that earlier models missed. They did not name the product, the vendor, or a case number. Wiz says Scan for Good is using Argon next to Gemini 3.8 Flash Cyber. That confirms use. It does not confirm that finding.

Inside Google, the blog says Argon agents freed over 300 TiB of memory once the change rolled out, with an estimate of 500 TiB to 1 PiB if the rest lands. On libgav1, their open video decoder, agents replaced 32,000 lines of hand-tuned code and landed a Rust build 2.7x faster than the existing Rust port, with the same video out. That is not faster than the optimized C++. Those rewrites are still in review before production. A useful picture of what they are trying. Not a receipt.

What it is not

Not a model you can put in a product this week. Not first on every legal or coding row. Not a 1 million token answer in a single eval call. Not a published write-up of a hospital bug. The Fairwind partner count is the program, not a headcount of who has Argon.

Bottom line

Argon is the long-job Gemini. On the office benchmarks other people run, it is at or near the front, and the intro price is friendly until you count how much it writes. The catch is access. If your work is a multi-hour coding or office loop, it is the one to try when the API opens. Until then it is a Fairwind model, not yours.