Gemini 4 Argon is for cyber defenders only, and Gemini 3.1 Pro is still a preview

In mid-September, a developer posted an Ask HN titled "Did Google kill its enterprise workhorse model?" His company runs millions of cases through Gemini, each with hundreds of thousands of tokens of documents, and he needs a Pro-class model that gets the reasoning right the first time. He wrote that Google was sunsetting Gemini 2.5 Pro in October, before any Pro model was generally available to replace it.
On September 30, Google announced Gemini 4 Argon. The first comments on the Hacker News thread for it were about the price and about Google's record of shipping models. One commenter asked Google not to turn off its old generally available Pro model before a new one is generally available.
Both threads describe the same gap. Argon is open to a small group of cyber defenders, and Google's current Pro model is still a preview. Below: what Google's docs say, what developers are doing instead, and what Argon changes for them.
A developer with millions of long documents asked where Google's Pro model went
The Ask HN poster, waldrews, says Gemini had a niche in document comprehension: a thousand-page document takes about 300K tokens, and he says OpenAI and Anthropic have nothing quite like it, even at more than 10x the token-adjusted price. He says Google's 3.x Flash models beat the old Pro models on benchmarks, but aren't the same on reasoning-heavy work over very large documents. He also says Gemini 3.1 Pro is "bizarrely still not in General Availability status," so he can't run it for U.S.-restricted workloads.
His October sunset date for 2.5 Pro isn't on Google's Gemini API deprecations page, which lists no shutdown date for it. He may be reading a different Google schedule, such as Vertex AI's. I couldn't confirm it.
Google's docs list one Pro model, and it is a preview
This is what Google's Gemini API deprecations page said when it was last updated on September 30:
| Model | Status on Google's page |
|---|---|
| Gemini 3.1 Pro Preview | Released Feb 19, 2026. Still a preview. No shutdown date. |
| Gemini 2.5 Pro | Not deprecated, served until further notice. Limited to users who have used it before. |
| Gemini 3.5 Pro | Not listed. Axios reports it never came out. |
| Gemini 3.8 Flash | Stable, released Sep 2. Google's recommendation for new projects, along with 3.5 Flash-Lite. |
| Gemini 4 Argon | Not listed. Fairwind Program only. |
A developer starting a new project today is told to use Flash. The one Pro model open to new users has been a preview for seven months.
Recommended
Nvidia Paid 86 Times Revenue for Hugging Face

Nvidia just spend $12.93 billion on Hugging Face, a company producing roughly $150 million in annualized revenue. The multiple is about 86 times revenue. It shifts depending on which number counts. Nvidia's SEC filing pu…
Read nextDevelopers split on whether Flash replaces Pro
The Ask HN thread has the debate in miniature. Commenter ernsheong said "Flash is the new Pro, try it first." Waldrews answered that Flash is a great writer and handles large context, but complex reasoning with convoluted rules and low tolerance for hallucination is still larger-model territory. When another commenter suggested a stronger harness with multi-agent verification loops, he said that fits interactive development, but his work is high-volume with big inputs and low error tolerance, where "usually get it right the first time" is part of the cost.
Other commenters landed elsewhere:
- davedx is migrating carefully with evals. He has seen some regressions but finds 3.x Flash works well for his team's cases, and says fallbacks to other models cost engineering time.
- VirusNewbie says 3.8 Flash on high thinking is better than 3.1 Pro.
- torvin92 and kennywinker argue the sunset notice proves the case for open-weight models. torvin92 called an API a dependency you don't control, and kennywinker said building a business on anything else is a bad idea.
- trio8453 says Google's problem is direction, not research. He calls Gemini 3.8 Flash impressive and the model line confusing, down to the version numbering.
So the people closest to the problem have picked different fixes depending on their workload: a stronger harness, migration with evals, or leaving for open weights.
Gemini 4 Argon's price doubles after launch, and its context window isn't published
Google set Argon's price before opening access. Here are the list prices, from Google and MarkTechPost's comparison table:
| Model | Input (per 1M tokens) | Output (per 1M tokens) | Cached input (per 1M tokens) |
|---|---|---|---|
| Argon, introductory | $2 | $10 | $0.10 |
| Argon, after the introductory period | $4 | $20 | $0.20 |
| Claude Opus 5.5 | $4 | $20 | $0.20 |
| GPT-6 Astra (short-context tier) | $10 | $50 | $1.00 |
Google says cached input is 95% off the input price, which gives $0.20 after the increase. Google hasn't said how long the introductory period lasts.
Take waldrews's workload as an example. One 300K-token document per case and a million cases is 300 billion input tokens. That costs $600,000 at $2 per million and $1.2 million at $4. This is an illustration that ignores output, caching, and whatever he pays now.
It also assumes the documents fit. Google hasn't published Argon's input context window, and the announcement only states the output limit of 1M tokens, up from 64K. Nobody outside Google's partners can check yet whether a thousand-page document fits.
Argon leads 12 of 18 of Google's benchmarks and trails on three
MarkTechPost's read of Google's comparison says Argon leads outright on 12 of 18 tests and ties for first on one. It leads on DeepSWE v1.1 (77.9%, against 74.2% for Opus 5.5 and 74.1% for GPT-6 Astra), AutomationBench (51.3%) and LVBench (91.7%).
It trails on three tests:
- FrontierSWE v2: 55.0%, against 65.5% for GPT-6 Astra.
- Terminal-Bench 4.0: 57.4%, against 66.4% for Opus 5.5.
- OSWorld-2.0: 69.2%, against 72.6% for GPT-6 Astra.
Google chose the tests. Argon's coding lead is on one benchmark, and two coding-related tests go the other way.
People who use Gemini describe a good writer that sometimes misbehaves
Commenters on the Argon thread report mixed real-world experience:
- One says Gemini writes some of the nicest, easiest-to-read prose.
- One says it has been "truly incredible" for writing, frontend and sysadmin work, including configuring a NixOS system.
- One says Gemini has a lot of "overloaded" hiccups.
- One calls Gemini the model that is routinely "borderline psychotic," and says it scares them.
- One pays for Google's AI Plus plan and says the newest model in the Gemini app is 3.6 Flash-Lite.
These are individual reports, not measurements. Google says it monitors Argon's chain-of-thought and actions and stops execution when needed, and that Argon leads Gray Swan's prompt injection benchmark. None of that can be checked outside Google's partners yet.
Google's engineers used Argon agents to rewrite a video decoder
Argon agents took an existing Rust port of libgav1, Google's open source video decoder, and replaced 32K lines of hand-tuned SIMD code with safe Rust the compiler can speed up on its own. Google says the result runs 2.7x faster than the Rust port and produces identical video output. The identical output is a check the agents could run after every attempt.
Google is also migrating up to 800K+ lines of the Fuchsia Zircon kernel to Rust with Argon agents. Those rewrites are still in audit and review. On the Argon thread, one commenter remembered Google's C++ team passing on Rust for Carbon and Swift, and a reply said LLMs favor languages with more code to learn from.
Four things to do before Argon opens
- Budget at $4 input and $20 output. Treat $2 and $10 as a bonus, and wait for a published context window before planning long documents.
- Check the shutdown date for the exact product you use. If you're on 2.5 Pro, look at both the Gemini API and Vertex AI schedules. Google's Gemini API page lists none, but waldrews says otherwise.
- Write your own test. Pick five to ten real cases with a pass/fail check. Run Flash on high thinking, 3.1 Pro Preview, and later Argon, and compare cost per correct answer, not cost per token.
- Put the model call behind one function. Switching models should be a config change, which makes the fallbacks davedx describes cheaper to build.
The Ask HN poster's question was where Google's Pro model went. On September 30, Google answered with a model that only cyber defenders can use.
Sources
- Gemini 4 Argon: our next era of frontier intelligence, Google
- Gemini deprecations, Google Gemini API docs (read Sept 30, 2026)
- Gemini 4 Argon, Hacker News thread
- Ask HN: Did Google kill its enterprise workhorse model?, Hacker News
- Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens, MarkTechPost (benchmarks, competitor prices, context window note)
- Google unveils long-awaited Gemini 4, Axios (Gemini 3.5 Pro never came out)
- Proactive cyber defense for governments and enterprises, Google (Fairwind Program)
Comments
No comments yet
Be the first to share your thoughts