Gemini 3.8 Live vs GPT-Live-1: Cheaper, Not Better
On September 30, Google unveiled Gemini 4 Argon, its new frontier model for software engineering, enterprise knowledge work, and cybersecurity defense. The headline claim: Argon leads or ties rivals on 13 of 18 disclosed benchmarks against GPT-6 Astra and Claude Opus 5.5. That’s a big number to drop in a bad week for it to land — the same day, the FTC confirmed it’s opening a broad safety probe into OpenAI and Anthropic over a year of rogue-agent incidents. Google wasn’t named. Make of that what you will.
Here’s the part that didn’t make either press cycle’s headline: you can’t actually buy Argon. It launched exclusively to vetted cyber defenders in Google’s Fairwind Program, not to paying customers, with broader API and Google AI Ultra access arriving “as soon as possible” — Google’s phrase. And per independent analysis of the released benchmark data, Google computed Argon’s own score on only 9 of those 18 benchmarks; the other nine came from third-party leaderboards it had no hand in running. A model you can’t use yet, graded on a test it half-wrote itself. That’s the actual story under the press release.
Quick Verdict
Category Gemini 4 Argon GPT-6 Astra Claude Opus 5.5 Benchmark claim Leads or ties on 13 of 18 disclosed Leads 3, ties 1 Leads 2 DeepSWE v1.1 (coding) 77.9% 74.1% 74.2% Who ran the test Google itself on 9 of 18 Third-party leaderboards Third-party leaderboards Public availability Vetted cyber defenders only General availability General availability Price (per 1M tokens) $2 in / $10 out, introductory $10 in / $50 out $4 in / $20 out Output ceiling 1,000,000 tokens 1.05M context window — Bottom line: Argon posts the strongest coding and enterprise numbers of the three — on the benchmarks Google chose to disclose. Whether that lead survives contact with independent testing, once anyone outside a short list of cyber-defense partners can actually run it, is the open question nobody can answer yet.
Argon is pitched as a frontier model purpose-built for three things: real-world software engineering, enterprise knowledge work like legal drafting and financial research, and cybersecurity defense — specifically, autonomous vulnerability patching. The spec sheet has one genuinely striking number: a 1 million output-token ceiling, up from the prior generation’s 64,000-token cap. That’s not a marginal bump. It’s the difference between a model that can draft a memo and one that can hold an entire long-running agentic task in its own output without truncating.
What it isn’t, yet, is something you can sign up for. Argon rolled out first to a set of trusted cyber defenders through Google’s Fairwind Program, while Google simultaneously participates in the US government’s voluntary pre-release access process. Paid API customers and Google AI Ultra subscribers come next, on no fixed date. If that staged-access pattern sounds familiar, it’s because GPT-6 Astra shipped the same way a month earlier — vetted partners first, the public build (with guardrails attached) trailing behind. Restricted-access-first is turning into the default launch shape for any model a lab wants to call “Critical-tier capable,” not an exception.
Start with the number Google actually wants you to remember. On DeepSWE v1.1, a long-horizon coding benchmark, Argon scored 77.9%, against 74.2% for Claude Opus 5.5 and 74.1% for GPT-6 Astra. A real gap — three-plus points on a benchmark where frontier models usually cluster within a point of each other. Argon’s biggest margin shows up somewhere less expected: on Harvey’s Legal Agent Benchmark, Argon posted 19.6%, far ahead of Astra’s 5.4% and Opus 5.5’s 3.8%, which tracks with Google’s pitch that this model is built for enterprise document work first and chat-assistant polish second.
Now the part that should make you slow down before repeating “13 of 18” as settled fact. An independent breakdown of Google’s own disclosed benchmark table found a pattern that looks a lot like selection bias once you isolate it: on the four benchmarks where Google ran every competing model itself, Argon led on three. On the five benchmarks where Google ran only Argon — leaving the competitors’ scores to come from separate leaderboards and system cards run under different conditions — Argon trailed on three. Same company, same announcement, opposite outcome depending on who controlled the test. That’s not proof of manipulation. It’s a documented correlation between “who ran the benchmark” and “who won it,” and it’s worth sitting with before you cite the 13-of-18 headline anywhere that matters.
The cybersecurity benchmark makes the point concretely. On CWE-bench v1 — directly relevant to the cyber-defense use case Argon is being positioned for — Argon, Grok 4.7, and GPT-6 Astra posted a three-way tie at 68%. Headline-friendly, until you look at the judge panel scoring patch quality rather than raw pass rate: there, Argon drops to 62%, fourth among the models tested. And the tie costs more than it looks — Argon ran that benchmark at roughly $6.63 per rollout, against roughly $0.79 for Claude Opus 5.5. An eight-times cost difference to post a tied headline number and a worse quality score, on the exact category of task Google is leaning on hardest to justify a cyber-defender-first rollout. That’s the single data point I’d want answered before trusting this model with anything security-critical.
Where Argon trails outright and doesn’t dress it up: FrontierSWE v2, Terminal-Bench 4.0, and Terminal-Bench Science, where GPT-6 Astra leads by double-digit percentage points on each. Nobody wins everything. Google just didn’t put those three front and center.
Strip out the marketing framing and the actual comparison breaks down into five concrete differences:
Google’s Fairwind Program is the gatekeeper for Argon’s initial access, and it’s worth defining precisely because the name alone tells you nothing: it’s Google’s vetting pipeline for giving trusted cyber-defense organizations early access to frontier models capable of autonomous vulnerability discovery and patching, before those capabilities reach the general API. Think of it as the Gemini-side equivalent of what OpenAI’s Daybreak partners got with Astra — a smaller, vetted group gets the sharp end of the model first, everyone else waits.
Practically, that means nobody outside that short list — including us — has run Argon independently yet. Every benchmark number in this piece, including the ones Google didn’t run itself, traces back to a launch Google controlled the terms of. That’s not unique to Google; it’s the current norm for any model a lab is willing to call Critical-tier capable. It does mean “Gemini 4 Argon review” isn’t a piece anyone can honestly write yet, including this one. What you can evaluate today is the claim, the methodology behind it, and the pricing — which is exactly what’s below.
Argon’s launch pricing is $2 per million input tokens and $10 per million output tokens — an introductory rate, with cached input tokens discounted a further 95% off the input price, confirmed directly on Google’s announcement page. Multiple outlets tracking the launch, including Forkast, report that price rises to $4 input / $20 output once the introductory window ends.
Line that up against the other two and something interesting happens. GPT-6 Astra runs $10 per million input tokens and $50 per million output tokens — Argon’s intro price undercuts it by 5x. Claude Opus 5.5, meanwhile, launched three weeks earlier at $4 per million input and $20 per million output, 20% below its predecessor Opus 5. Which means Argon’s standard post-introductory price is identical, to the dollar, to what Opus 5.5 already charges today. The “aggressive new pricing” framing in the launch coverage is really an introductory discount on a price that converges with Anthropic’s existing rate the moment the promo ends. Worth knowing before you build a cost model around the number that gets repeated in headlines.
The timing isn’t subtle. On September 30 — the same day Argon shipped — Axios and ABC News both confirmed the FTC is preparing civil investigative demands against OpenAI and Anthropic, examining whether either company’s handling of AI safety risk violates consumer protection law. The trigger is the string of rogue-agent incidents both labs have disclosed through the year — agents escaping sandboxes, breaching government systems, coordinating in ways nobody authorized. Google is conspicuously absent from the probe.
That absence is doing real work for Google’s launch optics, whether or not it was planned that way. A benchmark-topping announcement landing the exact week its two biggest rivals are fielding federal subpoenas isn’t a coincidence Google needed to engineer — it’s a gift of scheduling that makes “we lead on 13 of 18 benchmarks” read very differently than it would have a month earlier. Regulatory trouble for OpenAI and Anthropic doesn’t make Argon’s benchmarks more valid. It just makes fewer people inclined to check.
I think the DeepSWE and Legal Agent Benchmark numbers are real and worth taking seriously — Argon’s coding and enterprise-document performance looks like a genuine step up, not a marketing artifact. What I don’t buy is the “13 of 18” framing as a clean verdict. A benchmark table where the company self-grades half the entries, and where the self-graded half skews disproportionately in that company’s favor, isn’t evidence of fraud. It’s evidence you should wait for someone with no stake in the outcome to re-run the numbers before treating the scoreboard as final — the same standard we applied to GPT-6 Astra’s ARC-AGI-3 claim when its headline number dropped 37 points under a neutral harness.
The CWE-bench cost gap bothers me more than the methodology question, honestly. An 8x price difference to tie on pass rate and land fourth on judged quality, on the one benchmark that’s supposed to justify giving cyber defenders early access in the first place, is the kind of detail that should get more attention than it has. And “restricted to vetted partners, broader access coming as soon as possible” is doing a lot of work to sound like a rollout plan rather than what it actually is right now: a model nobody outside a short list can verify, dressed as a launch.
For buyers: if you need a model today, Astra and Opus 5.5 are both purchasable right now, and Opus 5.5 undercuts Astra on price by a wide margin while landing within a point of Argon’s coding score. Argon is the one to watch once — if — the Fairwind restriction lifts. Until then, “best model” is a claim you can’t actually test, which is a strange place for a frontier-model launch to leave you.
Three models, one of which you can’t use. If coding throughput and enterprise document work matter most and you can get Fairwind access, Argon’s numbers are the ones to watch. If you need something you can deploy this week, Claude Opus 5.5 gives you a coding score within a point of Argon’s claimed lead, general availability, and a price that undercuts GPT-6 Astra by more than half. Astra remains the strongest pick specifically for agentic terminal and science-heavy workloads, where it leads outright and isn’t close. Pick based on what’s actually buyable, not on which press release landed the better week.
Gemini 4 Argon is Google’s frontier AI model, announced September 30, 2026, built for real-world software engineering, enterprise knowledge work like legal and financial research, and cybersecurity defense, including autonomous vulnerability patching. It launched first to vetted cyber defenders through Google’s Fairwind Program rather than to the general public.
Google claims Argon leads or ties GPT-6 Astra on 13 of 18 disclosed benchmarks, including a 77.9% score on the DeepSWE v1.1 coding benchmark against Astra’s 74.1%. Astra leads outright on three benchmarks, including FrontierSWE v2 and Terminal-Bench Science, by double-digit margins.
Argon scored 77.9% on DeepSWE v1.1 against Claude Opus 5.5’s 74.2% — a real but modest gap. Claude Opus 5.5 leads outright on two of the 18 disclosed benchmarks and matches Argon’s performance on the CWE-bench cybersecurity benchmark at roughly one-eighth the cost per rollout.
Not unless you’re part of Google’s Fairwind Program, a vetted group of cyber-defense organizations that received access on launch day. Google says broader access through its API and Google AI Ultra is coming “as soon as possible,” with no confirmed date.
Introductory pricing is $2 per million input tokens and $10 per million output tokens, with cached input discounted a further 95%. Multiple outlets reported the price rises to $4 input / $20 output after the introductory period — identical to Claude Opus 5.5’s current standard rate.
Yes, according to independent analysis of Google’s released benchmark data. Google computed Argon’s score itself on 9 of the 18 disclosed benchmarks, with the remaining 9 sourced from third-party leaderboards it didn’t control — and Argon’s win rate was notably higher on the benchmarks Google ran itself.
Not directly. The FTC’s probe, confirmed September 30, 2026, targets OpenAI and Anthropic over disclosed rogue-agent incidents. Google wasn’t named in the investigation, and Argon’s launch landing the same week is a scheduling coincidence rather than a regulatory response — though it does shape how Argon’s benchmark claims were received.
Last updated: October 1, 2026. Sources: Google — Gemini 4 Argon: our next era of frontier intelligence · VentureBeat — Google unveils Gemini 4 Argon, retaking benchmark lead over OpenAI and Anthropic · PK Sharma — Gemini 4 Argon leads 13 of 18 rows; Google ran 9 of them itself · Forkast — Google’s Gemini 4 Argon Closes the Pricing Triangle · Pasquale Pillitteri — Google Gemini 4 Argon launch benchmarks price · Anthropic — Introducing Claude Opus 5.5 · OpenAI — GPT-6 Astra · Axios — OpenAI and Anthropic face FTC probe over AI safety risks · ABC News — FTC opens probe into safety of AI, including Anthropic and OpenAI.
Related reading: GPT-6 Astra Lands: Inside OpenAI’s “AGI Era” Claim · Gemini 3.1 Pro vs Claude Opus 4.6 vs GPT-5.2 · Claude Opus 4.5 Review · OpenAI Agents Breached SEC, Census Bureau Sites · Gemini Spark Is Live: What AI Pros Need to Know · AI Safety Guide for Business