Grok 4.6 Is 80% Cheaper. Here's the Catch.
Google just cut the price of a realtime voice conversation by more than half. Gemini 3.8 Live, the newest entry in Google’s voice-model line, prices audio input at $0.005 per minute and output at $0.018 per minute — a structure that lands an hour of back-and-forth conversation at roughly $1.38. GPT-Live-1, OpenAI’s answer and the successor to GPT-Realtime-2, charges a flat $0.05 per minute. Run the same hour through it and you’re at $3.00 or more.
That’s the headline, and it’s real. It’s also not the whole story. On Artificial Analysis’ Speech-to-Speech Quality Index, Gemini’s premium Extended Thinking tier edges out GPT-Live-1’s mid-tier configuration — barely. Standard tier against standard tier, the race flips, and GPT-Live-1’s base configuration wins the composite score outright. Add in the architecture gap — GPT-Live-1 runs full-duplex, Gemini 3.8 Live’s standard tier is turn-based — and “cheaper” and “better” stop being the same question. They’re not even measuring the same thing.
Quick Verdict
Aspect Gemini 3.8 Live GPT-Live-1 Best For High-volume, cost-sensitive voice apps; multimodal sessions Natural, interruptible conversation; premium support and sales lines Standard Pricing $0.005/min input, $0.018/min output $0.05/min flat Cost per hour (conversation) ~$1.38 ~$3.00+ Premium tier ($/hour) Extended Thinking: $3.50 Sol: $4.47 Speech-to-Speech Quality Index 82.6% (Extended Thinking) 81.5% (Astra, Medium) Architecture Turn-based (standard tier) Full-duplex Languages 97+ Fewer at launch Extras Background API calls, visual input while talking Native interruption handling Bottom line: Gemini 3.8 Live is the cheaper voice API by a wide margin and it wins one narrow benchmark at the top tier. GPT-Live-1 is still the better conversation partner for most standard-tier use cases, because full-duplex audio beats turn-based audio in the moments that actually annoy users — interruptions, overlapping speech, the parts of a real conversation nobody scripts.
Google’s pitch for Gemini 3.8 Live is volume economics. Per-minute pricing that splits input and output — $0.005 and $0.018 respectively — rewards exactly the workload every contact center and voice-app team is trying to build: long sessions, lots of listening, shorter bursts of talking back. A support call that’s mostly the customer explaining a problem costs less than one where the model does most of the talking. That’s a deliberate pricing shape, not an accident.
OpenAI’s GPT-Live-1 keeps it simple instead: one flat rate, $0.05 per minute, no separate meter for direction. Easier to budget. Also, on a typical two-way conversation, meaningfully more expensive — about 2.2x the Gemini rate on the same hour, based on the published per-minute figures.
Both companies also shipped a second, premium tier aimed at harder reasoning during a live call — tool use, multi-step lookups, the kind of thing that used to force a handoff to a human agent. Here the pricing structure converges: both charge a flat hourly rate instead of splitting input and output.
| Tier | Rate |
|---|---|
| Gemini 3.8 Live Extended Thinking | $3.50/hour |
| GPT-Live-1 Sol | $4.47/hour |
| Grok Voice Think Fast 2.0 High | $4.80/hour |
Gemini undercuts both competitors here too. Against GPT-Live-1 Sol — built on the same Sol-line reasoning behind GPT-5.6 Sol — Extended Thinking runs about 22% cheaper per hour. Against Grok Voice Think Fast 2.0 High, it’s about 27% cheaper. Route meaningful volume through a reasoning-heavy tier and that gap compounds fast.
The math isn’t close. An hour of standard-tier conversation on Gemini 3.8 Live costs about $1.38. The same hour on GPT-Live-1 costs $3.00 or more. For a voice app running thousands of hours a month — appointment scheduling, order status lines, first-tier support triage — that difference is the line item that decides whether the product is profitable at scale, not a rounding error on an invoice.
On Artificial Analysis’ Speech-to-Speech Quality Index — a composite that scores naturalness, latency-adjusted comprehension, and task completion in a live voice session — Gemini 3.8 Live’s Extended Thinking tier posts 82.6%, narrowly ahead of GPT-Live-1 Astra at the Medium setting, which scores 81.5%. A one-point gap isn’t a rout. It is, at minimum, proof that Google’s reasoning tier isn’t just cheaper — it’s competitive at the top of the market, not just the bottom.
Gemini 3.8 Live can accept visual input mid-conversation — point a camera at something while you’re still talking and the model responds to what it sees without you having to stop, upload, and resume. It also supports background API calls, meaning the model can trigger a tool call or a lookup without breaking the audio stream to do it. And it covers 97+ languages, a wider net than GPT-Live-1 has at launch. None of that shows up in the per-minute price. All of it matters if your product needs to see as well as hear.
Here’s the part the pricing table hides. GPT-Live-1 is full-duplex — it listens and speaks on the same continuous audio stream, the way a phone call actually works. You can interrupt it. It can register a “wait, no” mid-sentence and adjust instead of finishing a thought nobody wants to hear anymore. Gemini 3.8 Live’s standard tier is turn-based. It waits for you to stop talking, then responds. That’s a walkie-talkie model wearing a phone-call price tag.
For scripted, single-intent tasks — “check my order status,” “book me a 2pm” — turn-based is invisible. Nobody notices the handoff because there isn’t much to interrupt. For anything closer to an actual conversation — sales calls, complex support escalations, therapy-adjacent coaching apps, any use case where a customer talks over the agent because that’s what humans do — the turn-based lag is the first thing users complain about. It doesn’t show up in a benchmark score. It shows up in call recordings and abandonment rates.
Reasoning tiers aside, when you compare GPT-Live-1’s standard configuration against Gemini 3.8 Live’s standard configuration on the same Speech-to-Speech Quality Index, GPT-Live-1 comes out ahead. The composite score rewards exactly the full-duplex behavior above — lower perceived latency, fewer awkward pauses, better handling of overlapping speech. Gemini’s win at the top of the market doesn’t carry down to the tier most developers will actually ship on.
Teams already built on OpenAI’s Realtime API — the line GPT-Realtime-2 established back in May — inherit GPT-Live-1 without a rebuild. Function calling patterns, tool-use conventions, and integration code carry forward. That’s a real cost saving even if it never shows up on a per-minute pricing chart, and it’s the same switching-cost logic that keeps teams on GPT-5.5 even when a cheaper model scores higher on a benchmark.
“Cheaper” assumes your traffic matches the pricing shape. Gemini’s per-minute math looks best on listening-heavy calls. An app where the model does most of the talking — narration, walkthroughs, long explanations — shifts the blended cost closer to GPT-Live-1’s flat rate than the headline $1.38 suggests. Run your own traffic mix before assuming the discount holds.
A one-point benchmark win is not a category win. Gemini’s 82.6% against GPT-Live-1’s 81.5% is a margin, not a gap. Citing “Gemini beats GPT-Live-1 on quality” without naming the tier means quoting the one comparison that favors Google and skipping the one that doesn’t.
Turn-based isn’t a bug Google is unaware of. It’s a known trade-off, and it’s reasonable to expect Google to narrow it — the same way GPT-Realtime-2 closed the reasoning gap the original Realtime API left open. Buying today means buying today’s architecture, not the roadmap.
Neither premium tier is cheap in absolute terms. $3.50 to $4.80 an hour sounds trivial until a contact center runs reasoning-tier calls around the clock. At that point the standard-tier discount matters more, because most call volume shouldn’t need the expensive tier at all.
Google won the pricing argument decisively. An hour of standard-tier conversation costs less than half as much on Gemini 3.8 Live, and the Extended Thinking tier undercuts both GPT-Live-1 Sol and Grok Voice Think Fast 2.0 High while narrowly beating the former on quality. If your product is a high-volume, scripted voice workflow, that math is hard to argue with.
But “cheaper” and “better” are answering different questions here, and the title of this piece isn’t being cute about it. GPT-Live-1’s full-duplex architecture wins the standard-tier comparison that most developers will actually ship on, and it wins the interaction pattern that matters most in any conversation that isn’t fully scripted. Gemini’s one-point benchmark win lives at the top tier, where fewer teams will spend.
The honest recommendation: default to Gemini 3.8 Live for high-volume, low-complexity voice workloads where the price gap does real work on your margins. Default to GPT-Live-1 where the conversation itself is the product and users will notice — and complain about — a model that can’t be interrupted. Test both on your actual call recordings before committing either way. Benchmarks measure averages. Your users notice the calls that go wrong.
Gemini 3.8 Live is Google’s realtime voice AI model, priced at $0.005 per minute of audio input and $0.018 per minute of output on its standard tier, with a premium Extended Thinking tier at $3.50 per hour for reasoning-heavy calls. It supports 97+ languages, visual input during a live conversation, and background API calls without interrupting the audio stream. Its standard tier uses a turn-based conversational architecture.
GPT-Live-1 is OpenAI’s realtime voice AI model and the successor to GPT-Realtime-2, available through the OpenAI Realtime API. It charges a flat $0.05 per minute on its standard tier and offers a premium Sol-based reasoning tier at $4.47 per hour. Unlike Gemini 3.8 Live’s standard tier, GPT-Live-1 runs full-duplex — it can listen and speak on the same continuous stream, allowing natural interruptions.
Yes, substantially, on standard-tier pricing. An hour of two-way conversation costs roughly $1.38 on Gemini 3.8 Live versus $3.00 or more on GPT-Live-1 — more than double. The gap narrows at the premium tier, where Gemini’s Extended Thinking ($3.50/hour) still undercuts GPT-Live-1 Sol ($4.47/hour), but by a smaller percentage.
GPT-Live-1, on the standard tier that most developers will use. Its full-duplex architecture handles interruptions and overlapping speech the way a real phone call does. Gemini 3.8 Live’s standard tier is turn-based, meaning it waits for a pause before responding — fine for scripted, single-intent tasks, noticeably less natural in open-ended conversation.
It depends on the tier compared. On Artificial Analysis’ Speech-to-Speech Quality Index, Gemini 3.8 Live’s Extended Thinking tier scores 82.6% against GPT-Live-1 Astra (Medium) at 81.5% — a narrow win for Google. Compare the two standard tiers instead, and GPT-Live-1 wins the composite score, largely on the strength of its full-duplex architecture.
Gemini 3.8 Live’s Extended Thinking tier at $3.50/hour is cheaper than Grok Voice Think Fast 2.0 High at $4.80/hour — about a 27% discount. Both sit below GPT-Live-1 Sol’s $4.47/hour, making Gemini the cheapest of the three premium reasoning tiers currently on the market.
If your workload is high-volume and listening-heavy — scheduling, status checks, triage — the price difference likely justifies testing a migration. If your product depends on natural, interruptible conversation where users talk over the agent, GPT-Live-1’s full-duplex architecture is doing real work that a lower price tag doesn’t replace. Run both against your actual call recordings before deciding; benchmark averages won’t catch the specific failure modes your users will.
Full-duplex audio lets a model listen and speak simultaneously on one continuous stream, so it can be interrupted mid-sentence and adjust in real time — the way a phone call works. Turn-based audio requires one party to finish speaking before the other responds, closer to a walkie-talkie. GPT-Live-1 runs full-duplex; Gemini 3.8 Live’s standard tier is turn-based.
Last updated: September 17, 2026. Sources: Google Gemini product blog · OpenAI — Introducing GPT-Live-1 · OpenAI Realtime API documentation · Artificial Analysis · xAI.
Related reading: OpenAI’s GPT-Realtime-2 Rewires Voice AI · ElevenLabs $500M ARR: Voice AI Goes Institutional · Mistral Voxtral: Open-Source Voice AI · Grok 4.6 vs GPT-5.6 Sol vs Claude Fable 5 · Google AI Pro vs Ultra Pricing · Gemini Spark Is Live