GPT-Live-1 is OpenAI's new voice model, live in the API since 10 September 2026 at $0.05 a minute. The pitch is simple: it listens and speaks at the same time, in one model, instead of the stitched-together speech-to-text, reasoning, text-to-speech pipeline every voice agent has run on until now. It launched yesterday. There's already a real open-source demo built on it, an official integration tutorial from Twilio, and one company's actual test numbers. This is what it does, what the price doesn't include, and who's really using it.
01. What it actually is
one model, not three stitched together
Every voice agent built before this ran a cascade: speech-to-text turns your audio into words, a reasoning model decides what to say, text-to-speech turns that back into audio. Three models, three handoffs. Each handoff adds latency, and each one is a place to lose the natural rhythm of a conversation, like when someone should have noticed you said "um" and paused, not talked over you.
GPT-Live-1 collapses the listening and speaking into a single full-duplex model. It hears you while it's still talking, so a quiet "mm-hmm" doesn't stop it, background noise doesn't confuse it, and if you add a detail mid-sentence it can fold that in without waiting for its own turn to end. It deliberately does not try to reason deeply itself. That work is delegated to a separate backend model and your own tools, which is why it needs a backend model paired with it to actually do anything.
Redrawn from OpenAI's own description of the architecture. The backend model box is not optional: GPT-Live-1 needs one paired with it to do anything beyond talk.
02. The real numbers
OpenAI's own reported benchmarks, not independently audited
These are OpenAI's own launch-day figures, cross-checked across two independent write-ups. Treat them as the vendor's own numbers, not a third-party audit.
| Metric | GPT-Live-1 | Prior generation |
|---|---|---|
| Full Duplex Bench v1.5, Interactivity | 80.10% | 45.4% / 47.8% |
| Artificial Analysis Conversational Dynamics | 97.3% | 95.7% / 95.3% |
| Turn-taking latency | 0.798s | 1.41s |
| Customer-service task success | 86.2% | 45.7% |
The two "prior generation" columns are GPT-Realtime-2 and GPT-Realtime-2.1, the cascaded-style models this replaces. The interruption-handling jump is the number that matters most in practice: it's the difference between an agent that talks over you and one that doesn't.
03. What $0.05/min doesn't cover
the headline price is the voice layer, not the whole bill
$0.05 a minute, billed per second, is the real, confirmed price for the GPT-Live-1 voice layer. It is not the total cost of running an agent on it. The backend model that does the actual reasoning, and any tools it calls, are billed separately on top. A voice agent that also needs GPT-5-class reasoning will cost noticeably more than a nickel a minute once you add that in.
A few more real constraints worth knowing before you plan around this: it's unavailable on the Free tier, it only works through the new v1/live/sessions endpoint (it is not a drop-in swap for an existing Realtime API integration, that's a real migration), rate limits scale by usage tier from 25 up to 500 concurrent sessions, and the model ships with 12 built-in voices, with a custom voice requiring a sales conversation.
04. Real builds already running on it
not rumors or renders, shipped code and public demos
The strongest proof a new model is actually usable isn't a benchmark, it's someone outside the company shipping something with it in the first 24 hours. Two did.
Twilio published an official tutorial, written by their own engineer Xinghao Huang, showing a working GPTLiveProvider that bridges Twilio's phone-call infrastructure straight to GPT-Live-1. It handles both inbound and outbound calls, keeps tool execution separate from the real-time conversation loop so a slow function call doesn't freeze the call, and ships as a real package (twilio-agent-connect[server,gpt-live]) you can install today.
Separately, Bin Liu, VP of Product & Agent Engineering at HeyGen, posted a demo of a real-time poker coach built by combining GPT-Live-1 with two other tools: HyperFrames for generating the on-screen cards and UI live, and LiveAvatar for the visual presence. He open-sourced the whole thing.
we were legit amazed by GPT-live-1
— Bin Liu (@liu8in) September 10, 2026
and when we combined it with @HyperFrames_ and @TryLiveAvatar - and it's magical
here's a demo of a real-time poker coach:
- all cards & ui fully generated on the fly w hyperframes
- GPT-live is the brain and voice
- live avatar makes it all come to life
and we open sourced everything, so anyone can build on top (and yes, the repo is agent-friendly)
Real post from a verified account, checked live. Source: x.com/liu8in.
05. What OpenAI's launch partners are seeing
named companies, and one honest anonymous case
OpenAI's own launch materials name four companies already running GPT-Live-1: Yelp reports meaningful improvements in call handling for reservations and food orders, Speak cut interruptions by almost 80% during language learners' thinking pauses, Fin says its AI voice support flows more naturally, and Cognition paired it with Devin for collaborative ideation. One more case is real but deliberately anonymous: a healthcare customer says switching off their cascaded build cut their codebase by 80%, about 23,000 lines removed. No company name is attached to that one anywhere, so it stays anonymous here too.
The most detailed public number so far comes from Genspark, who posted their own test results rather than a launch-partner quote:
Congrats to the @OpenAI team on launching GPT-Live. 🎉
— Genspark (@genspark_ai) September 10, 2026
It sounds more natural. We ran it through 80 real restaurant booking calls first: task completion more than doubled over the previous generation, with 92% perfect comprehension.
Interruptions are smoother. A quiet "mm-hmm" no longer stops it, and if you ask something new mid-sentence, it switches without a pause.
It's live in Genspark today. Now in Call for Me, Realtime Voice Agent, and something new: you can call your agents in GenTeam. Hand off work, ask where things stand, review results. Just like calling a teammate.
See the demo 👇
Real post from a verified account, checked live. Source: x.com/genspark_ai.
06. If you're building voice agents for real
what this changes if this is your actual job, not just news
I build AI voice agents for real estate, healthcare and solar clients, mostly on cascaded-style platforms today. So this one isn't just interesting to read about. It's a real architectural alternative to what I currently sell. The honest read: the interruption-handling and latency numbers are a genuine jump, not a marginal one, and that specific pain (agents that talk over the caller, or freeze mid-sentence) is the single most common complaint I hear from clients running voice agents in production.
But it's not a free upgrade. Migrating an existing integration means touching a new endpoint, not flipping a config flag, and the real total cost is the $0.05/min voice layer plus whatever your backend model and tools already cost you. For a new build, I'd genuinely evaluate this before defaulting to a cascaded stack. For an existing one already working in production, the migration cost needs to earn its keep against a real, current pain point, not just a better demo. I've written about a real voice AI build before, an e-commerce case where the cascaded approach was the right call at the time: the Voice AI Employees case study.
07. The catch
what one day of launch coverage can't tell you yet
It launched yesterday. Every real-world number in this piece, outside OpenAI's own benchmarks, comes from launch partners and early builders who had access before the rest of us. That's a real signal, but it's not the same as independent developers reporting back after weeks of production use, because that data doesn't exist yet. It also can't see, only hear and speak, so anything involving a screen or an image is out of scope by design. And because it deliberately offloads reasoning, a bad backend model or a badly designed toolset will still produce a bad agent, just one that talks over you less while doing it.
GPT-Live-1 is a real architectural shift, not a repackaged Realtime API, and the first 24 hours of independent builds back that up. Whether it's worth migrating to depends entirely on whether interruption-handling and latency are your actual bottleneck, or just a nice-to-have.
Sources and further reading
- OpenAI's official API docs for
gpt-live-1: pricing, modalities, rate limits, endpoint support. - OpenAI's own launch posts: Introducing GPT-Live and Build more natural voice experiences with GPT-Live-1 in the API.
- Unite.AI's writeup, source for the launch-partner cases and the anonymous healthcare figure.
- Twilio's official engineering tutorial, the real GPTLiveProvider build.
- Real builds checked live on X: Bin Liu's poker coach demo, Genspark's 80-call test, and OpenAI Developers' launch-day posts.
// Talk to me directly
Join the Discord
I'm most active there and I actually reply. Drop your build in #models-and-tools, or ask in #help-and-tips if you're stuck on something.
#models-and-tools#help-and-tips// Free newsletter
I send out guides like this every week
Real setups, real sources, no hype. Drop your email and I'll send you the next one.