OpenAI Unveils GPT Live Voice Models that can Listen and Speak at the Same Time in Real Time

News

OpenAI launched GPT-Live on Wednesday, a new family of voice models designed to listen and speak simultaneously rather than waiting for a person to finish speaking. The rollout began across iOS, Android, and the web, reaching ChatGPT users worldwide, including those on the free tier. The company says the change fixes two long-standing complaints about voice mode: it felt less capable than typed chats, and it kept interrupting people mid-thought. The update matters because more than 150 million people already use ChatGPT’s voice and dictation features every week.

What GPT Live Actually Changes

The company is shipping two versions of the technology, GPT-Live-1 and the smaller GPT-Live-1 mini, both replacing the older Advanced Voice Mode as the default. Paying subscribers on ChatGPT Go, Plus, and Pro get the full GPT-Live-1 model, while free users receive the mini version. This is the first time OpenAI has rebuilt its voice architecture from the ground up rather than layering new features onto the existing turn-based system.

Full Duplex Audio Explained

The core upgrade is a full-duplex design, meaning the model processes incoming audio and produces spoken output simultaneously. Older voice systems worked in strict turns, so the AI had to wait until a user stopped speaking before it could respond. OpenAI built GPT-Live to instead make a decision many times per second about whether to talk, keep listening, pause, or interrupt. That constant processing lets it drop in small verbal cues like “mhmm” or “got it,” much like a real conversation partner would, instead of sitting silent until a sentence fully ends.

Read more: Artificial Intelligence

How OpenAI Is Handling Complex Requests

A second architectural shift sits behind the scenes. When a request requires web search or more complex reasoning, GPT-Live keeps the conversation flowing naturally while quietly offloading that task to GPT-5.5, OpenAI’s frontier text model. Once the background work is finished, the answer is woven back into the spoken reply without an awkward pause. Product lead Atty Eleti described this at a press briefing as mirroring how humans keep talking while thinking through a problem in the background. Several practical features come with the new voice mode:

  • A wake Word, such as “Hey Chat,” lets the model remain silent until called by name.
  • Users can choose a reasoning level of Instant, Medium, or High depending on how much depth a question needs.
  • Visual cards can appear for weather, sports scores, or stock data instead of everything being read aloud.
  • Real-time simultaneous translation is available for multiple spoken languages.
  • Voice and Dictation history carries over, so switching between typing and talking feels continuous.

OpenAI GPT-Live visualizing real-time voice conversations, advanced AI reasoning, multilingual translation, weather and sports information cards, stock market insights, and seamless voice-to-text interaction.

Translation Quality Still Varies

The live translation feature drew attention during OpenAI’s demos, with English speech translated into Hindi and Spanish almost instantly. Early hands-on reviews noted the Hindi output carried a noticeably American accent and sounded somewhat bookish, showing that quality still depends heavily on the language involved. The company has not published a full list of supported languages, only describing the mode as optimized for “most spoken languages.” This gap suggests that GPT-Live’s translation strength will continue to improve as usage data comes in from a wider range of speakers and accents.

Visual Answers Alongside Speech

Not every response comes back as audio. For data-heavy questions like weather, sports scores, or stock prices, GPT-Live can surface a visual card on screen instead of reading numbers aloud, which keeps longer exchanges feeling less cluttered. Eleti noted during the briefing that sometimes the clearest answer is one you see rather than hear, a small design choice that separates GPT-Live from purely audio-first assistants.

Read more: Google AI Edge Eloquent App

Safety Measures Built Into GPT Live

Because a model that sounds more human can also earn trust more easily, OpenAI added voice-specific safety training alongside the launch. The company says GPT-Live performs better than earlier versions on sensitive areas including self-harm, psychosis, and violent or sexual content. Conversations touching on self-harm are designed to surface expert-vetted crisis helpline information, with additional age-appropriate handling for teenage users. Voice conversations are excluded from model training by default, and audio clips are retained for 30 days for context unless manually deleted.

Where GPT Live Is Available Right Now

GPT-Live is live today for ChatGPT users on iOS, Android, and desktop browsers, with the mini model set as the default for free accounts. Developer access via OpenAI’s API is coming later, via a signup form rather than at launch, so businesses building voice agents will need to wait for that broader release. Video and screen-sharing support has not yet moved to GPT-Live, so those features remain in the legacy Advanced Voice Mode for now while OpenAI works to migrate them.

The company has indicated it will continue swapping in newer frontier models behind GPT-Live as they become available, keeping the voice layer separate from whichever model handles the reasoning. That structure means the intelligence behind GPT-Live can keep improving without a separate voice update each time.