MDL-4353EST.2026 · IDX.239
AudioIn production

GPT-Realtime-2.1

OpenAI · USA · 2026

A refresh aimed at the parts of a phone call that break voice agents: codes, silence, noise and interruptions.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

GPT-Realtime-2.1, released on 6 July 2026, is an incremental update to OpenAI's reasoning voice line, and the list of what changed says a lot about where voice agents actually fail. The company cites better alphanumeric recognition — order numbers, postcodes, card digits read aloud — plus improved handling of silence and background noise and better behaviour when a caller talks over the model. None of that shows up in a benchmark score; all of it decides whether a support line works. The architecture is unchanged from GPT-Realtime-2: speech to speech with configurable reasoning effort, tool use, a 128,000-token context and up to 32,000 tokens of output, text, audio and images in, text and audio out. Prices are unchanged too: 32 and 64 dollars per million audio tokens, 4 and 24 for text, 5 for image input, with cached input at 40 cents. Knowledge ends on 30 September 2024 and the model runs only on the Realtime endpoint.

#speech to speech#voice agents#realtime#reasoning#api only
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review