MDL-2320EST.2024 · IDX.474
AudioIn productionretires 1/20/2027 — in 154 days

GPT-4o Realtime

OpenAI · USA · 2024

The model behind the first live voice API: speaks over WebRTC, can be interrupted mid-sentence — and remembers only 32,000 tokens of the conversation.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

GPT-4o Realtime, launched with the Realtime API on 1 October 2024, was the model that made spoken conversation with a machine feel like a phone call rather than an exchange of recordings. It keeps a connection open over WebRTC or a WebSocket, starts speaking before the sentence is finished and stops when the human cuts in. The price of that immediacy is memory: it holds 32,000 tokens of context and returns at most 4,096 — a quarter and a quarter of what its turn-based sibling GPT-4o Audio manages — while audio tokens cost exactly the same, $40 in and $80 out per million. Text runs at $5 and $20, with cached input at $2.50. Input and output are text and audio, knowledge ends on 1 October 2023, and the model never left preview. OpenAI announced its retirement on 20 July 2026; it goes silent on 20 January 2027, replaced by gpt-realtime-2.1.

#audio#speech#realtime#api only#retiring
Official website

News

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review