GPT-4o Realtime
OpenAI · USA · 2024
The model behind the first live voice API: speaks over WebRTC, can be interrupted mid-sentence — and remembers only 32,000 tokens of the conversation.
GPT-4o Realtime, launched with the Realtime API on 1 October 2024, was the model that made spoken conversation with a machine feel like a phone call rather than an exchange of recordings. It keeps a connection open over WebRTC or a WebSocket, starts speaking before the sentence is finished and stops when the human cuts in. The price of that immediacy is memory: it holds 32,000 tokens of context and returns at most 4,096 — a quarter and a quarter of what its turn-based sibling GPT-4o Audio manages — while audio tokens cost exactly the same, $40 in and $80 out per million. Text runs at $5 and $20, with cached input at $2.50. Input and output are text and audio, knowledge ends on 1 October 2023, and the model never left preview. OpenAI announced its retirement on 20 July 2026; it goes silent on 20 January 2027, replaced by gpt-realtime-2.1.
▸News
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!