GPT-Realtime-2
OpenAI · USA · 2026
A voice that thinks before it answers — GPT-5-class reasoning inside a live conversation.
GPT-Realtime-2, announced on 7 May 2026, is OpenAI's most capable speech-to-speech model and the first voice model the company describes as reasoning at the level of its GPT-5 text line. Earlier voice models were fast but shallow: they answered quickly because they barely thought. This one carries an adjustable reasoning effort, so a developer can trade latency for depth on a per-request basis, and it keeps the conversation going while it works rather than falling silent. It handles text, audio and images on the way in, speaks and writes on the way out, holds a 128,000-token context and calls tools in parallel. The trade-off is price and scope: audio costs 32 dollars per million input tokens and 64 per million output, and the model runs only on the Realtime endpoint — it cannot be used through chat completions, batch or the assistants API. Its knowledge stops at 30 September 2024, so anything newer has to arrive through tools.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!