GPT-Realtime-2.1 mini
OpenAI · USA · 2026
The same live conversation for roughly a third of the price — a distilled version of the reasoning voice model.
GPT-Realtime-2.1 mini, released on 6 July 2026 together with the full model, is a distilled version of OpenAI's reasoning voice model, and the reason to care about it is arithmetic. Audio costs 10 dollars per million input tokens and 20 per million output, against 32 and 64 for the full model — a live voice minute that reasons before it answers becomes roughly three times cheaper. Text drops even harder, to 60 cents in and 2.40 out. Everything structural is kept: a 128,000-token context, up to 32,000 tokens of output, reasoning tokens, function calling, text, audio and images on the way in, text and audio out, over WebRTC, WebSocket or SIP. What is traded away is depth, in the usual way of a distilled model — it is the cheap tier for high-volume voice traffic, not the model to reach for when an answer has to be right the first time. Knowledge ends on 30 September 2024 and the model runs only on the Realtime endpoint.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!