GPT-Realtime mini
OpenAI · USA · 2025
OpenAI's budget voice model — a replacement that lasted eight months before being scheduled for shutdown itself.
GPT-Realtime mini is the cut-price tier of OpenAI's speech-to-speech API: the same WebRTC, WebSocket and SIP connections as the full model, the same 32,000-token context and 4,096-token output limit, but roughly a sixth of the text price at $0.60 and $2.40 per million tokens. It arrived with the snapshot dated 6 October 2025 and, from 7 May 2026, was the model OpenAI told developers to move to when gpt-4o-mini-realtime-preview was switched off. That recommendation held for eight months: on 20 July 2026 the mini itself was put on the deprecation list, with removal from the API on 20 January 2027 and gpt-realtime-2.1-mini named as its successor. The squeeze started earlier — the original October 2025 snapshot was already shut down on 23 July 2026, leaving only gpt-realtime-mini-2025-12-15 running. Feature support is deliberately thin next to the full model: function calling and prompt caching, no MCP servers. OpenAI no longer publishes audio-token rates for it; the current price list carries only the successor, at $10 and $20 per million audio tokens.
▸News
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!