GPT-4o mini Realtime, available since 17 December 2024, brought live spoken conversation within reach of applications that could not justify the flagship rate: $10 per million audio tokens in and $20 out, against $40 and $80 for GPT-4o Realtime. Text costs $0.60 and $2.40, cached input $0.30. What the buyer gives up is room to remember. The model holds 16,000 tokens of context — half of the full realtime model and an eighth of the turn-based mini — and answers with at most 4,096, which in a long voice conversation is the constraint that bites first. It speaks and listens over WebRTC or a WebSocket, takes text and audio in and out, and its knowledge ends on 1 October 2023. Like the rest of the legacy voice line it was retired on 20 July 2026 and stops responding on 20 January 2027; the successor is gpt-realtime-2.1-mini.