GPT-4o mini TTS
OpenAI · USA · 2025
The first OpenAI speech model that can be told how to say something, not just what. Built on GPT-4o mini, priced per token rather than per character.
GPT-4o mini TTS shipped on 20 March 2025 in the same release as the gpt-4o-transcribe pair, and it marks the point where OpenAI's speech synthesis stopped being a standalone converter and became a language model with an audio output. The practical consequence is steerability: a developer can instruct the model on delivery — tone, pace, character — and not only on the words. OpenAI's own framing runs from customer service to creative storytelling. The voices remain artificial presets, monitored by OpenAI to ensure they match the synthetic set; this model does not clone real voices. Input is capped at 2,000 tokens per request, which rules out feeding it a whole chapter in one call and makes it a per-utterance tool rather than a batch renderer. Billing changed with it. The older TTS-1 pair charges per character; this model charges $0.60 per million text input tokens and $12 per million audio output tokens. Because a token of text produces far more than a token's worth of audio, the two schemes are not directly comparable — which is precisely why the older, simpler models still have a constituency. The current default snapshot is gpt-4o-mini-tts-2025-12-15.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!