GPT-4o Transcribe Diarize
OpenAI · USA · 2025
The only OpenAI model that says who spoke. That single ability keeps it in service even though the vendor now recommends a cheaper model that lacks it.
GPT-4o Transcribe Diarize is a speech-to-text model with speaker diarization built in: it does not merely write down what was said, but splits the transcript by who said it, returning a diarized_json response with speaker labels and start and end timestamps for each segment. It reached the Audio API in late October 2025 and is available through the transcription endpoint alone — no Chat Completions, no Assistants, no batch or fine-tuning. Technically it is the narrowest member of the 4o audio family: 16,000 tokens of context, 2,000 tokens of output, knowledge to 1 June 2024. It bills $2.50 per million input tokens and $10 per million output tokens, which works out at roughly $0.006 per minute of audio — the same rate as plain gpt-4o-transcribe and whisper-1. Its position changed in July 2026 without the model itself changing at all. OpenAI released GPT Transcribe at $0.0045 a minute and made it the recommended starting point, describing the 4o transcription models as suitable for existing integrations but no longer the place to begin. Speaker labels, however, did not move: OpenAI's own transcription guide still sends anyone who needs to know who was talking to this model. Meetings, interviews, call-centre recordings and anything destined for a court-style transcript therefore run on a model that costs a third more than the recommended one and carries no recommendation. As of August 2026 it appears on no deprecation list.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!