Octave 2
Hume AI · USA · 2025
A speech model that reads the script like an actor — you direct it in prose, not in markup.
Octave 2 is Hume AI's speech-language model, and the distinction from ordinary text-to-speech is the company's central claim: rather than converting characters into sound, Octave reads the script the way an actor would, inferring from the words themselves when to whisper, shout or explain calmly. A user can also direct it in plain prose — describe the character and the delivery, and the model builds a voice to match, instead of picking from a fixed roster. The second generation, released on 1 October 2025, made that approach practical for live conversation: under 200 milliseconds to audio, a 40 percent speed gain over Octave 1 at half the price, and coverage of eleven languages — Arabic, English, French, German, Hindi, Italian, Japanese, Korean, Portuguese, Russian and Spanish — with the company promising at least twenty. It added multi-speaker conversation, voice conversion that swaps one voice for another while preserving timing and phonetics, and phoneme-level editing for fixing how a name or an emphasised word is pronounced. Hume's origin in emotion research is what shapes the product: the company sells expressive intent rather than fidelity, which places it against ElevenLabs on interpretation rather than on latency alone.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!