Sonic 3.5
Cartesia · USA · 2026
The voice model that is not a transformer — state-space architecture bought it sub-90 ms speech in 42 languages.
Sonic 3.5 is Cartesia's flagship speech-synthesis model and the most visible commercial product built on state space models rather than the transformer architecture that dominates the rest of the field. That choice is the point: SSMs process sequences with a constant-size internal state instead of re-reading a growing context, which suits streaming audio and is how Cartesia reaches sub-90-millisecond time-to-first-audio while still ranking at the top of independent naturalness comparisons. The model covers 42 languages, Polish among them, and its 2026 refresh concentrated on the failures that make synthetic agents unusable in practice rather than on headline quality: alphanumerics — order numbers, phone numbers, IDs and email addresses — are pronounced correctly across all languages without preprocessing, and English heteronyms such as read, bass and bow are resolved from context. Cartesia announced it alongside Ink-2, its streaming transcription model, giving the company both halves of a real-time voice loop.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!