Kimi-Audio 7B Instruct
Moonshot AI · China · 2025
One open model for everything audio: transcription, sound and emotion recognition, and spoken conversation that answers in speech, trained on 13 million hours.
Kimi-Audio 7B Instruct, published on 25 April 2025, is Moonshot AI's attempt to replace a stack of separate audio tools with a single model. The same weights handle speech recognition, audio question answering, audio captioning, speech emotion recognition, sound event and scene classification, and end-to-end spoken conversation — tasks that a typical production system assembles from three or four specialised components. Pre-training used more than 13 million hours of audio covering speech, music and general sound, alongside text. The architecture is what makes the range possible. Audio enters in two forms at once: a continuous acoustic representation preserving how something sounds, and discrete semantic tokens carrying what was said. A language-model core then generates text and audio tokens through parallel output heads, and a chunk-wise streaming detokenizer based on flow matching turns those audio tokens back into a waveform while the answer is still being produced, which is what keeps spoken replies from arriving late. Two notes for anyone reading the specification. First, the name understates the model: 7B refers to the language backbone, while the published weights come to 9.8 billion parameters once the audio tokenizer and the generation heads are counted. Second, Moonshot's claim of state-of-the-art results across audio benchmarks is not backed by numbers on the model card itself — the figures live only in a technical report PDF in the project repository, so the card cannot be checked on its own terms. Weights for the instruction-tuned and base checkpoints are on Hugging Face under an MIT licence; the inference code ships as a separate repository with a Docker image and no licence file of its own. There is no paid Moonshot endpoint for this model.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!