MDL-2075EST.2023 · IDX.778
Language modelRetiredretired 3/30/2025

Mixtral 8x7B

Mistral AI · France · 2023

The first freely licensed mixture-of-experts model to match GPT-3.5: 46.7B parameters in memory, only 12.9B used per token.

wujec.ai score

–/10

Community score

no votes yet
Sign in to rate

Mixtral 8x7B was announced on 11 December 2023 and it changed what an open model was expected to cost to run. Instead of one dense network, each feed-forward block holds eight separate groups of parameters; for every token, at every layer, a small router picks two of them. The model therefore stores 46.7 billion parameters but spends only 12.9 billion per token — it answers at the speed and the cost of a 13B model while scoring like something far larger. Mistral's own comparison put it at or above Llama 2 70B on most benchmarks with roughly six times faster inference, and at or above the GPT-3.5 base model — the model behind the free tier of ChatGPT at the time. It handled a 32,000-token context, English, French, Italian, German and Spanish, and generated code competently. The licence was again Apache 2.0, which is the part that mattered commercially: a company could match a paid frontier product on its own hardware without a negotiation. That claim is worth reading row by row, because the table is closer than the summary sounds. Of its seven rows Mixtral takes four — MMLU 70.6% against 70.0% for GPT-3.5 and 69.9% for Llama 2 70B, ARC Challenge 85.8%, MBPP 60.7% and GSM-8K 58.4% — and three of the four are won by under 1.4 points; only MBPP is a clear gap, 8.5 points over GPT-3.5. In the remaining three rows Mixtral is behind: HellaSwag goes to Llama 2 70B, 87.1% against 86.7%; WinoGrande goes to both rivals, 83.2% and 81.6% against 81.2%; and MT-Bench — the row for the instruct models — goes to GPT-3.5 by two hundredths, 8.32 against 8.30, the only figure in the whole table printed in bold for GPT-3.5. Mistral's own wording, "comparable to GPT3.5", is exact. Saying the instruct model "scored 8.3" without the neighbouring column, as this profile previously did, is not. The release was also a piece of theatre that the industry copied afterwards. Mistral put the weights out first as a bare magnet link on X, with no paper, no benchmark table and no blog post, and let the open-source community measure the model for three days before publishing its own numbers. Sparse mixture-of-experts designs went on to become the default for large models across the field. Mixtral 8x7B no longer appears in Mistral's API line-up and no shutdown date was ever announced; under Apache 2.0 the weights remain downloadable.

#open weights#Europe#historical#Apache 2.0#mixture of experts
Official website ↗

▸Videos

No videos yet.

▸Reviews

No reviews yet. Be the first!

Sign in to write a review