All newsReleases

Two flagship models launched on the same day. One promises to spend more tokens, the other to spend fewer

Published: 9/10/2026 · Source: Google / Meta AI Research / Meta Model API docs

On 2 September 2026 Google released Gemini 3.8 Flash and Meta released Muse Spark 1.3. Both are pitched at the same job — long-horizon coding and autonomous agents — and both were added to our catalogue this week. Read side by side, their launch materials argue for opposite things. Google keeps the price list untouched: USD 0.75 per million input tokens and USD 3.75 per million output, the same introductory rate 3.7 Flash opened, with both figures doubling on 1 January 2027. The context window, the output limit and the three thinking levels are unchanged too. What changed is how hard the model works. Google states it plainly: on complex tasks 3.8 Flash executes extra reasoning steps and calls tools iteratively, and "might use more tokens to maximize performance, especially at higher effort levels". The company then tells cost-sensitive developers to lower the effort level or simply stay on 3.7 Flash, which remains fully supported. Meta's announcement leads with the opposite number. Muse Spark 1.3 is described by Meta's own engineers as significantly faster than Muse Spark 1.2 while using around 20% fewer tool calls and around 25% fewer tokens — fewer turns where they are not needed, less verbose code. The published scorecard backs a real jump in coding (DeepSWE v1.1 rises from 55.0 to 75.4) and a very large one in long-context retrieval (MRCR 512K–1M from 55.5 to 98.1), though in most agent rows the model still finishes behind Anthropic's Opus 5. The practical lesson for anyone budgeting an agent is that the per-token rate has stopped being the whole story. A frozen price list can still produce a bigger invoice if the model decides to think longer, and a model with a higher headline rate can be cheaper if it stops calling tools it does not need. Neither company publishes the figure that would settle it — the average token cost of a completed task — so the comparison has to be run on your own workload. One asterisk on Meta's table: version 1.3 is measured at max reasoning while 1.2 is measured at xhigh, so the two columns are not set to the same effort. A postscript that sharpens the point. Meta put no price in its launch post, and the figures that spread afterwards — USD 0.10 per million input tokens and USD 0.20 per million output, widely reported as the lowest rate on the market — turn out to belong to a separate model id, muse-spark-1.3-contributor. That tier is cheap precisely because you grant Meta permission to train on your prompts and completions. The Standard tier, where Meta states your data is not used for training, costs USD 1.25 and USD 4.25. In other words the cheapest published rate for this model is not a discount on compute at all: it is a price for your data.