Doubao Seed 2.1 Turbo
ByteDance Seed · China · 2026
Half the price of ByteDance's flagship, and on the maker's own benchmark table it beats that flagship at competition maths and office work.
Doubao Seed 2.1 Turbo is the cheaper of the two models ByteDance released on 23 June 2026 as the Seed 2.1 family, alongside the flagship Seed 2.1 Pro. It carries the same 256,000-token context window, the same multimodal input including video, the same deep-thinking mode and the same tool calling — the difference the buyer sees is the invoice: 3 yuan per million input tokens and 15 yuan per million output tokens on Volcano Engine, exactly half the Pro rate, or roughly 0.41 and 2.07 US dollars. What makes the model interesting is how little that halving costs in measured quality. On ByteDance's own comparison table Turbo trails Pro by around two to four points on most tests — SuperGPQA 67.4 against 70.8, NL2Repo-Bench 43.7 against 47.0, Terminal Bench 2.1 67.6 against 71.0 — and on two of them it is simply ahead: BeyondAIME 88.0 against 87.0 for competition mathematics, and Workspace Bench 54.7 against 53.0 for high-economic-value office tasks. Against the Western models in the same table, Turbo posts higher scores than Claude Opus 4.7 on visual reasoning (MathVision 90.1 versus 83.1), visual knowledge (WorldVQA 48.6 versus 35.9), perception (BabyVision 62.9 versus 22.2) and spatial reasoning (ERQA 71.3 versus 52.5). The gap that does open is agentic autonomy. On Agent Startup Bench, which measures a model running an unattended long chain of work, Turbo scores 54.0 against Pro's 68.8 — a 14.8-point drop, by far the largest in the table, and a fair summary of where the cheaper distillation was allowed to give ground. On debugging (SWE-Atlas 30.6 against 35.2) and on visual puzzles (ZEROBench 11.0 against 18.0) the loss is also visible. Weights are closed and access runs through the Volcano Engine API, with batch inference at half the online rate again (1.5 and 7.5 yuan) and a low-latency tier at double it. ByteDance does not disclose the parameter count, the architecture or the relationship between the two models — whether Turbo is a distillation, a smaller sibling trained alongside, or the same network at reduced compute is not stated anywhere in the public material.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!