Qwen3.8-27B
Alibaba Cloud · China · 2026
Alibaba's compact open-weights flagship: a 27-billion-parameter dense vision-language model under Apache 2.0, with a 262k-token context that stretches to a million.
Qwen3.8-27B is the open-weights half of the Qwen3.8 generation, published on 14 August 2026 alongside the much larger closed flagship Qwen3.8-Max. Where Max is a 2.4-trillion-parameter mixture-of-experts served only through Alibaba's cloud, the 27B model is a dense checkpoint of 27.78 billion parameters that anyone can download from Hugging Face or ModelScope and run on their own hardware under the Apache 2.0 licence. It is a native vision-language model: images and video go in alongside text, from STEM diagrams and scanned documents to hour-scale footage. The context window is 262,144 tokens natively and, according to Alibaba, extensible to one million; a hosted version with the million-token window enabled by default is announced for Qwen Cloud but was not yet live at publication. Thinking mode is on by default and can be switched off per request, with reasoning depth tuned through a `reasoning_effort` setting and reasoning kept across turns via `preserve_thinking`. The architecture is unusual for a dense model of this size: 64 layers arranged as sixteen repetitions of three Gated DeltaNet blocks followed by one gated-attention block — a hybrid of linear and full attention that keeps memory use down over long inputs. Hidden dimension is 5,120, the vocabulary 248,320 tokens, and the model was trained with multi-token prediction. Alibaba's own benchmark table is the reason this release drew attention. On SWE-bench Pro the company reports 61.7 for the 27B model against 57.6 for its own closed Qwen3.7-Plus, and on DeepSWE 1.1 it reports 42.2 against 13.3 for the previous 27B generation. On Terminal Bench 2.1 it reports 73.0, below Opus 4.6 Max at 78.2 but well above the 63.4 of Qwen3.6-27B. All of these are vendor-reported figures and have not been independently reproduced; what is not in dispute is that a model small enough to fit on a single high-end accelerator now carries a specification sheet that would have described a frontier system eighteen months ago. The footnotes under that table matter as much as the numbers. Every baseline in the coding block was re-run by Alibaba itself inside the Claude Code harness, and the SWE-bench Pro task set was corrected before the re-run, so most of the comparison is between the vendor's own executions rather than published results; the single exception is Opus 4.6 Max, which keeps its officially reported SWE-bench Pro score of 53.4. QwenSWEBench and CoWorkBench are in-house benchmarks with no external definition. HLE is graded by GPT-4o. The multimodal block leans the other way: on MathVision Alibaba fixes one prompt for its own model and reports the better of two prompt variants for every competitor, and the model still loses that row to its own Qwen3.7-Plus, 90.0 to 90.3. Two more rows go against it in the same direction: Humanity's Last Exam 30.8 against 34.7, and GPQA Diamond 89.2 against 90.3, both to the closed in-house flagship of the previous generation.
▸News
Alibaba gives its rivals the better of two prompts and keeps one for itself — then loses that row to its own previous flagship
9/13/2026 · Qwen3.8-27B model card, Hugging Face
Alibaba opens both ends of Qwen3.8 — but only the small model gets Apache 2.0
8/15/2026 · Qwen — karty modeli i plik licencji na Hugging Face
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!