Tencent Hy4 preview
Tencent · China · 2026
Tencent's fourth-generation flagship: 770B parameters under Apache 2.0, and a vendor chart on which every rival still scores higher.
Hy4 preview is Tencent's fourth-generation flagship language model and the largest set of weights the company has ever published. It went up on Hugging Face, ModelScope, GitCode and CNB on 27 August 2026 under the Apache 2.0 licence, in a full BF16 edition and an FP8 quantised edition, and appeared on OpenRouter the following morning. Tencent had confirmed during its second-quarter earnings call on 12 August 2026 that Hy4 was still in training; two weeks later the weights were public. The model is a sparse mixture of experts with 770B total parameters and 49B activated per token. The backbone has 78 layers: the first uses a dense feed-forward block, the remaining 77 are MoE layers with 256 routed experts and one shared expert each, and every token activates the top eight routed experts plus the shared one. A native multi-token-prediction layer of 10B parameters (0.7B activated) is built in for speculative decoding, which brings the published weight file to 780B parameters — roughly 1.6 TB in BF16. Attention uses gated DeepSeek Sparse Attention with cross-layer index reuse, and the residual path uses identity hyper-connections with four residual streams. The context window is 1 million tokens, with output capped at 64,000. Tencent presents Hy4 preview as a productivity model rather than a benchmark champion: it says the training data was built around the work of its own software engineers, game developers, financial analysts and security staff, and that the model is co-designed with the CodeBuddy and WorkBuddy products. In a blind internal comparison, 163 Tencent experts scored 203 engineering tasks and rated Hy4 preview 2.99 against 2.92 for GLM 5.3 and 2.94 for Kimi K3 — a narrow lead, with roughly four in ten tasks lost in each pairing. The published benchmark chart is worth reading carefully, because it does not show a leader. In all twelve tests Tencent chose, the best score belongs to a rival — usually Claude Opus 5 or GPT-5.6 Sol. Hy4 preview scores 85.4 on Terminal Bench 2.1 against 88.3 for GLM 5.3 and GPT-5.6 Sol, 43.4 on the text-only Humanity's Last Exam against 53.2 for Claude Opus 5, and 17.5 on ProgramBench against 39.5 for Claude Opus 5. What the chart does show is the size of the generational step: on the same three tests Hy3 scored 70.8, 34.4 and 3.0. Tencent calls this an early version, admits the model spends longer reasoning than it needs to and over-verifies its own work, and says it would rather ship early than wait — the same approach it took with Hy3 preview. Unlike the multimodal successor trailed in press reports, this preview takes text in and returns text out.
▸News
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!