MDL-8790EST.2026 · IDX.014
Robotics AIIn production

Tencent Hy-Embodied-VLM-1.0

Tencent · China · 2026

Tencent's open vision-language model for robots: thirty billion parameters of which only three work on any given token, so it can drive a machine that has to answer in real time.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Hy-Embodied-VLM-1.0 is the model a robot uses to look at a scene and decide what to do next. Tencent Robotics X published the weights on 15 July 2026, built on the company's Hy3-A3B language backbone and its Hy-ViT2 vision encoder. The engineering point is the ratio. The model holds about 30.5 billion parameters, but routes each token through only 8 of its 128 experts plus one shared expert — roughly 3 billion parameters actually computed. Tencent's own comparison is with its previous generation, which activated 32 billion to reach nearly the same scores: about one tenth of the working parameters for comparable results. That matters more here than in a chatbot, because a robot that pauses to think has already dropped the cup. What the model is trained to do is narrower than general image understanding. Tencent organises it around three levels: reading the state of the machine and its surroundings, reasoning about what an action would change, and planning across many steps with room to notice a mistake and recover from it. In practice that covers naming where an object sits relative to another, judging whether a grip has succeeded, flagging a risky move before it is made, and re-planning when a step fails. Results are the vendor's: first place on 19 of 38 embodied benchmarks and second on 11 more, ahead of Qwen3.6-A3B by 4.4 percent on average, plus a stated best result on R2R-CE vision-and-language navigation using camera input alone. No independent evaluation of this model was found. The technical report is published as arXiv 2607.12894, verified as belonging to this model and not to a neighbouring number. The licence is plain Apache 2.0. That is worth stating explicitly, because the sibling model HY-Embodied-0.5-X carries Tencent's community licence, which does not apply in the European Union, the United Kingdom or South Korea. Within one product family the company ships both, so the licence has to be read per model. Context is 32,768 tokens — short next to Tencent's translation models, and a reminder that this is a perception-and-action model rather than a document reader. It accepts up to 128 images in a single prompt at their native aspect ratios.

#embodied AI#robotics#vision-language#open weights#Apache 2.0#China
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review