MDL-6170EST.2026 · IDX.562
Robotics AIIn production

Tencent Hy-Embodied-0.5-VLA-UMI

Tencent · China · 2026

Tencent's action model: it does not describe a scene, it outputs the movements of two robot arms. Trained on 10,000 hours of recorded human hands, and released together with a fifth of that recording.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Hy-Embodied-0.5-VLA-UMI is the point in Tencent's robot stack where thinking turns into movement. The model underneath it, HY-Embodied-0.5, looks at a scene and decides what should happen. This one takes that decision and emits the actual trajectory of two robot arms: position, rotation and gripper opening, fifty steps ahead at ten per second. The published file holds about 4.5 billion parameters. Roughly 3.8 billion of that is the borrowed vision-language backbone; the part that is genuinely new here is a 370-million-parameter dual-tower flow-matching transformer that Tencent calls the action expert. Flow matching means the model does not pick the next move from a list of options but shapes a continuous trajectory, which is what a physical arm actually needs. The actions are stored as a change relative to the first frame rather than as joint angles, so the same policy can be moved onto a robot built differently — Tencent reports transfer across four real platforms. What makes this release unusual is the data. Most robot models arrive as weights alone; the recordings that produced them stay inside the company, because collecting them is the expensive part. Tencent trained this checkpoint on more than 10,000 hours of two-handed demonstrations captured with a custom fingertip rig and optical motion capture, and then published 2,163 hours of it — 250,304 episodes, 233 million frames, 18.8 TB, under CC BY 4.0. That is roughly a fifth of the training corpus, and it is a fifth more than the field usually sees. This checkpoint is deliberately not finished. It is the generalist starting point, trained for 200,000 steps on 64 GPUs with a single camera frame and no history, meant to be fine-tuned onto a target robot. The version tuned for a benchmark is a separate model, Hy-Embodied-0.5-VLA-RoboTwin. One detail matters for a European reader. The vision-language backbone this model is built on ships under a Tencent licence that expressly does not apply in the European Union, the United Kingdom or South Korea. This model does not: it is plain Apache 2.0, as is the released dataset. Two files, two different legal positions, in one family. Technical report: arXiv 2606.14409.

#embodied AI#robotics#vision-language-action#open weights#Apache 2.0#China
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review