Pelican1.0-VL-72B
Beijing Innovation Center of Humanoid Robotics (X-Humanoid) · China · 2025
A 73-billion-parameter vision-language model meant to serve as a robot brain rather than a chatbot: it reads a scene, works out what has to be moved first and produces the step-by-step actions.
Pelican1.0-VL-72B is the flagship of Pelican-VL 1.0, a family of open-weight embodied brain models published by the Beijing Innovation Center of Humanoid Robotics — the state-backed centre that also builds the Tiangong humanoids and trades internationally as X-Humanoid. The 7B and 72B checkpoints went up on 13 November 2025, a 3B followed on 1 December 2025 and a 235-billion-parameter mixture-of-experts build with function calling on 7 February 2026. All of them are Apache 2.0. The model is not a general assistant. It takes images and video alongside text and answers the questions a robot has to answer before it moves: where the object is in space, which item blocks which, what a plausible grasp point looks like, and in what order a multi-step task should be carried out. The 72B build carries 73.41 billion parameters in bfloat16, uses the Qwen2.5-VL architecture as its base and accepts a 128,000-token context, so a long video sequence fits in one pass. The training method is what the authors put forward as the contribution. They call it DPPO, Deliberate Practice Policy Optimization: instead of a single fine-tuning pass, the model runs a loop of reinforcement learning, refinement, diagnosis and supervised fine-tuning, in which each cycle generates new hard examples aimed at whatever the model just got wrong. The raw corpus behind it held over 4 billion tokens before distillation, and the reported cost is a cluster of more than 1,000 A800 GPUs and over 50,000 A800 GPU-hours per checkpoint. The headline results come from the team's own report and have not been independently reproduced: a 20.3% uplift over the base model, and 10.6% above open-source models in the 100-billion-parameter class on established embodied benchmarks. Taken at face value they would place the model level with proprietary systems on this narrow class of tasks. The practical caveat is adoption. Despite the Apache 2.0 licence and the size record the family claims — the largest open-source embodied multimodal brain model at the time of release — the official checkpoints attract only tens of downloads a month on Hugging Face, while third-party quantised repackagings of the same weights draw more traffic than the originals.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!