News
Xiaomi throws depth bombs: 100,000 hours of real data, the first system verification of robot Scaling Law
2 min read
Source: zhidx.com
Wisdom of Things Author | Sanbei Editor | Mo Ying At the beginning of 2026, Huang Renxun predicted that the "ChatGPT moment" of physical AI is coming. However, it is much more difficult for robots to replicate the Scaling Law of large models than imagined. Robot training data needs to be obtained from the real world, and the industry has long been stuck in the "manual workshop" stage of small data, single tasks, and repeated parameter adjustment. Today, Xiaomi has just thrown out a "depth bomb" - the Xiaomi-Robotics-1 embodied base model, in an attempt to change this situation. Xiaomi-Robotics-1 is pre-trained based on 100,000 hours of real-world operation trajectories, and then uses approximately 11,000 hours of cross-body data to complete post-training. It is reported that this is the first time in China that a relatively complete system verification of Scaling Law has been carried out in the robot strategy model. Experimental results show that when the pre-training data is expanded from 2,500 hours to 20,000 hours, the model's action prediction loss on the verification set continues to decrease; when the parameter size is increased from 2 billion to 5 billion or 10 billion, the action prediction ability also steadily improves. The robot's success rate in completing tasks such as shoe storage and school bag packing in unseen home environments has also increased. Embodied intelligence is moving from the 1.0 era, which relied on single-task data and experience parameter adjustment, to the "industrialization" 2.0 era, which is driven by data and model scale. ▲Xiaomi-Robotics-1 floor-standing robot video demonstration project homepage: https://robotics.xiaomi.com/xiaomi-robotics-1.html 1. 100,000 hours of data to verify the robot scaling law. The robot industry has never lacked the consensus that "the more data, the better", but the difficulty is that the data is too expensive and too fragmented. Traditional data mainly comes from real machine remote operation. Operators need to complete tasks such as grabbing, sorting, and transporting in a real environment, as well as handle failed retries and equipment maintenance. This kind of data is not only slow to collect, but also naturally bound to specific machines.