News
Zhiyuan’s full-mode large model ranks first in the world! Overwhelm Google and Nvidia, outperform major manufacturers
2 min read
Source: zhidx.com
Intelligent Stuff Author | Jiang Yu Editor | Mo Ying When a robot stands in an exhibition hall where many people are talking, can it hear who is talking to it from the noise and determine whether a sentence requires a response? Can it take over when the other person pauses and continue naturally after being suddenly interrupted? Can it still answer and make appropriate gestures and expressions to match the tone? These seemingly ordinary communication details are the core capabilities of embodied intelligence and are also the most difficult part for robots to complete when they cross the threshold of natural interaction. For a long time, the robot industry has been trapped in an embarrassing paradox: continuous breakthroughs in individual abilities, and the integration of "seeing, listening, speaking, and moving" into a coherent, human-like whole, has always been short of breath. The essence of this gap is that the interaction ability cannot keep up - perception, reasoning, voice, and movement are all running separately, and they have never been coordinated on the same timeline. Recently, Wisdom's self-developed full-modal large model WITA-Omni Preview gave a weighty answer. On DailyOmni, the industry-recognized authoritative list of embodied full-modal understanding, it topped the list with a comprehensive score of 85.21 points, surpassing leading domestic and foreign models such as Qianwen, Gemini, Doubao, and NVIDIA, and ranked first in six out of eight subdivision indicators. On the main line of core interactive capabilities of "understanding the environment, understanding sounds, listening and thinking, and speaking and moving", a company focusing on embodied intelligence has moved ahead of general large model manufacturers. Previously, Genie Envisioner-Sim 2.0, a self-developed world model developed by Zhiyuan, has won the first place in the WorldArena "World Model Perception and Action Response" track. This time WITA-Omni has topped the all-modal understanding list. The two results correspond to the robot's key capabilities of understanding the physical world, simulating the physical world, and interacting with the physical world. This has also allowed the outside world to observe that in addition to building the robot itself, Zhiyuan is simultaneously building a self-developed AI landscape that integrates "interaction-operation-movement". 1. A big test of "sound and picture counterpoint": How can DailyOmni test its true full-modal capabilities?