News
Dialogue with Om AI Zhao Tiancheng: Perseverance for many years, betting on the "streaming" future native to physical AI
2 min read
Source: 36氪
A multi-modal model that has never seen surveillance footage understands surveillance better than a small model "veteran" who has been practicing on surveillance data for many years. This is not a science fiction movie. This is an "unintentional intervention" by Om AI in 2023. It is also a key node for CEO and chief scientist Dr. Zhao Tiancheng to become more convinced that "multimodal training methods can bring generalization to the open world of physics." At that time, the AI industry was pursuing generative AI with large language models as the core. Three years later, this multi-modal model evolved into VLX, the world’s first end-side flow multi-modal model series for physics AI. For the first time, a new model architecture of "end-side native streaming multi-modality" was proposed. Through this architecture designed from Day 1 for the end-side computing power constraints, the VLX model opened up a complete closed loop of physical AI of "continuous perception + precise positioning + action decision-making" on the end side for the first time. VLX Overview Chart There is no doubt that as of the summer of 2026, the interest in the AI industry has begun to shift from digital AI to physical AI. Different from digital AI in 2023, although the popularity of physical AI is very high, the industry has not yet converged. Multiple routes coexist and are still far away from production: language-centered VLA, pixel-centered video generation, 3D structure-centered simulation, visual representation-centered JEPA... The industry's choice of the two most mainstream conceptual routes of VLA and world model is still wavering. Is it replacement, coexistence, or fusion? While the industry is still debating who is the ultimate answer to physical AI, Om AI's VLX series models have completed the commercial closed loop of physical AI from simulation experiments to industrial implementation. By giving robots, drones, wearable devices, security cameras, AI PCs and other types of physical terminals the "cerebellum" for autonomous perception and the "brain" for cognitive decision-making, these physical terminals can achieve a leap from "passive execution of instructions" to "active adaptation to scenarios". VLX Architecture Diagram This is not an accident, but the inevitable result of Om AI Lianhui’s long years of focus. In 2019, when Zhao Tiancheng graduated from the CMU Institute of Language Technology with a Ph.D., his resume was dazzling enough: he worked in the field