News

Alibaba Qianwen releases Qwen-Audio-3

2 min read
Alibaba Qianwen released Qwen-Audio-3.0-ASR-Flash, speech recognition to overcome the "last mile" of professional scenarios. On July 31, Alibaba Tongyi Qianwen officially released the large-scale speech recognition model Qwen-Audio-3.0-ASR-Flash, which focuses on "no loss of words in long audio and no need to teach industry words". It has been opened for use on the Alibaba Cloud Bailian platform. The new model has been upgraded in five major dimensions: First, long audio context memory, you can refer to the previous semantics when transcribing, and the names and technical concepts of several hours of meetings remain consistent throughout, saying goodbye to cross-segment "fragments"; second, the built-in multi-industry vocabulary library, the recall rate of professional words in medical scenarios reaches 95.36%, IT programming reaches 91.87%, and unpopular terms can be accurately identified without manual configuration; The third is hierarchical hot word customization, and enterprise-specific vocabulary takes effect immediately. The recall rate in most scenes exceeds 99%, and there is no false trigger due to the increase of hot words; the fourth is integrated voice polishing, which can complete the removal of catchphrases, clean up repetitions, process self-correction and semantic reorganization in one step, and the output readability is close to the two-step solution of "ASR + large model polishing"; Fifth, a set of models covers more than 30 languages, with an average semantic error rate of 17.09% in seven languages, which is better than peer models such as Azure and Gemini, and is suitable for multinational conferences and overseas customer service scenarios. The Qwen-Audio-3.0-ASR-Streaming launched simultaneously is oriented to real-time scenarios, with a theoretical character output delay of 300 milliseconds, a typo rate of 7.80% in Chinese industrial scenarios, and 11.52% in English, leading the industry. This series has previously ranked first in the world with a typo rate of 1.7% in the Artificial Analysis evaluation, and is widely used in meeting minutes, real-time subtitles, educational recording and broadcasting, intelligent customer service and other fields. via AI News (author: AI Base)