News

Google DeepMind releases SL2T, sign language AI entering consumer-grade mobile phones for the first time Google's DeepMind releases a new multi-lingual sign language-to-text model SL2T and installs it on the Pixel 11 mobile phone system, promoting sign language AI to enter general consumer electronics products for the first time

2 min read
Google DeepMind releases SL2T, sign language AI entering consumer-grade mobile phones for the first time. DeepMind, a subsidiary of Google, releases a new multilingual sign language-to-text model SL2T and installs it on the Pixel 11 mobile phone system, promoting sign language AI to enter ordinary consumer electronics products for the first time. This model is the first to be connected to the Gboard keyboard and Live Transcribe. Users can complete sign language actions through the front camera to edit messages, search web pages and talk to Gemini, replacing traditional keyboard input. SL2T is trained on more than 50 sign languages ​​and a total of 100,000 hours of sign language material, about a quarter of which comes from the American Sign Language data set. Cross-language joint training helps the model learn the common action logic between different sign languages, and refreshes the record of the sign language transliteration model in the FLEURS-ASL evaluation. Unlike traditional solutions that rely on fixed vocabulary labels, SL2T can directly parse human movements to generate text, while processing non-hand semantic information such as facial expressions, body spatial positions, and lip movements. In terms of privacy, MediaPipe Holistic on the mobile phone tracks a total of 130 key points on the hands, face and torso in real time, and only transmits the action coordinates outward. The original video will be deleted immediately and will not be uploaded to the cloud. Currently, this feature only supports American Sign Language to English conversion, and Google plans to expand to more sign language and Android models in the future. DeepMind previously open sourced the sign language translation model SignGemma in May 2025, providing an algorithm basis for the implementation of SL2T. In addition, its barrier-free technologies such as WaveNet, Euphonia and Live Transcribe have also accumulated voice and visual multi-modal experience. As Google, iFlytek, Huawei and Apple continue to deploy, AI barrier-free interaction is accelerating from scientific research verification to consumer-grade products. via AI News (author: AI Base)