Google DeepMind Introduces SL2T Sign Language Translation Model
Key Takeaways
- •Google DeepMind's SL2T model is the first sign language translation technology to move from research prototypes into consumer products, launching initially on Pixel 11 devices.
- •The model was trained on over 100,000 hours of data across more than 50 sign languages and achieved a record 70 BLEURT score on the FLEURS-ASL benchmark.
- •SL2T processes sign language as pose landmark coordinates rather than raw camera feeds, discarding original video immediately to protect user privacy.
- •The technology initially supports American Sign Language to English translation in Gboard and Live Transcribe, with additional languages and devices planned.
- •Google DeepMind established the AI Sign Language Advisory Committee and involved Deaf community members throughout development to ensure responsible deployment.

Google DeepMind Introduces SL2T Sign Language Translation Model
Google DeepMind Sign Language Team
Google DeepMind is introducing sign-language-to-text (SL2T), a new model designed to power sign language features for Deaf and hard of hearing users.
Over recent decades, AI systems have made major advances in processing spoken languages, enabling automatic translation, dictation, and conversational interfaces for hearing users. However, that progress has not extended to the world's more than 200 sign languages or the estimated 70 million Deaf and hard of hearing people who use them.
Google DeepMind said it is now launching a massively multilingual sign-language-to-text translation model that it describes as a breakthrough in quality and generality. For the first time, the technology is moving from the lab into consumer products. While academic researchers and technology companies have explored sign language recognition for years, real-time sign-to-text translation has remained largely confined to research prototypes, making SL2T's deployment in consumer applications a notable milestone in accessibility AI. SL2T powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, beginning with American Sign Language (ASL) to English. More devices are expected soon, and additional languages will follow.
Like spoken dictation for hearing users, the feature allows Deaf users to sign to their phone wherever they would normally type. Users can sign to search the web, draft messages or documents, and ask Gemini to answer queries or carry out tasks. In Live Transcribe, users can sign responses during conversations instead of typing back and forth. According to testers, signing in ASL is faster, more natural, and more delightful than typing in English.
Why sign languages matter
Sign languages are the primary languages of Deaf communities around the world and a central part of Deaf cultural identity. Deaf people vary widely in their proficiency in signing, speaking, reading, and writing, so support across all modalities is important. Sign language processing can benefit Deaf people in the same way spoken language processing benefits hearing users, while also creating new possibilities for communication between Deaf and hearing communities.
Despite that potential, progress in sign language AI has been slow. Google DeepMind said this reflects both the technical difficulty of the work and common misconceptions about how sign languages function.
Compared with spoken-language transcription, sign language translation presents two major challenges. First, speech transcription is a sequential mapping from sound to text in the same language, while sign languages are independent natural languages with their own grammars and vocabularies. That means they require true machine translation, not a simple sign-to-word conversion. Second, the model must learn to see and interpret physical movement. Sign languages use simultaneous movements of the hands, arms, torso, head, and face, which makes accurate tracking at high frame rates a demanding computer vision task.
Google DeepMind said this is why early sign language technologies such as sign language gloves were fundamentally limited. Sign languages are not simply "English on the hands." They require detailed visual perception of fine-grained whole-body movement and full language translation. SL2T is designed to do both.
SL2T treats sign language input as points on the signer's body and translates those inputs into streaming text output. Example from the FLEURS-ASL benchmark.
How SL2T works
Google DeepMind said it built SL2T by combining a user-centric, culturally informed approach with large-scale data training. The model was trained on more than 100,000 hours of data across more than 50 sign languages, with roughly a quarter of the data in ASL. This scale is notable in a field where data scarcity has been a persistent bottleneck, as collecting sign language video is far more resource-intensive than gathering speech or text. Training jointly on diverse languages, dialects, and proficiency levels helped the model learn shared structures and outperform single-language models in the company's experiments.
To protect user privacy, SL2T processes sign language as a sequence of pose landmark locations rather than a raw camera feed. An on-device model, MediaPipe Holistic, tracks the location of points on the signer, and only those geometric coordinates are sent to the server for translation, allowing the original video to be discarded immediately. This design reflects a broader industry shift toward privacy-preserving AI architectures, where sensitive inputs such as camera feeds are processed locally whenever possible to reduce data exposure.
The model translates the coordinate sequence directly into text, bypassing intermediate annotations called glosses that are widely used in earlier sign language translation research. Google DeepMind said glosses do not capture non-linear aspects of sign languages such as non-manual markers and spatial constructions. Translating directly from landmarks removes artificial vocabulary limits and allows translation quality to scale with data.
According to key benchmarks such as FLEURS-ASL (sd-test), which measures ASL-to-English translation quality, SL2T is the most capable sign language translation model to date. Google DeepMind said the model achieved a zero-shot score of 70 BLEURT, higher than any previously reported score.
The company said it also focused on practical deployment issues beyond benchmark performance, including minimizing streaming latency, preventing hallucinations on non-signing inputs, ensuring fairness for the 10% of signers who are left-handed, and improving performance for one-handed signing, which is used when holding a smartphone in the other hand.
Examples from the FLEURS-ASL benchmark show SL2T translating complex ASL into fluent English. Google DeepMind said some errors remain, including rare signs, rapid fingerspelling ("prey" → "grey"), passive constructions, classifier depictions (dropping "claws"), and tense without context ("kicked off" → "start").
Building with the community
Google DeepMind said it believes in building with the Deaf community, not just for it. Deaf perspectives influenced every stage of the project, from conceptualization by Sam Sepah, a Deaf Googler, to data collection with Deaf partners, evaluation in Deaf user studies, and impact assessment with Deaf experts.
To support responsible real-world deployment, the company established the AI Sign Language Advisory Committee (AISLAC), which brings together global Deaf organizations and subject-matter experts. Through this participatory governance model, the communities most affected by the technology can directly shape development priorities.
Google DeepMind also co-authored a joint impact report for the release of SL2T 1.0 in Gboard and Live Transcribe, describing the technology's capabilities and current limitations. The company said it plans to continue that approach for all major sign language releases.
Looking ahead
SL2T builds on decades of foundational research across academia and industry, but Google DeepMind said bringing ASL input to users' phones is only the beginning. The company said its mission is to organize the world's information and make it universally accessible and useful, and that universal accessibility requires parity with spoken and written languages.
The team said it is working to expand the technology to additional sign languages, sign language generation, and frontier AI capabilities. Google DeepMind said it aims to share progress responsibly so that access through sign languages becomes standard across the digital landscape.
The Pixel-first rollout follows Google's established pattern of debuting AI-powered features on its flagship hardware before extending them to the wider Android ecosystem. Users can try SL2T in Gboard and Live Transcribe first on Pixel 11, with more devices coming soon, at no additional cost.
Acknowledgements
The work was carried out jointly by teams from Google DeepMind and Android. The core team that developed the SL2T model includes Garrett Tanzer, Benoit Brard, Elizabeth Clark, Tim Dozat, Sebastian Ebert, Dan Garrette, Manfred Georg, Vicky Holgate, Shankar Kumar, Mohammad Saboorian, Miloš Stanojević, Megh Umekar, John Wieting, Andy Zhang, and Chris Dyer.
The Android team that integrated the model into Gboard and Live Transcribe includes Ausmus Chang, Sai Aditya Chitturu, Dayle Chiu, Anna Chou, Ajay Dudani, Angana Ghosh, Alex Huang, Joanne Kim, Ed Lee, Thomas Lin, James Su, Yanchao Su, and Sharlene Yuan.
Additional support came from Anelia Angelova, Abhishek Bapna, Sara Basson, Glenn Cameron, Scott Crowell, Trevor Cohn, Noah Fiedel, Zoubin Ghahramani, Raia Hadsell, Tom Hudson, Alexander Hauerslev Jensen, Kazuya Kawakami, Peike Li, Liam McCafferty, Caroline Pantofaru, Abhinav Parashar, Christopher Patnoe, Laura Rimell, Sam Sepah, Thad Starner, Dave Uthus, and Biao Zhang.
The team also thanked those who participated in early-stage testing of the models.