Google introduces sign-language-to-text AI model in Pixel 11: How it works and what are the challenges – CNBC TV18

Google introduces sign-language-to-text AI model in Pixel 11: How it works and what are the challenges

The technology will power a new sign-to-text feature in Gboard and Live Transcribe on the Pixel 11 family.

By CNBCTV18.com  August 13, 2026, 11:29:43 AM IST (Updated)

Google introduces sign-language-to-text AI model in Pixel 11: How it works and what are the challenges
Google is bringing sign-language translation to its latest Pixel phones, giving hearing-impaired people a new way to communicate with apps and other people without typing.

At its Made by Google event on Wednesday, Google DeepMind introduced sign-language-to-text (SL2T), a multilingual AI model designed to translate sign language directly into written text. The technology will power a new sign-to-text feature in Gboard and Live Transcribe on the Pixel 11 family.

The feature uses a phone’s camera to capture signing, allowing users to sign wherever they would normally type. That includes searching the web, drafting text messages and emails, writing documents, or asking Google’s Gemini AI questions. In Live Transcribe, users can sign their responses during face-to-face conversations rather than typing them.


The rollout begins with American Sign Language (ASL) translated into English, with Google describing it as a sign-language equivalent of voice dictation.

Why sign-language AI is difficult

Google says sign-language translation presents challenges fundamentally different from speech recognition. Sign languages are independent natural languages with their own grammars and vocabulary, meaning they cannot simply be converted into spoken-language words one sign at a time. Secondly, the model must learn to ‘see’ and understand physical movement.

Meaning is also conveyed through coordinated movements involving the hands, arms, torso, head and face. Accurately tracking these at high frame rates makes sign-language processing a demanding computer-vision and machine-translation problem.

These challenges have limited earlier approaches, including sign-language gloves, which attempted to translate hand movements without capturing the full linguistic and visual context of signing.

SL2T is designed to address both sides of the problem: understanding complex physical movements and translating them as a language.

How Google’s SL2T model works

Trained on more than 100,000 hours of data spanning over 50 sign languages, SL2T has about a quarter of the training data coming from ASL. Training across multiple languages, dialects and levels of signing proficiency helped the model learn structures shared across sign languages, according to the company.

Privacy is another key part of the system. Rather than sending raw camera footage to a server, an on-device MediaPipe Holistic model tracks the location of points on the signer and converts it into geometric landmark coordinates. Those coordinates are then sent for translation, while the original video can be discarded immediately.

The model also translates directly from the landmark sequence into text instead of relying on earlier intermediate system of annotations known as ‘glosses.’ Google says gloss-based systems can miss elements such as facial expressions, spatial constructions and other non-manual markers that carry meaning in sign languages.

On the FLEURS-ASL benchmark, Google says SL2T is the most capable sign language translation model to date and has achieved a 70 BLEURT zero-shot score, which it describes as substantially higher than previously reported results.

Beyond benchmark performance, Google says it has optimized the model for real-world use, including reducing streaming latency, avoiding hallucinations when users are not signing, supporting left-handed signers, and improving one-handed signing for people holding the phone in their other hand.


Original source: https://www.cnbctv18.com/technology/

Leave a Reply

Your email address will not be published. Required fields are marked *