844-ai.ro
Everything that matters in AI, in one place.
News← Citește în română

Google Research Speeds Up Gemini Nano Models on Pixel with Innovative Multi-Token Prediction Technique

30 July 2026

Google Research has announced a new technical approach designed to boost the processing speed of Gemini Nano language models running locally on Pixel phones. The solution, called "frozen Multi-Token Prediction," marks an important optimization for on-device artificial intelligence that operates without relying on cloud servers.

What is multi-token prediction

Traditionally, language models generate text one token at a time, meaning each word or word fragment is produced sequentially, in a process that demands considerable computing time. The technique unveiled by the Google Research team changes this paradigm, enabling the model to anticipate multiple tokens simultaneously through an additional mechanism attached to the base architecture.

According to Google Research, the distinctive feature of this method lies in the fact that the model's core layers remain "frozen," meaning unchanged, while only the additional components responsible for multi-token prediction are trained separately. This strategy significantly cuts training costs, since there's no need to retrain the entire model from scratch.

Benefits for mobile devices

Applying this technique to Gemini Nano models—designed specifically to run efficiently on limited hardware such as smartphones—brings tangible advantages for Pixel device users. Faster text generation translates into quicker responses for AI-powered features built into the operating system, ranging from notification summaries to conversational assistance.

A key point highlighted by the researchers is that this speed boost doesn't come at a major cost to the quality of generated text. The models continue to produce coherent and relevant responses while retaining the benefits of local processing, which offers enhanced data privacy for users since information no longer needs to be sent to external servers for processing.

Implications for the future of on-device AI

This advancement from Google Research fits into a broader industry trend aimed at shifting more artificial intelligence capabilities directly onto users' devices, reducing dependence on cloud infrastructure. Optimizations like frozen multi-token prediction could become standard for future generations of compact models designed for phones, tablets, and other resource-constrained devices.

It remains to be seen how this technology will be implemented in future versions of Android and the Pixel ecosystem, and whether other makers of compact AI models will adopt this approach as well.

Source

Google Research

844-ai.ro reports based on the source above. Editorially synthesized article, with attribution.

Subscribe to our newsletter

Get the most important AI news once a week, straight to your inbox.