Google outlines a verifiable federated learning system built around trusted execution environments
Published: 6 October 2026
Google Research has described a new federated learning system that it says is designed to make privacy protections more externally verifiable while moving more computation from participating devices to servers. The system uses trusted execution environments, or TEEs, as the foundation for processing encrypted training data under defined access policies. Google says Gboard has already adopted the approach for English and Japanese next-word prediction models, reporting stronger privacy guarantees, improved accuracy and substantially faster compute times than in its previous federated learning system. The announcement is significant because it focuses not only on limiting data exposure, but also on making the relevant server-side processing available for independent inspection.
Federated learning is presented as a setting in which multiple clients collaborate on a machine-learning task under a service provider’s coordination while retaining control over their data and permitted workloads. Google links its work to four guiding principles: data minimization, anonymization, transparency and control, and verifiability and auditability. Earlier systems used approaches including differential privacy and Secure Aggregation. Yet Google says that, in prior arrangements, outside observers could not verify that uploaded device data was never logged or inspected. Secure Aggregation protected uploads cryptographically, but Google states that it was not compatible with state-of-the-art central differential privacy guarantees. The new architecture is intended to reduce the degree of trust placed in the server operator by making key privacy-relevant processing inspectable.
Under the described design, encrypted data collected from devices can be decrypted and processed only inside TEEs, using Python training programs represented in access policies, and only for a limited period after upload. Workload operators can see metrics and differentially private model weights, according to Google. Participating devices are said to know the complete set of server workloads that could access the data they upload. Those potential workloads are represented by policies published to Rekor, a public transparency log that external auditors can monitor. Google also says that the key-management and data-processing binaries can be reproducibly built from open-source code in its Confidential Federated Compute repository. Together, these elements are meant to allow third parties to examine what server-side workloads may operate on data, rather than relying solely on the operator’s assurances.
The performance claims are tied to a change in where and when training work occurs. Google says the system can collect device uploads before executing server-side training, avoiding training delays caused by changes in device availability and allowing the program to calculate participation schedules and tune other differential-privacy parameters at execution time. In earlier systems, Google says, training could take one to two months and was constrained by device availability, on-device computation and competition among workloads for device resources. Shifting client-gradient computation to the server allows parallelization across many machines, although Google says current speed is limited by available TEE resources. The company also sees the design as a possible route toward federated training of larger models, particularly if TEEs are integrated with accelerators.
The announcement does not claim that all privacy or security questions have been eliminated. Google explicitly qualifies the confidentiality and integrity properties of TEEs as being subject to limitations of current-generation hardware. It also notes that sideloaded serialized information can protect proprietary architectures and preprocessing logic while preserving auditability only so long as privacy-relevant logic remains hardcoded in the published Python program. Side-channel observations remain an area for further work, especially for dynamically loaded workloads and malicious server-side attacks. Google is experimenting with workloads beyond federated training, including synthetic data generation and combinations with TEEs specialized for functions such as LLM inference, but these are described as explorations rather than deployed outcomes. Google expects future hardware and research to offer deeper protections and anticipates that systems like this may one day come with fuller proofs of correctness for differential-privacy software and system components, but presents those as future possibilities rather than present guarantees.
Primary source
https://research.google/blog/toward-provably-private-learning-from-federated-data/ →Source publication date: 2 October 2026
Comments
Loading discussion…
Checking your session…
Subscribe to our newsletter
Get the most important AI news once a week, straight to your inbox.