NewsStocksGoogle DeepMind Unveils SynthID Bio to Watermark AI-Generated Proteins

Google DeepMind Unveils SynthID Bio to Watermark AI-Generated Proteins

Author: Google DeepMind Blog·

Key Takeaways

  • •SynthID Bio embeds imperceptible watermarks into the amino acid sequences and predicted 3D structures of AI-generated proteins, and the signatures remain verifiable in the synthesized physical proteins without compromising their biological function.
  • •In wet-lab tests against VEGF-A, the SARS-CoV-2 spike protein RBD, and PD-L1, watermarked protein binders matched unwatermarked versions in hit rate, binding affinity, and sequence diversity, producing the first biologically functional watermarked binders.
  • •For protein folding, the watermarking capability is built directly into AlphaFold 3's model weights, maintaining prediction accuracy while providing near-perfect detectability and resilience to digital noise or minor coordinate changes.
  • •The system is designed to strengthen biosecurity by giving DNA synthesis providers an automated signal that an order originated from a trusted model, and to help properly label or flag synthetic entries in databases such as the Protein Data Bank, UniProt, and GenBank.
  • •In collaboration with the Hie lab at Stanford University and the Arc Institute, DeepMind integrated SynthID Bio into the Evo 2 genomic model to watermark a designed bacteriophage genome, and early laboratory testing in bacteria cultures confirmed the watermarked phages are functional.
Google DeepMind Unveils SynthID Bio to Watermark AI-Generated Proteins

Google DeepMind has introduced SynthID Bio, a proof-of-concept effort to bring watermarking technology to synthetic biology. Announced in a blog post published on September 30, the system embeds an imperceptible signature directly into the biological code of AI-generated proteins, ensuring the watermark is verifiable not only on the digital model but also on the synthesized, physical protein itself — all while preserving its biological function in laboratory testing.

The work, authored by Pushmeet Kohli, David Stutz, Ali Cowen-Rivers and Jeremy Ratcliff, arrives as generative AI increasingly helps scientists address critical biological challenges: predicting the structure of proteins with AlphaFold, designing entirely new proteins with AlphaProteo and ProteinMPNN, and, more recently, developing new bacteriophages — viruses that infect bacteria. Yet these tools also present new challenges. Novel AI designs can bypass traditional DNA synthesis screening, while mislabeled synthetic 3D structures risk polluting public databases and misleading downstream research.

How SynthID Bio Works

SynthID Bio is a family of watermarking methods developed specifically for synthetic biology to strengthen biosecurity and scientific integrity. The approach adapts to the type of data involved: for sequences, it subtly guides the choice of amino acids; for predicted 3D structures, it adjusts atomic coordinates. In both cases, the modifications create a reliable signal for detection.

In experiments, these adjustments did not compromise the protein's biological function — a property Google DeepMind described as essential to effectively treating disease and advancing scientific research.

The team verified its approach for watermarking protein binders — molecules built to selectively latch onto other proteins — by combining its binder design method AlphaProteo with a SynthID Bio-enabled version of ProteinMPNN, a commonly used protein sequence generation method. In wet-lab testing across three target proteins — VEGF-A, the SARS-CoV-2 spike protein RBD, and PD-L1 — the watermarked designs matched the hit rate, binding affinity and natural sequence diversity of unwatermarked versions, successfully creating the first-ever watermarked and biologically functional protein binders. Binding affinity was measured as KD, with lower values indicating stronger binders.

For protein folding, SynthID Bio fine-tunes a small part of AlphaFold 3's diffusion network, building the ability to watermark directly into the model's weights. This ensures that predicted 3D coordinates inherently carry a detectable signature regardless of who runs the model. According to the company, the technique preserves AlphaFold 3's prediction accuracy while offering near-perfect detectability, maintaining key structural feature distributions and holding up against digital noise or minor coordinate changes. The team illustrated the results on the protein 7PPA, presenting the AlphaFold 3 predicted structure alongside the ground-truth structure and the watermarked structure.

Strengthening Biosecurity and Information Integrity

Biosecurity relies on layered defenses — a “Swiss cheese” model in which multiple independent safety measures work in tandem to cover one another's blind spots. Safeguards such as model-level mitigations and customer vetting each represent critical layers with potential gaps. As part of Google DeepMind's broader vision for bioresilience, SynthID Bio's watermarking approach serves as an important, tangible verification layer embedded in the biological design itself.

SynthID Bio is an important piece of the puzzle for tracking the provenance of biological designs,” said Sarah Carter, a biosecurity policy expert and Principal at Science Policy Consulting, who reviewed the work. “By linking designs to the model developer, these watermarks empower developers to lead on safety and allow synthesis providers to streamline screening for customers who have used those models.”

That layer is especially vital to DNA synthesis screening, which sits on the frontlines of biosecurity. Converting digital protein designs into physical molecules requires placing an order with DNA synthesis providers, who screen requests against databases of known threats. Historically, an unfamiliar sequence could be safely assumed to be an undiscovered natural organism. But because AI can create entirely new sequences with little resemblance to known hazards, screeners can no longer make that assumption. Verifying that an unfamiliar order is not an engineered threat requires exhaustive manual reviews that can stall vital research. SynthID Bio can provide an automated verification signal in this setting, proving an order originated from a trusted model with built-in safeguards.

“AI is expanding what scientists can design, and DNA synthesis companies have an important role in helping that innovation scale responsibly,” said James Diggans, Vice President, Policy and Biosecurity at Twist Bioscience, who provided early feedback on the paper. “For Twist, watermarking offers a promising new addition to the biosecurity toolbox that could strengthen screening, focus resources on sequences that warrant closer review and make biosecurity more efficient as AI-designed biology continues to advance.”

SynthID Bio could similarly help maintain the integrity of databases such as the Protein Data Bank, UniProt and GenBank. These databases, many of which are open to public submission, play a vital role in scientific advancement — but mislabeled entries can have an outsized negative impact on biosecurity decision-making, a challenge that may only grow with the addition of AI-generated biological data. As part of the submission process, SynthID Bio could help ensure synthetic entries are properly labeled or flagged for further review.

Looking Ahead

While no single biosecurity intervention is a silver bullet, SynthID Bio brings SynthID — Google DeepMind's tried and tested watermarking tool — to synthetic biology, marking an important first step toward reliably identifying and tracking AI-generated biological sequences and structures.

Moving forward, key challenges include making the watermark more robust against deliberate tampering. SynthID Bio can also be paired with provenance metadata approaches — similar to C2PA for digital media — or with central repositories of AI-generated biological data to better identify and track AI-generated proteins.

To match the growing capabilities of frontier AI technology, Google DeepMind is also researching how to apply watermarking to more complex biological objects. In ongoing work with the Hie lab at Stanford University and the Arc Institute, the team integrated SynthID Bio into Evo 2, an advanced genomic model, to watermark the genome of an Evo 2-designed bacteriophage. Early laboratory testing in bacteria cultures has confirmed these watermarked bacteriophages are functional. Google DeepMind said the work has the potential to address some of the biosecurity risks associated with genome design and that further details will be shared in a technical manuscript soon.

Realizing the full biosecurity benefits of this work will require community collaboration and further research. As part of its commitment to responsible innovation, the company is publishing its methods paper, open-sourcing the code and in vitro data, and releasing the weights to the research community. Google DeepMind said it aims to work openly with partners across biosecurity, gene synthesis and policy to ensure safety and responsibility keep pace with AI-driven discovery. Parties interested in partnering on the topic are asked to contact the team at synthidbio@google.com with a high-level proposal, and not to disclose confidential or proprietary information.

Acknowledgements

The project was initiated by Pushmeet Kohli. Research and technical development were led by Alexander I. Cowen-Rivers and David Stutz, with Kohli as adviser. Key engineering and research contributions were made by Guillermo Ortiz-Jimenez, Jeremy Ratcliff, Vinicius Zambaldi, Lindsay Willmore, Josh Abramson, Harshnira Patani, Christina Kouridi, Florian Stimberg, Mel Vecerik, Alex Chu, Sukhdeep Singh, Sumanth Dathathri, Eliseo Papa, Valentin De Bortoli, Arnaud Doucet, Jue Wang and Sven Gowal. The team thanked Adaptyv Bio for help with in vitro validation.

The extension of the work to watermarking bacteriophage DNA is a collaboration between Google DeepMind and the Hie lab at Stanford and the Arc Institute, with key contributions from Jeremy Ratcliff, Aleks Petrov, Alexander I. Cowen-Rivers, David Stutz, Elisa L. H. Wong, Victor Martin Palacios, Francesca Pietra, Alfred Piccioni, Tristan Oliver Kwan, Tor Lattimore, Sumanth Dathathri and Pushmeet Kohli of Google DeepMind, and Brian Hie, Samuel King and Aditi Merchant of the Arc Institute and Stanford.

Google DeepMind also thanked Rudy Bunel, Anna Cupani, Rob Fergus, Thomas Frerix, Sahra Ghalebikesabi, John Jumper, Jacob Kelly, David La, Victor Martin, Sebastian Nowozin, Stig Petersen, Aleks Petrov, Uchechi Okereke, Sylvestre-Alvise Rebuffi, Rosalia Schneider, Armin Senoner, Richard Shuai, Ashok Thillaisundaram, Elisa L. H. Wong, Zachary Wu and Augustin Žídek for contributions to the research article, and Julien Bergeron for contributions to the 3D rendering. The team closed by thanking Demis Hassabis for his encouragement and support of the project.