Google DeepMind has developed a watermarking system designed to identify proteins created by artificial intelligence, as AI models become increasingly capable of generating biological sequences that may not exist in nature.
The system, called Protein Watermarking, is intended to embed identifiable patterns into AI-generated protein sequences without significantly changing their predicted structure or function. The approach could help researchers determine whether a protein sequence originated from an AI system while preserving its potential scientific utility.
Unlike watermarks used for AI-generated images or text, identifying synthetic biological designs presents a different challenge. Proteins are built from sequences of amino acids, and changes to those sequences can affect how a protein folds and functions. Any watermark therefore needs to remain detectable without disrupting the biological characteristics researchers are trying to design.
DeepMind researchers developed the method to introduce statistical signals into generated amino acid sequences. These signals can later be detected computationally, providing a way to distinguish watermarked sequences from naturally occurring proteins or designs produced without the system.
The research comes as generative AI is increasingly being applied to biology and drug discovery. AI models can now help researchers predict protein structures and design new proteins for potential applications ranging from medicine and biotechnology to industrial processes.
Google DeepMind has played a major role in this field through AlphaFold, its AI system for predicting protein structures. The expansion from predicting existing biological structures to generating new ones has also raised questions around provenance, transparency and the ability to trace AI-created biological material.
A watermarking mechanism could provide researchers, laboratories and other organisations with an additional method for identifying the origin of computationally generated protein designs. It could also support future efforts to establish provenance standards as AI-designed biological sequences become more widely used.
However, watermarking is not the same as preventing the creation or misuse of a biological design. Its role is primarily centred on identification and traceability. The effectiveness of such systems can also depend on whether watermarks remain detectable after sequences are modified or processed further.
The work reflects a broader push across the AI industry to develop methods for identifying machine-generated outputs. Technology companies have already explored watermarking and provenance systems for text, images, audio and video as synthetic content becomes more difficult to distinguish from human-created material.
DeepMind's research extends that discussion into computational biology, where the output is not simply digital content but a biological sequence that could potentially be synthesised and studied in the physical world.
Disclaimer: This article may include information derived from interviews, press releases, public statements, research, company communications and other publicly available or third-party sources. Such material may be summarised, paraphrased or contextualised for journalistic and editorial purposes. All rights in third-party content remain with their respective owners.