SynthID embeds digital watermarks directly into AI-generated images, audio or video. The watermarks are added across generative AI consumer products from the above companies.
SynthID Detector: Digital Watermarking And Identification Tool Developed By Google DeepMind
SynthID is a digital watermarking and identification system developed by Google DeepMind to establish media provenance and protect information integrity. Google DeepMind researchers Sven Gowal, Rudy Bunel, Florian Stimberg, David Stutz, and Pushmeet Kohli led the development of the technology, with project sponsorship from Demis Hassabis. The system embeds imperceptible signals directly into AI-generated media to enable verification across digital distribution channels. It operates across images, video, audio, text, and synthetic biology. For images, a neural encoder alters pixel values across spatial color bands to embed a hidden pattern directly into the image.
Upload A File For Detection
Drag & drop, or click below to select an image from your device
Detects Media Generated By
Core Detection Capabilities
Watermarks are imperceptible to humans, but can be detected by SynthID's technology. It is robust against common transformations. When uploading files for detection, please upload the highest quality file possible to ensure maximum detection accuracy.
This is not a general AI detector. We can only detect media from the above companies, who have adopted SynthID technology. We are continuously working on extending our partnerships to allow detecting more generated media.
The Google DeepMind research paper: complete analysis of arXiv:2510.09263
In October 2025, DeepMind researchers published SynthID-Image: Image watermarking at internet scale (authored by Sven Gowal, Rudy Bunel, Florian Stimberg, David Stutz, Pushmeet Kohli, and colleagues). The study provides technical documentation on deploying deep learning watermarks across production internet infrastructure.
The core desiderata: quality, effectiveness, robustness, payload, security, and efficiency
The authors defined six requirements for internet-scale watermarking:
- Quality and diversity: The watermark must remain invisible and preserve generation diversity. Ad-hoc systems can restrict output variety, while post-hoc schemes protect diversity.
- Effectiveness and robustness: The detector must achieve high True Positive Rates at an ultra-low False Positive Rate (0.1 percent) across daily image transformations.
- Payload recovery: The watermark must carry multi-bit messages, measured through Bit Accuracy and Code Accuracy.
- Security: The architecture must resist watermark removal, forgery, model extraction, and secret extraction attacks.
- Computational efficiency: Encoding must add minimal latency (single-digit percentage overhead). Decoding must sustain high throughput across batch processing workloads.
- Deployment flexibility: The system must support internal infrastructure, external APIs, and open-model ecosystems.
The step-by-step verification guide
Forensic examiners follow a standard three-step workflow to verify digital assets using the SynthID checking interface.
Preparing And Submitting Content
Examiners prepare media assets to meet minimum signal requirements before scanning: Image assets (maintain native resolution, recommended 512x512 pixels), Text passages (input continuous sample of 80-100 words), and Audio recordings (submit clips of at least 3 seconds).
Running Multi-Pass Verification
The system initiates sequential verification: 1. Metadata pass (reads file headers to locate C2PA provenance, validates digital certificates), and 2. Neural tensor pass (passes the raw pixel or audio array into the convolutional decoder).
Interpreting Forensic Metrics
The report provides quantitative metrics: Conformal p-values, Payload extraction (for images containing multi-bit identifiers, displays recovered bits), and Spatial localization heatmaps (highlights specific pixel regions containing the signature).
Evaluating Invisibility
Quality is the most difficult requirement to evaluate. SynthID heavily relies on human evaluation side-by-side with unwatermarked content.
Empirical Robustness Findings
DeepMind forensics evaluation data regarding watermark survivability across aggressive transformations.
| Watermarking Model | Native Resolution | Payload Bits | Aggregated Random TPR | Aggregated Worst TPR | Combination Worst TPR |
|---|---|---|---|---|---|
| SynthID-O | 512x512 | 136 bits | 99.98% | 99.72% | 98.06% |
| WAM | 256x256 | 32 bits | 90.62% | 83.37% | 55.96% |
| VideoSeal-1.0 | 256x256 | 256 bits | 88.75% | 76.74% | 31.00% |
| TrustMark-Q | 256x256 | 100 bits | 78.61% | 68.73% | 21.22% |
| VideoSeal-0.0 | 256x256 | 96 bits | 78.35% | 61.22% | 27.71% |
| StegaStamp | 400x400 | 100 bits | 72.41% | 66.51% | 27.07% |
| InvisMark | 256x256 | 100 bits | 70.06% | 60.11% | 6.83% |
| TrustMark-P | 256x256 | 100 bits | 66.25% | 54.15% | 4.16% |
Theoretical Threat Models & Core Defenses
To formally evaluate the security guarantees of SynthID, DeepMind researchers modeled the ecosystem against four distinct theoretical attack vectors. These threat models represent the absolute worst-case scenarios where highly motivated, well-resourced adversaries attempt to subvert the watermarking infrastructure.
Watermark Removal
Creating a false negative to obscure content and claim ownership. The adversary mathematically subtracts the low-amplitude watermark signal, often utilizing gradient-descent attacks against a public verification API or applying brutal photometric degradation (such as aggressive Gaussian blurring paired with heavy JPEG compression) until the 136-bit payload falls below the verifiable True Positive Rate threshold.
Watermark Forgery
Creating a false positive to wrongly attribute ownership of a potentially incriminating piece of content to an innocent party. An adversary attempts to extract the cryptographic watermark payload from an AI-generated image and perfectly transplant that localized spatial pattern onto a human-captured photograph, attempting to falsely flag the real photo as a deepfake in forensic systems.
Model Extraction
Stealing the model architecture or localized weights through repeated black-box API querying. By submitting millions of synthetically generated noise matrices to the SynthID detector and observing the conformal p-value output distributions, adversaries aim to train a local "proxy model" that accurately mimics the proprietary Google DeepMind encoder, allowing offline adversarial optimization.
Secret Extraction
Finding the highly secured cryptographic keys utilized for the Pseudorandom Function (PRF) in tournament sampling (SynthID-Text) or the geometric alignment keys (SynthID-Image). If these central server keys are extracted, the adversary gains the ability to generate undetectable "zero-watermark" text directly from the API, or decrypt and read the internal metadata signatures of any Google generative asset.
Robust Training
The neural encoder f and decoder g are not just trained on clean imagery. DeepMind incorporates massive adversarial training loops where the network is forced to decode watermarks through aggressive data augmentations, including simulated JPEG artifacts, 50% random cropping, and affine rotations. This forces the neural network to encode the payload redundantly across robust, low-frequency structural features rather than fragile, high-frequency pixel noise.
Randomness
The encoder injects high-entropy stochastic elements into the embedding process. This ensures that if a user prompts an AI generator 100 times to create identical images of a "red apple", the SynthID system will generate 100 mathematically unique watermark perturbation masks. This stochastic diversity completely prevents "collusion attacks", where adversaries average multiple watermarked images together to identify and subtract the static watermark pattern.
Specificity
The watermark generation is highly content-dependent. The spatial frequency perturbations are mathematically bound to the semantic geometry of the host image (e.g., following the contours of a face or the structural lines of a building). Because the watermark shape is entirely specific to the generated content, an adversary who extracts the pattern and pastes it onto a different image will fail the decoder's geometric correlation check, completely neutralizing forgery attacks.
Content Filtering
The system utilizes dynamic entropy analysis to automatically refuse watermarking on specific corner-case content. For example, applying watermarks to completely flat, uniform color blocks (like a pure white background) would cause visible artifacting. Similarly, in SynthID-Text, generations with zero entropy�such as regurgitated code snippets, famous historical quotes, or rigid mathematical formulas�are bypassed to prevent altering factual correctness or introducing syntax errors.
Can SynthID Be Removed, Stripped, Or Bypassed From Media?
SynthID watermarks resist everyday digital transformations such as social media compression, image format conversion, and cropping. However, no watermarking system is immune to deliberate, destructive adversarial attacks. An adversary can strip an invisible watermark if they are willing to degrade the underlying media. Any attack strong enough to fully remove a SynthID watermark degrades the perceptual or semantic quality of the file to the point where its utility is lost.
Adversarial Image Attacks: Diffusion Re-Generation And Filtering
Attackers use three primary techniques to target image watermarks:
- Diffusion re-generation attacks: An attacker inputs a watermarked image into a generative diffusion model or variational autoencoder (VAE) and re-synthesizes the scene using an img2img pipeline. High denoising parameters wash out high-frequency watermark patterns but alter facial features, textures, and fine details.
- Universal adversarial perturbations: Attackers optimize small, calculated noise masks designed to reduce decoder activation logits.
- DeepMind defenses: DeepMind trains decoders with adversarial training and data augmentations. The encoder applies content-conditioned patterns, binding the watermark to the geometry of each image to prevent pattern extraction.
Adversarial Video Attacks: Temporal Decimation And Transcoding
Video watermark evasion attempts focus on temporal structure:
- Frame rate decimation: Dropping frame rates from 30 fps to 8 fps disrupts motion-vector continuity.
- High-quantization transcoding: Compressing video with Constant Rate Factors above 38 in H.264 or AV1 removes subtle inter-frame signals.
- Spatial letterboxing and aspect-ratio distortion: Forcibly squeezing videos into non-native aspect ratios weakens spatial detector alignment.
Adversarial Audio Attacks: Pitch Shifting And Filtering
Audio evasion attacks target frequency representations:
- Steep notch filtering: Applying frequency cuts across narrow acoustic bands attempts to excise watermark energy.
- Extreme pitch and tempo changes: Shifting audio pitch by more than 25 percent alters spectrogram alignments.
- Acoustic re-recording: Playing audio through distorted speakers into low-grade microphones introduces ambient noise that weakens signal clarity.
Adversarial Text Attacks: Paraphrasing And Prompt Laundering
Text watermarks are vulnerable to linguistic rewording:
- Automated paraphrasers: Running text through rewriting tools replaces watermarked words with synonyms, disrupting the sequence of pseudorandomly keyed tokens.
- Translation round-tripping: Translating text through multiple languages (for example, English to German to Japanese and back to English) replaces vocabulary choices.
- Watermark dilution: Splicing human-written paragraphs into watermarked text lowers the average text statistic T, pulling the score below detection thresholds.
The Security Dilemma: Visual Degradation Versus Watermark Evasion
Digital watermarking operates under a fundamental security balance. An adversary can always erase an imperceptible watermark by applying aggressive filters, heavy compression, or random noise. However, these destructive edits ruin the visual quality and commercial value of the asset. The technical measure of a watermarking system is forcing the adversary's required editing distortion beyond the threshold of acceptable quality.
Deployment Architecture Security
The security of a digital watermark depends heavily on who has access to the decoder. DeepMind evaluates SynthID across three distinct deployment models.