A new paper from researchers including Saifur Rahman Tamim delivers a uncomfortable truth bomb to anyone who's been cheering for mandated AI watermarks: the technology isn't ready for courtroom prime time. The study, posted to arXiv on July 17, 2026 (abs/2607.16010), puts three major watermarking schemes through rigorous forensic testing—and every single one fails to meet evidentiary standards that courts require.

Why This Matters Now

Governments are betting big on watermarks as a transparency solution. The EU AI Act demands markings that are "sufficiently reliable and robust." California's SB 942 requires disclosure that's "permanent or extraordinarily difficult to remove." These mandates assume watermark detection produces evidence admissible in court—but that assumption has never been empirically tested. Until now. The researchers evaluated three representative methods: KGW, Unigram, and the MarkLLM implementation of Google's SynthID-Text. They structured their analysis around a custom Forensic Readiness Score (FRS) framework with 12 criteria, three mandatory gates, and a 60-point scoring system. The attack vector? Meaning-preserving paraphrase—which is both legally realistic and hard to dismiss as evidence tampering.

The Numbers Are Brutal

Here's where it gets ugly for watermark proponents. Across 846 valid paraphrase runs using 15 diverse prompts per method, every initially-detected KGW and Unigram text lost its watermark after paraphrasing—100% conditional removal rate. SynthID fared only "slightly better" at 98.3%. These aren't edge cases; this is systematic failure under basic adversarial conditions. But wait, there's more. Even without any attack, false-negative rates were already alarmingly high: 70% for KGW, 83% for Unigram, and 80% for SynthID. The SynthID configuration also falsely flagged 5.4% of paraphrased human-written controls as AI-generated—a significant false positive problem. Perhaps most damning, SynthID showed an 18.6% paradox rate, with 80% of its own pristine watermarked output landing in the uncertainty deadband where detection is inconclusive.

Legal Standards and Technical Reality

The researchers tested against the Daubert admissibility criteria—the legal standard US courts use to evaluate expert testimony and scientific evidence. None of the three methods satisfy more than two of five Daubert factors. When your watermark scheme can't clear a federal evidentiary bar, "mandatory disclosure" starts looking like regulatory theater.

Key Takeaways

  • All tested watermarks (KGW, Unigram, SynthID) fail forensic readiness standards—100% removal under paraphrase attack for KGW and Unigram
  • False-negative rates range from 70-83% even without any adversarial manipulation
  • Zero methods satisfy more than two of five Daubert legal admissibility factors
  • California's SB 942 and EU AI Act mandates rest on assumptions that this data contradicts

The Bottom Line

This research confirms what anyone who's spent time in security knows: fragile detection mechanisms fail under real-world adversarial conditions. Regulators pushing watermarks as a transparency silver bullet are building policy on sand. If courts can't trust watermark evidence, mandatory disclosure becomes compliance theater—useful for appearances, worthless for accountability.