OpenAI’s October 5 announcement of textGrain watermarking is the latest reminder that AI provenance is a data modeling problem, not just a detection problem. The feature injects a statistical signal into eligible model output, with API customers able to opt in for select models and EU ChatGPT/Codex users receiving it over the coming weeks. Crucially, detector access is currently limited to approved researchers and expert organizations, not a public endpoint. This means most builders must design workflows that function without direct detector access, treating provenance as one observation among many rather than a definitive verdict.

Separate Capture, Signal, and Decision

Harshith Vaddiparthy argues that collapsing these three stages into a boolean written_by_ai field destroys necessary uncertainty. A robust architecture requires distinct records: capture logs (version hashes, author declarations, permitted AI policies), signal records (what a check actually observed, including not_run or inconclusive states), and decision logs (human review outcomes and rationale). For instance, a watermark result of not_detected is a statement about the check’s failure to find a signal, not proof of human authorship. Your data schema must explicitly support states like unknown, unsupported, and pending to prevent downstream dashboards from mislabeling missing evidence as human-written content.

Make the Limits Visible in the Interface

OpenAI explicitly states that textGrain cannot identify users, measure human contribution, or verify accuracy. The company’s own evaluation data shows significant fragility: for 400-token English passages, replacing just 10% of words with synonyms dropped detection rates from approximately 92% to 66%, while a 25% replacement rate plummeted detection to 17%. These are company-reported results under specific conditions, not universal benchmarks. Interfaces must display these constraints alongside results. A simple red/green badge encourages overconfident decisions; instead, show the input version, eligibility status, and method version. If your tool returns a probability score, preserve its calibration rather than converting it into a binary verdict via a hidden threshold.

Route Consequential Cases to People

Automated adverse actions based solely on detector callbacks are dangerous. A detected watermark might justify asking how a document was drafted, but it cannot verify whether the sources support the recommendations. Review queues should aggregate the document version, declarations, source links, and edit history, allowing humans to check claims that matter. If a result affects a person’s reputation or opportunity, the system must allow for explanation and error correction. Notably, you can and should build this review infrastructure without a watermark detector. Capture and review workflows remain valuable even when signal access is limited, ensuring continuity if detector availability expands later.

Key Takeaways

  • Design schemas that explicitly support unknown, not_run, and inconclusive states to prevent false positives/negatives.
  • Treat watermark detection as one input among many; never let it automatically trigger adverse actions without human review.
  • Display method limitations and eligibility status in the UI to combat overconfidence in binary results.
  • Build review workflows now that function without detector access, focusing on capture and human decision logging.

The Bottom Line

Stop trying to force AI provenance into a boolean. The real engineering win is building systems that can honestly say 'we don't know' and preserve the evidence trail for human review.