The web increasingly reads like it was written by nobody.
Synthetic text detection is framed as a binary classification problem. Reality is a provenance problem with missing data, adversarial incentives, collaborative authorship, and a public desperate for a little badge that says REAL.
Existing detectors tend to turn model-specific statistical tendencies into universal claims about authorship. This creates false accusations, especially for non-native English writers and people whose prose already resembles an annual report. A useful system should expose evidence, uncertainty, and limitations—not issue a metaphysical judgment about who typed each word.
What if we measured the nobody?
A detector changes the thing it detects.
Any public signal becomes a target. Human authors use grammar tools. Models copy humans. Humans copy models. Publishers flatten metadata. The benchmark leaks into training data before the procurement officer finishes adding the watermark to the PDF.
- RISK 01The extension confidently accuses the Declaration of Independence of using ChatGPT.
- RISK 02Users interpret “62% synthetic characteristics” as “62% of these words are fake.”
- RISK 03Deployment creates an evolutionary pressure toward more tedious human prose.
Provenance is a bundle of weak signals.
A detector need not be an oracle to be useful. A calibrated, inspectable ensemble can combine document metadata, cryptographic credentials when present, revision patterns, source citations, stylometry, and multiple model-likelihood estimators. The interface can say what it observed, what it did not, and how fragile the result is.