I was on a call when a lawyer’s voice-mail replayed a perfect imitation of his client asking for money. He mouthed a safeword and the recording hesitated—a tiny, telling break. That pause turned a routine evening into a lesson about how fast trust can vanish online.
A late-night phone that sounded real.
I tell you that because it’s how Dr. Hany Farid lives now: noticing the small breaks and training people to notice them too. Farid, a Dartmouth computer scientist who founded GetReal the day ChatGPT became public, treats digital evidence like a detective treats fingerprints. He watches our screens with the skepticism of someone who has seen scams evolve from crude fakes into near-perfect imitations.
Can I trust AI detection tools to stop fake content?
Short answer: they help, but they don’t stop everything. Tools from companies like Pangram and Originality.ai analyze dozens or hundreds of tiny signals across a text to say whether it’s machine-made. Pangram claims near-perfect accuracy—99.98%—and a University of Chicago study ranked it highly. Yet CEOs like Max Spero and Jon Gillham admit these systems produce probabilities, not guarantees.
You should think of detection more as probabilistic forecasting than absolute proof. One useful metaphor: detection tools are like a weather map predicting rain; the forecast helps you decide whether to take an umbrella, but it doesn’t change the sky. That mindset keeps you practical instead of panicked.
A startup dashboard that promises certainty.
Three years after ChatGPT, scores of detection tools appeared. Pangram, Originality.ai and others promise to clear the static. They do find patterns humans miss—subtle punctuation patterns, distribution of rare words, strange placement of symbols—and they combine those signals into a confidence score.
But generative models change fast. Engineers tune models to avoid obvious tells—em dashes were once a reliable giveaway, then they weren’t. When detection systems lean on a set of signals, model builders alter outputs and the tells shift. That chase is constant and expensive, and it’s why even confident-seeming tools still hedge.
Small-but-happy win:
If you tell ChatGPT not to use em-dashes in your custom instructions, it finally does what it’s supposed to do!
— Sam Altman (@sama) November 14, 2025
A watermark requirement on paper, not always on the page.
Here’s a visible change: the EU’s AI Act requires labeling and invisible watermarks for AI-generated content intended to inform the public. In practice, that’s progress. Anthropic added invisible watermarks to Claude to comply with the law. Google embeds SynthID and C2PA credentials into metadata. Yet developers and hobbyists find ways around marks, and some tools allow visible tags to be stripped.
How reliable are watermarks and labels?
Watermarks are useful starting points. They answer the narrow question, Did an AI touch this? What they can’t tell you is how much the AI altered the piece. Pangram and others are pushing to score degrees of AI involvement so you can tell if a human wrote most of an essay or if a model wrote the first draft and a human edited it. If you want a single answer, you won’t get one—watermarks are one instrument in a larger toolkit.
A designer on a Slack channel who removes the visible tag.
That example is real: developers already share tricks for stripping visible badges from images and videos. Companies will keep patching holes; others will find new ones. You and I live inside that cycle. Detection vendors improve, models adapt, and the cycle repeats.
Another metaphor: detection systems pull together hundreds of weak signals the way a weaver finds meaning in scattered threads—alone they’re thin, together they form a pattern.
Will detection ever outpace generative AI?
No one I spoke with promised a permanent lead. Max Spero told me you won’t reach 100% accuracy. Jon Gillham compared detection to forecasting: reliable often, flawless never. The models get better because they are trained on vast amounts of data; detectors get better by digesting more examples and refining math. It’s an arms race of patterns.
A professor telling students to stop pretending they can be experts overnight.
Farid’s advice is sharp and practical. He tells his students to stop trying to become forensic experts in a semester and to instead learn what a trusted source looks like. The people behind serious outlets—major newsrooms, established journals—have editorial standards, fact-checking routines, and reputational risk that act as a real filter. That matters more than a single detection score.
So what do you do? Combine tools and judgment. Use detection tools as guidance, not gospel. Check metadata when you can. Favor reporting with named sources and transparent methods. Treat sensational claims with healthy doubt. Ask: who benefits if this is believed?
A small habit that costs you nothing and buys you clarity.
Farid and his wife use a safeword on remote calls. It sounds simple because it is simple. You can adopt small rituals—verify via a second channel, ask for specifics only the person would know, cross-check a suspicious clip against reputable outlets. Those habits are your daily defenses.
Companies like GetReal, Pangram, Originality.ai, and standards from Anthropic, Google, and regulators will move the needle. They’re part of the solution, not the whole of it. Detection will improve; deception will too. The real skill is learning to weigh signals, not to expect a single magic answer.
Who will you trust the next time a flawless fake shows up on your timeline?