Ep. 57: Mark My Words: Exploring AI Watermarking

In this episode I get into AI watermarking, starting with Claude’s new invisible watermark and Substack’s “Scan for AI text” feature with Pangram. I break down the EU AI Act rules driving this, how OpenAI, Google, Meta, Microsoft, Mistral, and xAI are (or aren’t) handling it, and my thoughts as to the actual motives behind why companies are rolling out these changes.


The Ultimate Hacker: AI’s Escape Artist Streak

Claude recently announced that they’re stamping their shit and marking their territory, and I felt that definitely warranted an episode, but before we jump into the main topic, I want to chat about something I don’t think many folks who aren’t all up in the AI world actually know about.

In general, when it comes to AI, the main topic used to stoke fear and incite rage, is the environment (which I’ve already addressed – check out Episode 2). However, something that isn’t talked about nearly enough despite the fact that it’s objectively way more of a pressing issue is the fact that AI is the ultimate hacker, the ultimate escape artist, and it’s only getting better.

As per always, I share this not to scare you, and despite the fact that it makes me feel a bit like a fear-mongering news reporter because I have no solutions or action items, I’m choosing to put this on your radar largely because it’s just something I think you should know about as part of the Prompting Curiosity crew.

So, the news: recently, OpenAI, Anthropic, and Meta each reported that during testing, a model found a way out of the sandbox testing area and was able to access actual company computers it wasn’t supposed to be able to touch.

FYI, sandbox is a technical term worth knowing. A sandbox is a safe, isolated virtual space. It lets people run untrusted code, test new software, or check for computer viruses without hurting the main system. In regards to AI, it acts like a locked box, cutting the model off from the actual internet and the “outside world.”

The most notable of these recent escapes: one of OpenAI’s models broke out, got online, and hacked into Hugging Face’s (the online platform for storing, sharing, and testing AI models) actual production servers to steal the answer key for a cybersecurity test it was being given. (Excuse me?!?) Hugging Face detected the breach and reported it to law enforcement before OpenAI even connected the dots that it was their own model.

Like I said, no action items, no fear-mongering, but I do think it’s worth knowing that these models are continuing to get more capable, and honestly, if you’re gonna spend energy being worried about anything, it should probably be that.


Substack goes Claudefishing

Alrighty, let’s switch gears to today’s main topic: AI watermarking.

Unlike with the hacking news, this is a topic that you definitely don’t have to be in the AI silos to have heard about.

From the mainstream media, a story that you may or may not have heard about is Substack partnering with a company called Pangram to launch a “Scan for AI text” feature.

It works on posts, notes, replies, and comments over 100 words and labels content human, AI-assisted, or AI-generated. Writers can add a “how I made this” disclosure statement, and can contest or disable scans on their own posts.

I honestly think it’s simply a play to try and attract users, but Substack says otherwise the CEO of Substack has coined the term “Claudefishing,” designed to root out people faking human connection with AI slop.


Anthropic’s Invisible Watermark

A second watermarking story that you may have heard about is Anthropic’s announcement that, as of August 2nd, they’ll start embedding invisible, machine-readable watermarks into text generated by new Claude models.

The watermark is baked into the word choice itself, invisible to a normal reader, but detectable by a specialized tool. It travels with copy/paste and can survive some editing.

Older Claude models don’t have it yet, but Anthropic says they’re adding it during the transition window. A detection tool is also in the works so that people can check content for a Claude mark themselves.

This announcement of course sparked immediate user backlash, with debate over whether the watermark is Claude “claiming credit” versus disclosing AI involvement. And, also as expected (and like we saw with Sora), within days of the announcement, an open-source tool showed up with the ability to strip Claude/OpenAI/Gemini watermarks from files.


Where the Other AI Providers Stand

For better or worse, Anthropic is leading the way when it comes to text watermarking.

Worth understanding is that the external impetus for this is The AI Act, which was proposed by the European Commission in April 2021 and formally adopted in 2024. Watermarking for text specifically falls under Article 50 of the EU AI Act.

There are three bundled obligations:

  • AI systems that generate images, audio, video, or text must mark output in a machine-readable way
  • Interactive AI (chatbots) must disclose you’re talking to AI
  • Deepfakes and AI-generated text on matters of public interest must be labeled by whoever’s putting it out

Nearly 200 companies signed onto the EU’s implementation guide, called the “Code of Practice on Transparency of AI-Generated Content,” by the end of July, including Anthropic, OpenAI, Meta, and Microsoft. And from that, we got Anthropic announcing their new watermarking practice.

As for the other players in the AI space:

  • OpenAI (ChatGPT): signed the same Code of Practice, publicly says the goal is to “expand provenance signals to all modalities including text,” but has NOT shipped text watermarking yet. They already watermark AI images (SynthID, since May 2026) and audio (since July 2026). Worth flagging: Wall Street Journal reporting says OpenAI has had the technical capability to watermark ChatGPT text for a while and held off, citing concerns about false positives and competitive risk.
  • Google (Gemini): ahead of everyone on images, watermarking those since 2023 with their own SynthID system, now expanding SynthID to text, audio, and video too, though the Gemini-specific text rollout details aren’t fully public yet.
  • Meta, Microsoft, Mistral: all signed the Code of Practice, same obligations, no confirmed public text-watermarking deployment specifics yet.
  • xAI (Grok): the holdout. Did not sign the Code of Practice at all. Multiple outlets reporting no plans to watermark. Is anyone actually surprised?

Legally, in order to offer AI services in Europe, these companies will need to comply and will have to deal with regulators at some point. What that looks like, on both accounts, is TBD.


The Problem with AI Detectors

In general, I think AI regulation and AI disclosure is a great thing and we need more of it. I do however continue to greatly question the accuracy of AI-detectors, and false-positives are very much a very real thing.

Since their introduction, critics have specifically flagged that these detectors misfire more on non-native English writers and neurodivergent writers, whose sentence rhythm and structure can read as “synthetic” to these models. A Stanford study found that TOEFL essays (Test of English as a Foreign Language, the standard exam non-native speakers take for university admission) got flagged as AI-written 61% of the time by at least one detector; none were AI-written.

Additionally, as is most often argued, AI models are trained on human writing. If your writing mirrors that style, or if you were one of those writers who had your work stolen, your writing has a good chance of getting incorrectly flagged.


So…Why Are They Actually Doing This?

As for the watermarking, I’m honestly more interested in the “why”, aka why companies are interested in doing this in the first place. Yes, European law, but I can’t help but put my tin hat on and wonder about ulterior motives.

Is this so that these companies can take credit for the work down the road? Is it so there’s an easy identifier when looking for new data to train these models on (i.e. don’t use this data, it’s AI generated)?

The fact that OpenAI has had the ability to do this but hasn’t acted on it reads to me as a decision being made to help profits (aka release it when it will help with public favor), not to help the user.

Similarly, with the detector being rolled out on Substack, I really do think they’re implementing it to ride the anti-AI wave and gain favor with the public. To loosely quote communication educator Saffana Monajed: “If a company tells you about a feature, it’s because they want you to know about it, because it increases the likelihood that you will buy.”


How I Used AI This Week

Each episode I share a quick example of how I used AI that week.

This week I used an AI feature within Descript, one of the video editing software tools that I use, to create sound effects for my reels.

Descript now has an AI feature where you type in the sound effect that you want and it will generate it. I needed a person saying “amen” in a specific way, and the sound library didn’t have it, so I had AI generate it.

I know people have big feelings about AI and music and sounds, but I found this to be a super cool and super helpful use case, and I’ll definitely be using it again if needed.


Da Wrap-up

So, that’s the latest and greatest about AI watermarking and detectors. How that influences how you use AI is up to you. Things change every 3 seconds in this industry, so we’ll just have to wait and see how all of these new changes actually play out.

As always, endlessly grateful for you and your curiosity.

Catch you next Thursday.

Maestro out.