Why is voice watermarking so hard?

As shocking as it sounds, almost no voice AI company seems to be keeping up with the regulatory wave coming its way. Our product and our long-term vision depend on long-term partnerships. Those depend on trust. So I feel a responsibility to get the right information to the right people.
If you haven’t heard, 2 December 2026 is the EU’s marking deadline for covered generative AI systems placed on the market before 2 August. The rules already apply to newer systems. California also has marking requirements in force, and other countries are moving in the same direction.
For AI voice, there are two separate duties to understand. One is disclosure, something like “Hey, I’m an AI voice agent calling from…” The other is marking the generated content in a way a machine can detect. Saying that you are an AI agent does not replace the second duty.
Why should you care? Fines for these EU transparency violations can reach €15 million or 3% of global annual revenue, whichever is higher. For SMEs, including startups, and small mid-cap companies, the lower ceiling applies. These are maximum penalties, but they are still serious enough to matter to any voice AI company.
Even among those who know, the reaction is often the same. We’ll figure it out. It’s just watermarking. That problem was solved 30 years ago.
To be honest, the word “watermarking” does sound old. Our first marketing move was going to be retiring it. We gave up when Anthropic brought it back.
Traditional watermarking, and many of the methods available today, were not built for real-time phone calls. Why?
Two reasons. Real time. And phone calls.
Every millisecond matters
A conversation has a latency budget of just a few hundred milliseconds. Companies building TTS models know this well. They know what every millisecond means.
Many watermarking models are built around fixed audio windows, mostly a second long. Your TTS model produces chunks just a few tens of milliseconds long and streams them continuously to the user. The watermarking step has to fit into that flow. Waiting to collect a second of audio changes the whole conversation.
So the algorithm needs to mark very short pieces of audio, carry enough bits to be useful, and finish its work in just a few milliseconds.
Look through the literature. Finding something that meets all of those requirements at once is much harder than finding something described as “real time.”
After all that, your voice still has to travel through a phone call.
Your TTS model may run at 24 kHz or higher. Traditional narrowband telephony uses 8 kHz audio. Then come compression, lost information, noise, delays, and missing packets. The watermark has to survive all of it.
Combine those channel conditions with the speed and model-size constraints, and things get complicated very quickly.
Metadata looks like a great alternative. Instead of a watermark, we attach some information to the file and use that as its provenance.
But where is the file in a phone call?
There is no audio file travelling intact from the voice company to the listener. Call protocols can carry metadata, but you cannot assume that your provenance data will survive every network boundary, speaker, or recording. What reaches the person on the other end is the sound.
That brings us back to putting the mark inside the audio.
And after all the work needed to make it fast and robust, there is still one requirement we cannot forget. Companies do not want the watermark to damage the voice.
How do you make something inaudible and still strong enough to survive a channel as harsh as a phone call?
That is the hard, hard part.
The roughly 3.4 kHz upper limit of a traditional phone channel was not chosen by accident. That band carries much of the information we need to understand speech. We can lose the rest and still understand what someone is saying, even if the voice no longer sounds perfect.
So the watermark has to live in the part of the spectrum where it is hardest to hide. That is also the part most likely to survive the call.
Even modern approaches based on deep learning can take shortcuts here.
Meta’s AudioSeal concentrates its watermark in a narrow band around 1,100 Hz, visible as an unnatural line on a spectrogram. Filter that band, and you can break detection.
The problem is not just whether the mark can be removed. It is whether the system has learned a pattern that is easy to find and easy to target.
A watermark that uses the same pattern regardless of the audio is content-agnostic. I think that is a naive way to approach this problem.
It can be the result of setting up the training problem badly. Deep-learning models will find the easiest route to the goal. If the objective rewards hiding the mark but does not sufficiently punish fragile shortcuts, the model can learn exactly that.
Where the content goes, the watermark follows
The alternative is a content-dependent watermark. Where the content goes, the watermark follows.
On a spectrogram, it can look like the sound itself. Like its shadow.
Subtract the original audio from the watermarked audio, and you get what we call the residual. With this approach, the residual can follow the speech so closely that you can hear it.
Read that again. You can hear speech in the watermark itself.
Incredible, right?
“Reinvented watermarking,” as we call it, is still a relatively young field. There are many problems left. This is far from solved, and only a small number of research teams are working on this particular combination of constraints.
For systems covered by the EU transition period, the deadline is 2 December. Two months away.
We have been working on this for six years.
If you think your team can solve it by 2 December, we will not stand in your way or try to convince you otherwise.
But you are taking a serious risk. Two months from now, you could have a solution that damages audio quality or fails to meet the marking requirements.
One can cost you customers. The other can expose you to serious fines.
Get in touch before time runs out.
Mark what you generate. Prove what you ship.
Compliance you can demonstrate today, and a path to verified AI voice your customers actually trust, both built into the media you already produce.
