DeepMark has joined Y Combinator's Fall 2026 batch. We're live on Launch YC: upvote us
07/10/2026

DeepMark Vigil vs Meta AudioSeal for real-time voice AI

SK
by Slavko Kovačević7 min read
DeepMark Vigil vs Meta AudioSeal for real-time voice AI

When a company puts a voice agent on the phone, the identity of that agent should survive the call. That sounds simple. The audio may be compressed, resampled, interrupted by packet loss and passed between different networks before it reaches the person listening.

We built Vigil, DeepMark’s proprietary audio watermarking model, for that setting. It works with streaming speech, supports chunks as short as 80 milliseconds and runs inside the customer’s infrastructure.

Meta’s AudioSeal is an open-source audio watermarking system with publicly available code and model weights. It is a useful reference for anyone evaluating voice watermarking. The question for an enterprise buyer is how each option fits the system they actually need to run.

For us, that comes down to three things. Can the watermark keep up with the voice agent? How does it behave when the audio goes through a call? And how closely does it follow the speech itself?

Vigil and AudioSeal at a glance

Vigil and AudioSeal features
Feature DeepMark Vigil Meta AudioSeal
Availability Proprietary model, commercially provided by DeepMark Open-source code and weights under the MIT license
Streaming Native streaming with chunks as short as 80 ms Streaming implementation available in the official repository
Identifier payload 32 bits by default; 64-bit models available Optional 16-bit message
Embedding time on an L40S 1.45 ms median per batch of 32 one-second clips at 24 kHz, using CUDA graphs 9.32 ms median per batch of 32 one-second clips at 16 kHz, using torch.compile and the standard checkpoint
Integration and deployment SDK supporting multiple programming languages; runs inside your infrastructure Open-source code and model weights for integration and deployment in your infrastructure

Streaming that fits a live voice agent

Vigil can watermark a stream in chunks as short as 80 ms. A voice agent can send its speech through the model as it generates it, without waiting for the complete sentence or recording.

AudioSeal also has an official streaming implementation. For a production system, the useful comparison is the processing budget at your chunk size, sample rate and concurrency. Those are the conditions we want to evaluate with a customer.

Embedding speed on an NVIDIA L40S

Embedding is the operation on the outgoing audio path, so that is the speed we care about here.

In our L40S benchmark, Vigil embedded a batch of 32 one-second clips in 1.45 ms, compared with 9.32 ms for AudioSeal. That is approximately 6.4× higher embedding throughput, using each model’s fastest tested execution mode. Vigil processed 24 kHz audio with a 32-bit payload, while AudioSeal processed 16 kHz audio with a 16-bit payload. These are median embedding times, not total voice-agent latency.

Our October 8, 2026 test used the same NVIDIA L40S and PyTorch 2.11.0 for both models, with five seeds and 100 iterations per seed. Vigil ran in FP32 with CUDA graphs, while AudioSeal 0.2.0 used torch.compile.

We timed AudioSeal’s standard audioseal_wm_16bits checkpoint. Its separate streaming checkpoint was not part of this test. Total deployment latency also depends on buffering and the rest of the speech pipeline.

A watermark that follows speech like a shadow

This is the part I find easiest to explain by showing it.

Vigil’s watermark follows the speech. Its pattern changes with the voice’s structure over time and frequency. Subtract the aligned original speech from the watermarked version and you can see that relationship in the remaining signal. It is like a shadow of the speech. That is what we mean by content-dependent watermarking.

In the example below, the isolated watermark is amplified by 21 dB so you can inspect it. The watermarked recording contains the mark at its original level.

AudioSeal also uses the input waveform to generate its watermark. The important question is how strongly the resulting mark follows the host audio and where its energy ends up. A published structural analysis found concentrated AudioSeal watermark energy around 1.25 kHz and showed that targeted filtering could substantially reduce presence detection while preserving high objective speech quality in the configuration it tested.

The demonstration below shows the relationship between Vigil’s watermark and the speech carrying it.

Two spectrograms compare original speech with Vigil’s isolated watermark over five seconds. The watermark follows the timing and frequency patterns of speech. Its level is increased by 21 dB for visibility.
Vigil follows speech like a shadow. The lower plot shows watermarked audio minus the original, amplified by 21 dB for visibility. The dashed line marks 3.4 kHz, the conventional upper edge of narrowband telephone speech. Speech source is LibriVox, public domain.

Where AudioSeal falls short for production

I see AudioSeal as a good research baseline. I would not rely on the configuration we evaluated, as-is, to carry an agent’s identity through production voice calls. Our published DeepMark Benchmark shows why.

Sign inversion is one example. Multiply every audio sample by −1 and the waveform flips upside down. Its timing, level and frequency magnitudes stay the same. The speech remains intact, and the change is largely inaudible. Yet in our benchmark, this simple operation disrupted AudioSeal’s recovery of the embedded bits.

Cropping the beginning is another. In the published test, removing the first 10% of a recording reduced AudioSeal’s bit accuracy to roughly 60%. Removing the same proportion at a random position did not produce the same breakdown. Where the audio started mattered.

These results concern recovery of the embedded bits. For an agent’s identity, we need the correct identifier back, consistently. Knowing that some watermark may still be present does not give us that identity.

That is the gap between a useful research baseline and the reliability we require in production. We expect a watermark to survive simple signal operations like these. A phone call adds further transformations, which makes that requirement even more important.

The phone call is part of the problem

Speech leaving a voice agent and speech arriving on a customer’s phone are different signals. A watermark has to deal with the transformations in between. Phone-call survival is a core requirement for Vigil.

We use PhoneSim’s voip_to_cellular_narrowband profile as a reference path for calls that travel from VoIP into a narrowband mobile network. It combines an Opus hop with narrowband resampling and filtering, an AMR-NB mobile codec, packet loss with receiver concealment, adaptive playout, and clock drift. It also models input level and background noise.

We prioritize this path because it puts the watermark through the kinds of changes that matter when a voice agent calls an ordinary mobile number without end-to-end HD voice. PhoneSim uses real codecs and repeatable simulations, which makes it useful for comparing how a system behaves before and after the channel.

A simulated path is one part of an enterprise evaluation. The next step is the customer’s actual route, with its codecs, carrier connections, and receiving devices. That is the environment the watermark ultimately has to work in.

Open source, proprietary models and the deployment decision

AudioSeal’s open-source release gives an engineering team code and weights it can inspect, run and adapt under the MIT license. The team can build its own integration and evaluate it against the requirements of its product.

Vigil is a proprietary SDK that supports multiple programming languages. It runs inside the customer’s infrastructure. Its default payload is a 32-bit identifier, and we also have 64-bit models. The identifier is carried in the audio itself and can be used to associate a stream with an agent or source in the customer’s system.

For an enterprise, the decision includes the engineering work around the model. The outgoing speech pipeline, verification workflow, call channels, and operational requirements all need to fit together.

That is what we want to discuss with you. Bring the voice system you are building and the call path it needs to survive. We can evaluate Vigil in that context.

Talk to DeepMark about Vigil for your voice agents.

Mark what you generate. Prove what you ship.

Compliance you can demonstrate today, and a path to verified AI voice your customers actually trust, both built into the media you already produce.