When Numbers Aren't Enough: Why We Built a Listening Tool

At DeepMark, we spend a lot of time thinking about what makes a good watermark. Robustness is the obvious starting point - can the watermark survive compression, noise, re-encoding, even an analog replay through a laptop speaker and back? We've built an entire benchmark around that question, with more than 40 different attacks designed to push our models to the breaking point.
But there's another side to the equation, one that's arguably even harder to measure: quality. For our clients, audio quality isn't just a nice-to-have - it's the foundation their entire business is built on. Whether they're producing music, podcasts, audiobooks, or any other form of audio content, the quality of what they deliver is what their reputation depends on, and it's what keeps their listeners coming back. This is why we treat audio quality as something we will never compromise on. A watermark that survives everything but makes the audio sound worse defeats the whole point of being invisible, and that's simply not something we're willing to accept.
For a long time, we relied on a suite of well-established audio quality metrics to tell us whether our watermarked audio sounded as good as the original, and for a long time, those metrics told us what we expected to hear. Until they didn't.
The Problem with Measuring What You Hear
Our benchmark evaluates audio quality using metrics like PESQ, ViSQOL, SI-SDR, STOI, and NISQA - each designed to capture a different dimension of audio fidelity. PESQ estimates perceptual speech quality. ViSQOL produces a score modeled after human listening panels. SI-SDR measures signal distortion. STOI evaluates intelligibility. NISQA offers a non-intrusive quality estimate across noisiness, coloration, discontinuity, and loudness.
On paper, this sounds comprehensive. In practice, we started noticing something troubling.
These metrics would sometimes give high scores to audio that, to our ears, clearly had artifacts. Other times, they'd penalize changes that no listener would ever notice. A slight shift in spectral energy might tank a PESQ score while being completely inaudible. Conversely, a subtle but annoying tonal coloration might slip right past ViSQOL without raising a flag.
The root of the problem is that these metrics are stand-ins for human perception, not replacements for it. Each one was built with specific assumptions and constraints - narrow frequency ranges, fixed alignment between signals, averaging over time - that make them blind to certain kinds of artifacts while overly sensitive to others. They measure what they were designed to measure, but that's not always the same thing as what a person actually hears.
We found ourselves in a strange situation: our metrics said the audio was fine, but we could hear that it wasn't. Or they said something was wrong, and we couldn't hear any difference at all.
That's when we asked ourselves a simple question: if the goal is audio that sounds indistinguishable from the original to a human listener, why not just ask human listeners?
Introducing the Listening Tool
That question sat with us for a while, and the more we thought about it, the more obvious the answer became. We needed a way to put our audio directly in front of real people and let them tell us, in the simplest possible terms, whether they could hear anything different. Not a score on a scale, not a number generated by a formula, but a genuine human reaction - someone pressing play, paying attention, and sharing what they noticed.
So we built exactly that. The Listening Tool is a dedicated evaluation platform that lives inside our web application, designed from the ground up to collect honest, structured feedback from human listeners about the quality of our watermarked audio. It bridges the gap between what automated systems tell us and what people actually experience when they listen, giving us something that no metric can provide: a direct window into how our watermarks sound to the human ear.

The way it works is simple and deliberate. Someone on our team - typically a researcher or project lead - creates what we call a listening task. They upload a collection of audio recordings, where each one exists in two versions: the clean original and the same recording with our watermark embedded. They then choose a group of people to evaluate those recordings and send out the task. From that point on, each tester can open the tool whenever it suits them, work through the recordings at their own pace, take breaks, come back later, and pick up right where they left off. There's no pressure, no time limit, and no requirement to finish everything in one sitting, because we've found that relaxed, unhurried listening produces the most honest and useful results.
The tool offers two different ways for testers to evaluate the audio, and each one is designed to answer a slightly different question about how well our watermarks hide in plain sight.
Comparing two versions side by side
The first mode puts two recordings in front of the tester at the same time - one is the original, and the other carries our watermark, though the tester doesn't know which is which. They can switch back and forth between the two as many times as they like, replaying specific sections, jumping to different parts of the recording, and really taking their time to listen closely.

The first thing we ask is simply: do you hear any difference at all between these two? Many testers say no - and that's a great result for us, because it means our watermark is doing its job. But when someone does hear something, we want to understand as much as possible about what they noticed. We ask them to point out which of the two recordings they think contains the watermark, and then we ask them to describe how noticeable the difference is on a scale that ranges from barely perceptible all the way up to clearly annoying. They can even drag their cursor across the waveform to highlight the exact moment in the recording where they heard something unusual, which gives our researchers a precise timestamp to investigate rather than a vague sense that "something sounded off somewhere."

This mode is particularly valuable because it tells us not just whether our watermark can be detected, but exactly where and how strongly it reveals itself. When multiple testers independently highlight the same half-second of audio in the same recording, we know with confidence that there's a real perceptual issue in that spot - and we can go back to our model and figure out what's happening there.
Judging a single recording on its own
The second mode works differently, and it's designed to answer a question that's harder to pin down. Instead of giving testers two versions to compare, we play them just one recording - sometimes it's the original, sometimes it's the watermarked version - and we ask: does this sound like it might have been watermarked, or does it sound completely natural to you?

This might sound like a strange question to ask someone who has never worked with audio watermarks before, but that's exactly the point. What we're really measuring here is whether our watermark introduces any subtle feeling that something isn't quite right - a faint sense of processing, a slight unnaturalness in the sound, anything at all that makes a listener pause and think "this doesn't sound completely clean." If testers consistently identify watermarked recordings as watermarked, it means something about our embedding is leaving a perceptible trace, even without a clean reference to compare against.
When a tester suspects that a recording has been watermarked, we also ask them to rate the overall audio quality, which helps us understand not just that something was noticed but how much it bothered them. And because we know whether each recording actually contains a watermark or not, we can calculate exactly how often listeners are right, how often they miss our watermark entirely, and how often they incorrectly suspect a watermark in a perfectly clean recording - each of which tells us something different and important about our model's performance.

Together, these two modes give us a complete picture of how our watermarks are perceived. The side-by-side comparison tells us about precision - where exactly the watermark shows itself when you're actively looking for it. The single-recording evaluation tells us about the overall impression - whether watermarked audio simply sounds natural on its own, without any reference to compare it to.
Anyone Can Be a Tester
One of the most important decisions we made early on was to open participation far beyond our own engineering team, because we realized that the people best suited to judge whether our watermarks are truly invisible are not the ones who built them but the ones who will eventually hear them without ever knowing they're there. The tool can be used by anyone we choose to invite - colleagues from non-technical departments, friends, family members, and people with no background in machine learning or signal analysis of any kind.
This choice comes down to a simple truth: our watermarks don't live inside a controlled testing environment - they live in audio that real people hear during their everyday lives, whether that's a podcast listener on a morning commute, a musician reviewing a mix through studio monitors, or a journalist playing back an interview. If any of these people can sense that something about the audio feels slightly off, even without being able to say exactly what it is, then we know we still have work to do.
What makes this broad participation so valuable is the variety of ears and listening habits it brings into our process. Some testers turn out to be surprisingly sharp - picking up on differences that even members of our own team might miss - and they become our most demanding quality gatekeepers. Others rarely hear any difference at all, and their feedback is just as important, because it confirms that whatever traces our perceptive listeners catch are still well below the threshold where an average person would ever notice. Together, they give us an honest picture of how our watermarks actually sound when they leave our hands and enter the real world.
Built-In Rigor
Of course, subjective evaluation needs structure to be meaningful. We built several safeguards directly into the tool to ensure the results we collect are reliable.
One of them is what you might call an honesty check - we occasionally slip in pairs where both recordings are actually identical, with no watermark present in either one. If a tester consistently reports hearing differences in these identical pairs, we know their responses may not be trustworthy, and we can weight them accordingly when analyzing the overall results.
We also track how long each tester actually listened before responding, because a judgment made after half a second of playback tells us something very different from one made after careful, repeated listening.
And because testers can mark the exact time regions where they perceive a difference, we don't just learn that a watermark was noticed - we learn where. This is incredibly useful for our research team. When multiple independent listeners flag the same two-second window in the same recording, that's a clear signal pointing us toward a specific weakness in our embedding approach.
From Feedback to Better Models
The real value of the Listening Tool isn't in any single evaluation. It's in the feedback loop it creates.
Every listening task produces a detailed breakdown: per-item detection rates, per-tester reliability scores, severity distributions, region-level analysis. Administrators can compare results across tasks on a model leaderboard, tracking how successive versions of our watermarking system perform against human listeners over time.

When a new model version reduces the heard-difference rate from 30% to 12%, that's not a number improving on a spreadsheet - that's real people, listening carefully, genuinely unable to tell the difference. And when a particular audio sample gets flagged consistently across multiple testers, we know exactly where to focus our next round of improvements.
This is what continuous improvement looks like when your benchmark is human perception itself.
Why This Matters
We could have continued relying solely on automated metrics - they're fast, they're cheap, and they produce nice numbers for comparison tables - but we've seen too many cases where those numbers told a story that didn't match what we were actually hearing with our own ears.
The truth is that audio quality is ultimately a human judgment, and no formula, however sophisticated, can fully capture the experience of listening to something and deciding whether it sounds natural or whether something about it feels slightly wrong. Metrics can guide us, and they can flag obvious regressions, but when it comes to the subtle, almost hard-to-pin-down question of whether watermarked audio sounds the way it should - whether it sounds right - only a human ear can give the final verdict.
For a company that stakes its reputation on producing watermarks that are both unbreakable and inaudible, that verdict is everything, and the Listening Tool gives us a direct line to it.
It's Also Become a Bit of a Competition
Something we didn't expect when we launched the Listening Tool internally was just how much our team would enjoy using it. What started as a quality evaluation process has turned into a genuine source of enthusiasm around the office - people actually look forward to new listening tasks landing in their queue. There's a friendly rivalry that's developed over time, with testers comparing their accuracy scores, debating whether a particular recording really has a noticeable difference, and quietly trying to outperform each other on the leaderboard. It's the kind of healthy competition that makes everyone sharper, and it's pushed the overall quality of our evaluations to a level we didn't anticipate when we first built the tool.
Want to Help Us Listen?
If you've read this far and you're curious about what it's like to be a tester, we'd love to hear from you. We're always looking for fresh ears - people who enjoy listening to audio and are willing to spend a few minutes here and there helping us figure out whether our watermarks are truly invisible. You don't need any technical background, you don't need any special equipment beyond a decent pair of headphones, and you don't need to commit to anything long-term. If you're interested in participating and helping us make our models better, reach out to us at team@deepmark.me - we'd be happy to have you on board.
Mark what you generate. Prove what you ship.
Compliance you can demonstrate today, and a path to verified AI voice your customers actually trust, both built into the media you already produce.
