← Antonio LozanoSubly · Audio Description

Subly / Nov 2025 – Feb 2026

Audio description for people who had never made one

Audio description is the spoken track that describes what happens on screen. Subly sold accessibility to companies working to a legal deadline, and this was the last piece missing. I owned it end to end.

Role
Product owner and designer
Team
Me and two engineers
Scope
Interface and prompts
Shipped
Three releases in four months
Audio description in Subly, after the third release. Turn the captions on and the dialogue appears in the gaps each description was written into.

The problem

Audio description is the one accessibility feature you cannot check by looking at it.

The product already had captions, transcripts, and an analyser that checked a video against the standard and listed what was still missing. It covered most of it.

CoveredCaptions
CoveredTranscription
CoveredDescriptive transcript
CoveredSummary
The gapAudio description

The analyser. Checks a video against the standard and lists what is missing. For the customers who needed the whole standard, it kept listing the same thing.

The suite as it stood. The product was already naming its own gap to the customer.

Every other piece can be checked. A caption is right or wrong against the audio, and anyone can hear that. Audio description is a writing job. Someone decides what matters on screen, in what words, in the gaps the dialogue happens to leave. Two people describing the same shot write different things and both can be right.

So the person buying it had no way to tell whether they had bought something good.

The timing was not an accident. The European Accessibility Act had just come into force, with ADA Title II, the ACA and AODA doing similar work elsewhere. Subly had closed self-serve, so every customer arrived through a sales call.

Doing it by hand means a writer, a voice artist and a mix. Subly’s own published figure for the automated version is up to seven times cheaper.

Who it was for

Small teams handed a regulation,
and none of them were experts.

The move to B2B changed who was on the other end, and it was not who the product had been built for.

A company of several thousand people would run accessibility compliance with a small team, one to four people, and the work usually landed on one or two of them. They were not accessibility specialists and they were not video editors. They had been handed a regulation and a date.

Subly’s own published case studies say the same from the customer side. A university working to ADA Title II: “We needed a partner who could show us exactly what was required and help us fix it at scale.”

Sometimes the work went to an agency instead. An agency knows video. It does not know the regulation any better than the customer does.

How we worked

We knew who it was for. We did not know what they would do, and we shipped instead of waiting to find out.

There was research on the customer. There was none on how anyone would use this, and we chose not to wait for it. The company was too small for analytics to conclude anything and small enough to talk to users constantly, so the beta was the research.

Two or three interviews per release, with the people who would actually run it. Hotjar for where they stopped. And sales demos, which were unmoderated first use with a stranger.

What we tried, and what happened to each of them.

01Script firstNovemberWe write the script, you approve it, then we spend on the voice. Rejected in two weeks. Nobody could judge a script they had not heard.
02The journey, made explicitDecemberName the three steps so nobody waits for audio that has not been asked for. Temporary, and we knew it. It made the wait less bad while the real fix was built.
03Result firstFebruaryGenerate everything, hand back the described video, put the script beside it. Shipped. The one that answered the problem.

Release one / November

Replaced

Script first, because it was the fastest thing we could build

Generate the script. Show it. Let the user approve it. Then spend on the voice and the mix. In that order it was three separable jobs, which is why it was the version that fitted the time we had.

The reasoning on top of that was ours as a team, and we approved it without much research on purpose. Audio description was the most expensive thing Subly did, so we assumed people would want a say before firing it off.

One sentence on what audio description is, one button, and the order of the work underneath.
The stall. A list of lines with timecodes and no way to hear any of them.

What we found

A week and a half from shipping to knowing we were wrong

Day 3

First beta calls. People were not editing the script. They were sitting on it.

Day 5

Different customers, and the same behaviour again.

Week 1.5

Hotjar confirmed the pattern. They opened the script, read it, and stopped.

The reason was the same every time, and it is obvious once you hear it. You cannot judge a description without hearing where it lands. A line reads perfectly well on the page and then arrives on top of dialogue, or two seconds late.

They needed the audio and the video to check the words. We were asking them to check the words first.

So we had the answer in week two: give them the finished video and put the script next to it. It was not a screen change. It was the whole pipeline running in a different order.

Release two / December

Temporary

A patch, shipped knowing it was not the answer

Rebuilding around the finished video was weeks of engineering. Leaving release one alone meant a month of people staring at a script they had no way to judge.

So I built the middle option in the frontend myself: make the journey visible, so at least nobody is waiting for audio that has not been asked for yet. Three steps, named, with the current one marked.

The whole change is the strip across the top: Script, Review, Voice, with step two lit. Same generation, same script, same button.

It helped and it did not fix it, which is what we expected. Nobody had been confused about the steps. They had been unable to do step two.

Release three / February

Kept

The finished video first, and the script second

Everything runs. Transcription, description, voice, mix. What comes back is the described video playing, with the script beside it, and a wrong line is fixed in place.

Three fields and a button, with the type explained where it is chosen.
No percentage and no estimate, because neither would have been true. It says the credits come back if it fails.
What comes back: the described video, and the script beside it to fix a line.

Nobody has to judge a script any more. The thing that stalled release one is gone: the person who could not tell whether a line worked now hears it land.

What it cost the user. The script cost credits and the voice cost credits. Running them together spends both at once, and the cheap first step is gone.

We did it anyway, and the reason came out of the interviews, not the spreadsheet. The person running a video is usually not the person who put the budget on the account. Putting the arithmetic in front of them at the moment of action made them hesitate over a decision that was not theirs.

It also suited the business, and that is worth saying out loud. More generation is more consumption. We would have chosen it anyway, because a review step nobody can complete is not worth the saving. But the interests lined up.

There is a wall before any of it runs. Without the credits nothing starts, and nobody tops up their own account at Subly. Credits are arranged with someone on the team, so Book a call is the mechanism and not a sales gate. That is also why the screen carries no figures: none of them change which of the two buttons gets pressed.

Not enough credits. Nothing generated, nothing charged, and the settings are still there behind the modal.

The two modes

The first version paused every video.
The last one only pauses when it has to.

There are two ways to fit a description into a film. Standard puts it in the pauses that are already there. Extended stops the video to make room.

StandardThe description fits between two lines of dialogue

Dialogue Dialogue Dialogue A jaguar emerges from the shadows He leaps into the river

Nothing to decide. The gaps in the dialogue are long enough to hold what needs saying.

The problemContinuous dialogue leaves nowhere to put it

Dialogue Dialogue A caiman glides past, only its eyes above the surface

The whole difficulty of audio description in one picture. There is no room, and every way out costs something: cut the description, speed the voice past comprehension, or stop the video.

ExtendedThe video pauses to make the room

Dialogue Dialogue A caiman glides past, only its eyes above the surface

It solves the problem and changes the artefact: the viewer no longer gets the video they were given. That is worth saying at the moment of choosing, which is what the escalation does.

Why there are two kinds at all. Drawn in the page rather than screenshotted.

The first version only made Extended. Every video got paused, whether it needed it or not, so every customer got back a film that no longer ran the way the one they uploaded did.

Then we built Standard and made it the default. Extended now appears only when the dialogue leaves no gap big enough, and when it does, the product says so and explains what it changes.

Later we put the choice back. Some customers wanted Extended from the start, and once someone knows the difference there is no reason to keep deciding for them.

Extended, chosen on purpose. The screen says what it costs you before you pick it: “The film ends up a little longer.”

A first-time user never has to hear about the second mode. Someone who already knows can ask for it.

The escalation. No gap was big enough, and the warning says so plainly: “Your film gets longer.”
The film runs 01:09.28 instead of 01:04.29, and the red marks say where it stops. The captions show why each pause was needed.

What we cut

Two things I dropped to make the dates

Scope · cut

No per-cue control

Fixing one description without touching the rest. It would mean learning what a cue is, and it worked against what the business needed that quarter.

Scope · cut

English only

Other languages stayed in the backlog. A language picker on day one is one more decision for someone who does not want to make any.

Per-cue control is the one that still bothers me. I cut a feature that would have saved people money, because the company needed them to spend more that quarter. What I did instead was push the quality of the generation, so they would need it less often.

It held better than it deserved to, for a reason I did not design. Most people barely edited. The output was for compliance, not for broadcast, and once it was accurate they were done with it.

When it failed

It broke in parts, so a failure did not cost the whole run

The script, the voice and the mix run in order and fail separately. If the voice failed the script survived, and the error was about the voice alone.

When a run failed we said so and gave the credits back. Someone spending credits against a compliance deadline will not keep using a product where a failure also costs money.

Nothing came back, and the credits did. The refund is stated on the screen, so nobody has to chase it.
The voice failed and the script did not. “Your script is safe.” The retry is for the voice alone.

Where it ends

The download, and what Extended does to it

What shipped ordered the download by which part of the backend produced each file. The one thing most people came for, the video with the description already in it, was a checkbox in the far column that stayed off until an aspect ratio was picked first.

The order here is the other one. The video is what gets published and everything else is a file, so in Standard audio description sits inside the video’s options: same length, captions still fit, costs nothing else.

Extended is not an option in the same sense. The video stops for each description, so it runs 01:09.28 instead of 01:04.29 and every file with timecodes moves with it. It leaves the video’s options and goes up beside the language, as the version, and it is a dropdown rather than two large cards on purpose. The question at this point is not whether the person wants audio description. It is which of their two videos they are taking.

The same modal, both versions. Standard keeps audio description under the video. In Extended it moves up to the version row, and every file underneath is timed to the longer cut.

The prompts

I wrote the prompts, not just the screens

Because there is no per-cue editing, a wrong line costs a full regeneration and the user pays for it. The output had to be right the first time, and that is not a screen problem.

What gets described and what gets left out. How much detail. When to stop so the description does not talk over the dialogue. Someone who cannot see the screen is on the other end of every one of those calls. They are design decisions. They get filed as engineering because they are written in prompts instead of pixels.

The first version ran in two stages and it made things up. It described objects that were not in the video. Splitting it into three, with one rule each stage has to obey, is what made the output good enough to sell against a standard.

Stage 01

Context

The transcript, the extracted frames, and an explicit boundary on what counts as knowable.

The ruleDescribe only what is visually present.
Stage 02

Generation

Timed descriptions written to fit the natural gaps in the speech, not written first and squeezed after.

The ruleFit the gap, or say it does not fit.
Stage 03

Refinement

Every line is cross-checked against the source before it is allowed out.

The ruleIf it is not in the frame, it does not ship.

The third stage is the one that stopped the hallucinations.

Outcome

What changed, and the one customer I got to ask

The change is what the person on the other end has to do. In November they read a list of lines and had no way to tell whether any of it was right. In February they press one button, wait, and play the video.

The only feedback I have from a customer is one interview, after the third release shipped.

Works quite well. Still a bit lengthy to run, but way faster than making it by hand.
A national broadcast channel, in the only interview run on release three. This is my recollection of what they said, not a transcript, so it is set without quotation marks.

The wait is the fair part of that. Everything runs once they press the button and it grows with the length of the video, around five minutes. What those five minutes replace is a writer, a voice artist and a mix.

There was never a second interview. Subly was acquired a few weeks after the third release and I left that month, so there is no month of usage to point at and the user base was too small for the analytics to mean much before that. I could quote a number from the pipeline and it would not survive a follow-up question, so I would rather say this: the feature is live, the two versions are real, and the difference between them is the output below.

The clearest measure is the output itself. Run the same video through both versions and version one writes seven descriptions, version two writes twenty-three. That is not padding. Version one skipped long stretches of the video because it only wrote for the gaps it found easy, and it described what it saw in four or five words. Version two covers the video.

7 → 23Descriptions, same videoVersion one against version two. More of the video actually covered, not more words per line.
3 → 1Steps before a resultScript, review, generate became one action.

Version one7 descriptions

Version two23 descriptions, first 10 plotted

00:0000:4801:36

The same video, both versions, on the same scale. The hatched stretch is thirty-two seconds where version one said nothing at all: the jaguar enters the water and takes the caiman, and the viewer who cannot see it is told none of it. Version two also opens the video, which version one left silent for its first eight seconds. Only ten of version two’s twenty-three are plotted, so its real coverage is denser than drawn here, not sparser.

Version one7 descriptions

00:00:08 → 00:00:11A spotted jaguar stalks the riverbank.
00:00:16 → 00:00:18He moves through foliage.
00:00:24 → 00:00:28A caiman glides through the murky river.
00:00:30 → 00:00:32Caiman floats.
00:00:38 → 00:00:40Scarface watches from bank.
00:00:48 → 00:00:52Scarface watches as a caiman approaches.
00:01:24 → 00:01:27Scarface climbs the bank, dragging caiman.

Version two23 descriptions

00:00:00 → 00:00:02Tangled roots and dense green foliage overhang a riverbank.
00:00:08 → 00:00:11A spotted predator emerges from the shadows, walking slowly along the bank.
00:00:16 → 00:00:18The jaguar, its coat a pattern of dark rosettes, moves through the undergrowth.
00:00:24 → 00:00:26Scarface continues his patrol, his mouth slightly open.
00:00:26 → 00:00:28A caiman swims in the green water, its head and armoured back visible.
00:00:30 → 00:00:32A reptile on its back rights itself in the water.
00:00:38 → 00:00:40A third glides past, only its eyes and snout above the surface.
00:00:42 → 00:00:43From behind a screen of leaves, Scarface watches the water.
00:01:11 → 00:01:13He leaps from the bank, launching himself into the river.
00:01:20 → 00:01:23Scarface’s head emerges from the water, the caiman gripped in his jaws.

Ten of the twenty-three shown.