Generate automatic captions, correct them by hand, and attach them as a caption file. Twenty minutes per video for most content, and the corrections are what make them worth having.

Who is actually reading them

The obligation is usually framed around deaf and hard of hearing viewers, and that is a real and sufficient reason on its own.

It substantially understates the audience.

A large share of video on phones is watched with the sound off: in public, at work, in bed, on a bus, in a waiting room. Captions are how those people watch anything at all.

People watching in a second language use them. People with attention or processing difficulties use them. People in noisy environments use them.

Which means captions are one of the few accessibility measures whose main beneficiaries are the general audience, and it makes the business case straightforward rather than principled.

Automatic captions are a first draft

Every major platform will generate captions automatically, and the quality is far better than it was a few years ago.

It is not good enough to publish unedited.

Automatic recognition handles ordinary speech well and fails predictably on the things that matter most in a business video: product names, technical terms, place names, people's names, and anything said with an accent or over background noise.

Those are precisely the words a viewer needs, so the errors are concentrated in the useful parts.

Punctuation is the other weakness, and a wall of unpunctuated text is difficult to follow even when every word is correct.

The workflow

Faster than people expect once it is a routine.

For a three-minute video this is around twenty minutes. For an hour of recorded talk it is a couple of hours, which is when paying somebody starts to make sense.

Keep the caption file

A practical point that saves the work being repeated.

Once corrected, download the caption file rather than leaving it inside one platform.

It is a small text file with timings, and it can be attached to the same video wherever else it is published, uploaded to a different platform, or used if you move hosting.

It is also the source for a transcript, which is worth having separately.

Businesses that skip this end up captioning the same video three times because it appears on the site, on a social platform, and in a newsletter.

Burned-in captions are not the same

Worth distinguishing, because the two are frequently confused.

Captions burned into the picture are part of the image. They cannot be turned off, resized, restyled, or read by assistive technology, and they cannot be translated.

They are common on social platforms because they survive being reposted, and for a short promotional clip that is a reasonable choice.

For anything instructional or substantial, a proper caption file is better, because the viewer controls it and the text exists as text.

Where a video will appear in both places, do both: a burned-in version for social and a captioned file for your own site.

A worked example

A training provider had about forty short videos published over three years, none captioned, on the reasoning that it was a large project.

They tested one. Automatic captions were generated in a few minutes, and correcting them took eighteen minutes, most of which was the product names and one recurring technical term.

Forty videos at that rate was around twelve hours, which they spread over three weeks.

Two things happened that they had not anticipated.

Watch time rose on the captioned videos, which they attributed to people watching without sound.

And the corrected caption text, published as transcripts beneath each video, began attracting search traffic, because forty videos had become forty pages of relevant text where previously there had been none.

Publish the transcript too

The step that produces most of the unexpected benefit.

A transcript is the caption text without timings, laid out as readable paragraphs beneath the video.

It helps anybody who would rather read than watch, which is a substantial group and includes almost everybody trying to find one specific detail.

It gives search engines text to work with, where a video on its own is close to opaque.

And it makes the content usable by somebody who cannot watch video at all, for whatever reason.

It costs almost nothing once the captions are corrected, which is why it is worth doing at the same time rather than as a separate project.

Where to start with a backlog

Do not work chronologically.

Order the videos by views over the last year and start at the top, since the benefit is proportional to who is watching.

Then take anything embedded on a page that matters: the homepage, a main service page, or anything used in a sales conversation.

Then anything instructional, where somebody is trying to follow along and a missed word has consequences.

Leave the rest indefinitely. A backlog of forty is a project; the eight that matter is an afternoon.

The counter-case

Not every video needs the full treatment.

A silent atmospheric clip on a homepage with no speech has nothing to caption, and adding a caption track saying music playing is not useful.

What such a video does need, if it conveys information, is that the information exist somewhere in text on the page.

There is also a version of this that becomes disproportionate: paying for professional captioning of an internal video watched by six people is money that would do more elsewhere.

The test is whether anybody is watching and whether the audio carries meaning. Where both are true, caption it. Where neither is, leave it.

The routine

  1. Order your videos by views over the last year.
  2. Generate automatic captions on the top eight.
  3. Correct names, terms, numbers and punctuation.
  4. Download and keep the caption file.
  5. Publish the transcript beneath the video.
  6. Caption new videos before publishing, not after.
  7. Leave the tail uncaptioned without guilt.

Step six is what stops the backlog rebuilding, and it adds twenty minutes to a process that already took hours.

The original case for captions is covered in a video with no captions.


Frequently asked questions

Who actually uses captions?

Far more than deaf and hard of hearing viewers. A large share of phone video is watched with the sound off, and people watching in a second language or in noisy places rely on them.

Are automatic captions good enough?

As a first draft. They handle ordinary speech well and fail on product names, technical terms, place names, and accents, which are exactly the words that matter.

How long does it take?

About twenty minutes for a three-minute video: generate, play through, correct names and punctuation, check timing. An hour of recorded talk is a couple of hours.

Should I burn captions into the video?

For short social clips, reasonably. For anything instructional use a caption file, since burned-in text cannot be resized, restyled, translated, or read by assistive technology.

Why publish a transcript as well?

It helps people who would rather read, it gives search engines text where a video is otherwise opaque, and it costs almost nothing once the captions are corrected.

Where do I start with a large backlog?

Order by views over the last year and do the top eight, plus anything embedded on a page that matters. Leave the tail indefinitely.

West Coast Media Solutions Inc. provides web design, web development, hosting, digital marketing, and business consulting to organisations across Canada, drawing on more than twenty-five years in the field.

Forty uncaptioned videos?

Order them by views and do the top eight. Twenty minutes each, and publish the transcript underneath.

Start a Conversation