✓ ISO 17100 aligned | ✓ Native translators | ✓ Certified & notarised

Get a Free Quote

How to Make Videos Accessible in Multiple Languages

How to Make Videos Accessible in Multiple Languages

Make videos accessible in multiple languages using subtitles, dubbing, audio description, and sign language — a UK workflow with ISO 17100 quality control.

Request a Free Quote

Contact Form

20% Discount

for all NEW CLIENTS

Problems with the form? Send an email to info@translationservicesuk.co.uk

What we do

Our services

What does it mean to make a video accessible in multiple languages?

Making a video accessible in multiple languages means layering five localisation assets — transcript, translated subtitles, dubbed audio, audio description, and sign-language interpretation — so viewers follow the content regardless of the language they speak or a sensory disability.

What are the 5 layers of the multilingual video accessibility stack?

The five layers are a source-language transcript, translated subtitles or SDH, dubbed or voice-over audio tracks, audio description for visual information, and sign-language interpretation — each solving a distinct barrier and mapped to a specific WCAG 2.2 success criterion.

Captions vs subtitles vs SDH — which one does your video need?

Choose closed captions when the audio and audience share a language, translated subtitles when the audience speaks a different language than the audio, and SDH when the audience is both Deaf or hard-of-hearing and reads a different language.

How do you produce multilingual subtitles in an ISO 17100 workflow?

Producing ISO 17100 multilingual subtitles takes 4 sign-offs: transcribe the source audio, translate the transcript into each target language, revise the translation against the video’s visual context, and QA the timed subtitle file before final export.

Dubbing vs voice-over vs subtitles — what does each cost per finished minute in the UK?

In the UK, professional subtitles cost roughly £15 to £30 per finished minute per language, UN-style voice-over £40 to £80, and lip-synced dubbing £150 to £400, with dubbing carrying the longest lead time because it requires studio recording and mix.

How do you add audio description for blind and low-vision viewers?

Add audio description by writing a description script that fits the natural pauses in dialogue, recording it with a professional voice artist, and delivering it as a separate audio track (extended AD) or a mixed master where the description ducks the original audio.

Making a video usable across languages and disabilities requires more than uploading a caption file. It requires a five-layer localisation stack — transcript, translated subtitles, dubbed audio, audio description, and sign-language interpretation — produced under ISO 17100 quality control and mapped to WCAG 2.2 clauses that UK and EU regulators enforce.

What does it mean to make a video accessible in multiple languages?

Making a video accessible in multiple languages means layering five localisation assets — transcript, translated subtitles, dubbed audio, audio description, and sign-language interpretation — so viewers can follow the content regardless of the language they speak or a sensory disability they live with. A single caption track solves one barrier; the full stack addresses language, hearing, vision, and cognitive barriers simultaneously, which is why accessibility-first productions plan every asset before the source cut is locked.

A transcript is the indispensable foundation of the entire workflow. It provides a full plain-text record of everything spoken in the video, acts as the source document for translation into every target language, and stands alone as an accessibility asset for users who cannot play video at all. Without an accurate transcript, every downstream deliverable — subtitles, dubs, audio description scripts — carries forward whatever errors the transcription introduced.

A practical multilingual accessibility project touches four disciplines: transcription, translation, voice production, and platform delivery. Each discipline carries its own quality gate, and each output file — SRT, WebVTT, TTML, WAV, MP4 — must conform to the destination player’s ingest specification. Building accessibility in from the start of production, rather than retrofitting it after the video is published, consistently reduces cost, shortens delivery time, and produces higher-quality results across all language versions.

What are the 5 layers of the multilingual video accessibility stack?

The five layers are a source-language transcript, translated subtitles or SDH, dubbed or voice-over audio tracks, audio description for visual information, and sign-language interpretation — each solving a distinct barrier and mapped to a specific WCAG 2.2 success criterion.

LayerDeliverableBarrier solvedWCAG 2.2 clause
1. TranscriptPlain-text.txt or.docxAudio-only / SEO / deaf-blind refreshable braille1.2.1 Audio-only and Video-only (Prerecorded)
2. Translated subtitles / SDHSRT, WebVTT, TTMLLanguage barrier and hearing loss1.2.2 Captions (Prerecorded), Level A
3. Dubbed / voice-over audioWAV or AAC track, or muxed MP4Language barrier for low-literacy or child viewers1.2.3 Audio Description or Media Alternative (Level A)
4. Audio descriptionSeparate AD track or mixed masterBlind and low-vision viewers1.2.5 Audio Description (Prerecorded), Level AA
5. Sign-language interpretationPicture-in-picture insetDeaf viewers using BSL/ASL as first language1.2.6 Sign Language (Prerecorded), Level AAA

Layer 1 — What is a transcript and why start there?

A transcript is the full plain-text record of a video’s audio and it is the source asset from which every other layer is produced, which is why the workflow starts here. A verbatim transcript also acts as a standalone accessibility asset for deaf-blind users on refreshable braille displays and improves search indexing, because crawlers read text rather than watching video.

Our transcript pack is priced from £30 per page and delivered as a translated document under the same discipline as our Document Translation Services, with the source-language transcript feeding directly into the translation queue for every downstream language.

Layer 2 — How do translated subtitles and SDH differ?

Translated subtitles render spoken dialogue in a second language, while SDH (Subtitles for the Deaf and Hard-of-Hearing) additionally carry speaker labels and non-speech cues such as [door slams] or [music swells]. Standard subtitles assume the viewer hears the audio and only needs a language bridge; SDH assumes the viewer does not hear at all and needs the full acoustic scene rendered in text.

  • Translated subtitles: dialogue only, one language other than the audio.
  • SDH: dialogue + speaker IDs + sound effects + music cues, in the target language.
  • Closed captions: same-language transcription including non-speech audio, toggled on by the viewer.

Layer 3 — When should you dub the audio instead of subtitling?

Dub the audio when the target audience reads slowly, when the video is under 3 minutes and text would obscure the frame, or when platform norms expect a native voice track — children’s content, broadcast TV in France, Germany, Italy, Spain. Dubbing replaces the original voice track with a lip-synced target-language recording; voice-over (UN-style) leaves the original audio audible under a narrated translation.

Layer 4 — What does audio description add for blind viewers?

Audio description is a narrated track inserted between dialogue that describes visual information — actions, settings, on-screen text, and body language — so blind and low-vision viewers receive the same story information as sighted viewers. It sits on a secondary audio channel and slots into the natural pauses in the original mix. WCAG 2.2 clause 1.2.5 requires it at Level AA when significant visual information is not already carried by the dialogue.

Layer 5 — Why is sign-language interpretation a separate deliverable?

Sign-language interpretation is a separate deliverable because BSL, ASL, and other sign languages are full natural languages with their own grammar, not English on hands, and many Deaf viewers use sign as a first language even when they can read captions. British Sign Language has its own word order, spatial grammar, and idiom, and a qualified interpreter is filmed as a picture-in-picture inset (commonly bottom-right) so the sign track runs in parallel with the video.

Captions vs subtitles vs SDH — which one does your video need?

Choose closed captions when the audio and audience share a language, translated subtitles when the audience speaks a different language than the audio, and SDH when the audience is both Deaf or hard-of-hearing and reads a different language. Each format serves a distinct need, and confusing them leaves part of your audience without the access they require.

Closed captions are a text version of all audio — speech and non-speech — presented in the same language as the video and toggled on or off by the viewer. They satisfy WCAG 2.2 clause 1.2.2. Subtitles, by contrast, carry only translated spoken dialogue in a language different from the original audio; they are primarily a localisation tool rather than an accessibility one, though they are vital for making content useful and shareable across language markets. SDH — subtitles for the Deaf and hard-of-hearing — combines both functions: it delivers translated text in a different language while also including speaker identifications and non-speech audio cues such as [music], [applause], or [door slams], meeting the needs of Deaf viewers in a foreign-language market in a single track.

FeatureClosed captionsSubtitlesSDH
Language relative to audioSame languageDifferent languageDifferent language
Includes non-speech cuesYesNoYes
Includes speaker IDsYesNoYes
Primary audienceDeaf / hard-of-hearing / noisy environmentsForeign-language viewersDeaf viewers in a foreign market
WCAG clause satisfied1.2.2Not a WCAG requirement (localisation)1.2.2 for the target language

Open captions share the same content as closed captions but are permanently embedded in the video frame itself and cannot be edited or toggled off by the viewer — a distinction that matters when you need to guarantee caption display on a player that does not support subtitle tracks. UK licensed broadcasters must caption up to 80% of programming, provide signing on 5%, and audio-describe 10% under regulatory targets, while online content published by public-sector bodies is bound by WCAG 2.2 AA. From 28 June 2025, the European Accessibility Act extends mandatory accessible audiovisual media requirements to products and services sold into the EU.

How it works

How do you produce multilingual subtitles in an ISO 17100 workflow?

1

Which subtitle file format should you export — SRT, WebVTT, or TTML?

Export SRT for YouTube, Vimeo, and social uploads, WebVTT for HTML5 web players and styling control, and TTML/IMSC for BBC iPlayer and broadcast delivery. Burn subtitles into the MP4 only when the destination player cannot ingest a sidecar file — burned-in captions cannot be toggled and fail viewers who need a different language.

FormatFull nameBest forStyling support
SRTSubRip TextYouTube, Vimeo, LinkedIn, MetaMinimal
WebVTT (.vtt)Web Video Text TracksHTML5 <track> players, custom web playersPositioning, CSS classes
TTML / IMSCTimed Text Markup LanguageBBC iPlayer, broadcast, OTT platformsFull — fonts, colours, regions
Burned-in MP4Rendered pixelsSocial feeds with autoplay mutedN/A (not toggleable)

2

How long does subtitle translation take per video minute?

Human-quality subtitle translation runs about 8 to 12 minutes of translator work per finished video minute, so a 10-minute video is delivered in 24 to 48 hours under our ISO 17100 pack, or same-day when the recording is under 1,000 source words and placed London before 11:00 GMT. Reading-speed constraints (15–17 cps for adults, 12 cps for children) drive the tighter end of that range because translators have to condense rather than translate word-for-word.

Pricing

Dubbing vs voice-over vs subtitles — what does each cost per finished minute in the UK?

In the UK, professional subtitles cost roughly £15 to £30 per finished minute per language, UN-style voice-over £40 to £80, and lip-synced dubbing £150 to £400, with dubbing carrying the longest lead time because it requires studio casting, recording, and final mix. The right format for your project depends on budget, audience expectations, and the level of immersion the content demands — dubbing replaces the original voice track entirely with a lip-synced target-language recording, while voice-over lays a translated narration over a ducked original and subtitles keep the original audio intact.

Deliverable£ per finished minuteTurnaround (10-min video)What you get
Translated subtitles (SRT/VTT)£15 – £3024–48 hoursTimed subtitle file, ISO 17100 reviewed
SDH subtitles£20 – £3524–48 hoursSubtitles + speaker IDs + sound cues
Voice-over (UN-style)£40 – £803–5 working daysNarrated track over ducked original
Lip-synced dubbing£150 – £40010–15 working daysFull replacement voice cast, mix, master
Audio description£25 – £603–5 working daysScript + voiced AD track
BSL/ASL interpretation inset£90 – £2005–10 working daysFilmed interpreter, keyed and composited

Subtitles are the most cost-efficient entry point into multilingual accessibility and are particularly effective in environments where video is frequently watched without sound — a significant proportion of online video consumption happens in silent or sound-restricted settings, making subtitled content accessible to audiences who would otherwise skip it entirely. Dubbing is the premium option: it provides a fully immersive experience but requires a qualified voice cast and studio time, which is why quality-assured dubbing projects follow ISO 17100 — the translation-services standard that mandates a qualified translator, an independent reviser, and a formal QA sign-off — before any audio is committed to tape.

How do you add audio description for blind and low-vision viewers?

Audio description is a narrated account of the significant visual information in a video — actions, scene changes, on-screen text, and relevant body language — inserted into the natural pauses between dialogue so that blind and low-vision viewers receive the same informational content as sighted audiences. WCAG 2.2 clause 1.2.5 requires audio description at Level AA whenever meaningful visual information is not already conveyed by the spoken audio; a talking-head interview with descriptive narration needs far less AD than a wordless demonstration or an action sequence.

To achieve the highest standard of accessibility, audio description should be commissioned as part of the original production schedule rather than added after the edit is locked, because retrofitting AD into a densely edited sequence often requires cutting additional pauses into the timeline — a process that lengthens delivery and increases cost. A clear, well-lit video showing the speaker’s face also supports viewers who lip-read, making considered production choices a multiplier across multiple accessibility outcomes.

  1. Spot the video for description gaps — identify every silent or low-dialogue moment longer than approximately two seconds where visual information is being conveyed.
  2. Draft the AD script — describe actions, settings, on-screen text, and body language in concise, objective language without overlapping existing dialogue.
  3. Time-code each description line — align every AD cue to the exact frame at which it should enter to avoid clashing with speech.
  4. Record with a matched voice artist — select a voice whose tone, pace, and register suit the video’s style; record in a treated studio environment.
  5. Deliver the correct stems — provide either a separate AD audio track (extended AD) or a mixed master in which the original audio ducks by 6–10 dB beneath the description, depending on the destination player’s capability.

Sign-language interpretation is a complementary access service: a qualified interpreter is filmed signing the full audio content in BSL, ASL, or another sign language and composited as an on-screen inset, ensuring Deaf viewers who use sign language as their primary language receive the content in the most natural form for them.

When do you need sign-language interpretation and how is it filmed?

Sign-language interpretation is required when serving Deaf audiences whose first language is BSL, ASL, or another sign language, and it is filmed as a picture-in-picture inset with a qualified interpreter working from the finalised script under the same ISO 17100 review discipline. For UK broadcast, Ofcom requires licensed broadcasters to sign 5% of programming; for online content, sign-language provision is Level AAA under WCAG 2.2 clause 1.2.6.

  • Source the interpreter from a qualified register — for BSL, that is Signature or NRCPD in the UK.
  • Brief from the locked script, not the raw audio, so terminology and names are agreed in advance.
  • Film against a plain background (mid-tone blue or grey) with even lighting on hands and face.
  • Composite as a PIP inset sized at 25–33% of the frame width, positioned bottom-right by convention.
  • Peer-review the recording with a second qualified Deaf reviewer before delivery.

Which UK and EU laws require multilingual video accessibility?

Three overlapping regulatory frameworks govern video accessibility for UK and EU audiences, and each framework imposes distinct obligations on different categories of publisher. Understanding which regime applies to your organisation is the essential first step before scoping any multilingual video project.

  • Ofcom — UK licensed broadcasters are required to caption up to 80% of qualifying programming, provide audio description on 10%, and carry sign-language interpretation on 5%. These are hard targets against which channel compliance is measured each year. Failure to meet them is a regulatory matter, not merely a best-practice shortfall.
  • Public Sector Bodies Accessibility Regulations 2018 (PSBAR) — UK public-sector bodies must ensure their video content meets WCAG 2.2 AA. That standard includes clause 1.2.1 (audio-only and video-only alternatives), clause 1.2.2 (captions at Level A), clause 1.2.3 (audio description or media alternative at Level A), and clause 1.2.5 (audio description at Level AA). A public-sector video published without captions is already in breach of domestic law.
  • European Accessibility Act (EAA) — from 28 June 2025, the EAA mandates accessible audiovisual media services for products and services sold into the EU. Organisations that distribute video content commercially in any EU member state must ensure that content meets accessibility requirements as of that date, regardless of where the organisation is headquartered.
  • Equality Act 2010 — the reasonable-adjustment duty under this Act covers accessible video for any service offered to the public in Great Britain, giving individuals a potential avenue for challenge where accessibility barriers are not addressed.

ISO 17100 is the international quality standard for translation services and, while not a law, is required by many public procurement frameworks and by regulated industries. It mandates a qualified human translator, an independent reviser, and a formal QA sign-off for every translation — including subtitle and dubbing scripts — which means machine-only outputs do not satisfy ISO 17100 compliance. For regulated content such as compliance training, evidence videos, or public communications, meeting the translation quality standard is as important as meeting the technical accessibility specification.

For regulated content where subtitles and transcripts need to accompany certified documentary evidence — an evidence video paired with a certified transcript, for example — our Legal Translation Services handle the certification requirements in parallel with the accessibility deliverables.

Why are automatic YouTube captions not enough for multilingual accessibility?

Automatic captions carry a word-error rate that can reach 5 to 15 percent under normal production conditions, and that error rate rises sharply with accented speech, technical vocabulary, proper nouns, and code-switching between languages — the exact contexts that multilingual audiences encounter most. They mis-punctuate speaker changes, fail to insert non-speech cues, and cannot reliably translate idiom, all of which means they fall short of WCAG 1.2.2 and cannot satisfy the ISO 17100 requirement for a qualified human translator plus an independent reviser.

Some video platforms do allow a single video to carry multiple audio tracks in different languages, enabling viewers to switch to their preferred language from within their account settings. However, this multi-language audio feature does not automatically create those tracks — each dubbed audio track still has to be recorded by a qualified human voice artist and uploaded manually. The platform’s own automatic dubbing tools remain limited in scope and are not available to all publishers, which means most multilingual video projects cannot rely on automation for the voice production stage at all.

Auto-translated captions compound the problem further: they inherit every error from the auto-transcription layer beneath them, then introduce a second layer of errors through machine translation. The result is caption text that misrepresents both the original speaker and the target language, undermining the trust and comprehension of the audience segments those captions are meant to serve. Professional multilingual captions — built from a human-reviewed transcript, translated under ISO 17100, and timed to broadcast-standard reading speeds — are the only reliable route to genuine compliance and audience reach.

What is a practical checklist for shipping a multilingual, accessible video?

Shipping a multilingual accessible video is a sequenced workflow where each stage depends on the quality of the one before it. The transcript is the source of truth for every downstream deliverable, so a transcription error propagates into subtitles, dubs, and audio description scripts simultaneously — which is why locking an accurate, time-coded transcript before any translation begins is the most important quality gate in the entire process.

  1. Lock the source-language master — picture and audio must be final before any translation or timing work begins; changes to the edit after subtitles are timed require full retiming.
  2. Transcribe the audio — produce a full plain-text record of all speech, including speaker IDs and non-speech audio cues, time-coded to the frame.
  3. Translate the transcript into each target language under ISO 17100 — a qualified translator produces the first draft; an independent reviser reviews it; a QA sign-off completes the cycle before any subtitle timing or voice recording begins.
  4. Time subtitles to broadcast reading speeds — target 15–17 characters per second, maximum two lines per card, maximum 42 characters per line; subtitle timing is a specialist skill distinct from translation.
  5. Record voice-over or dubbing audio in a treated studio — deliver labelled stems (dialogue, music, effects) to allow post-production flexibility; lip-synced dubbing replaces the original voice track entirely and requires a full voice cast and final mix.
  6. Write, time, and record the audio description track — description lines must fit the natural pauses in the original dialogue without overlapping speech; deliver either a separate AD stem or a mixed master.
  7. Film the BSL or ASL interpreter and composite as a picture-in-picture inset — the interpreter must be a qualified signer for the specific sign language required; inset size and position must not obscure critical on-screen information.
  8. Embed subtitle tracks with correct BCP-47 language codes — use en-GB, es-ES, ar-SA, and equivalent codes so players surface the correct language to the correct viewer automatically.
  9. Apply player-level accessibility settings — disable autoplay-with-sound; confirm no content flashes more than three times per second (a WCAG seizure-safety requirement).
  10. Test on the accessible player — verify keyboard navigation, caption toggle, screen-reader control announcements, and correct audio-track switching across the destination platforms, including mobile browsers.

Building this checklist into your production schedule before the shoot — rather than treating accessibility as a post-delivery task — keeps every deliverable on the critical path and avoids the cost and delay of retrofitting access assets into a locked edit.

How do you get a UK-based multilingual video translation quote?

Request a quote by sending the finalised video file or an existing transcript, specifying the target languages, and identifying the destination platform. With that information we can return a fixed-price breakdown covering every access layer your project requires — transcription, translated subtitles, SDH, voice-over, lip-synced dubbing, audio description, and BSL or ASL interpretation — all produced under ISO 17100 quality assurance, meaning every translation goes through a qualified translator, an independent reviser, and a formal sign-off before delivery.

Our per-minute rates for subtitles start from £15 per language, with same-day options available for orders placed before 11:00 GMT. We deliver across more than 200 languages, and our team handles the full range of subtitle formats — SRT, WebVTT, and TTML — with correct BCP-47 language codes embedded so your video player surfaces the right track to the right viewer automatically.

When your multilingual video also needs to sit alongside certified documentary evidence — an evidence video paired with a certified transcript, or compliance training accompanied by a certified translation of the source script — our Online Certified Translation Services in the UK handle the certification requirement in parallel, so both the accessibility deliverables and the legal documentation arrive on the same timeline.