# Captions, Audio Description, and Intelligibility for Spoken Audio

Map where speech lives in a finished recording, transcribe it into captions in the format the player expects, align caption timing to the speech, mix in audio description with program ducking, and prove the message survives degraded listening.

> Canonical page: https://elysiatools.com/en/hubs/audio-accessibility-captions-and-description

- **Keywords:** audio accessibility workflow, generate captions from audio, SRT VTT caption timing, audio description mixing, ducking for narration, assistive listening enhancement, emergency alert intelligibility, voice activity detection

## Frequently asked questions

### Why must the edit be locked before captioning?

Caption timing is attached to the cut. Move one sentence and every caption after it sits wrong, which means redoing alignment across the whole program. Transcribe once the picture and the audio are final, and the caption work survives.

### My captions drift progressively — doesn't one more offset fix that?

A constant offset fixes a constant delay. Drift that grows across the program points at a frame-rate or pull-down mismatch between the caption file and the media — re-derive the captions against the correct source rather than nudging timing further.

### How do I choose ducking depth for description?

Set depth by the loudest program passage that runs under narration, not the average — if the narration reads clearly there, it reads clearly everywhere. Keep the release quick so music and dialogue bounce back instead of sinking.

### What does the alert intelligibility test actually simulate?

The channels a real warning meets — narrow phone-speaker bandwidth, lossy compression, and background noise. It reports whether the heuristic speech clarity of the alert survives each degradation, so you fix the message before the emergency, not during it.

## Related content

- [Podcast Recording, Editing, and Delivery](https://elysiatools.com/en/hubs/podcast-recording-editing-delivery): Rescue remote voices, clean speech artifacts, shape the edit, master loudness, add show assets, and validate podcast files before delivery.
- [Audio Analysis and Visualization Tools](https://elysiatools.com/en/hubs/audio-analysis-and-visualization-tools): Visualize audio, inspect stereo, diagnose signal and timing, and check files before delivery.
- [Audio Restoration, Noise Repair, and Cleanup Tools](https://elysiatools.com/en/hubs/audio-restoration-repair-workflows): Diagnose and repair clicks, hum, clipping, crackle, room sound, hiss, and harshness without promising lossless restoration.
