What are captions?
Captions are synchronized text that represents the audio content of a video. Unlike subtitles, which only translate dialogue, captions include all meaningful audio information: spoken words, speaker identification, sound effects, music descriptions, and other relevant audio cues.
Captions can be either open (always visible, burned into the video) or closed (can be toggled on and off by the viewer). Closed captions are generally preferred because they give users control over their viewing experience.
Who benefits from captions?
Captions provide access for a wide range of users:
- Deaf and hard of hearing users – Captions provide the primary means of accessing audio content in video.
- People in noisy environments – Viewers in loud spaces such as airports, gyms, or open offices can follow video content through captions.
- Non-native speakers – Reading along with spoken content supports comprehension for people learning a language or who are more comfortable reading English than hearing it.
- People without audio devices – Users who don’t have headphones or speakers available can still access video content.
- Viewers who prefer to watch without sound – Many people choose to watch videos on mute, especially on mobile devices and social media.
- People with cognitive disabilities – Captions reinforce spoken content and support comprehension for users with processing differences.
Types of captions
Understanding the different types of captions helps you choose the right approach for your content:
- Closed captions (CC) – Can be toggled on and off by the viewer. This is the most common and preferred format for online video.
- Open captions – Permanently visible in the video. These are useful for social media platforms that may not support closed caption files, or when you want to ensure captions are always seen.
- Real-time captions (CART) – Generated live during events, meetings, or broadcasts. For more information, see Live Captions.
- Auto-generated captions – Created automatically by platforms like YouTube, Zoom, or Kaltura. These are a good starting point but always require review and editing for accuracy.
What makes a good caption?
Accurate, well-formatted captions make a significant difference in the user experience. When creating or reviewing captions, keep these quality standards in mind:
- Accuracy – Captions should match the spoken content exactly. Names, technical terms, and acronyms must be correct.
- Synchronization – Caption timing should align closely with the audio. Captions that appear too early or too late can be confusing.
- Speaker identification – When multiple speakers are present, identify who is speaking, especially when the speaker is not visible on screen.
- Sound descriptions – Include meaningful non-speech sounds in brackets, such as [applause], [music playing], or [phone ringing].
- Readability – Use proper punctuation, capitalization, and line breaks. Avoid displaying too much text on screen at once.
- Completeness – Do not paraphrase or omit spoken content. Caption all meaningful audio.
Techniques for recorded video
The following platform-specific guidance applies to pre-recorded video. For live or real-time video, see Live Captions.
Most video platforms support uploading caption files in SRT or VTT format. If you have an existing caption file, you can typically upload it directly to any of these platforms.
YouTube
YouTube has built-in caption support and automatic captioning:
- Edit existing captions
- Add new captions to a video
- Use YouTube’s auto-caption feature as a starting point
- Always review and edit auto-generated captions for accuracy
- Add speaker names and sound descriptions
YouTube’s automatic captions have improved significantly but still make errors with technical terms, proper nouns, and accented speech. Always review and correct auto-generated captions before publishing.
Vimeo
Vimeo provides caption support for both uploaded and linked videos:
- Vimeo captions and subtitles guide
- Upload SRT or VTT caption files directly to your videos
- Vimeo also offers auto-captioning on certain plans – review for accuracy before publishing
Kaltura
Kaltura is UGA’s primary video platform (UGA MediaSpace) and supports captions. Video owners can upload caption files or use Kaltura’s built-in captioning tools. Kaltura also offers machine-generated captions that can serve as a starting point, but these must be reviewed and edited for accuracy.
For assistance with Kaltura captioning, submit a request through the UGA Captioning Service request form.
Zoom
Zoom meetings and webinars can be recorded with automatic transcription:
- Enable Zoom transcription
- Enable live transcription during meetings for real-time captions
- Save transcription when recording ends
- Review transcription for accuracy before sharing the recording
Zoom’s auto-generated captions are helpful for live meetings but should be reviewed and corrected before distributing recorded content. For more on live meeting accessibility, see Online Classes and Meetings.
eLC
If hosting video in UGA’s learning management system eLC (Desire2Learn/Brightspace), ensure captions are present in the video platform you use. Videos embedded in eLC from Kaltura, YouTube, or other platforms should already have captions applied at the source. For more information, see E-Learning Accessibility.
Social media
Video on social platforms has specific captioning needs. Many social media users watch videos with sound off, making captions especially important. Open captions (burned into the video) are often the most reliable option for social media, since not all platforms support closed caption files consistently.
For platform-specific guidance on accessible social media content, see Videos.
Prioritizing your videos for captioning
If you have a large library of uncaptioned videos, prioritize based on these factors:
- Audience demographics – If the target audience is likely to include individuals who are deaf or hard of hearing, the video should be a top priority.
- Traffic – Your most popular and widely viewed videos should be a high priority.
- Publication date – Newer videos should generally be prioritized over older content.
- Instructional use – Videos used in courses or training should be captioned promptly to ensure equal access for students.
In general, focus initial efforts on high-impact videos such as videos available to the public on a high-use website, videos used multiple times in a course, and videos developed by multiple faculty members for use across several classes.
Common mistakes to avoid
- Relying solely on auto-generated captions: Machine-generated captions are a starting point, not a finished product. They frequently misidentify technical terms, proper nouns, and accented speech.
- Missing speaker identification: When multiple people are speaking, failing to identify who is talking makes captions confusing and less useful.
- Omitting non-speech sounds: Sound effects, music, and other audio cues that are important to understanding the content must be included in captions.
- Poor synchronization: Captions that appear significantly before or after the corresponding audio create a disorienting experience.
- Paraphrasing or summarizing: Captions should represent the actual spoken content, not a condensed version of it.
- Forgetting to caption embedded videos: Videos embedded from external platforms still need captions. Verify captions are present at the source before embedding.
Additional resources
- Captioning Key – Best practices for creating high-quality captions
- W3C WAI: Captions/Subtitles – W3C guidance on caption requirements and techniques
Related DASH resources
WCAG criteria
Captions support these WCAG 2.1 success criteria:
- 1.2.1 Audio-only and Video-only (Prerecorded) (Level A)
- 1.2.2 Captions (Prerecorded) (Level A)
- 1.2.4 Captions (Live) (Level AA)
Getting help
If you have questions or need assistance, the Digital Accessibility Services Hub (DASH) is available to support colleges, schools, and administrative units as they build sustainable accessibility practices.
- Contact DASH at [email protected] for consultations, training, or technical assistance.
Accessibility is a shared responsibility, and every step you take makes UGA’s digital environment more inclusive.