The <track> Element
Technical Summary
track associates an external timed text track with an audio or video element. Its kind identifies subtitles, captions, descriptions, chapters, or metadata intended for scripts.
Adding a track alone does not establish that captions give users complete information or that a media resource meets accessibility criteria. Keep the normative track model, file and cue behavior, browser UI, assistive-technology access, and the accessibility of the media content as separate questions.
Definition / Categories
| Item | Definition |
|---|---|
| Meaning | Specifies an explicit external timed text track for a media element |
| Categories | None |
| Context | As a child of an audio or video element, before flow content |
| Content model | Nothing; it has no text or child elements |
| End tag | None |
| DOM interface | HTMLTrackElement; its track property returns the associated TextTrack |
Purposes identified by kind
The kind attribute tells the user agent what a timed text track is for. Subtitles and captions have different purposes.
kind | Normative purpose | Display and use boundary |
|---|---|---|
subtitles | Transcription or translation of dialogue when the audio can be heard but is not understood | Intended to be overlaid on video; requires srclang |
captions | Dialogue plus relevant sound effects, music cues, and other audio information | For when sound is unavailable or not clearly audible |
descriptions | Text descriptions of the video component | Intended for audio synthesis when the visual component is not usable; support needs implementation evidence |
chapters | Timed information that identifies chapters in the media | Chapter navigation UI and behavior also need user-agent evidence |
metadata | Timed information intended for use from scripts | Not intended as ordinary displayed captions |
If kind is omitted, its normative default is subtitles. An invalid value defaults to the metadata state. Attribute defaults and the track that a browser actually selects or displays are separate questions.
Attributes and content model
| Attribute | Normative role | Condition or boundary |
|---|---|---|
src | URL of the text track data | Required; must be a valid non-empty URL |
kind | Identifies subtitles, captions, descriptions, chapters, or metadata | Defaults to subtitles; an invalid value defaults to metadata |
srclang | Language of the text track data | Must be a valid BCP 47 language tag; required when kind="subtitles" |
label | User-readable name for selecting a track | User agents use it when listing subtitle, caption, and audio-description tracks |
default | Candidate to enable if user preferences do not indicate a more suitable track | Does not override user preferences; there are per-kind limits on default tracks |
A media element may contain track children whether or not it has a src attribute. Without src, child source candidates precede the track elements. In both forms, tracks are children of the media element. See the source element page for candidate selection and parent-specific attribute rules.
<video controls>
<source src="/media/lesson.webm" type="video/webm">
<source src="/media/lesson.mp4" type="video/mp4">
<track kind="captions" src="/media/lesson.en.vtt"
srclang="en" label="English captions" default>
</video>
WebVTT cues and APIs
WebVTT is a timed text format with cues. This cue is active for the first two seconds of the video. For captions, include meaningful sounds when people need that information to understand the media.
WEBVTT
00:00:00.000 --> 00:00:02.000
Door closes.
00:00:02.000 --> 00:00:04.000
Speaker: Hello.
HTMLTrackElement.readyState reports the state of text-track data retrieval.
| Value | State |
|---|---|
NONE (0) | Track data has not loaded |
LOADING (1) | Track data is loading |
LOADED (2) | Track data has loaded |
ERROR (3) | Track data failed to load |
HTMLTrackElement.track provides the associated TextTrack. A media element's textTracks collection can also include script-added tracks and tracks from the media resource; these sources should be distinguished when observing behavior. Cue access and processing are separate API questions.
Accessibility and conformance boundary
WCAG 2.2 SC 1.2.2 requires captions for in-scope prerecorded synchronized-media audio. SC 1.2.3 allows a time-based media alternative or audio description for prerecorded video; SC 1.2.5 requires audio description for prerecorded video. Applicability depends on conditions such as synchronized versus standalone media, live versus prerecorded content, the information conveyed, and whether the resource is a clearly labeled text alternative.
track is one mechanism for providing captions or descriptions. It does not automatically guarantee accurate and complete cues, descriptions of visual information, or a text alternative for audio-only content. How browsers render cues, expose controls and tracks in accessibility trees or Windows APIs, and support them in tools such as Narrator or NVDA requires separate observation.
Fact / Evidence
| Type | Claim | Conditions / scope | Status | Source |
|---|---|---|---|---|
| SPEC | track is a child of a media element and associates external timed text; it has no end tag or content. | HTML Living Standard track element: HTML syntax and content model. | Reviewed | HTML Standard: track |
| SPEC | kind identifies five purposes, and srclang, label, and default have distinct conditions. | Distinguish subtitles from captions, BCP 47 language tags, and preference-sensitive default track selection. | Reviewed | HTML Standard: track attributes |
| SPEC | HTMLTrackElement.readyState has four states: not loaded, loading, loaded, and failed. | Values range from 0 to 3. Fixture observations of a value are kept separate from its normative definition. | Reviewed | HTML Standard: HTMLTrackElement |
| SPEC | WebVTT defines timed cues, and WCAG time-based media criteria apply according to content conditions. | The presence of a track element or file does not establish conformance of the media content. | Reviewed | WebVTT · WCAG 2.2 SC 1.2.2 |
Evidence
- HTML Living Standard: The track element — content model, attributes, kind states, selection, and HTMLTrackElement
- HTML Living Standard: Timed text tracks — text-track model, cues, loading, and APIs
- WebVTT: The Web Video Text Tracks Format — cue format and parsing rules
- WCAG 2.2 SC 1.2.2, SC 1.2.3, and SC 1.2.5 — criteria for prerecorded synchronized media
- HTML-AAM 1.0 Working Draft — an entry point for the pending mapping review; this page makes no track-specific platform-mapping claim
- Web Platform Tests: media-elements/track — test discovery entry point; this page does not claim that the directory was run
Implementation Evidence
Normative claims are recorded separately from implementation observations. This page records limited Chrome observations using video-v1 and two track-related WPT results, while preserving the unverified scope for other browsers, Windows APIs, and assistive technology.
| Type | Observation scope | Conditions to record | Status |
|---|---|---|---|
| IMPL | Loading, readyState, cue, and caption rendering for one captions track in the video-v1 Chrome fixture |
2026-09-24 / Chrome 153.0.8010.50 / Windows NT 10.0.26200.0. In the video-v1 fixture, one track had kind=captions, label=English captions, language=en, mode=showing, readyState=2, and one cue; the cue rendered on the video. The fixture set mode to showing after load. Other browsers and conditions were not observed. For the fixture and detailed Chrome observations, see video Implementation Evidence. | Partial observation (Chrome 153 / video-v1) |
| WPT | Selected WPTs for track file loading / readyState and the default attribute |
2026-09-24 / Chrome 153.0.8010.50 / Windows NT 10.0.26200.0 / wpt.live. track-load-from-src-readyState.html (1/1 pass) and track-default-attribute.html (1/1 pass) were run. track-mode.html and error-sequence.html produced no final harness summary and are uncounted. The full directory and other browsers were not run.The test version used was not recorded. The linked test may change, so this result cannot be repeated with certainty against the same version. |
Partial run (2/2 pass) |
| AAM | Caption cue rendering in video-v1 and exposure in the Chrome Accessibility Tree | 2026-09-24 / Chrome 153.0.8010.50 / Windows NT 10.0.26200.0. The cue rendered on the video but did not appear as a separate node in the captured snapshot. This applies only to that snapshot. Windows platform APIs, assistive technology such as Narrator / NVDA, and other browsers were not observed. For the full snapshot scope, see video AAM evidence. | Partial observation (Chrome AX tree) |
IMPL and AAM use limited Chrome 153 observations from video-v1; WPT covers two track-related files. Do not generalize these results to other browsers, Windows platform APIs, assistive technology, or full WPT directories.
Coverage / Open Issues
- Specification reviewedElement context and content model, five
kindvalues, attribute conditions, and HTMLTrackElement - Specification reviewedWebVTT cue scope and WCAG 2.2 time-based media applicability boundaries
- Partial / openIMPL: one caption cue loaded and rendered in the Chrome 153 video-v1 fixture; other browsers, selection, and error conditions remain unverified
- Partial / openWPT: two selected files passed 2/2 assertions. Unreported tests such as track-mode, the full directory, and other browsers remain unverified
- Partial / openA Chrome Accessibility Tree snapshot did not expose the cue as a separate node. Windows platform APIs, assistive technology, and other browsers were not observed
This is an initial specification review with limited evidence. It does not claim cross-browser interoperability or that caption files and media content conform to accessibility requirements.
Related surface
For a beginner-friendly caption and WebVTT example, see track in Yugien. The full video model is covered by video; audio playback and text alternatives are covered by audio.