The Captions That Drifted Nine Seconds After the Edit: Timeline Repair for VTT Files in the Session Archive

Nine seconds is not a rounding error. It is long enough for a viewer to read a punchline before the speaker delivers it, long enough for a sponsor mention to land on the wrong slide, and long enough for an accessibility reviewer to flag the archive as non-conforming. When a VTT file drifts by nine seconds after an edit, the cause is almost never the captioner. It is a timeline mismatch introduced between the caption file and the final media asset.

This article is a repair procedure for that specific failure: a WebVTT file whose cues are internally consistent but globally offset from the published video. It assumes you have the final media file, the VTT file, and a deadline. Every step is reversible and testable before the archive goes public.

What WebVTT actually guarantees

The W3C WebVTT specification defines a cue as a text segment associated with a time interval. Each cue has a start time, an end time, and a textual payload. The format is designed for time-aligned text tracks referenced from an HTML <track> element. The specification is explicit that the file is a sequence of cues, each with its own interval. It does not define a global offset, a drift correction, or a synchronization anchor. That is the root of the problem: WebVTT tells you what each cue says and when it should appear relative to the media timeline, but it has no mechanism to detect or correct a mismatch between that timeline and the media file it accompanies.

The MDN WebVTT API documentation reinforces the same model. A text track is a container for time-aligned text data played in parallel with a video or audio track. Cues are the individual units. The API allows you to add, remove, and inspect cues programmatically, but it does not provide a built-in drift detector. If the media element’s timeline and the cue timestamps disagree, the browser will render the cues at the times written in the file. It will not compensate.

This means the repair must happen outside the player, in the file itself, before publication.

Why nine seconds, and why after the edit

A nine-second offset is a suspicious number. It is not the kind of drift that accumulates from clock skew over a two-hour session. It is the kind of offset introduced by a discrete edit: a title card added at the head, a sponsor bumper inserted before the first session, a cold-open trimmed, or a slate removed. If the caption file was generated against the raw recording and the published video has a different head, the entire file shifts by the duration of that difference.

The same logic applies to HLS packaging. RFC 8216 describes HTTP Live Streaming as a protocol for delivering continuous streams via playlists of media segments. Each segment has a duration declared in the playlist. If the caption file was timed against a different segmentation or a different playlist version, the cue timestamps will not align with the segment boundaries the player actually uses. The RFC does not define a caption synchronization mechanism; it defines media segments and playlists. Caption alignment is the operator’s responsibility.

So the first diagnostic question is not “how do we fix the drift?” It is “what changed between the caption source and the published media?” The answer determines whether you apply a constant offset or rebuild the file.

Diagnostic: confirm the offset is constant

Before you shift anything, verify that the drift is a constant offset and not a progressive drift. A constant offset means every cue is early or late by the same amount. Progressive drift means the offset grows over time, which points to a frame-rate mismatch or a timestamp discontinuity.

Use three checkpoints: the first cue, a cue near the midpoint, and the last cue. For each, note the cue’s start time in the VTT file and the actual time the corresponding speech occurs in the final media. If all three differences are within 200 milliseconds of each other, treat it as a constant offset. If the difference grows by more than 500 milliseconds between the first and last checkpoint, treat it as progressive drift and rebuild the file from a fresh transcription against the final media.

This threshold is a working recommendation, not a specification. The WebVTT spec does not define an acceptable synchronization tolerance. The 200-millisecond figure is a practical target for caption readability; the 500-millisecond figure is a practical trigger for rebuild. Document your own thresholds in the runbook so the decision is not made under pressure.

Repair: apply a constant offset

If the offset is constant, the repair is a timestamp shift. You are adding or subtracting the same duration from every cue’s start and end time. The cue text, identifiers, and settings remain unchanged.

Do not edit the file by hand. A two-hour session can contain 1,200 or more cues. Manual editing introduces transcription errors and is not reversible. Use a script that reads the VTT, parses the timestamp lines, applies the offset, and writes a new file. Keep the original file untouched.

The WebVTT timestamp format is HH:MM:SS.mmm or MM:SS.mmm. The specification allows both. Your script must handle both forms. It must also preserve the cue identifier lines, the cue settings, and any NOTE blocks. A NOTE block is a comment that starts with the word NOTE and ends at the first blank line. If your script strips comments, you lose the provenance information that tells you which captioner or tool produced the file.

Here is the logic in plain terms:

  1. Read the VTT file line by line.
  2. Identify timestamp lines by the presence of the string -->.
  3. Parse the start and end timestamps.
  4. Add or subtract the offset in milliseconds.
  5. Reformat the timestamps to the original precision.
  6. Write the new file with a distinct name, such as session-archive-offset.vtt.

If the offset would push a cue’s start time below zero, clamp it to zero and log the cue. A negative start time is invalid in WebVTT. If the offset would push a cue’s end time beyond the media duration, clamp it to the media duration and log the cue. These edge cases usually affect only the first and last cues, but they must be handled explicitly.

Verification: test before you publish

After applying the offset, verify the result against the final media file. Do not rely on the player’s default rendering. Load the media and the new VTT file in a controlled environment and check the same three checkpoints you used for diagnosis. The first cue should appear within 200 milliseconds of the corresponding speech. The midpoint cue should be within 200 milliseconds. The last cue should be within 200 milliseconds.

If the checkpoints pass, the offset is correct. If they fail, the offset is wrong or the drift is not constant. Revert to the original file and re-diagnose.

Also verify that the file still conforms to the WebVTT format. The specification requires the file to begin with the string WEBVTT. It requires a blank line between the header and the first cue. It requires each cue to have a start time, an end time, and a payload. A malformed file may be rejected by the player or rendered incorrectly. If you have a validator, run it. If you do not, check the first 10 lines and the last 10 lines manually.

Prevention: anchor the caption file to the final media

The repair is straightforward once you know the offset. The harder problem is preventing the offset in the first place. The root cause is almost always a mismatch between the media used for captioning and the media published. The fix is to caption against the final media, or to record the exact difference between the caption source and the final media.

In a live or hybrid event, the captioner may be working from a live feed while the archive is assembled from a separate recording. The two feeds may have different start times, different pre-roll, or different post-roll. If the caption file is exported from the live captioning system and paired with the archive recording without adjustment, the offset is inevitable.

The operational fix is to add a synchronization checkpoint to the post-event workflow. After the archive media is finalized, play the first 30 seconds and the last 30 seconds. Note the timecode of a distinctive spoken phrase in each. Compare those timecodes to the corresponding cue timestamps in the VTT file. If the difference is more than 200 milliseconds, apply the offset before publishing. This takes less than five minutes and catches the error before it reaches the audience.

For hybrid events, the room-to-chat handoff is a common source of timeline confusion. If the archive includes both the room feed and the chat replay, the two timelines may not start at the same moment. The caption file must be aligned to the primary media timeline, not the chat timeline. If your workflow links the two, verify which timeline the captions reference. The article on why hybrid events fall apart at the room-to-chat handoff covers the broader failure pattern; the caption alignment is a specific instance of it.

What about TTML and other formats?

WebVTT is not the only timed text format. TTML2 is a W3C Recommendation for timed text interchange. It defines a system model for authoring, transcoding, and presentation. Like WebVTT, it associates text with time intervals. Unlike WebVTT, it is designed for interchange among legacy distribution systems. If your archive pipeline uses TTML internally and exports VTT for the web player, the offset may be introduced during transcoding. Check the transcoding step for any time-shifting options. The TTML2 specification does not define a drift correction mechanism either; the responsibility is with the processor.

If you are working with HLS, RFC 8216 defines WebVTT as a supported media segment format. The RFC does not define caption synchronization. If your HLS packaging introduces a discontinuity, the caption file must be adjusted accordingly. The EXT-X-DISCONTINUITY tag signals a discontinuity in the media timeline. If a discontinuity is present and the caption file does not account for it, the cues will drift. Check the playlist for discontinuity tags and verify that the caption file’s timeline matches the media timeline after each discontinuity.

FAQ

How do I know if the drift is constant or progressive?
Check three cues: first, midpoint, last. If the offset is the same at all three, it is constant. If it grows, it is progressive. Progressive drift usually means a frame-rate mismatch or a timestamp discontinuity. Rebuild the file.

Can I fix the VTT file in a text editor?
You can, but you should not. A two-hour session can have over 1,200 cues. Manual editing is error-prone and not reversible. Use a script and keep the original file.

What is the acceptable synchronization tolerance?
The WebVTT specification does not define one. A practical target is 200 milliseconds. If the offset exceeds 500 milliseconds, rebuild the file. Document your own thresholds.

Does the player compensate for drift?
No. The browser renders cues at the times written in the file. The MDN WebVTT API documentation describes the cue model but does not define automatic drift correction.

What if the offset pushes a cue before zero?
Clamp the start time to zero and log the cue. Negative start times are invalid in WebVTT.

How do I prevent this in the future?
Caption against the final media, or record the exact difference between the caption source and the final media. Add a synchronization checkpoint to the post-event workflow. Verify the first and last 30 seconds before publishing.

Copyable artifact: VTT offset repair checklist

For the post-production operator or archive engineer:

VTT OFFSET REPAIR CHECKLIST

1. IDENTIFY THE OFFSET
   - Play the final media. Note the timecode of the first spoken phrase.
   - Open the VTT file. Note the start time of the first cue.
   - Calculate the difference. Record it in milliseconds.
   - Repeat for a midpoint cue and the last cue.
   - If the difference is constant within 200 ms, proceed.
   - If the difference grows by more than 500 ms, rebuild the file.

2. BACK UP THE ORIGINAL
   - Copy the original VTT file to a backup location.
   - Do not edit the original in place.

3. APPLY THE OFFSET
   - Use a script to parse timestamp lines (lines containing "-->").
   - Add or subtract the offset from start and end times.
   - Clamp start times to zero if negative.
   - Clamp end times to media duration if beyond.
   - Preserve cue identifiers, settings, and NOTE blocks.
   - Write to a new file: session-archive-offset.vtt

4. VERIFY
   - Load the new VTT file with the final media in a controlled player.
   - Check the first, midpoint, and last cues.
   - Confirm each is within 200 ms of the corresponding speech.
   - Confirm the file begins with WEBVTT and has a blank line before the first cue.

5. PUBLISH
   - Replace the original VTT file with the offset file.
   - Keep the backup until the archive is confirmed.
   - Log the offset value and the reason in the session record.

6. PREVENT
   - Caption against the final media, or record the source-to-final difference.
   - Add a 5-minute synchronization check to the post-event workflow.
   - For hybrid events, verify which timeline the captions reference.
   - For HLS, check for EXT-X-DISCONTINUITY tags and verify alignment.

This checklist is a working procedure, not a specification. The thresholds are practical recommendations. Adjust them to your workflow and document your own.