A line can be on time while one word is wrong
Cue-level subtitle editors treat a whole caption as one block. That works until the line starts correctly but a name highlights late, a short word flashes too long, or the final word lingers after the speaker stops. Moving the entire cue fixes one edge and breaks the other.
The problem becomes obvious with active-word captions. Viewers can see the highlight miss the voice even when the cue’s overall start and end look reasonable. An editor needs access to the timing inside the cue, not just around it.
CutCaption exposes a word-level timing strip for tracks that contain word timing. Each word keeps its own start and end boundary while remaining connected to the subtitle cue and the video preview.
A focused correction workflow
- Play the problem cue and listen for the exact word that feels early or late.
- Open the cue’s word timing strip.
- Adjust that word’s start or end boundary.
- Use the neighboring words and cue edges as the safe timing limits.
- Replay the short section at normal speed.
- Turn on the intended active-word or karaoke treatment and check the visual rhythm.
- Save the project and review the same section through the libass preview before export.
The goal is not to make every waveform edge mathematically perfect. It is to make the reading rhythm follow the speech without creating flicker, overlap, or words that disappear too quickly.
Timing and styling are separate controls
Word timing decides when a word becomes active. A selected-text or single-word style override decides how it looks. Keeping those concepts separate is useful:
- correct the word boundary without changing its color or font;
- emphasize one technical term without inventing a different timing window;
- apply a global active-word treatment while keeping one local visual exception;
- switch between karaoke modes without rebuilding the transcript.
CutCaption’s five karaoke modes include sweep/fill, instant activation, word flash, word build, and per-word styling behavior. They share the underlying timed words but produce different reading patterns.
When word timing is unavailable
Not every transcription language or imported subtitle track contains word-level timestamps. SRT and VTT imports normally describe cues, not a timed record for every word. In those cases, line-level timing remains editable and effects should degrade to the timing data that exists.
Do not imply precision the source does not contain. If a project needs exact active-word synchronization, start with a transcription result that includes word timing and review it before styling.
What to inspect before export
Check quick phrases, contractions, names, and the last word before a pause. Look for zero-length or nearly invisible activations, and avoid squeezing a correction past its neighboring word. Then compare the browser preview with the final delivery path.
To turn corrected timings into a visual sequence, see the online karaoke subtitle editor guide. For the renderer behind the preview, read how libass subtitle preview works.
