Skip to content

Feature / Captions

Captions built from the same transcript.

When captions are enabled, they are timed word by word from the transcript that placed the cut and burned into the file. The word being spoken flips to a highlight colour. 12 presets are the starting points, and every field behind them is yours to set - including the entrance each line arrives on.

The clips workspace, rendered by the app's own components with a sample run.

What are burned-in captions?

Burned-in captions are drawn into the video's own pixels rather than shipped beside it as a subtitle file, so when enabled they appear on every platform and in every player. cutspool burns them in word by word, timed from the same transcript that placed the cut; while the source is retained with READY status, the clip editor can hide them or change their styling before re-rendering and downloading.

The actual file this feature produces.

Timing
Word level
Render
Burned in
Presets
12
  • 0:34

    The cheapest thing on this desk is the one I would replace last

    Score
    8.2
    Length
    0:34
  • 0:35

    Stop buying a second screen before you fix the first one's height

    Score
    7.7
    Length
    0:35
  • 0:29

    Cable management is a one-hour job you will do four times

    Score
    7.1
    Length
    0:29

What a caption style is made of

Face

One of six fonts. You cannot upload your own.

Size and weight

Set as a proportion of the frame height, so 720 x 1280 and 1080 x 1920 read the same.

Colours

Base, highlight and outline, each any six-digit hex value.

Outline or box

One or the other. The box replaces the outline rather than sitting behind it.

Words per line

One to five, three by default, broken early on a noticeable pause.

Position

One of three vertical positions in the frame.

Output

Presets
12 named starting points
On a free plan
2 presets and their highlight colour
On a paid plan
All 12 presets and every field
Timing
Word level, from the transcript that placed the cut
Render
Burned into the picture when enabled, never a separate file

How it works

Timings

Take the word timings

The caption track is built from the same word-level timings that placed the cut, so a caption cannot drift out of sync with the audio.

The same workspace while a run is still cutting.
Lines

Group the words into lines

One to five words per line, three by default, broken early when there is a noticeable pause between words.

0:34
01 / 12

White words, yellow on the beat, hard black outline.

Font
Open Sans
Words per line
3
Highlight
#FFE500
Behind the words
4px outline
Entrance
None

The 12 presets as the renderer burns them in, on a sample run.

Highlight

Highlight the spoken word

Each word takes the highlight colour for exactly as long as it is being said, then hands it to the next word.

The clips workspace, rendered by the app's own components with a sample run.
Render

Burn them into the render

When enabled, captions are drawn in the same encode pass as the reframe, so they are part of the video rather than a separate subtitle file.

The clips workspace, rendered by the app's own components with a sample run.
Brand

Add your handle and logo

If you set them, your handle is drawn under the captions and your logo above centre, on every clip in the run.

One finished clip as it appears in the list.

What you control

Pick a preset, then change any field on it. Free plans get 2 of the 12 presets and their highlight colour; paid plans get all 12 and every field.

Caption preset

One of 12 named starting points

Caption style

Font, size, weight, caps, colours, outline or box, position, words per line

Output resolution

1080 x 1920 or 720 x 1280

Brand handle

Text line burned under the captions

Brand logo

Image overlaid above centre

Captioned the same way, whatever the recording.

Podcasts

Turn a 60-minute episode into a queue of reviewable vertical cuts.

Interviews

Find the compact answer without replaying the full recording.

Webinars

Pull reusable teaching moments out of a long session.

Creator content

Get a week of short posts out of one long upload.

Business content

Cut the recordings your team already has into short clips.

Where it works, and where it does not

Works best with

  • Clear speech
  • Conversation-led video
  • Distinct speaker turns
  • Source footage with enough resolution to crop to vertical

What it does not do

  • A style carries an outline or a background box, never both. The renderer's box replaces the outline rather than sitting behind it.
  • The box is drawn around each rendered line, not as a full-width banner.
  • 7 font faces, no upload of your own.
  • One to five words per line, and one of three vertical positions.
  • Captions are burned into the MP4 when enabled, and can be hidden before re-rendering and downloading. A finished clip's corrected transcript can also be downloaded as SRT, VTT or TXT; the full source transcript is not downloadable.
  • Which of these are on the roadmap

Questions about caption styling

Yes, while the source is retained with READY status. Start from one of 12 named presets, then set the font from 7 faces, the size, the weight, uppercase, the base and highlight colours, the outline colour and width, the shadow, a solid background box, the vertical position and the words per line. A style carries an outline or a background box, never both, because the renderer's box replaces the outline instead of sitting behind it. Free plans get 2 presets and their highlight colour; paid plans get all 12 and every field.

Yellow, #FFE500, on the Broadcast preset. Every preset sets its own, and any six-digit hex colour can replace it before the run.

No. When enabled, they are burned into the picture during the encode, so they travel with the MP4 and cannot be switched off by the platform you post to. You can hide them in the clip editor before re-rendering and downloading.

No. Every size, outline width, shadow and margin is a proportion of the frame height, so the same style renders the same at 720 x 1280 and 1080 x 1920.

Ready when the recording is

Your next clip is already inside the recording.