Skip to content

Feature / AI clipping

One recording in, a list of clips out.

The recording is transcribed, then a model reads the transcript and returns the ranges worth cutting. Each clip comes back with a title, a description, hashtags and a score from 0 to 10, sorted highest score first.

The clips workspace, rendered by the app's own components with a sample run.

What is AI video clipping?

AI video clipping is software choosing the short clips inside a long recording instead of an editor scrubbing for them. The recording is transcribed with word-level timings, a model reads that transcript and returns candidate ranges with a title, a description, hashtags and a score from 0 to 10, and each range is then cut, reframed to vertical and rendered with captions burned in. In cutspool a clip is one to four ranges from the same recording, and the clips arrive sorted highest score first.

The actual file this feature produces.

Output
1080 x 1920 or 720 x 1280
Ratio
9:16
Format
MP4, H.264 + AAC
  • 0:41

    The discount that quietly trained your best customers to wait

    Score
    8.6
    Length
    0:41
  • 0:49

    Raising prices twice in one year without losing a single account

    Score
    7.9
    Length
    0:49
    Ranges
    2
  • 0:28

    Nobody reads your pricing page the way you think they do

    Score
    7.4
    Length
    0:28

What comes back with every clip

Title

One line, up to 80 characters.

Description

One to three sentences, up to 300 characters.

Hashtags

Three to eight, lowercase, each with a leading #.

Score

0 to 10, sorted highest score first.

Ranges

One to four ranges from the same recording.

Output

Ratio
9:16
Resolution
1080 x 1920 or 720 x 1280
Frame rate
30 FPS or the source's own rate, whichever is higher
Container
MP4
Video codec
H.264
Audio codec
AAC

How it works

Transcript

Transcribe the recording

Audio is pulled out as 16 kHz mono and cut into five-minute chunks, each transcribed with word-level and segment-level timings. Every chunk carries padding on both ends and the stitch keeps whichever chunk owns a word's midpoint, so a sentence running across a boundary is not transcribed twice. The finished transcript is cached against the upload, so clipping the same recording a second time skips transcription entirely.

One file goes in.

MP4 / MOV / WEBM

Upload a podcast, interview, webinar or long-form recording. Credits are charged on the length of the file you upload, and returned if the run fails or you cancel it.

58:42

episode-114-pricing.mp4

1.2 GB · 58:42

Uploading0%
Charged
0 credits
Rate
1 credit per source minute
Refund
Full, on failure or cancel
A representative run, played back at a readable pace. The frames and the caption overlay are drawn by the app's own components; the stage lengths here are not a claim about how long a real run takes.
Model

Score candidate clips

A model reads the transcript and returns each candidate as one to four ranges from the same recording, with a title, a description, hashtags and a score from 0 to 10. Because the ranges come from what was said rather than from a fixed interval, one clip can stitch two moments that belong together and leave out the minutes between them.

The clips workspace, rendered by the app's own components with a sample run.
Ranking

Sort by score

Clips come back sorted highest score first, so the strongest moment is the first one in the list. The score orders the batch rather than filtering it: every candidate the model returned is there, and a lower-scoring clip is further down rather than thrown away.

0:41
0:52
0:34
Framing

Reframe to vertical

Each range is cut from the source and framed to 9:16. By default the frame is filled by blurring and padding the recording behind it; you can switch to a crop that fills the frame edge to edge instead. The crop holds one position for the whole clip - it does not follow a speaker around the frame.

What you upload · 16:9
0:38
What you get · 9:16
Captions

Burn in the captions

When enabled, captions are burned into the picture and light up word by word as they are spoken, off the same word-level timings the transcript returned. 12 presets are the starting points; behind them sit the individual fields - face, size, weight, colours, an outline or a solid box, position, letter spacing and how many words are on screen at once.

0:34
01 / 12

White words, yellow on the beat, hard black outline.

Font
Open Sans
Words per line
3
Highlight
#FFE500
Behind the words
4px outline
Entrance
None

The 12 presets as the renderer burns them in, on a sample run.

Delivery

Download

The finished files land in your clip list as MP4s - H.264 video, AAC audio, 9:16, at 720x1280 or 1080x1920. The recording itself stays in your source library for your plan's retention window, so a second batch from it costs no upload and no transcription.

0:41
0:52
0:50
0:34
0:55

What you control

Run inputs are fixed once a job is queued. After it finishes, a generated clip can be shortened or re-extended only inside its original generated timeline while the source is retained with READY status.

Number of clips

Up to your plan's cap

Clip length

One of five ranges, 10s to 180s10-35s · 15-60s · 30-90s · 60-120s · 90-180s

Pause removal

Off, or cut silences over 1.5s, 2s or 3s

One long recording, many short ones.

Podcasts

Turn a 60-minute episode into a queue of reviewable vertical cuts.

Interviews

Find the compact answer without replaying the full recording.

Webinars

Pull reusable teaching moments out of a long session.

Creator content

Get a week of short posts out of one long upload.

Business content

Cut the recordings your team already has into short clips.

Where it works, and where it does not

Works best with

  • Clear speech
  • Conversation-led video
  • Distinct speaker turns
  • Source footage with enough resolution to crop to vertical

What it does not do

  • The original ranges come from the model. There is no arbitrary source-range picker.
  • A generated clip's start and end can move only inside its original generated timeline, and only while the source is retained with READY status.
  • A clip stitches at most four separate moments from the same recording.
  • Source size and length are capped by your plan.
  • Which of these are on the roadmap

Questions about clip selection

The recording is transcribed with word-level timings, then a model reads the transcript and returns clip ranges. Each clip comes back with a title, a description, hashtags and a score from 0 to 10. A clip can stitch up to four separate moments from the same recording.

Not arbitrarily. You set the clip length before the run, and the model returns every strong qualifying moment up to your plan's limit. Once a generated clip exists, you can shorten it or re-extend it only inside its original generated timeline while the source is retained with READY status. Picking a new moment outside that timeline requires another run; once the source expires or is deleted, source-backed editing is unavailable.

Every strong qualifying moment the model finds, up to the cap your plan sets. It does not add weaker clips to fill the allowance.

The job is retried automatically. If it fails for good, the credits are refunded and the partial output is cleaned up. You can also cancel a queued job and get the credits back.

Ready when the recording is

Your next clip is already inside the recording.