Feature / AI clipping
One recording in, a list of clips out.
The recording is transcribed, then a model reads the transcript and returns the ranges worth cutting. Each clip comes back with a title, a description, hashtags and a score from 0 to 10, sorted highest score first.
What is AI video clipping?
AI video clipping is software choosing the short clips inside a long recording instead of an editor scrubbing for them. The recording is transcribed with word-level timings, a model reads that transcript and returns candidate ranges with a title, a description, hashtags and a score from 0 to 10, and each range is then cut, reframed to vertical and rendered with captions burned in. In cutspool a clip is one to four ranges from the same recording, and the clips arrive sorted highest score first.
The actual file this feature produces.





The discount that quietly trained your best customers to wait





Raising prices twice in one year without losing a single account





Nobody reads your pricing page the way you think they do
What comes back with every clip
Title
One line, up to 80 characters.
Description
One to three sentences, up to 300 characters.
Hashtags
Three to eight, lowercase, each with a leading #.
Score
0 to 10, sorted highest score first.
Ranges
One to four ranges from the same recording.
Output
- Ratio
- 9:16
- Resolution
- 1080 x 1920 or 720 x 1280
- Frame rate
- 30 FPS or the source's own rate, whichever is higher
- Container
- MP4
- Video codec
- H.264
- Audio codec
- AAC
How it works
Transcribe the recording
Audio is pulled out as 16 kHz mono and cut into five-minute chunks, each transcribed with word-level and segment-level timings. Every chunk carries padding on both ends and the stitch keeps whichever chunk owns a word's midpoint, so a sentence running across a boundary is not transcribed twice. The finished transcript is cached against the upload, so clipping the same recording a second time skips transcription entirely.
One file goes in.
Upload a podcast, interview, webinar or long-form recording. Credits are charged on the length of the file you upload, and returned if the run fails or you cancel it.

episode-114-pricing.mp4
Score candidate clips
A model reads the transcript and returns each candidate as one to four ranges from the same recording, with a title, a description, hashtags and a score from 0 to 10. Because the ranges come from what was said rather than from a fixed interval, one clip can stitch two moments that belong together and leave out the minutes between them.
Sort by score
Clips come back sorted highest score first, so the strongest moment is the first one in the list. The score orders the batch rather than filtering it: every candidate the model returned is there, and a lower-scoring clip is further down rather than thrown away.












Reframe to vertical
Each range is cut from the source and framed to 9:16. By default the frame is filled by blurring and padding the recording behind it; you can switch to a crop that fills the frame edge to edge instead. The crop holds one position for the whole clip - it does not follow a speaker around the frame.


Burn in the captions
When enabled, captions are burned into the picture and light up word by word as they are spoken, off the same word-level timings the transcript returned. 12 presets are the starting points; behind them sit the individual fields - face, size, weight, colours, an outline or a solid box, position, letter spacing and how many words are on screen at once.

White words, yellow on the beat, hard black outline.
- Font
- Open Sans
- Words per line
- 3
- Highlight
- #FFE500
- Behind the words
- 4px outline
- Entrance
- None
Download
The finished files land in your clip list as MP4s - H.264 video, AAC audio, 9:16, at 720x1280 or 1080x1920. The recording itself stays in your source library for your plan's retention window, so a second batch from it costs no upload and no transcription.


















What you control
Run inputs are fixed once a job is queued. After it finishes, a generated clip can be shortened or re-extended only inside its original generated timeline while the source is retained with READY status.
Up to your plan's cap
One of five ranges, 10s to 180s
Off, or cut silences over 1.5s, 2s or 3s
One long recording, many short ones.
Where it works, and where it does not
Works best with
- Clear speech
- Conversation-led video
- Distinct speaker turns
- Source footage with enough resolution to crop to vertical
What it does not do
- The original ranges come from the model. There is no arbitrary source-range picker.
- A generated clip's start and end can move only inside its original generated timeline, and only while the source is retained with READY status.
- A clip stitches at most four separate moments from the same recording.
- Source size and length are capped by your plan.
- Which of these are on the roadmap
Questions about clip selection
Ready when the recording is