Skip to content

Use case

Clips from podcasts without replaying the whole episode.

Upload the episode file you already have. It is transcribed with word-level timings, the moments come back as scored ranges, and each one is cut to vertical 9:16 with the captions in the picture.

Source

Start from the episode file.

An episode goes up as MP4, MOV or WebM, within the size and length ceilings your plan sets. The length is read from the file before the run, and that length is what the run is charged on.

The upload panel, with a source file staged and the run settings on the right.
Transcript

The episode is timed to the word.

Audio is pulled out, split into five-minute chunks and transcribed with word-level and segment-level timings. Those timings decide where a cut can land, so a clip opens on a sentence instead of mid-breath. There are no speaker labels: the transcript records what was said, not who said it.

The same workspace while a run is still cutting.
Candidates

Ranges come back scored.

A model reads the timed transcript and returns every strong qualifying candidate up to your plan's limit, ordered by score and carrying its ranges, title, description and hashtags. One candidate can stitch up to four separate moments from the same episode when they share a theme.

One finished clip as it appears in the list.
Output

Each clip lands vertical and captioned.

Every clip renders vertical at 9:16, 1080 x 1920 or 720 x 1280, with H.264 video and AAC audio. Word-by-word captions are burned in when enabled and can be hidden before re-rendering and downloading. The default framing keeps the whole episode frame and fills the space above and below with a blurred copy of it. While the source is retained with READY status, the caption style, framing, burned-in title and boundaries inside the clip's original generated timeline stay editable; a new source moment cannot be picked.

The clips workspace, rendered by the app's own components with a sample run.

Sample run

What one pricing episode came back with.

This is the run the frames on this page are drawn from: one episode file in, three scored candidates out, the strongest of them first. The ledger and the clip list beside it are the same run, not two illustrations of one.

Source file
episode-114-pricing.mp4
Length
58:42
Credits charged
59
Clips returned
3
Ranges cut
4 across 3 clips
Highest score
8.6
That clip's title
The discount that quietly trained your best customers to wait
Start clipping
The clips workspace, rendered by the app's own components with a sample run.

The parts that matter here.

AI clipping
Transcription
Captions

What a run costs.

An episode is charged once, on its own running time: the 58:42 file in the sample run above rounded up to 59 credits, and it would have cost the same if the run had handed back twelve clips instead of three. A run that fails puts the credits back.

See pricing

Other recordings this works on.

Interviews

Find the compact answer without replaying the full recording.

Webinars

Pull reusable teaching moments out of a long session.

Creator content

Get a week of short posts out of one long upload.

Business content

Cut the recordings your team already has into short clips.

Episode questions.

No. The transcript carries word and segment timings, not speaker names, and clips are picked on what is said instead of on who says it. Nothing in a finished clip identifies a voice.

Yes, up to four separate ranges from the same recording, stitched in the order they should play. The model does that only when the moments share one clear theme, and the joins are its own: you cannot add a range or move one.

Two stages. Transcription is split into five-minute chunks transcribed in parallel, so a 90-minute episode is not three times the wait of a 30-minute one. Building the clips is the longer stage: each clip is one encode that cuts, reframes and burns the captions in, and three are built at a time. The job reports which of its six stages it is in, and a percentage, while it works. None of that is a promised finish time.

Yes. The episode stays in your library for the window your plan sets, and starting another run from it skips the upload and skips transcription too, because the word-timed transcript is cached against that file. The second run still costs the episode's source minutes.

MP4, H.264 video and AAC audio, at 1080 x 1920 or 720 x 1280. Paid plans pick between the two and default to 1080 x 1920; Free renders at 720 x 1280. 9:16 is the only output ratio.

Ready when the recording is

The next episode is already carrying three of these.