Process
From source video to finished clip.
One run takes a recording you already have and returns vertical clips with the captions already in the picture. Here is every stage of it, including the parts you do not control.
The file goes straight to storage.
Pick an MP4, MOV or WebM. Your plan sets the size and length ceilings, and the pricing page lists both. The browser reads the file's length and size, then sends it directly to storage rather than through the API. Nothing is processed while it is still uploading.
Every word carries a timestamp.
The audio is extracted as 16 kHz mono, cut into five-minute chunks and transcribed. What comes back is word-level and segment-level timings for the whole recording. Those timings place the cuts and drive the captions, and a clip's own words can be corrected afterwards in the editor.
Candidates come back as ranges with a score.
A model reads the timestamped transcript and returns every strong qualifying clip it finds, ordered by score and capped by your plan. Each candidate names the ranges it is cut from and arrives with its own text. You set the length band before the run; the ranges themselves come from the model.
- Ranges per clip
- 1 to 4, stitched in playback order
- Score
- 0 to 10, highest first
- Also returned
- Title, description, 3 to 8 hashtags
Play what came out.
Finished clips land in the clip list with their title, description, hashtags, score and duration. You play them there and download the ones worth posting.
Adjust within the original moment
One output shape: vertical 9:16.
Every clip renders vertical at 9:16, 1080 x 1920 or 720 x 1280, at 30 FPS or the source rate if that is higher. There is no square or landscape output. Two framings exist; either can be chosen before the run and swapped afterwards while the source is retained, and whichever is chosen holds for the whole clip.
The source frame is scaled to the full width and the space above and below is filled with a blurred copy of it. Nothing is cut off. This is the default.
The frame is scaled up and cropped to its centre. A speaker sitting away from the centre can fall outside the crop: the crop is static and does not follow anyone.
Captions are burned in, word by word.
Each word is timed from the transcript and lights up as it is spoken, so the clip reads with the sound off. 12 presets are the starting points, and every field behind them is yours to set.
- Presets
- 12 starting points
- Font
- One of 7 faces
- Colours
- Base, highlight and outline, any hex value
- Outline or box
- One or the other, never both
- Words per line
- One to five, three by default
Download the MP4.
A finished clip is an MP4 at 1080 x 1920 or 720 x 1280. When captions are enabled they are already in the picture, and can be hidden before re-rendering and downloading. Playback and download links are signed and stay valid for seven days; opening the clip list issues fresh ones.
- Container
- MP4
- Video
- H.264, 1080 x 1920 or 720 x 1280, 9:16
- Audio
- AAC
Between runs
The upload stays, so a second run starts further along.
An uploaded recording is kept for as long as your plan says, and it stays listed with what has already been cut from it. Starting another run from one skips the upload entirely, and skips transcription too, because the word-level transcript is cached against that file. The run still costs its source minutes.
You can delete a stored upload whenever you want. Clips already rendered from it are not touched and stay downloadable in your list, but can no longer be edited or re-rendered.
A job reports the stage it is in.
There is no spinner standing in for progress. The job moves through six stages and reports a percentage as it goes. A queued or running job can be stopped, and its credits come back.
Upload received
The file lands in storage and its length is verified against what the browser reported.
Extracting audio
Audio is pulled out as 16 kHz mono and split into chunks.
Transcribing
Each chunk is transcribed with word-level and segment-level timings.
Finding clip candidates
A model reads the transcript and returns clip ranges, a title and a score.
Building clips
Each clip is cut, reframed to vertical 9:16 and has its captions burned in.
Ready
The finished MP4s appear in your clip list, ready to download.
What takes time, and what you get back.
Processing begins after the upload finishes. The file is checked in storage, the credits are reserved, and the job is queued for a worker.
The audio is split into five-minute chunks and many are transcribed at once, so a two-hour recording is not four times the wait of a 30-minute one.
Each clip is one encode: cut, reframe, captions burned in. Three clips are built at a time, and this is the longest stage of a run.
A failed job is retried automatically. If it fails for good, the credits are returned and the partial output is removed.
Limits and formats
- Supported input
- MP4, MOV, WebM
- Maximum file size
- 2 GB to 8 GB, by plan
- Maximum source length
- 60 minutes to 4 hours, by plan
- Clips per run
- 2 to 15, by plan
- Clip length
- 10 to 180 seconds, set before the run
- Output
- 1080 x 1920 or 720 x 1280, 9:16
- Output format
- MP4, H.264 video, AAC audio
- Captions
- Burned in, word-timed, 12 presets
- Usage unit
- 1 credit per rounded-up source minute
Ready when the recording is








