Skip to content

Process

From source video to finished clip.

One run takes a recording you already have and returns vertical clips with the captions already in the picture. Here is every stage of it, including the parts you do not control.

Upload

The file goes straight to storage.

Pick an MP4, MOV or WebM. Your plan sets the size and length ceilings, and the pricing page lists both. The browser reads the file's length and size, then sends it directly to storage rather than through the API. Nothing is processed while it is still uploading.

MP4 · MOV · WebM · size and length by plan

The upload panel while the file is transferring.
Transcribe

Every word carries a timestamp.

The audio is extracted as 16 kHz mono, cut into five-minute chunks and transcribed. What comes back is word-level and segment-level timings for the whole recording. Those timings place the cuts and drive the captions, and a clip's own words can be corrected afterwards in the editor.

The same workspace while a run is still cutting.
Find moments

Candidates come back as ranges with a score.

A model reads the timestamped transcript and returns every strong qualifying clip it finds, ordered by score and capped by your plan. Each candidate names the ranges it is cut from and arrives with its own text. You set the length band before the run; the ranges themselves come from the model.

Ranges per clip
1 to 4, stitched in playback order
Score
0 to 10, highest first
Also returned
Title, description, 3 to 8 hashtags
The clips workspace, rendered by the app's own components with a sample run.
Review clips

Play what came out.

Finished clips land in the clip list with their title, description, hashtags, score and duration. You play them there and download the ones worth posting.

Adjust within the original moment

The model chooses the source range; there is no arbitrary source-range picker. While the source is retained with READY status, you can shorten a generated clip or re-extend it inside its original timeline, and change its caption text and style, framing or title before re-rendering without spending credits. If that source expires or is deleted, those source-backed edits are unavailable.
One finished clip as it appears in the list.
Format

One output shape: vertical 9:16.

Every clip renders vertical at 9:16, 1080 x 1920 or 720 x 1280, at 30 FPS or the source rate if that is higher. There is no square or landscape output. Two framings exist; either can be chosen before the run and swapped afterwards while the source is retained, and whichever is chosen holds for the whole clip.

Blurred pad

The source frame is scaled to the full width and the space above and below is filled with a blurred copy of it. Nothing is cut off. This is the default.

Fill frame

The frame is scaled up and cropped to its centre. A speaker sitting away from the centre can fall outside the crop: the crop is static and does not follow anyone.

One finished clip as it appears in the list.
Caption

Captions are burned in, word by word.

Each word is timed from the transcript and lights up as it is spoken, so the clip reads with the sound off. 12 presets are the starting points, and every field behind them is yours to set.

Presets
12 starting points
Font
One of 7 faces
Colours
Base, highlight and outline, any hex value
Outline or box
One or the other, never both
Words per line
One to five, three by default
One finished clip as it appears in the list.
Export

Download the MP4.

A finished clip is an MP4 at 1080 x 1920 or 720 x 1280. When captions are enabled they are already in the picture, and can be hidden before re-rendering and downloading. Playback and download links are signed and stay valid for seven days; opening the clip list issues fresh ones.

Container
MP4
Video
H.264, 1080 x 1920 or 720 x 1280, 9:16
Audio
AAC
The clips workspace, rendered by the app's own components with a sample run.

Between runs

The upload stays, so a second run starts further along.

An uploaded recording is kept for as long as your plan says, and it stays listed with what has already been cut from it. Starting another run from one skips the upload entirely, and skips transcription too, because the word-level transcript is cached against that file. The run still costs its source minutes.

You can delete a stored upload whenever you want. Clips already rendered from it are not touched and stay downloadable in your list, but can no longer be edited or re-rendered.

The picker behind Your uploads: each stored file is ready to be clipped again.

A job reports the stage it is in.

There is no spinner standing in for progress. The job moves through six stages and reports a percentage as it goes. A queued or running job can be stopped, and its credits come back.

  1. Upload received

    The file lands in storage and its length is verified against what the browser reported.

  2. Extracting audio

    Audio is pulled out as 16 kHz mono and split into chunks.

  3. Transcribing

    Each chunk is transcribed with word-level and segment-level timings.

  4. Finding clip candidates

    A model reads the transcript and returns clip ranges, a title and a score.

  5. Building clips

    Each clip is cut, reframed to vertical 9:16 and has its captions burned in.

  6. Ready

    The finished MP4s appear in your clip list, ready to download.

Shown mid-run: stage 03 of 06

What takes time, and what you get back.

Start

Processing begins after the upload finishes. The file is checked in storage, the credits are reserved, and the job is queued for a worker.

Transcription

The audio is split into five-minute chunks and many are transcribed at once, so a two-hour recording is not four times the wait of a 30-minute one.

Building clips

Each clip is one encode: cut, reframe, captions burned in. Three clips are built at a time, and this is the longest stage of a run.

Failures

A failed job is retried automatically. If it fails for good, the credits are returned and the partial output is removed.

Limits and formats

Supported input
MP4, MOV, WebM
Maximum file size
2 GB to 8 GB, by plan
Maximum source length
60 minutes to 4 hours, by plan
Clips per run
2 to 15, by plan
Clip length
10 to 180 seconds, set before the run
Output
1080 x 1920 or 720 x 1280, 9:16
Output format
MP4, H.264 video, AAC audio
Captions
Burned in, word-timed, 12 presets
Usage unit
1 credit per rounded-up source minute

Ready when the recording is

Your next clip is already inside the recording.