Skip to content

Feature / Transcription

A word-timed transcript of the whole recording.

Audio is extracted at 16 kHz mono, split into chunks and transcribed with word-level and segment-level timings. Those timings are what place every cut and time every caption.

The same workspace while a run is still cutting.

What is word-level transcription?

Word-level transcription puts a timestamp on every single word, not only on each sentence or segment. cutspool extracts the audio as 16 kHz mono, splits it into chunks and transcribes each one with both word-level and segment-level timings. Those timings do two jobs: they place a clip's cut points, and they time each word of the burned-in captions, which is why the highlight lands on the word being spoken.

The actual file this feature produces.

Granularity
Word and segment
Audio
16 kHz mono
  • 0:42

    The first thing she changed was who answers the phone

    Score
    8.9
    Length
    0:42
  • 0:52

    Why she stopped measuring tickets closed per day

    Score
    8.1
    Length
    0:52
    Ranges
    2
  • 0:34

    Hire the person who asks what happens after the ticket closes

    Score
    7.2
    Length
    0:34
Segment range
Word timing
Caption source

What the transcript carries

Word

Every word carries its own start and end against the source timeline.

Segment

Each segment carries a start, an end and the text spoken inside it.

Coverage

The whole source file, not only the ranges that end up as clips.

Placement

The word timings decide where a cut can land, so a clip opens on a sentence rather than mid-breath.

Captions

The same word list times the burned-in captions, so the two cannot drift apart.

Output

Audio
16 kHz mono, video track dropped
Granularity
Word and segment timings
Speaker labels
None
Language selection
None
Reuse
Cached against the upload

How it works

Source

Extract the audio

The video track is dropped and the audio is written out as 16 kHz mono.

The upload panel, with a source file staged and the run settings on the right.
Chunks

Split into chunks

Long recordings are cut into fixed-length chunks so a two-hour source transcribes in parallel pieces rather than one request.

The same workspace while a run is still cutting.
Timings

Transcribe with timings

Each chunk comes back with both word-level and segment-level timings, not just text.

One file goes in.

MP4 / MOV / WEBM

Upload a podcast, interview, webinar or long-form recording. Credits are charged on the length of the file you upload, and returned if the run fails or you cancel it.

47:54

interview-nadia-rademaker.mov

1.9 GB · 47:54

Uploading0%
Charged
0 credits
Rate
1 credit per source minute
Refund
Full, on failure or cancel
A representative run, played back at a readable pace. The frames and the caption overlay are drawn by the app's own components; the stage lengths here are not a claim about how long a real run takes.
Stitch

Stitch the chunks back together

Every timing is offset by its chunk's start so the finished transcript maps onto the source timeline.

The clips workspace, rendered by the app's own components with a sample run.
Handover

Hand the timings on

The clip ranges are chosen against this transcript, and the captions are timed from the same word list.

The clips workspace, rendered by the app's own components with a sample run.

What you control

Nothing. Transcription runs on every job and has no settings of its own.

Every one of these gets the same transcript.

Podcasts

Turn a 60-minute episode into a queue of reviewable vertical cuts.

Interviews

Find the compact answer without replaying the full recording.

Webinars

Pull reusable teaching moments out of a long session.

Creator content

Get a week of short posts out of one long upload.

Business content

Cut the recordings your team already has into short clips.

Where it works, and where it does not

Works best with

  • Clear speech
  • Conversation-led video
  • Distinct speaker turns
  • Source footage with enough resolution to crop to vertical

What it does not do

  • Only a finished clip's own transcript reaches the browser, not the source's. While the source is retained, editing a word changes the burned-in text; trim controls separately shorten or re-extend the clip inside its original generated timeline.
  • There are no speaker labels: the transcript records what was said, not who said it.
  • There is no source-transcript search or download. A finished clip's corrected transcript can be downloaded from its editor as SRT, VTT or TXT.
  • There is no language selection. The audio is transcribed as it is.
  • Which of these are on the roadmap

Questions about the transcript

Yes, while the source is retained with READY status. The editor shows the clip's own transcript on its timeline, and any word can be corrected - useful when a name or a technical term came back wrong. A correction changes the text that is burned in; it does not choose a new source moment.

No. The transcript carries word and segment timings only. It does not identify who is speaking, and nothing in the pipeline separates speakers.

A finished clip's corrected transcript can be downloaded from its editor as SRT, VTT or TXT. The full source transcript is not downloadable.

Ready when the recording is

Your next clip is already inside the recording.