Feature / Transcription
A word-timed transcript of the whole recording.
Audio is extracted at 16 kHz mono, split into chunks and transcribed with word-level and segment-level timings. Those timings are what place every cut and time every caption.
What is word-level transcription?
Word-level transcription puts a timestamp on every single word, not only on each sentence or segment. cutspool extracts the audio as 16 kHz mono, splits it into chunks and transcribes each one with both word-level and segment-level timings. Those timings do two jobs: they place a clip's cut points, and they time each word of the burned-in captions, which is why the highlight lands on the word being spoken.
The actual file this feature produces.



The first thing she changed was who answers the phone



Why she stopped measuring tickets closed per day



Hire the person who asks what happens after the ticket closes
What the transcript carries
Word
Every word carries its own start and end against the source timeline.
Segment
Each segment carries a start, an end and the text spoken inside it.
Coverage
The whole source file, not only the ranges that end up as clips.
Placement
The word timings decide where a cut can land, so a clip opens on a sentence rather than mid-breath.
Captions
The same word list times the burned-in captions, so the two cannot drift apart.
Output
- Audio
- 16 kHz mono, video track dropped
- Granularity
- Word and segment timings
- Speaker labels
- None
- Language selection
- None
- Reuse
- Cached against the upload
How it works
Extract the audio
The video track is dropped and the audio is written out as 16 kHz mono.
Split into chunks
Long recordings are cut into fixed-length chunks so a two-hour source transcribes in parallel pieces rather than one request.
Transcribe with timings
Each chunk comes back with both word-level and segment-level timings, not just text.
One file goes in.
Upload a podcast, interview, webinar or long-form recording. Credits are charged on the length of the file you upload, and returned if the run fails or you cancel it.

interview-nadia-rademaker.mov
Stitch the chunks back together
Every timing is offset by its chunk's start so the finished transcript maps onto the source timeline.
Hand the timings on
The clip ranges are chosen against this transcript, and the captions are timed from the same word list.
What you control
Nothing. Transcription runs on every job and has no settings of its own.
Every one of these gets the same transcript.
Where it works, and where it does not
Works best with
- Clear speech
- Conversation-led video
- Distinct speaker turns
- Source footage with enough resolution to crop to vertical
What it does not do
- Only a finished clip's own transcript reaches the browser, not the source's. While the source is retained, editing a word changes the burned-in text; trim controls separately shorten or re-extend the clip inside its original generated timeline.
- There are no speaker labels: the transcript records what was said, not who said it.
- There is no source-transcript search or download. A finished clip's corrected transcript can be downloaded from its editor as SRT, VTT or TXT.
- There is no language selection. The audio is transcribed as it is.
- Which of these are on the roadmap
Questions about the transcript
Ready when the recording is