← All notes
On Post

On the cut: editing long-form from the transcript

The transcript-first workflow that moves long-form interview cuts from days to hours, and the caveats that keep it honest.

8 min

A long-form interview is a strange thing to edit. The footage is not the problem; the problem is that the good material is scattered through two hours of a person thinking out loud, and finding it means watching all of it, repeatedly, at the speed it was recorded. For years that was simply the cost. You sat with the footage until you knew it well enough to shape it, and knowing it took as long as it took.

The change that mattered was learning to edit the words before the pictures. We transcribe the interview to accurate, time-aligned text, and then we do the first structural pass in the transcript itself, reading rather than watching. You can read two hours of talk in a fraction of the time it takes to watch it, and reading is where the shape of an argument becomes visible. Deleting a sentence of text deletes the corresponding footage; rearranging paragraphs rearranges the cut. The rough assembly stops being an afternoon of scrubbing and becomes an hour of editing prose.

The honest gain here is speed on the tedious part, and it is a real gain. Finding the four sentences that matter in a rambling answer, spotting that the point made at minute ninety belongs next to the setup from minute ten, cutting the false starts and the circling: this is work the transcript makes fast because it is fundamentally about language, and language is easier to scan on a page than in a waveform. A first cut that used to take days genuinely takes hours.

Now the caveats, because a workflow sold without them is a trap. The first is that a transcript is not the film. It carries the words and throws away everything else: the pause that lands the point, the breath before the hard admission, the look away that says more than the sentence. A cut that reads perfectly on the page can be lifeless on the screen because every edit sits on a hard consonant and nothing is allowed to breathe. So the transcript pass gets us the structure, and then we go back to the footage and edit for the things text cannot hold. The words find the skeleton; the pictures put a person back on it.

The second caveat is accuracy. Automatic transcription is good and it is not perfect, and it fails in exactly the places that matter most: proper names, technical terms, the one number that had better be right, the negation that flips the meaning. An editor cutting purely on trusted text will confidently assemble a sentence the person did not say. So we treat the transcript as a map, not the territory: every kept line gets checked against what was actually spoken before it goes to the client, and the numbers and names get checked twice.

The third is subtler and it is about drift. When editing feels like word processing, it is easy to over-cut, to tighten a person into saying something crisper and less true than what they meant, because the crisp version reads better on the page. The transcript makes that frictionless, which is precisely the danger. We hold a rule against it: the cut has to be something the person would still recognise as theirs. Faster is only better if it is still honest, and the transcript does not know the difference. That part is still ours.

Used with those caveats, the workflow is one of the genuine improvements of the last few years, not because it edits the film for us, but because it clears the reading so we can spend our attention on the watching. The machine reads the two hours. The editor still makes the cut. That division has not moved, and we do not intend to let it.

Next note

On culling: narrowing thousands of frames to the keepers

On Editing · 6 min
Read →