Eighty hours of back catalogue, unsearchable

Transcription is the easy half. Knowing who said which line is what turns a wall of text into something you can actually use.

By Cory FechnerAugust 24, 20265 min readBeyond the chat window

The short answer

Transcribing eighty hours is fast and nearly free. The step that makes it useful is speaker separation, which labels who spoke when, and no chat tool does it. With speakers labelled you get chapters, a clip list with timestamps, and show notes for episodes that never had any.

A podcaster with four years of episodes. Two hosts, frequent guests, no transcripts, no chapters, and no real idea what is in there. This is the last of the four everyday jobs in what AI can do that a chat window cannot, and it is the cheapest and fastest of them by some distance.

What she wants: a clean transcript of every episode with the speakers labelled, chapter markers, a list of the moments worth clipping, and show notes for the episodes that never got any.

Why the chat window stalls

Upload limits. Most chat interfaces cap audio around thirty minutes. A ninety-minute episode has to be chopped into three by hand, and the splits land mid-sentence — usually mid-anecdote, which is precisely the material worth finding.

No speaker labels. This is the capability gap, and it is total rather than partial. Chat tools will return the words. They will not tell you who said them. A two-host transcript without labels is a wall of text where a question and its answer are indistinguishable, which makes it nearly useless for the thing she wanted it for.

And there are a hundred and sixty episodes. Even at ten minutes of handling each — upload, wait, download, name, file — that is a working week of nothing but administration.

The models that do the work instead

faster-whisper, on the large-v3 model, transcribes far faster than real time. Eighty hours of audio does not take eighty hours to read; it takes about eighty minutes. Speech-to-text is the mature end of this field and it shows.

pyannote separates the voices and labels who spoke when. This is the step that turns a transcript into something usable, and it is genuinely a different model doing a different task — it is not listening to words at all, it is listening to voices.

Qwen3 reads the finished, labelled transcripts and produces the chapters, the clip candidates with timestamps, and the show notes. A language model, doing the language-model part, at the end of a chain that was mostly not about language.

The run, step by step

The run $3.80 of $25.00 cap
  1. STEP 1Transcribe 80 hours1 hr 20 min
  2. STEP 2Separate the speakers55 min
  3. STEP 3Align the timestamps18 min
  4. STEP 4Chapters and clips40 min
  5. STEP 5Write show notes22 min
Machine time3h 35m
Your attention~15 min
Cost of the run$3.80

Transcription and speaker separation together are more than half the run. Aligning timestamps is quick and load-bearing: it is what makes a clip list usable, because a moment you cannot jump straight to is a moment you have to go find.

Her part is ninety seconds. This job is usually finished before anyone remembers to check on it.

What you will have when it is done

A transcript per recording, speakers named and timestamps throughout. Chapter markers you can paste straight into your player. And a clip list — the strongest moments with exact start and end times.

The clip list is the one that changes what people do with their archive. Four years of episodes is not a thing anyone re-listens to, so the good material inside it is functionally lost. A ranked shortlist with timestamps turns it back into inventory.

Show notes for the episodes that never had any come out of the same pass, and they are drafts. Publishable after a read-through, not before.

Accuracy, honestly

Transcription quality is not uniform and it is worth knowing where it degrades. Crosstalk — two people talking over each other, which is most of what makes a good conversation good — is the hardest case for both models at once. Strong accents, poor microphones, and specialist vocabulary each cost accuracy.

For finding things, that is fine. Search tolerates errors well: you are looking for the passage, and a misspelled word nearby does not stop you finding it. For publishing, it is not fine, and anyone promising otherwise has not checked their own output.

The same distinction runs through the document job, where extraction good enough to search is a much lower bar than extraction good enough to report from.

Why this job is the cheapest of the four

Under four dollars for eighty hours of audio is low enough that people assume something has been left out. Nothing has. Speech-to-text is simply the most mature capability in this whole field, and maturity shows up as speed.

Transcription runs many times faster than the audio plays, so the eighty hours in the archive is not eighty hours of anything. Speaker separation is slower per minute but still nowhere near real time. The language model at the end reads text, and text is cheap to read compared with anything involving pixels.

Compare that with the short video job, which costs roughly five times as much for a fraction of the source material, because generating video frames is the most demanding thing in the set.

The practical consequence is that audio archives are the easiest place to start. If you have never handed a pile to a machine before and you want to find out whether any of this is real, a back catalogue of recordings will tell you for the price of a coffee and about four hours of waiting.

The same job at fifty times the size

Everything here scales, and at a certain size it stops being the same job. Half a million recorded support calls is the enterprise version, and past a certain volume the sequence changes shape: rather than running five steps in order, the work splits across many machines at once, and the finish date becomes something you choose.

The models are largely the same ones. What changes is that transcription stops being the point. With half a million calls, converting them all to text just gives you the same unreadable pile in a different format, and the value is entirely in the four steps after.

How it works shows this job as a handover: three steps that are yours, five that are not.

Questions people ask

What is speaker labelling and why does it matter so much?

It works out how many distinct voices are in a recording and marks who spoke each line. Without it a two-host transcript is an undifferentiated wall of text, and the whole point was to find the moments worth clipping. Chat tools return the words but not who said them.

Do I need to split long episodes first?

No. Length is not a constraint here, which is one of the main differences from a chat tool. Upload limits in chat interfaces cap most files around thirty minutes, so every episode has to be chopped by hand and the splits land mid-sentence, breaking exactly the passages you wanted.

How accurate are the transcripts?

Good enough to search and clip from, not good enough to publish unread. Crosstalk, accents, poor microphones and technical vocabulary all cost accuracy. For a back catalogue where the goal is finding things, that trade is fine; for a published transcript, budget an editing pass.

What is a clip list?

The strongest moments across the archive with exact start and end times, produced by a model reading the finished transcripts rather than listening. It is a shortlist for a person to review, not a finished edit, and it is usually the single most useful output of the job.

Is under four hours realistic for eighty hours of audio?

It is illustrative rather than measured on our own runs. Transcription runs far faster than real time, so the audio length matters less than you would expect. Audio quality, the number of speakers and how much crosstalk there is move it more than total hours do.

EARLY ACCESS

Renting opens shortly. Want to know when?

We are opening rentals to a small group at a time. The first 100 accounts keep the rate they join at for twelve months, whatever our prices do afterwards. Leave an email and we will tell you when your turn comes. One message, no newsletter.

Own a GPU that sits idle? Put it to work. The average provider earns $180 to $420 per month per card.

List your GPU