Half a million recordings cannot be read by any reporting tool a company owns, cannot be uploaded to a third party under most contracts, and are not made useful by transcription alone. The work splits across many machines at once, so finishing in a day rather than seven months costs almost the same.
A company with half a million recorded calls going back to 2011. Stored, retained, backed up. Never read. This is the enterprise member of the set worked through in what AI can do that a chat window cannot, and it breaks the shape of the other four on purpose.
Every call was listened to once, by the person taking it, and then filed. The reason customers actually leave is somewhere in there. So is the product fault that started three years ago and never stopped. So is every promise a rep made that the company did not keep.
None of it is in the CRM. The CRM has the ticket category the rep picked from a dropdown at the end of a hard conversation.
Five hundred thousand calls, averaging seven minutes, is about fifty-eight thousand hours of audio. Listened to end to end by one person at eight hours a day, that archive takes twenty-eight years.
Why this has never been done
It is not that nobody wanted to. Three things stop it, and they stop it every time.
It is audio. There is no query you can run against a recording. Every reporting tool the company owns is useless here, because none of them can read the thing. The data has been sitting in the warehouse the whole time in a format the warehouse cannot see.
It is confidential. The calls contain customer names, account numbers, addresses, sometimes payment details. In most contracts and most jurisdictions, uploading them to a third-party service is not an option anyone will sign off on. That rules out the entire category of tools that would otherwise be the obvious answer.
Transcription alone does not help. A vendor will happily convert the archive to text. Then you have five hundred thousand transcripts, which is the same unusable pile in a different format. The work is what happens after.
That third point is the one that gets underestimated. Most attempts at this stop at transcription, declare victory, and produce a searchable corpus nobody searches.
Five models, each doing one job
The output of each becomes the input of the next. This is the part no single vendor sells, and it is where the answer actually comes from.
One: faster-whisper, large-v3. Turns every recording into text, far faster than real time.
Two: pyannote. Separates the voices and labels who spoke when. Without this a two-person call is a wall of text and you cannot tell a customer complaint from a rep's answer.
Three: Qwen3. Reads each transcript and classifies it — what kind of call was this, was it resolved, how did the customer sound by the end.
Four: Qwen3 again, structured. Pulls the specifics into consistent fields: which product, which fault, what was promised, what was refunded.
Five: BGE-M3 with clustering. Groups half a million calls by what they are actually about, then tracks each group across fifteen years to show what is growing, what is seasonal, and what quietly began and never stopped.
Step five is the one that produces something nobody has ever seen. The first four are preparation.
You do not wait longer. You run wider.
This is the part that does not apply to a smaller job. The work splits cleanly — no call depends on any other call — so it can run on as many machines at once as you are willing to pay for. The deadline becomes something you choose.
Notice that the cost barely moves. You are paying for the work, not for the calendar, so waiting six months saves nothing and there is no reason to wait.
This is why the other four jobs in this guide get a step-by-step ledger and this one does not. A ledger draws five things happening in order. What actually happens here is one thing happening two hundred times at once, and drawing it as a sequence would misrepresent the job.
Where the recordings go: nowhere
For most companies this is the question that decides the project, ahead of cost and ahead of accuracy.
The recordings go to machines started for this job and nowhere else. No third-party transcription service sees a call. No request leaves the session carrying audio.
The machines are destroyed when the job ends. Storage is wiped with them. There is no archive of your archive.
Personal details are removed inside the pipeline. Names, account numbers and card details are stripped between step one and step two, so everything downstream — and everything you keep — is already clean.
You get a record of what ran. Which models, which versions, on what, when. The thing an auditor asks for.
This is also why the work is possible at all. The same archive that cannot be uploaded to a service can be processed on machines you control.
The money, and the honest version of it
Read this part carefully before quoting it anywhere.
The saving on transcription alone is real, but it is roughly three to five times, not a hundred. Commercial speech-to-text has become genuinely cheap, and anyone in procurement will already know that. A managed transcription API would quote somewhere in the region of five figures for an archive this size; running the whole five-step chain on rented machines comes in around the low thousands. That is a good number. It is not a shocking one.
The case for doing it this way is not primarily price. It is that the recordings never leave machines you control, and that the four steps after transcription — the ones that turn text into an answer — happen in the same place, on the same data, in one pass. No transcription vendor sells you those. That is where the comparison stops being close, because on the other side of it there is no product to compare against.
What you actually get
Every call searchable in plain English — ask a question, get the calls that answer it, with the timestamp. Fifteen years of issues charted over time: what grew, what faded, what started quietly in year eleven. The reasons customers left, in their own words, grouped and counted rather than guessed at. Promises made and not kept, with the product and the date. And a cleaned, de-identified transcript archive you keep and can query again without rerunning anything.
The archive is already paid for. You are just not reading it.
If your version of this is smaller, the four everyday jobs in this guide are the same idea at a scale one person can start on a Tuesday. To talk about an archive rather than a batch, email us — enterprise work does not start with a job picker, it starts with a sample run on a few hundred of your own calls.
