Run Whisper large-v3

Ready in about 90 seconds. No install, no setup. From $0.45 per hour.

SOC 2 controls 142 verified providers 99.2% uptime Ready in about 90 seconds

What it does

Whisper listens to audio or video and writes down what was said, with timestamps.

It handles accents, background noise, crosstalk, and phone-quality recordings better than anything else available, and it works in roughly a hundred languages without being told which one it is hearing.

It will also translate as it goes, so a Spanish interview can come back as English text in one pass.

What you can make

Forty hours of interview tape transcribed overnight
Podcast episodes turned into searchable text
Subtitles generated for a video edit
Sales calls transcribed for a CRM
Lecture recordings turned into notes
A Spanish interview translated to English text

These are prompts, not fixed presets. The workspace opens with one of them already loaded so your first result is one click away.

Your setup

WE RUN THIS ON RTX 4090

We run this on an RTX 4090. It transcribes about an hour of audio in 2 minutes, so a forty hour archive is roughly 80 minutes of GPU time. $0.45 per hour, billed by the minute.

Why this card?

Whisper large-v3 only needs about 10GB of memory, so a cheaper card can run it. The RTX 4090 wins on throughput: it processes many segments of audio in parallel, and the batch size you can push through it makes a transcription job roughly three times faster than on an RTX 3090 for twice the hourly price. Faster and cheaper overall.

Memory
24GB
Speed
about 2 minutes per hour of audio
Rate
$0.45/hr

What will it cost you?

40 hours of audio 80 minutes $0.60

Billed by the minute on a RTX 4090 at $0.45 per hour. A batch job shuts the GPU off when the last hour of audio is done, so this is the whole cost.

Two ways to run it

OPEN THE APP

Click and use it

Drag a file in and read the transcript as it appears, with timestamps you can click.

Get early access Best for exploring.
RUN A BATCH JOB

Hand us the whole pile

The right mode for an archive. Connect a drive, transcribe everything, collect text and subtitle files.

Join the list Best for volume. The GPU shuts off automatically when it is done.
3,204 Whisper large-v3 jobs run on GPUVault in the last 30 days
Model facts
Parameters1.55B
LicenseApache 2.0
Memory required10GB minimum
Base modelTrained from scratch by OpenAI
PublisherOpenAI
Recommended hardwareRTX 4090, 24GB
Hugging Faceopenai/whisper-large-v3
COMING SOON

Want Whisper large-v3 the day it opens?

We are testing the rental flow with a small group first. Leave an email and we will tell you when Whisper large-v3 is ready to run. One message, no newsletter.

Whisper large-v3 questions

How much does it cost to transcribe 40 hours of audio?

About $0.60. Forty hours of audio takes roughly 80 minutes on an RTX 4090 at $0.45 per hour. Commercial transcription services charge between $0.25 and $1.50 per minute of audio, which would put the same job between $600 and $3,600.

Will it know who is speaking?

Not on its own. Whisper writes down what was said but does not separate speakers. Pair it with pyannote speaker diarization, which labels each segment by speaker, and run both in the same job. The transcription preset in the Job Runner does this for you when you tick "identify speakers".

Can I use the transcripts commercially?

Yes. Whisper is Apache 2.0 licensed with no restrictions on commercial use or on what you do with the output.

Do I need to install anything to use Whisper large-v3?

No. We start a machine with Whisper large-v3 already loaded and hand you a link. Everything runs in your browser, there is nothing to download, and nothing is left on your computer afterwards. It works the same on a Mac, a Windows laptop, or a Chromebook. A workspace is usually ready in about 90 seconds.

How much does it cost to run 40 hours of audio?

About $0.60. Whisper large-v3 takes about 2 minutes per hour of audio on the RTX 4090 we recommend, so 40 hours of audio is roughly 80 minutes of GPU time at $0.45 per hour. Billing is by the minute, and a batch job shuts the GPU off the moment the last item finishes. The calculator above works this out for your own numbers.

Can I run Whisper large-v3 on a cheaper card?

Sometimes, and the calculator will not always make it look worth it. Whisper large-v3 only needs about 10GB of memory, so a cheaper card can run it. The RTX 4090 wins on throughput: it processes many segments of audio in parallel, and the batch size you can push through it makes a transcription job roughly three times faster than on an RTX 3090 for twice the hourly price. Faster and cheaper overall. If you want to try a different card anyway, the advanced catalog lets you pick one and shows the estimated time before you commit.

Own a GPU that sits idle? Put it to work. The average provider earns $180 to $420 per month per card.

List your GPU