GPU rental pricing, card by card

Every card we run, what it costs an hour, and the reason the cheapest hourly rate is often the most expensive way to finish a job.

August 16, 20264 min readGPU rental

The short answer

GPU rental runs from about $0.22 an hour for an RTX 3090 to $3.10 for an H200. The RTX 4090 that suits most work is $0.45. What you pay depends far more on how long your job takes than on the hourly rate, so a faster card is often cheaper overall.

Pricing pages for compute tend to be either a wall of instance names that mean nothing, or a single number with an asterisk. This one is neither. Here is every card we run, what it costs, and — the part that actually decides your bill — which one your particular job should be on.

If you have not read how GPU rental works yet, start there. This piece assumes you know what you are renting and just want the numbers.

The whole price list

GPUVault hourly rental prices by card
CardMemoryClassPer hourAn 8-hour day
RTX 3090 24GB Value $0.22 $2
RTX 4090 24GB Standard $0.45 $4
RTX 5090 32GB Standard $0.69 $6
RTX A6000 48GB Large memory $0.79 $6
L40S 48GB Large memory $0.89 $7
A100 80GB 80GB Data center $1.35 $11
H100 80GB 80GB Data center $2.25 $18
H200 141GB 141GB Data center $3.10 $25
Billed by the minute. The RTX 4090 is the one most people end up on.

Billed by the minute, no minimum, nothing charged while the machine is off. The right-hand column is there because "an 8-hour day" is how most people actually think about a work session, and it makes the difference between the tiers concrete.

The hourly rate is the least useful number here

You are not buying hardware. You are buying the time it takes to finish something, and those are different purchases.

Take generating 500 images. Three models, three different cards, three very different hourly rates:

Cost of generating 500 images on three different models: SDXL Turbo $0.04, Stable Diffusion XL $0.50, FLUX.1 dev $2.23. SDXL Turbo RTX 3090 · 10 min $0.04 Stable Diffusion XL RTX 4090 · 67 min $0.50 FLUX.1 dev L40S · 150 min $2.23
500 images, start to finish, billed by the minute.
See the numbers
ModelCardTimeCost
SDXL TurboRTX 309010 min$0.04
Stable Diffusion XLRTX 409067 min$0.50
FLUX.1 devL40S150 min$2.23

The cheapest hourly rate does not win, and the most expensive one is not a disaster either. What decides it is seconds per image multiplied by the number of images, and only then multiplied by the rate.

This is why we put a calculator on every model page rather than a price list. The question "what does an A100 cost" has a tidy answer that will not help you. The question "what does my job cost" has a messy answer that will.

Which card for which work

RTX 3090 — $0.22 an hour. Same 24GB of memory as a 4090, meaningfully slower, half the price. The right choice for overnight batch work where nobody is sitting there waiting. If a job runs while you sleep, paying extra for speed buys you nothing.

RTX 4090 — $0.45 an hour. The default, and the card most jobs on this site end up on. Enough memory for every mainstream image model and for language models up to roughly 30 billion parameters. If you do not know what you need, this is it.

RTX 5090 — $0.69 an hour. 32GB rather than 24GB, and quicker. Worth it when a model is just slightly too big for a 4090, which happens more often than you would like.

RTX A6000 and L40S — $0.79 and $0.89 an hour. Both 48GB. This is the tier where video work, the larger image models like FLUX.1, and most fine-tuning live. The L40S is the faster of the two and usually the better buy despite the higher rate.

A100 80GB — $1.35 an hour. The previous generation of data-center card, and still the value pick when you need a lot of memory. Large language models fit here comfortably.

H100 80GB — $2.25 an hour. Fast, and priced accordingly. Worth it for training runs and for serving the biggest models, where the time saved is measured in hours rather than minutes.

H200 141GB — $3.10 an hour. The most memory you can get on one card. You will know if you need this, and if you are not sure, you do not.

The part nobody puts on a pricing page

A few things move the real cost that never appear in an hourly rate.

Idle time is the biggest one. A machine you left running is billed exactly like a machine doing work. Our sessions shut themselves off after 15 minutes of inactivity, which is a policy rather than a feature, and one worth checking on any platform before you use it. An RTX 4090 forgotten overnight is about five dollars. An H100 forgotten over a long weekend is over a hundred and sixty.

Startup time. Every provider has some. Ours is about 90 seconds because the model is already on the machine when you get it. Platforms that hand you a bare box and let you install things yourself are cheaper per hour and considerably more expensive per useful hour, because the first forty minutes are you reading installation notes with the meter running.

Failed jobs. We do not charge for work that did not complete, and failed sessions refund automatically. Not everyone does this. It matters more than it sounds, because early on you will fail a few jobs while you work out what you are doing.

How this compares elsewhere

Vast.ai will often be cheaper per hour than anyone, because it is a marketplace of individual machines with variable reliability and no hand-holding. RunPod sits in the middle and is genuinely good if you are comfortable with containers. Lambda is the enterprise end, priced for teams with a procurement process.

We are not the cheapest hourly rate on that list and do not claim to be. What we do is start the machine with your model already on it, cap your spend by default, and shut it down when you stop. If you know exactly what you want and enjoy configuring it, one of the others may suit you better, and we would rather say so here than have you find out afterwards.

Working out your own number

Pick the job, not the card. Every model page shows which card we run it on, why, and what your specific volume costs before you commit to anything.

Or start from the kind of work you do and let the model choose the hardware.

Questions people ask

How much does it cost to rent an H100 per hour?

{{price gpu=h100-80}} an hour on our fleet, billed by the minute. An H100 has 80GB of memory and is the card you want for training work and for the largest language models. For most image, video and audio jobs it is overkill, and an RTX 4090 at {{price gpu=rtx-4090}} finishes the same work for a fraction of the money.

How much does an A100 cost per hour?

{{price gpu=a100-80}} an hour for the 80GB version. The A100 is the previous generation of data-center card and it is still excellent value for anything that needs a lot of memory but does not need the very fastest training speed. If your model fits in 80GB and you are not in a hurry, it usually beats the H100 on total cost.

Why is the RTX 4090 so much cheaper than an A100?

Because it is a consumer card. It has 24GB of memory rather than 80GB and lacks some features that matter for large-scale training, but for running a model that already fits in its memory it is remarkably fast for the money. That is why it is the card most jobs on this site land on, and why we default to it unless a model needs more room.

Is a cheaper card always cheaper overall?

No, and this is the mistake worth avoiding. You pay for time, not for hardware. A card at half the hourly rate that takes three times as long costs you 50% more. Work out the cost of the whole job rather than comparing rates, which is what the calculator on every model page does before you start.

Is a GPU server the same thing as a graphics card?

Not quite, and the two words get used as if they were. The graphics card — the GPU, or just "the card" — is the component that does the work. The GPU server is the whole computer it sits inside, with its own processor, memory and disk, in a data center. When you rent from us you get a dedicated server with a card in it, not a share of somebody else's, so the card's full memory and full speed are yours for the length of the job.

Do prices change?

Yes. Hourly rates move with supply, and anyone publishing a price page that never changes is either not paying attention or not telling you. Every price on this site comes from one table that the whole site reads, so when a rate moves, every page moves with it on the same day. If you spot a stale number anywhere, it is a bug and we want to hear about it.

COMING SOON

Renting opens shortly. Want to know when?

We are testing the rental flow with a small group before we open it to everyone. Leave an email and we will tell you the day it opens. One message, no newsletter.

Own a GPU that sits idle? Put it to work. The average provider earns $180 to $420 per month per card.

List your GPU