GPU rental runs from about $0.22 an hour for an RTX 3090 to $3.10 for an H200. The RTX 4090 that suits most work is $0.45. What you pay depends far more on how long your job takes than on the hourly rate, so a faster card is often cheaper overall.
Pricing pages for compute tend to be either a wall of instance names that mean nothing, or a single number with an asterisk. This one is neither. Here is every card we run, what it costs, and — the part that actually decides your bill — which one your particular job should be on.
If you have not read how GPU rental works yet, start there. This piece assumes you know what you are renting and just want the numbers.
The whole price list
| Card | Memory | Class | Per hour | An 8-hour day |
|---|---|---|---|---|
| RTX 3090 | 24GB | Value | $0.22 | $2 |
| RTX 4090 | 24GB | Standard | $0.45 | $4 |
| RTX 5090 | 32GB | Standard | $0.69 | $6 |
| RTX A6000 | 48GB | Large memory | $0.79 | $6 |
| L40S | 48GB | Large memory | $0.89 | $7 |
| A100 80GB | 80GB | Data center | $1.35 | $11 |
| H100 80GB | 80GB | Data center | $2.25 | $18 |
| H200 141GB | 141GB | Data center | $3.10 | $25 |
Billed by the minute, no minimum, nothing charged while the machine is off. The right-hand column is there because "an 8-hour day" is how most people actually think about a work session, and it makes the difference between the tiers concrete.
The hourly rate is the least useful number here
You are not buying hardware. You are buying the time it takes to finish something, and those are different purchases.
Take generating 500 images. Three models, three different cards, three very different hourly rates:
See the numbers
| Model | Card | Time | Cost |
|---|---|---|---|
| SDXL Turbo | RTX 3090 | 10 min | $0.04 |
| Stable Diffusion XL | RTX 4090 | 67 min | $0.50 |
| FLUX.1 dev | L40S | 150 min | $2.23 |
The cheapest hourly rate does not win, and the most expensive one is not a disaster either. What decides it is seconds per image multiplied by the number of images, and only then multiplied by the rate.
This is why we put a calculator on every model page rather than a price list. The question "what does an A100 cost" has a tidy answer that will not help you. The question "what does my job cost" has a messy answer that will.
Which card for which work
RTX 3090 — $0.22 an hour. Same 24GB of memory as a 4090, meaningfully slower, half the price. The right choice for overnight batch work where nobody is sitting there waiting. If a job runs while you sleep, paying extra for speed buys you nothing.
RTX 4090 — $0.45 an hour. The default, and the card most jobs on this site end up on. Enough memory for every mainstream image model and for language models up to roughly 30 billion parameters. If you do not know what you need, this is it.
RTX 5090 — $0.69 an hour. 32GB rather than 24GB, and quicker. Worth it when a model is just slightly too big for a 4090, which happens more often than you would like.
RTX A6000 and L40S — $0.79 and $0.89 an hour. Both 48GB. This is the tier where video work, the larger image models like FLUX.1, and most fine-tuning live. The L40S is the faster of the two and usually the better buy despite the higher rate.
A100 80GB — $1.35 an hour. The previous generation of data-center card, and still the value pick when you need a lot of memory. Large language models fit here comfortably.
H100 80GB — $2.25 an hour. Fast, and priced accordingly. Worth it for training runs and for serving the biggest models, where the time saved is measured in hours rather than minutes.
H200 141GB — $3.10 an hour. The most memory you can get on one card. You will know if you need this, and if you are not sure, you do not.
The part nobody puts on a pricing page
A few things move the real cost that never appear in an hourly rate.
Idle time is the biggest one. A machine you left running is billed exactly like a machine doing work. Our sessions shut themselves off after 15 minutes of inactivity, which is a policy rather than a feature, and one worth checking on any platform before you use it. An RTX 4090 forgotten overnight is about five dollars. An H100 forgotten over a long weekend is over a hundred and sixty.
Startup time. Every provider has some. Ours is about 90 seconds because the model is already on the machine when you get it. Platforms that hand you a bare box and let you install things yourself are cheaper per hour and considerably more expensive per useful hour, because the first forty minutes are you reading installation notes with the meter running.
Failed jobs. We do not charge for work that did not complete, and failed sessions refund automatically. Not everyone does this. It matters more than it sounds, because early on you will fail a few jobs while you work out what you are doing.
How this compares elsewhere
Vast.ai will often be cheaper per hour than anyone, because it is a marketplace of individual machines with variable reliability and no hand-holding. RunPod sits in the middle and is genuinely good if you are comfortable with containers. Lambda is the enterprise end, priced for teams with a procurement process.
We are not the cheapest hourly rate on that list and do not claim to be. What we do is start the machine with your model already on it, cap your spend by default, and shut it down when you stop. If you know exactly what you want and enjoy configuring it, one of the others may suit you better, and we would rather say so here than have you find out afterwards.
Working out your own number
Pick the job, not the card. Every model page shows which card we run it on, why, and what your specific volume costs before you commit to anything.
Or start from the kind of work you do and let the model choose the hardware.
