Skip to main content

Pricing

Same silicon, less than half the bill

NexGPU connects directly to 1,175 compute nodes across 51 countries. An H100 costs $3.582 per GPU-hour here — 47% below the AWS on-demand rate. Billed hourly, no minimum spend, no contract.

Entry tier per GPU-hour
$0.083
H100 cheaper than AWS on-demand
47%
GPU models in stock
68
Live rentable nodes
1,175

01 — Rate card

Every GPU we carry

Datacenter and consumer classes listed separately. Rates are per GPU-hour — the lowest rentable price for that model.

Datacenter class

A100, H100, L40S and other professional accelerators — for large-scale training, high-concurrency inference and long stable runs.

Datacenter class
ModelVRAMNodesFrom / GPU-hr
Quadro P4000up to 2x8GB6$0.110
RTX A2000up to 2x6GB3$0.137
RTX A4000up to 8x16GB31$0.145
Tesla V100up to 4x32GB21$0.188
Tesla P40up to 2x24GB2$0.214
Q RTX 6000up to 4x23GB5$0.274
RTX 4000Adaup to 2x20GB4$0.274
Tesla T4up to 2x15GB3$0.298
RTX A5000up to 2x24GB5$0.355
RTX PRO 4000up to 4x24GB28$0.408
A10up to 2x22GB4$0.414
Q RTX 8000up to 2x48GB4$0.508
RTX PRO 4500up to 2x32GB5$0.625
L4up to 2x22GB6$0.629
RTX 5000Adaup to 2x32GB3$0.682
RTX A6000up to 4x48GB8$0.817
A100 PCIEup to 4x40GB22$0.824
RTX 5880Adaup to 4x48GB8$0.829
A100 SXM4up to 8x80GB22$1.088
RTX 6000Ada48GB8$1.109
RTX PRO 5000up to 2x48GB17$1.125
L40Sup to 4x45GB8$1.639
RTX PRO 6000 Sup to 8x96GB19$1.852
RTX PRO 6000 WSup to 4x96GB32$2.055
H100 SXMup to 8x80GB16$3.582
H100 PCIEup to 8x80GB8$4.182
H100 NVL94GB3$5.201
H200up to 8x140GB17$6.660
H200 NVLup to 2x140GB7$7.342
B200up to 8x179GB16$10.797
B300up to 8x269GB3$15.742

31 models

Consumer class

RTX 5090, 4090 and 3090 gaming cards — the cheapest compute per dollar, ideal for image generation, LoRA fine-tuning and small-to-mid model inference.

Consumer class
ModelVRAMNodesFrom / GPU-hr
GTX 1660 Sup to 8x6GB11$0.083
GTX 1080up to 6x8GB10$0.094
GTX 1070up to 2x8GB3$0.095
RTX 3060 Tiup to 2x8GB12$0.097
GTX 1070 Tiup to 4x8GB2$0.099
RTX 3060up to 6x12GB48$0.100
GTX 1080 Tiup to 7x11GB13$0.108
RTX 3070up to 8x8GB25$0.109
RTX 4060up to 4x8GB8$0.112
RTX 2060Sup to 2x8GB5$0.117
RTX 2060up to 2x12GB3$0.120
Titan Xpup to 7x12GB10$0.121
RTX 5060up to 2x8GB5$0.123
RTX 4060 Tiup to 4x16GB27$0.126
RTX 2080 Tiup to 8x11GB9$0.135
RTX 5060 Tiup to 8x16GB70$0.136
RTX 3060 laptopup to 2x12GB12$0.137
GTX 1660 Ti6GB4$0.138
RTX 4070Sup to 8x12GB20$0.139
RTX 4070up to 4x12GB19$0.149
RTX 3080up to 2x10GB19$0.161
RTX 3080 Tiup to 4x12GB24$0.163
RTX 4070 Tiup to 4x12GB12$0.174
GTX 1060up to 2x3GB2$0.176
Titan Vup to 2x12GB4$0.178
RTX 3070 Tiup to 2x8GB5$0.187
RTX 5070up to 2x12GB18$0.191
RTX 3090up to 8x24GB66$0.193
RTX 5070 Tiup to 4x16GB29$0.229
RTX 4080up to 4x16GB13$0.242
RTX 4080Sup to 8x16GB25$0.255
RTX 4070S Tiup to 8x16GB15$0.265
RTX 5080up to 8x16GB33$0.283
Titan RTX24GB2$0.304
RTX 3090 Tiup to 2x24GB4$0.389
RTX 4090up to 14x24GB130$0.540
RTX 5090up to 8x32GB107$0.723

37 models

Prices are the lowest available per-GPU hourly rate and move with supply and demand. Multi-GPU nodes rent whole, so the amount charged is the per-GPU rate times the GPU count.

02 — Head to head

Against three mainstream GPU clouds

Same card, same hourly billing, same on-demand terms — no reserved discounts or annual commitments. Competitor rates come from each vendor's public pricing page; the check date is noted below the table.

Against three mainstream GPU clouds
ModelNexGPUAWSCoreWeaveLambda
H100 SXM80GB$3.582$6.880−47%p5.48xlarge$6.160−41%HGX H100$3.990−10%8x H100 SXM
H200141GB$6.660No per-GPU rate published$6.310+6%HGX H200Full cluster only
A100 SXM80GB$1.088$3.431−68%p4de.24xlarge$2.700−59%A100$1.990−45%1x A100 SXM
RTX 509032GB$0.723Not offeredNot offeredNot offered
RTX 409024GB$0.540Not offeredNot offeredNot offered

None of the three mainstream clouds offer consumer GPUs at all. For image generation, LoRA fine-tuning and small-model inference, the RTX 5090 and 4090 have no equal on price-performance — in this tier there is no comparison, only availability.

Competitor figures are each vendor's published US-region on-demand rate, normalised to a per-GPU hour. AWS is us-east-1 on-demand; CoreWeave and Lambda are their listed pricing-page rates. A green − means we are cheaper; an amber + means they are. H200 supply is currently thin and we do not win that tier — it is listed as it stands. Percentages round down throughout: our advantage is never inflated and our disadvantage is never shrunk. AWS · CoreWeave · Lambda

03 — How billing works

Three charges, all of them stated

The most common surprise on a GPU cloud bill is discovering compute was not the only line item. All three are written out here — including the one nobody else mentions.

  • Compute — metered per second, priced per hour

    Starts when the instance reaches running state and stops the moment it is destroyed. Zero compute charge while stopped — we reconciled this against upstream billing second by second.

  • Storage — keeps running while stopped

    Disk is billed per GB-month from the moment the instance is created — including during image download — and it continues after you stop the instance. This is industry standard; most platforms just do not say so. The network-wide median is about $0.414 per GB-month, which means a 200GB instance still costs around $2.76 for every stopped day. If you are done, destroy it — do not merely stop it.

  • Network — egress metered

    Ingress and egress are metered separately and priced per GB, with a median around $0.0081 per GB. For normal training and inference this is typically under 5% of the bill; check the estimate before a large data export.

Why we lead with the storage charge

Because it is the only charge that keeps running when you are not using anything. The industry norm is to put it in clause 14 and never mention it again, letting the customer find it at month end. We put it in the middle of the pricing page, on the order form, on the instance card and in the low-balance email. A bill you can predict is worth more than a bill that merely looks cheaper.

04 — Run the numbers

What common jobs actually cost

Measured single-GPU run times multiplied by the starting rate. Substitute your own duration.

What common jobs actually cost
JobGPURuntimeCost
100 SDXL images at 1024pxRTX 4090$0.540/hr~15 min$0.14
One image-style LoRA fine-tuneRTX 4090$0.540/hr~2 hours$1.08
Transcribe 10 hours of audio, Whisper large-v3RTX 4090$0.540/hr~30 min$0.27
QLoRA fine-tune of a 7B model, single GPURTX 5090$0.723/hr~6 hours$4.34
Serve a 70B model for a full dayA100 SXM4$1.088/hr24 hours$26
One week of H100 training, single GPUH100 SXM$3.582/hr168 hours$602

Prices are the lowest available per-GPU hourly rate and move with supply and demand. Multi-GPU nodes rent whole, so the amount charged is the per-GPU rate times the GPU count. Runtimes are typical single-GPU magnitudes and vary with resolution, sequence length, dataset size and step count. Estimates only.

05 — Included

What we do not charge extra for

Line items that are commonly upsold elsewhere are inside the sticker price here.

  • Public direct ports

    Every instance ships with a public IP and a direct port range. SSH, Jupyter and your own services map straight through — no port fees and no mandatory bastion hop.

  • Web terminal and file manager

    Open a root shell in the browser, and optionally install a visual control panel for files, processes and cron with one click. No extra charge.

  • The whole template library

    PyTorch, vLLM, ComfyUI, Whisper and the rest are all available. There is no "enterprise image" tier.

  • No minimums, no setup fee

    No minimum rental period, no monthly floor, no instance creation fee. Run out of balance and it stops; top up and it resumes.

06 — FAQ

Pricing and billing

The six questions asked most often before a first order.

How is GPU rental billed — hourly or by the second?

Metered per second, priced per hour. Ninety minutes of runtime is charged as 1.5 hours; there is no rounding up to a whole hour. Compute charges begin when the instance reaches running state and end when it is destroyed.

Do I still pay after stopping an instance?

Compute charges stop, but disk storage continues to accrue per GB-month, because your data still occupies the host's drive. If you will not need it again soon, destroy the instance. If you need to keep the data, copy the results out first, then destroy.

Is $3.582/hr for an H100 a rate I can actually get?

It is the lowest per-GPU price across all H100 SXM nodes. Nodes of the same model vary by region, bandwidth and host reliability, so the marketplace lists every rentable node with its own live rate and you pay the price you see. Multi-GPU nodes rent whole: an 8-GPU node bills at the per-GPU rate times eight.

Why is this so much cheaper than AWS, CoreWeave or Lambda?

Different supply structure. Traditional clouds price against owned datacenters, so the rate carries facility construction, redundant standby capacity and long depreciation schedules. NexGPU aggregates distributed supply where utilisation is set by market matching rather than one vendor's capacity plan. The H100 is physically the same card — what differs is how the supply reaches you.

Is there a minimum top-up? Any hidden fees?

No minimum spend, no setup fee, no port fee, no image licensing fee. The bill has exactly three lines — compute, storage and network — and the instance detail page shows the running total and unit rate for each.

Can prices rise while my job is running?

No. The rate is locked at order time and holds until that instance is destroyed, even if the market moves up. New rates apply only to newly created instances.

Start building on NexGPU

Enterprise R&D team or solo developer — either way, your first job can be running in minutes.

Sign up to browse live network pricing. No payment method required.