Blog/GPU Guides

NVIDIA H200 GPU Price (August 2026): $3.99/hr, Spot $1.99

Vishnu Subramanian
Vishnu Subramanian

Founder @ JarvisLabs.ai

January 8, 2026·17 min read·Updated August 3, 2026
NVIDIA H200 GPU Price (August 2026): $3.99/hr, Spot $1.99

Remember when 32GB of GPU VRAM seemed like a luxury? I sure do. Back in my Kaggle competition days, I was desperately trying to get my hands on a V100 with its "massive" 32GB VRAM. AWS had them, but by the time I cleared their quota hurdles, the competition was long over. Fast forward to 2026, and those numbers feel almost laughable.

Here's the short answer if you're in a hurry...

The NVIDIA H200 GPU costs $30,000–$40,000 to buy outright and $3.82–$10.60 per GPU hour to rent on-demand (August 2026). The H200 offers 141GB of HBM3e VRAM and 4.8 TB/s of bandwidth. Jarvislabs runs H200s at $3.99/hr on-demand — below Nebius ($4.50) and RunPod ($4.39) — with 1,000+ GPUs available right now, as single GPUs or full 8-GPU nodes, on VMs or containers.

The latest LLaMA 4 models from Meta are pushing the boundaries, requiring a minimum of 80GB VRAM just to get started. It's a stark reminder of how quickly AI model requirements are evolving - what was once cutting-edge is now barely enough to run the smallest of today's foundation models.

I'll admit it - I was skeptical when NVIDIA announced the H200. The H200 price tag was steep, and I wasn't convinced that the 76% VRAM increase (from 80GB to 141GB) would justify the investment. But here's where I was wrong: the timing couldn't have been better. As models like DeepSeek and LLaMA 4 hit the scene, that extra VRAM became a game-changer. Let me break it down with a real-world H200 example: Running LLaMA 4's larger models (like Maverick 400B) on H100s requires two full nodes with 8 GPUs each for a higher context window. But with the H200? You can do it on a single 8-GPU H200 node. That's not just a convenience - it's a massive cost and complexity reduction that makes advanced AI more accessible to everyone.

NVIDIA H200 Price Snapshot (August 2026)

TL;DR — H200 on-demand pricing ranges from roughly $3.82 to $10.60 per GPU hour. Jarvislabs runs H200 at $3.99/hr on-demand, below Nebius ($4.50) and RunPod ($4.39), and is one of the very few providers that will rent you a single H200 instead of forcing you into an 8-GPU node.

Jarvislabs H200 Pricing

PlanPrice / GPU-hourBest for
On-demand$3.99Everything — fine-tuning, inference, interactive work. No commitment, per-minute billing.
Spot (interruptible)$1.99Batch inference, checkpointed training, evals
Longer commitments & larger clustersTalk to usSubstantially better rates for multi-month commitments and larger cluster reservations

Every plan is per-minute billed, available as a single GPU or a full 8×H200 node, and runs on either a VM or a container — containers when you want to iterate fast, VMs when you need full control of the machine.

Committing for a year or need a larger cluster? Discounts get significantly better with duration and scale. Email sales@jarvislabs.ai and we'll put a number in front of you.

H200 Cloud Pricing Comparison (verified August 2026)

ProviderTypeMin. GPUsOn-Demand $/GPU-hrSpot / PreemptibleSingle GPU?
JarvislabsDedicated cloud1$3.99$1.99✅ Yes
Vast.aiMarketplace1from $3.82 1✅ Yes
RunPodDedicated cloud1$4.31–$4.39 2✅ Yes
NebiusDedicated cloud1$4.50 3$2.45✅ Yes
AWSHyperscaler8~$10.60❌ 8-GPU minimum
AzureHyperscaler8~$10.60❌ 8-GPU minimum
OracleHyperscaler8~$10.00❌ 8-GPU minimum
Google CloudHyperscaler8Not published$3.72❌ 8-GPU minimum

Note: Neocloud rates were checked against each provider's public pricing page in August 2026. Hyperscalers only offer H200s in 8-GPU HGX bundles, so the per-GPU figure divides the node price by 8; those list prices vary by region and commitment, so check the source links before budgeting. 4 5 6

Key Takeaways

  • H200 rental prices are moving in both directions right now. Nebius raised its H200 on-demand rate to $4.50, while marketplace rates on Vast.ai drifted down to around $3.82. If you're comparing quotes, check the date on whatever table you're reading — including this one.
  • Single-GPU access is rare. AWS, Azure, Oracle and Google Cloud all require an 8-GPU node. If you want one H200 to prototype on, your options are narrow.
  • Hyperscalers cost roughly 2.5x more on-demand — about $10.00–$10.60 per GPU-hour versus $3.99.
  • Marketplace rates aren't like-for-like. A marketplace quote comes from whichever individual host is cheapest that hour, with variable hardware and no guarantee the capacity is there when your job restarts. Dedicated capacity at a similar price is a better deal than it looks.

H200 Availability: The Part Nobody Talks About

Price is the easy question. In 2026 the harder one is can you actually get one today. H200 supply has been genuinely volatile this year — rental rates have swung double digits inside a single week as supply and export policy shifted, and "available" on a pricing page often means a waitlist, a quota request, or a capacity block you reserve weeks out.

Where Jarvislabs sits right now:

  • 1,000+ GPUs available on the platform, not a waitlist number.
  • No quota approval process. No sales call, no capacity request form. Launch from the dashboard.
  • Single GPUs or full 8×H200 nodes, provisioned in about 90 seconds.
  • VMs or containers, your choice, at the same hourly rate.
  • Dedicated hardware, not a marketplace. You're not bidding against other tenants for whichever host happens to be cheapest this hour, and the machine you get tomorrow is the machine you got today.

That last point matters more than it sounds. Several of the cheapest H200 listings you'll find are peer-to-peer marketplaces where the headline rate comes from whichever individual host is undercutting everyone that hour — variable hardware, variable network, variable uptime, and no guarantee the capacity is there when your run restarts.

Hardware Purchase Price (MSRP vs Street)

If you're buying rather than renting:

ConfigurationList / MSRPStreet & ResaleNotes
Single H200 SXM module$40,000–$55,000$30,000–$40,000Module only, not a usable system on its own
4-GPU HGX board$180,000–$220,000$160,000–$190,000Includes NVLink fabric
8-GPU HGX server$400,000–$500,000$350,000–$420,000Full system with CPUs, networking, chassis

Two things people miss here. First, an H200 module isn't a computer — the 8-GPU server price includes CPUs, NVSwitch fabric, networking and chassis, which is why it's more than 8× the module price. Second, the sticker price is the small half of the bill: power, cooling, rack space, and depreciation on an asset that a newer generation is already competing with.

NVIDIA H200 vs H100: Specs & Price Comparison

For a detailed breakdown of H100 pricing and availability, check out our NVIDIA H100 Price Guide. You may also find our H100 vs A100 comparison helpful for understanding the GPU evolution. If you're looking for a more budget-friendly option, see our A100 Price Guide or our L4 GPU Guide for cost-efficient inference.

FeatureH100 (SXM)H200 (SXM)🔍 What Changes
Launch architectureHopper (2023)Hopper + HBM3e (2024)Same core silicon; memory subsystem upgraded
GPU memory80 GB HBM3141 GB HBM3e+76 % capacity lets you keep ≥ 70 B-parameter models on a single card (NVIDIA)
Memory bandwidth3.0 TB/s4.8 TB/s+60 % throughput cuts batch-size bottlenecks on inference (NVIDIA)
FP8 peak (Tensor Core)3.96 PFLOPS3.96 PFLOPSCompute parity—raw flops aren't the selling point
NVLink 4 speed900 GB/s900 GB/sMulti-GPU scaling unchanged (NVIDIA)
PCIe interfaceGen 5 ×16Gen 5 ×16No change
Max TDP (SXM)700 W (configurable)700 W (configurable)Same rack-power footprint
Jarvislabs price$2.69/hr$3.99/hrH200 carries a ~48% premium for +76% memory

Key takeaways — why the extra 61 GB matters

  • Single-GPU fits for giant models. A lone H200 can load Llama 4 70B or Mixtral-8×22B in full precision, eliminating tensor-parallel gymnastics (splitting model tensors across GPUs) across two H100s.
  • Bandwidth boosts context length. When you crank up sequence length (≥ 32 k tokens) the 4.8 TB/s pipe keeps attention kernels fed instead of stalling on HBM (High Bandwidth Memory).
  • No free compute lunch. Peak TFLOPs are identical, so training throughput only rises if you're memory-bound.
  • Power and networking stay flat. If your rack can cool an H100, it can cool an H200; NVLink fabric configs carry over 1-for-1.

When to choose which

Pick thisIf you…
H200 ($3.99/hr)Need to run ≥ 70 B models, push long-context inference, or want the simplest 8-GPU topology without sharding headaches.
H100 ($2.69/hr)Care more about tokens/sec per dollar on models that already fit in 80 GB, e.g., Mistral 7B-MoE or SD-XL training runs.

NVIDIA H200 vs B200: Is Blackwell Worth Waiting For?

Blackwell is shipping, which raises the obvious question: skip the H200 entirely?

FeatureH200 (SXM)B200 (SXM)What it means
ArchitectureHopperBlackwellGenerational jump, new instruction set
GPU memory141 GB HBM3e192 GB HBM3e+36% capacity
Memory bandwidth4.8 TB/s8.0 TB/s+67% — the headline gain
Max TDP700 W~1000 WBlackwell needs denser power and cooling
Precision supportFP8FP8 + FP4FP4 roughly doubles inference throughput on supported models
AvailabilityBroad, mature toolingTighter supply, newer stackHopper is the safe operational choice today

The honest answer for most teams is that the H200 is still the better buy in 2026, for three reasons:

  1. Cost per useful token. B200 rates run roughly $7–$10/GPU-hour across the clouds we checked in August 2026 (Nebius $7.15, Lambda $6.99, RunPod $8.64+) against the H200's $3.99/hr. That's a 75–150% premium, and unless you're memory-bandwidth-bound or can actually exploit FP4, you won't recover it.
  2. Software maturity. The Hopper stack — vLLM, TensorRT-LLM, FlashAttention — is battle-tested. Blackwell kernels are improving fast but you'll hit more sharp edges.
  3. Availability and power. B200 supply is tighter, and ~1000W per GPU means many data centers need retrofitting.

Where the B200 does win clearly: very long context inference, extremely large MoE models, and FP4-quantized serving at scale. If that's you, see our NVIDIA B200 specs breakdown.

The rental-market wrinkle worth knowing: H200 rates have at times traded above newer Blackwell parts during supply crunches, because the H200 is what everyone's production stack already runs on. Generation number and price don't move in lockstep.

NVIDIA H200 Specs (SXM and NVL)

The H200 ships in two form factors, and picking the wrong one is an expensive mistake.

SpecH200 SXMH200 NVL (PCIe)
GPU memory141 GB HBM3e141 GB HBM3e
Memory bandwidth4.8 TB/s4.8 TB/s
Form factorSXM module (HGX baseboard)Dual-slot PCIe card
Max TDP700 W (configurable)Up to 600 W (configurable)
InterconnectNVLink + NVSwitch, 900 GB/s all-to-allNVLink bridge, up to 4 GPUs
CoolingHigh-airflow or liquid, HGX chassisAir-cooled, standard servers
Best for8-GPU training nodes, large distributed jobsRetrofitting existing PCIe servers, smaller deployments

Which one you actually want: if you're renting cloud H200s for training or multi-GPU inference, you want SXM — the NVSwitch fabric gives every GPU 900 GB/s to every other GPU, and that's what makes 8-GPU jobs scale. NVL exists mainly so enterprises can drop H200s into standard air-cooled PCIe servers without buying an HGX chassis. Jarvislabs runs H200 SXM.

Both variants carry the same 141 GB of HBM3e and 4.8 TB/s of bandwidth, so single-GPU memory-bound workloads perform similarly. The difference shows up the moment you scale past one GPU. Full specifications are on NVIDIA's H200 datasheet.

What Does It Cost to Actually Train or Serve on an H200?

Hourly rates are abstract. Here's what they translate to:

WorkloadSetupTimeCost at $3.99/hrCost at $1.99/hr spot
LoRA fine-tune, 7B model1× H200~3 hrs~$12~$6
Full fine-tune, 13B model4× H200~12 hrs~$192~$96
Continuous 70B inference endpoint1× H200730 hrs/mo~$2,913/mo— (not suited to preemption)
Large training run8× H2001 week~$5,363~$2,675

For always-on inference, on-demand runs about $35,000/year per GPU at list — which is exactly the point where a commitment starts paying for itself. Multi-month and cluster rates come down substantially from there; talk to sales for a number. Either way you're not paying for the server chassis, the power, the cooling, or the depreciation on an asset a newer generation is already competing with.

Frequently Asked Questions About NVIDIA H200 Price

How much does an NVIDIA H200 GPU cost?

The NVIDIA H200 GPU costs $30,000–$40,000 at street price to purchase outright, or $40,000–$55,000 at list/MSRP. For cloud rentals, H200 prices range from $1.99 to $10.60 per GPU hour depending on provider and commitment. Jarvislabs offers H200 at $3.99/hour on-demand and $1.99/hour spot.

How much does it cost to rent an H200 GPU?

H200 rental prices vary widely by provider. As of August 2026, Jarvislabs offers H200 at $3.99/hour on-demand and $1.99/hour spot. Vast.ai marketplace rates start around $3.82/hour, RunPod is $4.31–$4.39/hour, and Nebius is $4.50/hour. AWS, Azure and Oracle charge roughly $10.00–$10.60 per GPU hour on-demand. Most hyperscalers require renting a full 8-GPU node, while Jarvislabs rents single H200 GPUs. Longer commitments and larger clusters get substantially better rates — contact sales@jarvislabs.ai.

Can I rent a single H200 GPU or do I need an 8-GPU server?

Hyperscalers ship H200 only in 8-GPU HGX nodes. Jarvislabs is one of the few platforms offering single H200 GPU access on-demand at $3.99/hr, so you can prototype without paying for a whole server. You can also rent full 8×H200 nodes when you need them.

Are H200 GPUs actually available right now?

Yes. Jarvislabs has 1,000+ GPUs available on the platform with no waitlist and no quota approval process — you can launch a single H200 or a full 8-GPU node in about 90 seconds, on either a VM or a container. This is worth checking with any provider, because H200 supply has been volatile through 2026 and "available" often means a capacity request or a multi-week reservation.

What is the NVIDIA H200 price vs H100?

The H200 costs roughly 15–20% more than the H100 to buy: an H100 runs $25,000–$30,000 against the H200's $30,000–$40,000. For cloud rentals on Jarvislabs, the H200 is $3.99/hr versus $2.69/hr for the H100 — a 48% premium for 76% more memory and 60% more bandwidth. The H200 wins whenever you're memory-bound; the H100 wins on tokens-per-dollar for models that already fit in 80GB.

Is the H200 or B200 better value in 2026?

For most teams the H200 remains better value. The B200 offers 192GB and 8.0 TB/s bandwidth against the H200's 141GB and 4.8 TB/s, but carries a significant price premium, needs ~1000W per GPU, and runs on a less mature software stack. The B200 wins for very long context inference, very large MoE models, and FP4-quantized serving. Otherwise the H200 at $3.99/hr delivers better cost per useful token.

What is the difference between H200 NVL and H200 SXM?

The H200 NVL is the dual-slot PCIe variant designed for air-cooled standard servers, drawing up to 600W and linking up to 4 GPUs over an NVLink bridge. The H200 SXM is built for NVIDIA's HGX platform, draws up to 700W, and connects through NVSwitch at 900 GB/s all-to-all. Both carry 141GB HBM3e and 4.8 TB/s bandwidth, so single-GPU performance is similar — but SXM scales far better past one GPU. Jarvislabs runs H200 SXM.

Is it cheaper to rent or buy an H200 for 24/7 use?

Renting is cheaper for nearly everyone. A single H200 on Jarvislabs at $3.99/hr costs about $35,000/year at on-demand list, against $30,000–$40,000 for the bare module — before the server to put it in, power, cooling and depreciation. Commit for a year or take a larger cluster and the rate drops substantially below that, which is where renting wins outright. Buying only makes sense at sustained multi-year utilization with data center capacity you already own.

Why is the H200 more expensive than the H100?

You're paying for memory: the H200 jumps from 80GB HBM3 to 141GB HBM3e and bumps bandwidth to 4.8 TB/s (+60%), letting it handle 70B+ parameter models on one card. The compute silicon is the same Hopper die, so FP8 peak throughput is identical.

Yes — NVLink/NVSwitch fabric is included in the instance rate at Jarvislabs. Unlike other providers, we don't charge extra for bandwidth or data transfer. What you see is what you pay.

Will NVIDIA H200 prices drop now that Blackwell has shipped?

Not as predictably as you'd expect. Previous-generation flagships historically see roughly 15% list-price cuts within six months of a new generation, but H200 rental rates have at times traded above newer Blackwell parts during supply crunches, because production stacks are already built on Hopper. Expect volatility rather than a steady decline.

Conclusion & Next Steps

Want NVIDIA H200 performance without the $30K+ price tag?

  • $3.99/hr on-demand with no commitment, or $1.99/hr spot for interruptible work
  • Better rates for year-long commitments and larger clusterstalk to sales
  • 1,000+ GPUs available now — no waitlist, no quota request
  • Single GPU or 8-GPU node, on a VM or a container
  • Skip the hardware headaches — power, cooling, depreciation

The math is simple: launch an H200 in 90 seconds, scale when needed, pay as you use.

Ready to try it? Rent an NVIDIA H200 GPU or see full GPU pricing. For multi-node setups or longer commitments, email support@jarvislabs.ai for a custom quote.

Once you have your H200, check out our guides on deploying Ollama for LLM inference or optimizing inference with vLLM quantization to get the most out of your GPU.

We keep this NVIDIA H200 price guide updated monthly. Last reviewed August 2026.


Need custom H200 pricing? Committing for a year, or need a larger cluster? Discounts scale with both duration and size, and go well beyond the on-demand rate above. Email sales@jarvislabs.ai for a quote.


Footnotes

  1. Vast.ai H200 pricing — marketplace rate, updates hourly. Checked August 2026.

  2. RunPod pricing — checked August 2026.

  3. Nebius pricing — checked August 2026.

  4. Google Cloud GPU Pricing

  5. Oracle Cloud Infrastructure Price List

  6. AWS EC2 Capacity Blocks Pricing

Get Started

Build & Deploy Your AI in Minutes

Cloud GPU infrastructure designed specifically for AI development. Start training and deploying models today.

View Pricing