Best Cloud GPU Providers for AI in 2026: Pricing Compared

Vishnu Subramanian
Vishnu Subramanian
Founder @ Jarvislabs

There is no single best cloud GPU provider. The right one depends on whether you need one GPU for an afternoon, an 8×H200 node for a month, an InfiniBand cluster for a training run, or an API endpoint that scales with traffic. Specialist GPU clouds (Jarvislabs, RunPod, Lambda, CoreWeave) price per GPU-hour and launch in seconds to minutes. Hyperscalers (AWS, Google Cloud, Azure) price whole instances and make sense when your data and tooling already live there. Jarvislabs is built for teams that want root-access VMs, ready-to-run Templates, Serverless and Managed Endpoints from one account, with published per-GPU rates, per-minute billing and committed capacity when a team needs it.

Updated September 2026. Competitor rates were verified against each provider's public pricing page on 2026-09-05. Jarvislabs rates are the US-tier on-demand rates from the pricing page; accounts billed in India see INR rates on jarvislabs.ai/in.

What changed since the last version of this guide

The GPU cloud market moved quickly in 2026, and so did Jarvislabs. Since this guide was first written:

  • H200 became the default large-model GPU. Single-GPU H200 SXM instances, multi-GPU nodes and InfiniBand clusters are all available on Jarvislabs. Our measured H100 vs H200 comparison shows where the extra 61 GB of memory changes the result.
  • Inference stopped being an afterthought. Jarvislabs now has three inference products: Serverless (your framework, your model, autoscaling workers), Managed Endpoints (a model catalog with the serving stack maintained by our team) and Model APIs (per-token, OpenAI-compatible).
  • Networking and storage caught up. VMs can join a VPC with Security Groups and Reserved IPs, and Shared Filesystems attach to multiple instances. Both are available on every Jarvislabs region.
  • Spot and committed pricing are public. H100 spot is $1.19/hr and a one-year commitment brings H100 to $2.47/hr.
  • The CLI arrived. One command, jl run, uploads a project, creates the GPU, runs the script and pauses the machine afterward. It was designed for coding agents as much as humans.

Several rates in the previous version were out of date, including our own H200 price. The table below replaces them.

Quick provider comparison

Single-GPU on-demand rates in USD per GPU-hour unless noted. H100 rows use the SXM variant where the provider publishes one.

ProviderNVIDIA H100 SXMNVIDIA H200NVIDIA A100 80GBBillingBest for
Jarvislabs$2.69$3.99$1.49Per-minuteVMs, Templates, Serverless and clusters from one account; committed capacity
RunPod (Secure Cloud)$3.29$4.59$1.59Per-secondBroad GPU catalog, Community Cloud, mature serverless
Lambda$4.29 (PCIe $3.29)Not listed single-GPU8×node only ($2.79/GPU)Per-minuteReserved InfiniBand clusters, B200 self-service
CoreWeave$6.16*$6.31*$2.70*Per-hourManaged Kubernetes at cluster scale
Paperspace$5.95Not listed$3.18Per-hourGradient Notebooks, DigitalOcean ecosystem
Vast.aiMarketplace quoteMarketplace quoteMarketplace quotePer-secondLowest headline prices, reliability depends on the host
AWS / Google Cloud / AzureWhole-instance pricingWhole-instance pricingWhole-instance pricingPer-second / per-hourEcosystem integration, compliance programs

*CoreWeave publishes 8-GPU node prices (H100 $49.24/hr, H200 $50.44/hr, A100 80GB $21.60/hr). The per-GPU figures are those totals divided by eight and are not single-GPU offers.

Hyperscaler GPU instances bundle CPU, RAM, local NVMe and networking, and the flagship H100 and H200 instances (AWS P5, Google Cloud A3, Azure ND) are 8-GPU nodes. Their list on-demand rate per GPU-hour is typically well above the specialist clouds in this table, with spot and committed-use discounts narrowing the gap. Use their calculators for a like-for-like number.

How to read the table

  • SXM vs PCIe matters. H100 SXM has higher memory bandwidth and NVLink between GPUs. A lower PCIe rate is not the same product.
  • Single GPU vs node. Some providers only publish 8-GPU node prices. Divide by eight for a rough comparison, then check whether you can actually buy one GPU.
  • Storage is separate almost everywhere. Jarvislabs bills retained storage at $0.00014/GB/hour (about $0.10/GB/month) whether the instance is running or paused. RunPod and Paperspace also bill storage separately.
  • Billing granularity compounds. Per-minute or per-second billing matters when jobs are short or bursty. Providers that bill in whole hours charge a 65-minute job as two hours.

Specialist GPU clouds

Jarvislabs

What it is: The AI Cloud for builders and teams. Four ways to run: On-Demand VMs with root access, Templates with a configured software stack, InfiniBand clusters when one node is not enough, and Serverless or Managed Endpoints when you want an API rather than a machine.

Numbers that matter:

  • Billing is per minute on VMs, Templates and Serverless workers. Pause an instance and GPU charges stop; retained storage keeps billing. Autopause triggers on duration or account balance, not idle detection.
  • Single-GPU instances launch in under 90 seconds; 8-GPU VMs can take up to about eight minutes. 27,000+ developers, 50M+ GPU hours served.
  • Seven GPUs in the public lineup: NVIDIA H200 SXM ($3.99/hr), H100 SXM ($2.69/hr), RTX Pro 6000 ($1.89/hr), A100 80GB ($1.49/hr), A100 40GB ($0.89/hr), L4 ($0.44/hr) and A30 ($0.41/hr).
  • Spot pricing on selected GPUs: H100 $1.19/hr, H200 $1.99/hr, RTX Pro 6000 $0.99/hr, A100 80GB $0.89/hr.
  • Commitments from one month to one year. H100 drops to $2.47/hr and H200 to $3.67/hr on a one-year plan.
  • Up to 8 GPUs per VM or Template. H200 SXM InfiniBand clusters beyond that, with deployments of 128, 256 and 1,024+ GPUs planned with sales.

Platform:

  • VPC, Security Groups and Reserved IPs for VMs. Private addresses survive pause and resume; Reserved IPs keep the same public address ($5/month).
  • Shared Filesystems up to 10 TB, mountable on several instances in the same region, surviving instance deletion.
  • Serverless: deploy with vLLM, SGLang or Ollama, get an OpenAI-compatible endpoint, set worker limits and idle timeout, scale to zero. Billed per GPU-minute of worker runtime.
  • Managed Endpoints: pick an open model from the catalog and deploy. Our team pins the GPU, image, precision, parallelism and kernels. Why we built it.
  • Model APIs: hosted open models behind an OpenAI-compatible API, billed on tokens used. See the catalog.
  • jl CLI and Python SDK: check availability, create, SSH, transfer files, run a script and pause from the terminal, or hand the same workflow to a coding agent.

The private network stays inside Jarvislabs: there is no VPN or private interconnect to AWS or other clouds, so external databases are reached over their public endpoint.

Best for: Engineering teams that want to move from one GPU to multi-node training or production inference without changing clouds, with published rates, per-minute billing and committed capacity when they need it.

Jarvislabs pricing | GPU Cloud India | Compare: RunPod · Lambda · CoreWeave · Vast.ai · Paperspace

RunPod

What it is: GPU cloud with a managed tier (Secure Cloud) and a peer-hosted marketplace tier (Community Cloud), plus Serverless workers and on-demand or reserved clusters.

Strengths: a broad GPU catalog, per-second billing, flex and active serverless workers, a large template community.

Limits: Community Cloud instances can be interrupted and vary by host. Secure Cloud H100 SXM ($3.29/hr) and H200 ($4.59/hr) rates sit above Jarvislabs. Storage is billed separately, and GPU and storage availability should be checked in the same region.

Best for: Teams that want GPU variety, an established serverless product, or budget compute via Community Cloud. Jarvislabs vs RunPod

Lambda

What it is: GPU cloud and hardware company with self-service instances, Lambda Stack and reserved InfiniBand clusters.

Strengths: Lambda Stack keeps drivers and frameworks in sync, B200 is available self-service, and reserved 1-Click Clusters are a well-trodden path for multi-node training. Per-minute billing.

Limits: Single-GPU H100 SXM is $4.29/hr, with PCIe at $3.29/hr. A100 80GB is only sold as an 8-GPU node. No serverless product, so inference means running your own stack. On-demand inventory is first come, first served.

Best for: Research labs planning reserved multi-node training and teams that value Lambda Stack. Jarvislabs vs Lambda

CoreWeave

What it is: Enterprise GPU cloud built around managed Kubernetes and HPC networking (InfiniBand and RoCE).

Strengths: large-scale cluster infrastructure, mature Kubernetes tooling, serverless products for RL and inference.

Limits: pricing is per 8-GPU node (H100 $49.24/hr, H200 $50.44/hr), so single-GPU experiments are not the target use case. The operating model assumes your team runs Kubernetes.

Best for: Organisations that already operate Kubernetes and need contracted cluster capacity. Jarvislabs vs CoreWeave

Vast.ai

What it is: A marketplace where hosts rent out GPUs at prices they set.

Strengths: marketplace offers that can undercut managed providers' list rates, a wide range of GPU types and host configurations, and on-demand, reserved and interruptible rental modes.

Limits: every host and offer must be evaluated individually. Reliability, storage persistence, transfer costs and data location depend on the host. Interruptible bids need checkpointing.

Best for: Fault-tolerant batch jobs where the operator can checkpoint often and price beats predictability. Jarvislabs vs Vast.ai

Paperspace

What it is: GPU Machines, Gradient Notebooks and Gradient Deployments inside the DigitalOcean portfolio.

Strengths: a notebook-first workflow, container deployments with autoscaling, DigitalOcean billing and support.

Limits: H100 at $5.95/hr and A100 80GB at $3.18/hr on demand, per-hour billing, and storage and public IPs that keep billing while machines are off.

Best for: Teams already on DigitalOcean who want notebooks and simple deployments. Jarvislabs vs Paperspace

Hyperscalers: AWS, Google Cloud and Azure

What they are: GPU instance families inside a general-purpose cloud. AWS P5 (H100) and P5e/P5en (H200), Google Cloud A3 (H100) and A3 Ultra (H200), Azure ND H100 v5 and ND H200 v5.

Strengths:

  • Integration with the storage, identity, networking and managed ML services you may already use (S3 and SageMaker, GCS and Vertex AI, Blob and Azure ML).
  • Compliance programs and audit tooling that enterprise procurement recognises.
  • Spot and preemptible capacity at steep discounts when your job can tolerate interruption.
  • Alternatives to NVIDIA: AWS Trainium and Inferentia, Google TPUs.

Limits:

  • The flagship H100 and H200 instances are 8-GPU nodes. Single-GPU SKUs exist in some regions, but usually sit behind a quota request rather than an instant launch.
  • Whole-instance pricing bundles CPU, RAM and NVMe you may not need, and data transfer and block storage add separate line items.
  • Instance startup takes minutes rather than seconds, and GPU quota often requires a support request before you can launch at all.
  • List on-demand rates per GPU-hour are well above the specialist clouds in this guide.

Best for: Teams whose data, compliance requirements or ML platform already live in one of these clouds, and workloads where the cost of moving data outweighs the GPU price difference.

How to choose

By workload

WorkloadRecommendation
Fine-tuning a 7B to 70B modelJarvislabs H100 or H200 VM or Template. Persistent storage and per-minute billing suit iterative runs. Use spot if you checkpoint.
Pretraining or multi-node trainingJarvislabs H200 SXM InfiniBand clusters, Lambda reserved clusters or CoreWeave, depending on whether you want VMs, Lambda Stack or Kubernetes.
Production LLM inference, sustained loadJarvislabs Managed Endpoints for catalog models, Serverless for your own stack. RunPod Serverless is the closest alternative. Keep at least one worker warm for large models.
Bursty or low-volume inferenceJarvislabs Model APIs or any per-token API. Pay for tokens, not idle GPUs.
Image and video generation (FLUX, SDXL, Wan)Jarvislabs RTX Pro 6000 (96 GB) for large image and video models, L4 for smaller ones. Both are datacenter GPUs; consumer cards are not licensed for datacenter use.
Notebook-based experimentationJarvislabs Templates (PyTorch, ComfyUI, JupyterLab preinstalled) or Paperspace Gradient.
Fault-tolerant batch jobsVast.ai interruptible offers or Jarvislabs spot. Both need checkpointing.
Data residency requirementsJarvislabs VMs inside your own VPC with Security Groups (India regions), or a hyperscaler region that matches the requirement.

By team

Individual engineers and small teams: Jarvislabs Templates or RunPod. Pick by the GPU you need and the region closest to your data.

Product teams (2 to 20 engineers): Jarvislabs. Move from a fine-tuning VM to a Serverless endpoint to a Managed Endpoint as load grows, on one bill, driven from the CLI.

Research lab: Lambda or Jarvislabs clusters for reserved multi-node capacity. Lambda if you want Lambda Stack; Jarvislabs if you also want Templates, Serverless and committed capacity.

Enterprise platform team: Jarvislabs for committed capacity, VPC-isolated VMs and dedicated endpoints billed per minute; CoreWeave if your team already runs Kubernetes; a hyperscaler if procurement requires it.

By region

Most specialist clouds concentrate GPU capacity in the United States and Europe. Jarvislabs runs H100, H200 and RTX Pro 6000 in India, with INR billing available at jarvislabs.ai/in. RunPod added an India region in July 2026. The hyperscalers cover more regions, but GPU quota is often the first thing to run out. Pick the region closest to your data, then confirm the GPU and quantity you need are launchable there.

The total cost checklist

The hourly GPU rate is the biggest line item, but rarely the only one. Before comparing providers, price the whole job:

  1. GPU hours at the on-demand, spot or committed rate you will actually use.
  2. Storage while the data is retained, including while compute is paused.
  3. Idle time. Per-second and per-minute billing waste less than per-hour billing. Autopause and scale-to-zero waste less than forgotten instances.
  4. Egress. Downloading checkpoints and datasets is free on some providers and expensive on hyperscalers.
  5. Engineering time. Setting up a serving stack for a new model can take days. A managed catalog or a preconfigured Template is a cost saving even when the hourly rate is not the lowest.
  6. Reserved capacity. If you know you will need 8×H200 for three months, ask for committed pricing. Jarvislabs publishes one-month to one-year rates on the pricing page, and Lambda, CoreWeave and Paperspace list reserved or commitment terms.

FAQ

Which cloud GPU provider has the lowest prices?

For a single GPU on a given afternoon, a Vast.ai marketplace offer or a RunPod Community Cloud instance can have the lowest headline price, with the reliability tradeoffs described above. Among managed providers publishing fixed rates, Jarvislabs lists H100 SXM at $2.69/hr and spot at $1.19/hr, and RunPod Secure Cloud lists H100 SXM at $3.29/hr. Jarvislabs storage is billed separately at about $0.10/GB/month. The lowest hourly rate is not always the lowest total cost. Compare total cost with the checklist above.

Which cloud GPU provider is best for beginners?

Jarvislabs Templates or Paperspace Gradient. A Template gives you PyTorch, CUDA, JupyterLab and SSH already configured, launches in under 90 seconds and bills per minute. RunPod's template library is also large and well documented.

Can I switch between providers easily?

Yes, for most workloads. Every provider in this guide runs standard CUDA and supports Docker, so code and models move. The friction is data: moving a 500 GB dataset or a set of checkpoints takes time and may incur egress fees. Keep data on a filesystem you control and script your environment setup so you can rebuild it anywhere.

Should I use a hyperscaler or a specialist GPU cloud?

Use a hyperscaler when your data, identity and ML platform already live there, or when procurement requires their compliance program. Use a specialist cloud when you want to rent one GPU without a quota ticket, pay per GPU-hour rather than per node, and start in seconds. Many teams do both: GPUs on a specialist cloud, object storage and services on a hyperscaler.

How do I estimate my monthly GPU cost?

Multiply hours per job by jobs per month by the GPU rate, then add storage. Example: 40 hours of fine-tuning per week on one Jarvislabs H100 SXM at $2.69/hr is about $470/month on demand, or about $210 on spot. A 500 GB filesystem adds about $50/month. An 8×H200 Managed Endpoint kept warm around the clock is about $23,000/month on demand, before committed pricing.

Which provider has the best H100 and H200 availability?

Self-service inventory changes hourly on every provider. Jarvislabs publishes live availability in the dashboard and through the jl gpus command. For a planned run, reserve capacity rather than assuming on-demand stock will be there on the day. Lambda, CoreWeave and Jarvislabs all offer committed or reserved capacity.

Do any of these providers offer GPUs in India?

Jarvislabs runs NVIDIA H100, H200 and RTX Pro 6000 in its India regions, with VPC networking, Shared Filesystems and INR billing. RunPod opened an India region in July 2026 with H100s. AWS, Google Cloud and Azure have Indian regions with limited GPU quota. See GPU Cloud India.

Is per-minute billing really different from per-hour billing?

For long training runs, no. For everything else, yes. Take a developer who launches ten separate 20-minute experiments a day. On Jarvislabs that is 200 billed minutes. On a provider that charges a one-hour minimum per launch it is 10 billed hours, a 3× difference at the same hourly rate. Serverless per-GPU-minute billing applies the same logic to inference workers.

Build & Deploy Your AI in Minutes

Get started with Jarvislabs today and experience the power of cloud GPU infrastructure designed specifically for AI development.

← Back to FAQs