On-Demand VMs · This page
You own the stack.
For custom kernels, your own Docker daemon, or a GPU node in your Kubernetes cluster.
- Root on the VM via sudo
- Your OS, drivers and runtime
- SSH access · typically about 90 seconds to launch
Run your AI workloads on 1 to 8 dedicated NVIDIA GPUs, with SSH access and control over your software.
NVIDIA H100 SXM $2.69/hr · Billed by the minute
$ jl ssh <instance-id>
$ sudo whoami
root
Measured performance
Six internal benchmarks on 8× NVIDIA H200 SXM. VM and container results were within 1.2% across compute, memory, communication and training.
+0.26% observed difference · Final 10-step mean. FP8, Flash Attention 3.
-0.17% observed difference · 1 GiB message across 8 GPUs. The NVLink and NVSwitch path, preserved.
Internal Jarvislabs benchmark, June 2026, run on separate 8× NVIDIA H200 SXM allocations. A Jarvislabs GPU container compared with a KVM VM using direct GPU passthrough. Training was nanochat with FP8 and Flash Attention 3; probe software versions differed slightly between runs. These figures describe the tested systems and workloads, not a universal guarantee.
| Workload | Container | VM | Difference |
|---|---|---|---|
| BF16 matrix multiply, mean of 8 GPUs | 694.1 TFLOP/s | 695.9 TFLOP/s | +0.25% |
| FP16 matrix multiply, mean of 8 GPUs | 623.0 TFLOP/s | 630.3 TFLOP/s | +1.16% |
| HBM copy bandwidth, mean of 8 GPUs | 4.286 TB/s | 4.266 TB/s | -0.47% |
| NCCL all-reduce bus bandwidth, 1 GiB message | 468.2 GB/s | 467.4 GB/s | -0.17% |
| nanochat training, 1 GPU | 205,133 tok/s | 206,131 tok/s | +0.49% |
| nanochat training, 8 GPUs | 1,594,318 tok/s | 1,598,385 tok/s | +0.26% |
Differences in either direction at this scale are run-to-run variation: clocks, temperature and software builds move results by fractions of a percent. Read the table as parity, not as the VM winning four rows of six.
Direct GPU passthrough lets CUDA execute on the GPUs. NVLink and NVSwitch connect the eight H200s in the tested configuration.
Choose your compute
Choose a GPU count to see supported configurations and the hourly price for the whole instance.
Typically ready in about 90 seconds.
| GPU | GPU memory | vCPU | System RAM | Instance / hour |
|---|---|---|---|---|
| NVIDIA H200 SXM | 141 GB | 28 | 300 GB | $3.99/hr |
| NVIDIA H100 SXM | 80 GB | 16 | 200 GB | $2.69/hr |
| NVIDIA RTX Pro 6000 Blackwell | 96 GB | 28 | 160 GB | $1.89/hr |
| NVIDIA A100 80GB | 80 GB | 16 | 112 GB | $1.49/hr |
| NVIDIA A100 40GB | 40 GB | 16 | 112 GB | $0.89/hr |
| NVIDIA A30 | 24 GB | 16 | 112 GB | $0.41/hr |
| NVIDIA L4 | 24 GB | 32 | 124 GB | $0.44/hr |
vCPU and system RAM scale with the number of GPUs in the instance.
Instance storage is $0.00014/GB/hour. A paused instance is charged storage only.
GPU availability and supported counts vary by region. Choose your region in the dashboard.
From launch to work
GPU drivers are ready. Install your runtime, use your own images, or join an existing Kubernetes cluster.
Select a GPU, count and region in the dashboard. Most VMs launch in about 90 seconds; allow up to ~8 minutes for eight GPUs.
Take root with sudo. Install Docker, configure your drivers, or bring your own orchestration.
Pausing stops GPU charges. Your data and environment stay on disk, billed at $0.00014/GB/hour.
Choose your level of control
VMs and Templates use the same GPUs at the same hourly rates. Choose who manages the software underneath your code.
On-Demand VMs · This page
For custom kernels, your own Docker daemon, or a GPU node in your Kubernetes cluster.
Templates
For notebooks, fine-tuning and image generation in an environment that is already configured.
Before you launch
The practical details, including launch times, storage charges and configuration limits.
Yes. SSH in and take root with sudo: install what you like, patch the kernel, load your own modules. It is a virtual machine, not a namespace on somebody else’s host.
Yes, your own daemon at your own version, pulling from your own registry. The NVIDIA container toolkit is yours to install or skip.
Yes. The VM behaves as a GPU node: install a kubelet and join an existing cluster, or run a single-node cluster on the box. We do not operate your control plane.
Not measurably, for the GPU paths we tested. In internal 8× H200 runs, VM and container results stayed within 1.2% of each other across tensor compute, HBM bandwidth, NCCL communication and nanochat training throughput, and five of the six results were within 0.5%. The GPUs reach the VM by direct passthrough, so CUDA runs on the card itself.
About 90 seconds to a root shell. If you want a working PyTorch environment in 1.8 seconds instead, that is what Templates are for.
GPU charges stop the moment you pause. You pay storage only, at $0.00014/GB/hour, and your disk is kept.
No. Your GPUs are dedicated to your instance for as long as it runs, and on a VM you control the kernel, the drivers and the OS on top of them.
Eight is the maximum in a single VM. For workloads that need more GPUs, use a multi-node cluster across several VMs. Talk to us about a cluster configuration for your workload.
Not every GPU is offered at every count. NVIDIA H200 SXM, NVIDIA H100 SXM, NVIDIA RTX Pro 6000 Blackwell, NVIDIA L4 run at 8 per instance. The rest are capped lower, and the configurations table shows exactly which.
A $10 minimum to add credit. Billing is per minute with no minimum rental period. There is no free trial.
Your next experiment starts here
Start with $10 in credit. Pay by the minute. No minimum rental period.