Open models.
Ready to deploy.

Deploy a model from our catalog as an API endpoint, with the GPU and serving configuration managed by Jarvislabs.

Browse models

You choose the model and scaling settings. We maintain the serving stack.

The serving configuration is included

We prepare the model.
You build the application.

Our team selects, tests and maintains the configuration for each model in the catalog.

GPU and serving image

We select the GPU configuration and serving image for the model, so you don’t have to assemble the stack yourself.

Precision and performance

We use the model authors’ recommended precision, choose parallelism and kernels, and benchmark the configuration before release.

Ongoing improvements

We incorporate serving optimizations into the image as our team validates them.

Read how we build Managed Endpoints

Scaling and billing

Choose how much
capacity stays ready.

Set worker scaling to match your traffic and response-time requirements.

Keep workers warm for critical workloads

Large models take time to load, and GPU availability can change. Keep at least one worker warm when your application needs to respond without waiting for a cold start.

Scale to zero when waiting is acceptable

Let workers stop when they are not needed if your workload can tolerate the time needed to start a worker and load the model.

Pay for worker runtime

Billing is per GPU-minute of worker runtime. Warm workers are billed even when idle; current rates are shown in the deployment flow.

Two ways to serve models

Choose who manages
the serving stack.

Both products support scaling, with different levels of control over the deployment.

Managed Endpoints

Choose a model from the catalog and use our maintained serving configuration.

  • Jarvislabs selects the GPU and serving image
  • Our team maintains model-specific optimizations
  • You configure worker scaling
Browse models

Serverless

Configure your own serving stack when your workload needs more flexibility.

  • Choose your serving framework and configuration
  • Control how your model runs
  • Scale workers with your workload
Explore Serverless

Before you deploy

Common questions.

Details about configuration, scaling and costs.

What does Jarvislabs manage?

We select and maintain the GPU configuration, serving image, model precision, parallelism and kernels for each model in the catalog. You choose the model and configure how the deployment scales.

Which models can I deploy?

The dashboard catalog shows the current models and deployment options. We add support as we prepare and validate serving configurations, so check the catalog for the latest selection.

Can I bring my own serving stack?

Choose Serverless when you need control over the serving stack. Managed Endpoints use the configurations maintained by our team for models in the catalog.

Can an endpoint scale to zero?

Yes. For large models serving critical workloads, we recommend keeping at least one worker warm. Starting a new worker depends on GPU availability and takes time to load the model.

How am I billed?

You pay for GPU worker runtime, billed by the minute, rather than per token. Warm workers incur charges even when they are not handling requests. Check the deployment details for current pricing.

Can you help with capacity or data residency requirements?

Talk to our team about your model, expected traffic and region requirements. We’ll work through the deployment options and available capacity with you.

Get started

Deploy your model.

Browse the catalog or talk to us about the capacity your application needs.

Browse models