Managed Endpoints
Choose a model from the catalog and use our maintained serving configuration.
- Jarvislabs selects the GPU and serving image
- Our team maintains model-specific optimizations
- You configure worker scaling
Deploy a model from our catalog as an API endpoint, with the GPU and serving configuration managed by Jarvislabs.
You choose the model and scaling settings. We maintain the serving stack.
The serving configuration is included
Our team selects, tests and maintains the configuration for each model in the catalog.
We select the GPU configuration and serving image for the model, so you don’t have to assemble the stack yourself.
We use the model authors’ recommended precision, choose parallelism and kernels, and benchmark the configuration before release.
We incorporate serving optimizations into the image as our team validates them.
Scaling and billing
Set worker scaling to match your traffic and response-time requirements.
Large models take time to load, and GPU availability can change. Keep at least one worker warm when your application needs to respond without waiting for a cold start.
Let workers stop when they are not needed if your workload can tolerate the time needed to start a worker and load the model.
Billing is per GPU-minute of worker runtime. Warm workers are billed even when idle; current rates are shown in the deployment flow.
Two ways to serve models
Both products support scaling, with different levels of control over the deployment.
Choose a model from the catalog and use our maintained serving configuration.
Configure your own serving stack when your workload needs more flexibility.
Before you deploy
Details about configuration, scaling and costs.
We select and maintain the GPU configuration, serving image, model precision, parallelism and kernels for each model in the catalog. You choose the model and configure how the deployment scales.
The dashboard catalog shows the current models and deployment options. We add support as we prepare and validate serving configurations, so check the catalog for the latest selection.
Choose Serverless when you need control over the serving stack. Managed Endpoints use the configurations maintained by our team for models in the catalog.
Yes. For large models serving critical workloads, we recommend keeping at least one worker warm. Starting a new worker depends on GPU availability and takes time to load the model.
You pay for GPU worker runtime, billed by the minute, rather than per token. Warm workers incur charges even when they are not handling requests. Check the deployment details for current pricing.
Talk to our team about your model, expected traffic and region requirements. We’ll work through the deployment options and available capacity with you.
Get started
Browse the catalog or talk to us about the capacity your application needs.