Blog/Perspectives

The $500 Billion Buys the Factories. Utilization Decides the Winners.

Vishnu Subramanian
Vishnu Subramanian

Founder @ JarvisLabs.ai

August 11, 2026·10 min read

On August 10, Nvidia announced partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR: financing platforms designed to mobilize over $500 billion of third-party capital for AI infrastructure. Jensen Huang's framing was direct: AI factories can now be financed as productive infrastructure, backed by long-term institutional capital. In his words, "in AI, compute is revenue."

The market's first reaction was doubt. Nvidia's stock fell nearly 3% on the day, and most of the coverage circled the same two words: circular financing.

I've spent the last six years building Jarvis Labs, a GPU cloud platform serving customers across the world. I sit at the layer where capital like this actually becomes productive infrastructure: buying inventory, pricing it, and living with the consequences when it sits idle.

From that seat, I think both the cheerleaders and the skeptics are arguing about the wrong thing. The $500 billion buys the factories. But factories don't produce revenue. Utilization does. Compute is revenue only for the companies that convert it into intelligence efficiently. Infrastructure alone won't decide the winners of this buildout. The companies that use it well will.

Here's what that looks like from the ground.

The pricing reversal nobody predicted

Historically, compute was considered a depreciating asset. Nvidia ships a new GPU generation every two to three years, and the standing question was whether the older generation would become obsolete before it paid for itself.

Two years back, the market behaved exactly that way. A huge wave of supply came in, and providers competed by cutting prices on H100s and H200s. It looked like a race to the bottom.

Over the last year, we've watched that trend reverse. Prices are increasing. Availability has gone down. And customers who used to rent by the hour now want to commit for much longer durations, to make sure the compute they need for their workloads is actually there.

The numbers on our own platform tell the story. Our on-demand price today is $2.69 per hour for an H100 and $3.99 for an H200, which is 20 to 30% higher than a year back. For committed capacity, pricing depends on cluster size and duration: a large H200 cluster runs anywhere between $3.20 and $4.00 per hour, and a two- or three-year contract can bring that toward the lower end of $3 to $4.

Here's the detail that surprises people: when customers want an instant cluster, capacity right now rather than next quarter, they usually end up paying more than the on-demand price. Scarcity has inverted the old discount logic.

Nvidia cited the same trend in its announcement: one-year H100 rentals rose from about $1.70 to $2.35 per GPU-hour between October 2025 and March 2026, and cross-provider on-demand medians moved from roughly $2.00 to $2.70. We saw the same movement in our own pricing well before the announcement. The asset is holding value.

The layers between compute and intelligence

But holding value and producing value are different things, and this is where the $500 billion question actually lives.

A lot of companies can buy raw compute, buy software from someone else, club it together, and sell it to end customers. That's one model. The simplest way to make money from a GPU cluster is to rent it out as bare metal, and it's a good place to start. But bare metal becomes a commodity very quickly. When Jensen talks about generating intelligence out of compute, he's not talking about selling the compute anymore. To get there, you have to move up several layers. At Jarvis Labs, we've spent six years building them.

The layers between compute and intelligence: the Jarvis Labs stack, from bare metal to Token Factory

The first layer is VMs and templates built on top of containers: single machines with up to eight GPUs (H100s, H200s), scaling up to large clusters connected over InfiniBand, with observability across them. Alongside this sits another revenue stream: spot, or preemptible, instances, priced lower for workloads that can tolerate interruption.

The second layer is burst compute. Let's say you want to run 1,000 operations at once, each lasting a few seconds or minutes. VMs don't fit that shape of work. Our containers cold-start in 1.8 seconds, which is what makes that workload economical.

The third layer is serving. Say you want an open-source model deployed for use inside your organization. Now you need a load balancer and intelligent routing across serving layers, and every company building that plumbing themselves is not making the best use of their time. Our serverless platform provides it out of the box.

On top of that sit custom endpoints. Let's say a company wants to move to an open-source model like GLM 5.2. The traditional path is renting or buying GPUs, hiring AI engineers and data scientists, and spending several weeks to months plumbing everything together before consuming a single token. With custom endpoints, they deploy that model in a few clicks, and our engineering and data science teams keep optimizing it as new techniques arrive, shipping them when they're production-ready. Customers benefit from that research without spending anything extra on it.

What does that optimization actually mean? Concrete things. Loading a very large model from disk to GPU memory can take anywhere from five to ten minutes out of the box; our optimizations bring that to two to three minutes. At scale, those minutes pile up into real cost. For smaller models, GPU snapshotting brings loading down to one or two seconds. To be honest about the limits: snapshotting works well when the model fits in a single GPU. For a model that has to fit across eight GPUs, it doesn't work well today. We are working on figuring that out, and as Nvidia makes it more possible, we will bring it to large models as well, which would make loading them much faster than it is today. Let's say a large model could load in under 30 seconds or a minute: that changes the unit economics for our end customers. We work on speculative decoding: sometimes off-the-shelf, and sometimes we train our own draft models to make sure customers get the right results.

The reason we can do this is that we operate the entire stack, from bare metal to the top layer. Companies that operate on just one layer of someone else's stack have a different, smaller set of optimizations available to them. In an inference market growing this fast, that difference compounds.

Efficiency is what gets financed

There's a second-order effect in Nvidia's announcement that I think most coverage missed. The financing partners will independently underwrite each opportunity on demand, utilization, cash flow and residual value.

Read that again: utilization is now underwriting criteria. In this model, efficiency isn't just an operational virtue. It's what makes a company financeable. Inefficient operators won't get capital.

Utilization is also why we've operated as a global company from day one. For a long time, close to 70% of our revenue came from outside India, and international revenue contributes to a major extent even today. The reason is simple mathematics. Many workloads are time-sensitive. If you sell only to Indian customers, you have roughly 12 hours a day to optimize revenue generation. The other 12 hours, your inventory sits idle. Sell globally, and the same GPUs generate revenue around the clock. Going global isn't just a growth strategy for a GPU cloud. It's a utilization strategy.

My hope is that Indian financiers see this shift too. The way the world underwrites this buildout is changing, and there is a real opportunity for Indian capital to participate in it early. And the direction is clear: as efficiency improves, fund managers will look for efficiently run companies to back.

The bubble question, honestly

Is this circular financing? This is probably one of the hardest questions, and at a high level I understand why people see it that way. But there's another perspective from where I sit.

In the last six months, the ask for inventory has been way more than what we could supply. That's demand we see directly: companies paying real money, committing to real durations.

And enterprises, historically the last to adopt any technology, are only just starting. My estimate is that enterprise adoption is under 2 to 3% today. When that reaches 10 or 15%, the demand will be far greater, and the infrastructure is not ready for it. Someone has to take that risk, and our existing financial systems may not be built to take it. Nvidia is betting on the larger vision, and the financiers will run independent evaluation on each company: what is their vision, their strategy, how far are they from achieving it, what have they achieved. Anyone who has taken even a small bank loan knows the due diligence involved. At billions of dollars, it will be far more than we anticipate.

Two more things give me conviction. First, even if AI research stopped across the world today, adopting the innovation that already exists would take years, and would need far more compute than currently exists. Second, the depreciation math is running backwards. Models and serving technology keep getting more efficient, which means you can squeeze more value out of inventory created two years back than you could when it was installed. Nvidia points to the A100, still commercially productive six years after launch. Our experience says the same. The asset improves with software while usage keeps growing. That's not what a bubble asset looks like, but it only holds for operators who actually do the squeezing.

Where this goes: tokens, and where they're generated

For the last six years, we've been building the fundamental blocks of what we call the Token Factory. Our serverless layer has been live for some time and is used by a lot of companies today; the Token Factory sits on top of it. The honest blocker to shipping it has been inventory availability for our own use, and that resolves in the next couple of months. When it does, we ship.

For a large country like India, there's one more dimension: sovereignty. India cannot rely on tokens generated outside the country as its primary source of intelligence. We want to be one of the primary places where people consume these tokens. As our software matures, we can leverage it across inventory available anywhere in the world.

So here's the one idea worth keeping from all the coverage of Nvidia's announcement. The $500 billion is real, the demand is real, and the underwriting will be more disciplined than the skeptics assume. But capital and chips are the entry ticket, not the game. Every industrial revolution was built on infrastructure, and in every one, the fortunes went to the operators who ran that infrastructure best. In AI, compute is revenue. For the companies that know how to use it.

Get Started

Build & Deploy Your AI in Minutes

Cloud GPU infrastructure designed specifically for AI development. Start training and deploying models today.

View Pricing