Scaling LLM Inference

From Fixed Reservations to Burst-Out Parallelism

October 9, 2026
Ted De Graaf
© 2026 Schmied Enterprises LLC. All rights reserved.

Introduction

Large language model (LLM) inference is often billed on a per-token basis, which can feel expensive when it adds up and you are challenged with changing per-token bills every month. Hiring a manager to deal with it adds to the costs.

Most providers also offer an option called “reserved instance” where you pay a flat annual fee that covers a fixed size server container in advance. The downside of that model is that the data-center is often left idle – you pay for space you never use, and the cost per token stays high because the service has to stay on standby even when traffic is low.

In this article we walk through how a simple scaling strategy changes those numbers. We start by estimating what part of the datacenter capacity is actually being used under the burst plan knowing the reserved market prices and costs. Then we look at what happens if you switch to burst-mode agents that share a single node, and finally we see how adding many parallel nodes (e.g., 50 sub-agents) turns the whole system into a high-throughput engine.

The calculations are useful well beyond a single bill. They help estimate financing needs, support the valuation of equity for CTOs, and even offer a way to check the real profit margins of hyperscalers.

Estimating Utilisation When Scaling Out

Our baseline service moves 100 tokens per minute in and out together. If we keep using only one agent instance at this rate, the total yearly volume is roughly 52 million tokens. At a price of $2 per million tokens that works out to about $105 a year for what you would have paid if every token were billed at market rates.

The reserved-instance plan costs $36 per year, which is far less than the $105 you’d pay in a pure “pay-as-you-go” model. The ratio of 36 to 105 tells us that only about one-third of the full capacity is being used of the per token on-demand partition – roughly a 34 % utilisation. This is the case of burst instances shared by per-token customers. Security and compliance is a reason why banks and hospitals are tied to reserved instances: a dedicated container stays on standby even when much of it sits idle, there to guarantee service rather than to fill every token of capacity.

Infrastructure teams are pushed to the input cost of the servers themselves, competing with each other on both the supplier side and the customer side. Their margin on reserved instances competes with a PC on a desk.

Speed-Up When Burst Agents Share One Node

If you replace the reserved agent with a burst-mode instance and run subagents in parallel, you can take advantage of available capacity. The slice stays proportional to customer use. Cloud nodes usually come with reserved output egress proportional to their core count. A single request can be generated faster, which reduces the latency and increases overall throughput. The effect is roughly three times faster due to the 34% average utilisation of the burst partition. Each token-generation unit still respects its own limit of 100 tokens per minute, and running those subagents in parallel multiplies the output by a factor near 3.

So on a single burst node, with subagents running in parallel, you can expect an effective throughput close to 290 tokens per minute – more than twice what the reserved model could deliver, with lower latency on each request.

Parallel Burst Nodes: Adding More Sub-Agents

Now imagine we spin up fifty independent burst agents on separate nodes. Each agent is still capped at 100 tokens per minute, but because they are not sharing a single node, the multiplier from step two no longer applies to each individual node; instead you simply add their rates together. The result is a system that can handle 5 000 tokens per minute – ten times the original reserved output.

The key insight here is linear scaling: every extra sub-agent adds another 100 tokens per minute, and there is no diminishing return as long as you have enough nodes to support them. Hyperscalers win. Nodes usually share a few pinned models making read-only requests scalable.

The cost remains $2 per million tokens, and customers who run sparse large queries can benefit. Reserved-instance customers can achieve the same result, but at a heftier price, since they do not share their capacity with other customers. Ten nodes cost $360 a year.

Bottom Line

By moving from a fixed reservation model to one that scales out, a single request is generated faster, latency falls, and overall throughput rises, while the used slice stays proportional to customer use. Capacity can still grow linearly. Burst customers keep paying $2 per million tokens, which favors sparse large queries, while reserved customers pay a heftier price for capacity they do not share. This end-to-end view shows how a modest change in architecture can turn a modest service into a high-performance solution.

Reserved nodes move. Shared burst servers stay near 34%. Two reserved nodes fluctuate from empty to full. Ten shared burst servers wobble, and one sometimes hits 100%. RESERVED · 2 NODES Utilisation fluctuates, 0–100% 100% NODE 1 NODE 2 100 tok/min $36 / node / year Not shared with other customers, randomly utilized, 100% paid BURST SHARED · 10 SERVERS Each moves a little. Sometimes one reaches 100%. Long-term average stays 34%. 34% 100% Shared by per-token customers $2 / million · long-term average 34%