Most inference GPUs are bought for the weights and starved by the cache. Llama 3 8B in bf16 needs about 16 GB for its weights, so it “fits” on a 24 GB card. At a 0.9 memory ceiling, roughly 5.5 GB is left over, which holds about five 8,000-token conversations at once. GPU cost optimisation is the practice of choosing hardware, precision and batching so that cost per served token falls, not just cost per GPU-hour. Those are different numbers, and teams usually optimise the wrong one.
Why does the KV cache decide GPU size?
The KV cache, not the weights, sets how many requests a GPU can serve at once. Every generated token keeps its attention keys and values in memory for the rest of the sequence. The size per token is 2 × layers × KV heads × head dimension × bytes per value. For Llama 3 8B that is 2 × 32 × 8 × 128 × 2 bytes, or 128 KiB per token, and 1 GiB for a single 8,192-token sequence. The attention mechanism behind that cache explains why it grows linearly with context and cannot be skipped.
How that memory is managed matters as much as how much there is. The vLLM team showed in 2023 that serving systems of the time wasted most of their KV memory to fragmentation and over-reservation.
The vLLM launch post attributes its throughput gains to exactly this: paging the cache lets far more sequences share one card. If your serving stack predates paged attention, switching engines is the cheapest GPU cost optimisation available, ahead of any hardware change.
| Setup for Llama 3 8B | KV room at 0.9 ceiling | Concurrent 8K sequences |
|---|---|---|
| 24 GB card, bf16 e.g. NVIDIA L4 |
About 5.5 GB | About 5 |
| 24 GB card, 4-bit weights | About 16 GB | About 14, at some accuracy cost |
| 40 GB card, bf16 | About 20 GB | About 18 |
| 80 GB card, bf16 e.g. A100 80GB, H100 |
About 56 GB | About 52 |
These figures ignore activation memory and CUDA overhead, so treat them as ceilings. The shape is what matters. Moving from 24 GB to 80 GB triples memory but multiplies concurrent capacity roughly tenfold, because the weights are a fixed cost paid once per card.
Is the biggest GPU always the cheaper option?
No, and the counter-argument deserves weight. A large card only wins when traffic keeps it full. Decoding is limited mostly by memory bandwidth, so batching fifty sequences costs little more per step than batching five, and cost per token drops sharply under sustained load. At 3 a.m., with four users online, that same card bills the full hourly rate for work a 24 GB part could handle.
Bursty internal tools, overnight batch jobs and low-traffic pilots all fall on the small-card side. So do workloads that can scale to zero on serverless infrastructure, where idle time costs nothing. Quantisation shifts the line as well: 4-bit quantisation nearly triples KV room on a 24 GB card, at an accuracy cost you must measure on your own evaluation set rather than assume.
Where does GPU cost optimisation leave your sizing decision?
Size from measured traffic, in this order. Take p95 concurrent sequences and p95 context length from production logs. Multiply them by the per-token KV size to get the cache you need. Add the weights, divide by 0.9, and pick the smallest card, or card count, that fits. Then load-test that configuration and compute cost per million output tokens at real utilisation. Compare configurations on that number alone, because cost per GPU-hour hides the batch size you actually achieved.
GPU cost optimisation is not a one-off exercise, so recheck after every model change. Moving from an 8K to a 32K context window quadruples the cache per sequence, and the card that was right last quarter becomes the bottleneck without a single alert firing.
- GPU cost optimisation for inference should minimise cost per served token, which depends on achieved batch size, not on the hourly GPU price alone.
- Llama 3 8B in bf16 needs about 128 KiB of KV cache per token, so one 8,192-token sequence occupies 1 GiB of GPU memory.
- The vLLM team reported that pre-PagedAttention serving systems wasted 60 to 80% of KV memory, against under 4% with paging.
- Larger GPUs win under sustained load because weights are a fixed cost, while smaller cards or scale-to-zero win for bursty or low traffic.
- Size GPUs from p95 concurrency and p95 context length, then compare configurations on cost per million output tokens at measured utilisation.
Conclusion
Self-hosting only pays when your cards stay busy, so run the sizing exercise above before signing any reservation. If utilisation comes out low, compare it against managed endpoints first. Generative AI as a service in the cloud covers that build-versus-buy trade-off.
Frequently Asked Questions
How much GPU memory does an 8B LLM need for inference?
About 16 GB for the weights in bf16, plus KV cache for every active sequence. For Llama 3 8B the cache is roughly 128 KiB per token, so each 8,192-token conversation adds about 1 GiB. A 24 GB card therefore serves only a handful of long concurrent conversations at full precision.
How do I calculate KV cache size per token?
Multiply 2 by the number of layers, the number of key-value heads, the head dimension and the bytes per value. For Llama 3 8B in bf16 that is 2 × 32 × 8 × 128 × 2 bytes, or 128 KiB per token. Multiply by context length and concurrent sequences for the total.
Is a bigger GPU cheaper for LLM inference?
Only under sustained load. A larger card holds many more concurrent sequences because the weights are a fixed cost, so cost per token falls when it stays busy. At low or bursty traffic it bills full price while idle, and a smaller card or scale-to-zero option usually costs less.
Does quantization reduce GPU cost for inference?
Usually, yes. Storing weights in 4 bits instead of 16 cuts weight memory by about three quarters, freeing room for KV cache and more concurrent requests on the same card. The trade-off is some accuracy loss, which you should measure on your own evaluation set before deploying.
What metric should I use to compare GPU configurations?
Use cost per million output tokens measured under realistic load. Cost per GPU-hour ignores how many requests the card actually batched, which is where most of the difference between configurations lies. Load-test each option with production-like prompt and output lengths before comparing.