The most common home-lab disappointment is not a slow GPU or a bad model. It is a model that loaded fine, answered the first three questions fine, and then at message twelve slowed to a crawl while the system monitor showed VRAM pinned at 100%. Nothing was broken. The user had budgeted for the file on disk and the machine was being charged for work that happens after loading.
Two numbers, not one
Memory during inference splits into two pools that behave nothing alike. The first is weights: the model itself, laid out in VRAM when it loads. The second is the KV cache: the key/value tensors scraped off every token the model has already processed, kept so it doesn't have to re-read the conversation each turn. Weights are a fixed cost. The KV cache is a variable cost that starts near zero and grows with every token of context — including the ones you generated yourself.
This is why a setup can be perfectly stable for a short chat and unusable at document-length context. The weights never changed. The cache did.
Sizing the weights honestly
The usual back-of-envelope is "quarter of the 16-bit size" for 4-bit quantization. That is close enough for a shopping decision and wrong for a memory plan, for three reasons. Quantization almost never covers the whole model — embedding layers, some normalization weights and the output head commonly stay at higher precision. The runtime holds activations, the compute graph and framework buffers on top of that. And a context window you allocate but don't use is usually still reserved.
A rule that has held up well in practice: budget file size plus 10–15%, then add your context allocation on top. So the map on a single 24 GB card looks like this:
- 7B–14B at 4-bit: roughly 5–9 GB of weights. Enormous headroom — this is where long-context work, RAG over a document pile, and multi-agent experiments live.
- 32B at 4-bit: roughly 19–20 GB with runtime overhead. This is the configuration people actually mean when they say "the card fits" — and it is also the one where context is tight.
- 70B at 4-bit: roughly 40 GB, meaning two 24 GB cards at minimum, with the caveat that multi-GPU splits introduce their own overhead.
Note what the 32B row implies: on 24 GB, the weights have already eaten the card. Everything the model needs for context has to come out of the remaining slice or out of system RAM. See our 3090 piece for why that specific card shows up in this conversation at all.
The KV cache is where the plan dies
The KV cache size, per token, is 2 × layers × KV heads × head dimension × bytes per element — the 2 because every token stores both a key and a value, and each element is typically 2 bytes at fp16. Head dimension is usually 128, and layers run into the dozens for modern mid-size models, so a single token can easily cost well over 100 KB before you multiply by context length.
Multiply by context length and the shape becomes clear. A model with a sizable per-token cache and an 8K window can consume several gigabytes; the same model at 32K can consume more than the weights do. Two details trip people up:
- Grouped-query attention changes everything. Modern models share KV heads across many query heads precisely to reduce this cost. Two models with similar layer counts can differ by 4× in cache size purely because one uses 8 KV heads and the other 32. If you only compare parameter counts, you will mispredict which one runs.
- Quantizing weights does not quantize the cache. These are separate settings. Weight quantization shrinks the fixed cost; KV cache quantization shrinks the growing one.
Practical settings that actually help
Quantize the KV cache
Running the cache at 8-bit instead of 16-bit roughly halves its footprint, and at typical quality thresholds the difference is hard to detect in general chat. This is the single highest-leverage toggle in the whole stack, and it costs you one flag. Quality-sensitive work — long code refactors, precise arithmetic, anything where you'll notice a subtle drift — is where you pay the 16-bit price back deliberately.
Set your context to what you need
Allocating 128K "just in case" on a card that can't serve it is not a plan; it is a slow-motion out-of-memory event. Pick a context you actually use, and treat long-document work as a separate configuration rather than the default.
Watch VRAM, not GPU utilization
The tell for cache spill is not a crash. It is a sudden collapse in tokens per second while the GPU sits near 100% VRAM, because the overflow is being shuffled over PCIe every token. In monitoring tools, prefer memory-residency numbers over percentage-utilization numbers; utilization will look happily busy while you're actually measuring the bus.
Spilling to system RAM is a valid fallback, at a price
Offloading layers to system memory will run a model that otherwise wouldn't fit, and on a machine with fast DDR5 and a wide memory bus it can be usable for single-user use. Expect a large, not marginal, throughput penalty — the difference between conversational and "come back in a minute." If your workload is occasional, that trade can be correct. If it's interactive, it usually isn't, and the honest answer is a smaller model rather than a bigger one that creeps.
Where this leaves the buying decision
The upshot of the arithmetic is that VRAM capacity, not raw compute, is the binding constraint on a home box — and it's binding a specific way: the ceiling is set by weights, while usability within that ceiling is set by the cache at your working context length. A machine that can "technically load" a 32B model is a different machine from one that can hold it plus a real conversation. Before upgrade season, run the numbers for the context length you actually use, with the cache set to 8-bit, and you'll usually find the honest answer is not a bigger card. It's a better-configured one.
Comments