
> GGUF VRAM Calculator: Check Before You Download▋
2 October 2026· 4 min read
One field note a week, no noise — get it by email
ggml_backend_cuda_buffer_type_alloc_buffer: failed to allocate 3412 MiB. The file downloaded fine. The card was never going to take it — and the arithmetic that predicted this would fit in the margin of a receipt. The usual workflow is backwards: hit download, watch the progress bar, let the out-of-memory error do the math. That ordering is what today's tool refuses. The GGUF VRAM calculator is live: model size, quantization and context in — weights, KV cache and a verdict for every common card out. Or run it in reverse from your VRAM budget to the largest quant that fits. It runs entirely in your browser, nothing is uploaded, and it joins the run-LLMs-locally hub next to the rest of this cluster.
It has two modes, because the question comes in two shapes:
- "How much VRAM does this model need?" — parameters, quant, context, architecture in; a breakdown and a card-by-card verdict table out (6 through 96 GB).
- "What fits in my VRAM?" — budget, model size, context in; every quant from Q2_K to F16 with its own verdict out.
How much VRAM does a GGUF model need?
Three parts, always the same three:
VRAM ≈ weights + KV cache + overhead
weights = params × bits-per-weight / 8
KV cache = 2 × layers × kv_dim × context × 2 bytes (f16)
overhead ≈ 0.7 GB CUDA/runtime + ~5% compute buffers
Worked example, the most common one: an 8B model at Q4_K_M with 32k context. Weights: 8 × 4.85 / 8 = 4.85 GB. KV cache: 2 × 32 layers × 1024 kv-dim × 2 bytes = 128 KB per token × 32,768 = 4.0 GB. With overhead: ~9.8 GB total — "Q4 8B", the model everyone calls an 8-GB-card model, misses an 8 GB card by 23% the moment you give it a long context. The quant was never the villain; the context was. That failure mode has its own field note, because it also eats system RAM when you offload.
Which GGUF quant fits my card?
Mode 2 exists because the reverse question is what people actually have: a fixed card, a model in mind. A 7B at 4096 context against an 8 GB budget returns exactly this:
| Quant | Weights | Total (est.) | Verdict |
|---|---|---|---|
| Q4_K_M | 4.24 GB | ~5.7 GB | fits with headroom |
| Q6_K | 5.77 GB | ~7.3 GB | tight |
| Q8_0 | 7.44 GB | ~9.0 GB | partial CPU offload, much slower |
One screen, and the age-old forum debate "Q4 or Q8" collapses into your hardware's answer instead of somebody else's.
Where the numbers come from
The bits-per-weight table comes from llama.cpp's ggml quant formats — the same numbers the files on Hugging Face are built with:
| Quant | bpw | Quant | bpw |
|---|---|---|---|
| Q2_K | 3.35 | Q5_K_M | 5.69 |
| Q3_K_M | 3.91 | Q6_K | 6.59 |
| IQ4_XS | 4.25 | Q8_0 | 8.50 |
| Q4_K_M | 4.85 | F16 | 16.0 |
The KV cache half uses each family's public geometry. Grouped-query attention is the whole story there: a Llama-3-class 8B keeps 8 KV heads × 128 dims × 32 layers, which is 128 KB per token — while a Llama 2 13B, with no GQA, burns 640 KB per token, five times more, from an older and nominally smaller-era model. The architecture dropdown carries the four common shapes plus a custom mode for anything you can read off a config.json.
The file size told you what you would download. It never told you what you could run.
What the estimates deliberately miss
Three things, on purpose. Mixture-of-experts routing: only the active experts get touched at inference, but this tool prices the whole weight set, so MoE totals read high. Mixed quantization: Q4_K_M is itself an average across tensors, so a real file lands a few percent off the table value. And CUDA-side padding varies by backend version — the flat 0.7 GB is a middle-of-the-road figure, not a constant of nature. When a decision is worth more than a few hundred megabytes, skip the estimates and read the file's header with llama.cpp's gguf_dump.py — header-reading calculators do exactly that. This one stays arithmetic on purpose: three inputs you can re-derive by hand, and a verdict you can argue with.
The tool is at github.com/mrsaynothing/gguf-vram-calculator — one HTML file, no dependencies, no build step. It sits alongside the local GGUF guide, llama.cpp vs Ollama and the best-local-LLM shortlist, with the serving-side counterpart in can vLLM run GGUF.
If a verdict disagrees with your load times, open a repository issue with the model, quant and context — the table should survive contact with real hardware, and anywhere it doesn't is a row to fix.
faq
+ How much VRAM does a 7B GGUF model need?
A Q4_K_M of a 7B model holds about 4.2 GB of weights; with 4k context and runtime overhead, plan for roughly 5.7 GB — a comfortable fit on an 8 GB card. Q8_0 of the same model at the same context is ~9 GB and will not fit.
+ Does context length affect VRAM?
Yes, through the KV cache. A Llama-3-class 8B with GQA burns about 128 KB per token at f16, so 32k context costs ~4 GB before a single weight is loaded. The context, more than the quant, is what breaks the budget.
+ Is the GGUF file size the VRAM I need?
No. File size approximates the weights only. KV cache grows with context, and runtime overhead (CUDA buffers, compute margins) adds a fixed chunk on top. Disk space and VRAM are different budgets.
+ How accurate is a VRAM calculator?
This one is deliberate napkin arithmetic from public quant tables — right for the fits-or-not decision within a few hundred MB. For byte-exact numbers, read the GGUF header itself with llama.cpp's gguf_dump.py.
— mrsaynothing
$ Related entries
We Deleted Our CI. The Machine Ships Anyway.
2026-10-01
GitHub CI deleted: 118 lines of workflow YAML gone, one 27-line local script in. What got safer, what got worse, and the receipts from both.
Can vLLM Run GGUF? Yes — on GPU Only
2026-09-23
Can vLLM run GGUF? Yes — via the official plugin, on GPU only. The serve syntax, the tokenizer trap, the hardware limits, and when llama.cpp still wins.
GGUF Quantization Levels: Q4_K_M vs Q8_0, Size and VRAM
2026-09-13
GGUF quantization levels compared: Q4_K_M vs Q8_0 on size, VRAM and quality, plus the one rule — default Q4_K_M, step up only for code and math.