mrsaynothingmrsaynothing's blog
← Back to blogGGUF VRAM Calculator: Check Before You Download

> GGUF VRAM Calculator: Check Before You Download▋

2 October 2026· 4 min read

One field note a week, no noise — get it by email

ggml_backend_cuda_buffer_type_alloc_buffer: failed to allocate 3412 MiB. The file downloaded fine. The card was never going to take it — and the arithmetic that predicted this would fit in the margin of a receipt. The usual workflow is backwards: hit download, watch the progress bar, let the out-of-memory error do the math. That ordering is what today's tool refuses. The GGUF VRAM calculator is live: model size, quantization and context in — weights, KV cache and a verdict for every common card out. Or run it in reverse from your VRAM budget to the largest quant that fits. It runs entirely in your browser, nothing is uploaded, and it joins the run-LLMs-locally hub next to the rest of this cluster.

It has two modes, because the question comes in two shapes:

  • "How much VRAM does this model need?" — parameters, quant, context, architecture in; a breakdown and a card-by-card verdict table out (6 through 96 GB).
  • "What fits in my VRAM?" — budget, model size, context in; every quant from Q2_K to F16 with its own verdict out.

How much VRAM does a GGUF model need?

Three parts, always the same three:

VRAM ≈ weights + KV cache + overhead
weights  = params × bits-per-weight / 8
KV cache = 2 × layers × kv_dim × context × 2 bytes   (f16)
overhead ≈ 0.7 GB CUDA/runtime + ~5% compute buffers

Worked example, the most common one: an 8B model at Q4_K_M with 32k context. Weights: 8 × 4.85 / 8 = 4.85 GB. KV cache: 2 × 32 layers × 1024 kv-dim × 2 bytes = 128 KB per token × 32,768 = 4.0 GB. With overhead: ~9.8 GB total — "Q4 8B", the model everyone calls an 8-GB-card model, misses an 8 GB card by 23% the moment you give it a long context. The quant was never the villain; the context was. That failure mode has its own field note, because it also eats system RAM when you offload.

Which GGUF quant fits my card?

Mode 2 exists because the reverse question is what people actually have: a fixed card, a model in mind. A 7B at 4096 context against an 8 GB budget returns exactly this:

Quant Weights Total (est.) Verdict
Q4_K_M 4.24 GB ~5.7 GB fits with headroom
Q6_K 5.77 GB ~7.3 GB tight
Q8_0 7.44 GB ~9.0 GB partial CPU offload, much slower

One screen, and the age-old forum debate "Q4 or Q8" collapses into your hardware's answer instead of somebody else's.

Where the numbers come from

The bits-per-weight table comes from llama.cpp's ggml quant formats — the same numbers the files on Hugging Face are built with:

Quant bpw Quant bpw
Q2_K 3.35 Q5_K_M 5.69
Q3_K_M 3.91 Q6_K 6.59
IQ4_XS 4.25 Q8_0 8.50
Q4_K_M 4.85 F16 16.0

The KV cache half uses each family's public geometry. Grouped-query attention is the whole story there: a Llama-3-class 8B keeps 8 KV heads × 128 dims × 32 layers, which is 128 KB per token — while a Llama 2 13B, with no GQA, burns 640 KB per token, five times more, from an older and nominally smaller-era model. The architecture dropdown carries the four common shapes plus a custom mode for anything you can read off a config.json.

The file size told you what you would download. It never told you what you could run.

What the estimates deliberately miss

Three things, on purpose. Mixture-of-experts routing: only the active experts get touched at inference, but this tool prices the whole weight set, so MoE totals read high. Mixed quantization: Q4_K_M is itself an average across tensors, so a real file lands a few percent off the table value. And CUDA-side padding varies by backend version — the flat 0.7 GB is a middle-of-the-road figure, not a constant of nature. When a decision is worth more than a few hundred megabytes, skip the estimates and read the file's header with llama.cpp's gguf_dump.py — header-reading calculators do exactly that. This one stays arithmetic on purpose: three inputs you can re-derive by hand, and a verdict you can argue with.

The tool is at github.com/mrsaynothing/gguf-vram-calculator — one HTML file, no dependencies, no build step. It sits alongside the local GGUF guide, llama.cpp vs Ollama and the best-local-LLM shortlist, with the serving-side counterpart in can vLLM run GGUF.

If a verdict disagrees with your load times, open a repository issue with the model, quant and context — the table should survive contact with real hardware, and anywhere it doesn't is a row to fix.

faq

+ How much VRAM does a 7B GGUF model need?

A Q4_K_M of a 7B model holds about 4.2 GB of weights; with 4k context and runtime overhead, plan for roughly 5.7 GB — a comfortable fit on an 8 GB card. Q8_0 of the same model at the same context is ~9 GB and will not fit.

+ Does context length affect VRAM?

Yes, through the KV cache. A Llama-3-class 8B with GQA burns about 128 KB per token at f16, so 32k context costs ~4 GB before a single weight is loaded. The context, more than the quant, is what breaks the budget.

+ Is the GGUF file size the VRAM I need?

No. File size approximates the weights only. KV cache grows with context, and runtime overhead (CUDA buffers, compute margins) adds a fixed chunk on top. Disk space and VRAM are different budgets.

+ How accurate is a VRAM calculator?

This one is deliberate napkin arithmetic from public quant tables — right for the fits-or-not decision within a few hundred MB. For byte-exact numbers, read the GGUF header itself with llama.cpp's gguf_dump.py.

— mrsaynothing

$ Share this post

$ Get the next build log by email

One email per post. Receipts and mistakes included.

self-hosted · no third parties · one-click unsubscribe

what is this?