
> vLLM GGUF FAQ: Ten Search Questions, Answered▋
Ten real search questions about vLLM and GGUF — CPU support, serve syntax, quant coverage, the tokenizer trap — answered from the docs and our tested posts.
mrsaynothing· 9 octobre 2026· 6 min de lecture
vllm serve unsloth/Qwen3-0.6B-GGUF:Q4_K_M --tokenizer Qwen/Qwen3-0.6BThe five-minute checklist
The failure modes have a table — work it top to bottom.
| Symptom | Cause | Fix |
|---|---|---|
404 or "model not found" on serve |
GGUF passed as a plain path — vLLM wants repo:quant_type |
Use the repo:quant_type form (next section) |
| Tokenizer error while loading | --tokenizer missing or pointing at the GGUF repo |
Point it at the original base model repo |
| Runs in Ollama/llama.cpp, refuses in vLLM | Same file, different engine — vLLM's GGUF path is GPU-only | CPU boxes stay on llama.cpp |
| Quant scheme rejected at load | Exotic quant outside the plugin's coverage | Drop to Q4_K_M or friends — see the quantization guide |
Can vLLM run GGUF?
Yes — through the official vllm-gguf-plugin, on NVIDIA and AMD GPUs. The plugin is vLLM's own, not a community fork. Both halves matter: the plugin is official, and the hardware table stops at GPU. The full serve walkthrough lives in Can vLLM Run GGUF? Yes — on GPU Only.
Does vLLM support GGUF on CPU?
No. vLLM has no CPU kernel behind its GGUF path — x86/Arm CPUs and Intel GPUs sit outside the supported hardware table. This is the trap behind the classic report: the same GGUF that 404s in vLLM on a laptop runs fine in Ollama, because GGUF's home turf is CPU-first engines. The file was never the problem.
How do you serve a GGUF model with vLLM?
One line — the repo, the quant type, and the tokenizer:
vllm serve unsloth/Qwen3-0.6B-GGUF:Q4_K_M --tokenizer Qwen/Qwen3-0.6B
The syntax is repo:quant_type. The trap is the --tokenizer flag: it names the original model repo, never the GGUF repo — GGUF files carry weights, and vLLM wants the real tokenizer from the source model.
verify: serve it, then nvidia-smi — the vLLM process should be holding VRAM, not system RAM.
Which GGUF quants does vLLM support?
The K-quants people actually download — Q4_K_M and friends — work; exotic schemes may not. llama.cpp remains the reference implementation for the format, so new quant schemes and new model architectures land there first and reach the plugin later. Before downloading anything, size it with the GGUF quantization guide.
Is vLLM's GGUF support production-ready?
The docs answer this one themselves: "GGUF support in vLLM is highly experimental and under-optimized." That sentence sits in the first paragraph of the official GGUF page — which is more honesty than most experimental features get. Treat the plugin as a way to reuse files you already have, and keep the exit to llama.cpp open.
Does vLLM load GGUF faster than llama.cpp?
No — and speed was never the point. llama.cpp memory-maps the file, so a GGUF starts lazily; vLLM loads it like any other checkpoint. GGUF inside vLLM is a memory-footprint feature: the docs frame it as a way to shrink VRAM use, not a throughput play. The throughput win somewhere else entirely is batching — and that bring-up cost repeats per restart, unlike kv-cache which pays rent while the model stays loaded.
vLLM or llama.cpp for GGUF — which wins?
Same file, opposite instincts: llama.cpp treats GGUF as the format, vLLM treats it as an option. One GPU serving several people → vLLM plus the plugin. One person, any hardware, widest quant choice → llama.cpp, one binary. The head-to-head table — CPU, concurrency, quant coverage, setup — is in the deep dive.
Why does my GGUF file fail in vLLM?
Three real causes, all in the table above: wrong serve syntax (plain path instead of repo:quant_type), a missing or wrong --tokenizer, or hardware outside the supported table. A fourth shows up rarer: a quant scheme the plugin doesn't cover. Work the table top to bottom — the first three rows cover most reports.
Does Ollama run GGUF?
Yes — natively; GGUF is Ollama's default format, not an add-on. Pull a GGUF from Hugging Face, point Ollama at it, done. If the GPU sits idle while it runs, that's a different problem with its own checklist: Ollama Not Using GPU? Fix It.
Can several users share one vLLM GGUF server?
That's the one job vLLM exists for: continuous batching queues concurrent requests against one loaded model. The GGUF plugin inherits that machinery — the honest caveat is the docs' own "under-optimized" line, so if you're sizing a shared server, measure your quant on your card before promising anyone a number. Everything this cluster assumes — the format, the engines, the hardware — is mapped in the run-LLMs-locally hub.
One more trap: it loads, then OOMs under load
A GGUF that serves fine for one user can still blow past VRAM when batching multiplies kv-cache per concurrent request. The weights shrank; the cache grows with each user. Drop concurrency, shorten context, or step the quant down — the sizing logic is the same as any vLLM deployment.
The ten answers above trace to two sources: the official vLLM documentation and our own serve logs. What nobody has published yet is a benchmark — Q4_K_M through the plugin versus the same file in llama.cpp on one card, tok/s side by side. When that bench runs, it lands in the local-LLM hub first.
$ Articles similaires · local llm
Ce que les modèles de décision ne savent pas faire : six limites honnêtes
Deux des six nouveaux modèles de décision locaux sont non commerciaux, aucun ne s'explique, et le chiffre de confiance est une affirmation tant que personne ne l'a pas banché. La vague, lue contre le courant.
2026-10-08 · 4 min de lecture

Can vLLM Run GGUF? Yes — on GPU Only
Can vLLM run GGUF? Yes — via the official plugin, on GPU only. The serve syntax, the tokenizer trap, the hardware limits, and when llama.cpp still wins.
2026-09-23 · 5 min de lecture

llama.cpp vs Ollama — which one should you run?
The record read, not run: Ollama pins its llama.cpp engine at b11351 while upstream ships b11443. One is an appliance, one is the engine room.
2026-10-06 · 5 min de lecture

GGUF VRAM Calculator: Check Before You Download
Model size, quant and context in — weights, KV cache and a per-card verdict out. The GGUF VRAM calculator answers will-it-fit before the download starts.
2026-10-02 · 4 min de lecture

$ Recevez le prochain how-to par e-mail
Un e-mail par article. Réparez et avancez.