> grep -r "GGUF"▋
5 entries
GGUF is the quantized model file format covered across the mrsaynothing wiki-in-blog: how to run it locally, which quant level to pick, whether vLLM can serve it, and how much VRAM it needs.[source][source][source][source]
Running GGUFs locally covers one-line Ollama pulls and llama.cpp loading straight from a Hugging Face URL.[source]
Quantization levels like Q4_K_M vs Q8_0 trade file size and VRAM against quality, with Q4_K_M as the default and step-ups only for code and math.[source]
vLLM serves GGUF via the official plugin on GPU only, with serve syntax, a tokenizer trap, and hardware limits documented.[source]
A VRAM calculator takes model size, quant and context and returns weights, KV cache and a per-card will-it-fit verdict before you download.[source]
see also: AI · local LLM · Ollama
auto-tended from 5 posts · every sentence cites its source · mortal on purpose
2026-10-06
llama.cpp vs Ollama — which one should you run?
The record read, not run: Ollama pins its llama.cpp engine at b11351 while upstream ships b11443. One is an appliance, one is the engine room.
2026-10-02
GGUF VRAM Calculator: Check Before You Download
Model size, quant and context in — weights, KV cache and a per-card verdict out. The GGUF VRAM calculator answers will-it-fit before the download starts.
2026-09-23
Can vLLM Run GGUF? Yes — on GPU Only
Can vLLM run GGUF? Yes — via the official plugin, on GPU only. The serve syntax, the tokenizer trap, the hardware limits, and when llama.cpp still wins.
2026-09-13
GGUF Quantization Levels: Q4_K_M vs Q8_0, Size and VRAM
GGUF quantization levels compared: Q4_K_M vs Q8_0 on size, VRAM and quality, plus the one rule — default Q4_K_M, step up only for code and math.
2026-09-07
How to Run GGUF Models Locally: Ollama, llama.cpp & vLLM
How to run GGUF models locally: one-line Ollama pulls, llama.cpp straight off a Hugging Face URL, and how to pick the right quant for your VRAM.