mrsaynothing.dev

> grep -r "GGUF"▋

5 entries

GGUF is the quantized model file format covered across the mrsaynothing wiki-in-blog: how to run it locally, which quant level to pick, whether vLLM can serve it, and how much VRAM it needs.[source][source][source][source]

Running GGUFs locally covers one-line Ollama pulls and llama.cpp loading straight from a Hugging Face URL.[source]

Quantization levels like Q4_K_M vs Q8_0 trade file size and VRAM against quality, with Q4_K_M as the default and step-ups only for code and math.[source]

vLLM serves GGUF via the official plugin on GPU only, with serve syntax, a tokenizer trap, and hardware limits documented.[source]

A VRAM calculator takes model size, quant and context and returns weights, KV cache and a per-card will-it-fit verdict before you download.[source]

see also: AI · local LLM · Ollama

auto-tended from 5 posts · every sentence cites its source · mortal on purpose