← home

$ ollama run everything

local LLMs, start to serving — the whole cluster, one hub

Your hardware already runs assistants — the hard part is the matrix of quantization, runtimes and VRAM math nobody explains end-to-end. These field notes cover the whole path: pick a model, quantize it, choose a runtime, keep it on the GPU, and understand where your ram actually went.