> grep -r "Ollama"▋
8 entries
Ollama is a local-LLM runner that wraps llama.cpp for convenience, trading raw speed and control for one-line model pulls and easy API serving.[source][source]
It serves models on port 11434; connection refused there means nothing is listening.[source]
Real-world friction: GPU fallback to CPU (drivers, VRAM, ROCm, Docker flags), RAM ceilings deciding what boots, and per-VRAM quant picks for coding models.[source][source][source]
Compared head-to-head with LM Studio for install, GPU use, speed, and API serving.[source]
see also: local LLM · AI · GGUF · Linux
auto-tended from 8 posts · every sentence cites its source · mortal on purpose
2026-10-06
llama.cpp vs Ollama — which one should you run?
The record read, not run: Ollama pins its llama.cpp engine at b11351 while upstream ships b11443. One is an appliance, one is the engine room.
2026-09-30
Ollama Connection Refused? The 60-Second Triage
Connection refused means nothing is listening on port 11434. The six real causes, ranked, plus the one curl command that tells you which one you have.
2026-09-19
Local-LLM Regrets Are RAM Problems. Nobody Talks About RAM.
A 4.9 GB model reserves 7.0 GB before it answers. VRAM gets the debates; RAM decides what boots, what fits, and what silently falls back to CPU.
2026-09-16
Ollama Not Using GPU? Fix It on Linux, Windows and WSL
Ollama ignoring your GPU and falling back to CPU? The five real causes — drivers, VRAM, pinned backends, ROCm, Docker flags — and the command that fixes each.
2026-09-10
llama.cpp vs Ollama: Which Should You Run in 2026?
llama.cpp vs Ollama: Ollama wraps llama.cpp for convenience, raw llama.cpp wins on speed and control. Benchmarks, GPU offload flags, and when each wins.
2026-09-07
How to Run GGUF Models Locally: Ollama, llama.cpp & vLLM
How to run GGUF models locally: one-line Ollama pulls, llama.cpp straight off a Hugging Face URL, and how to pick the right quant for your VRAM.
2026-09-04
Best Local LLM for Coding: 8GB to 24GB VRAM Picks
The best local LLM for coding by VRAM bracket: Qwen3 Coder vs DeepSeek at 8, 12, 16 and 24 GB, the right quant per card, plus a runnable Ollama setup.
2026-09-01
Ollama vs LM Studio: Which Local LLM Tool Should You Use?
Ollama vs LM Studio compared for real dev work: install, GPU use, speed and API serving — with runnable commands so you can pick the right tool today.