mrsaynothing.dev

> grep -r "Ollama"▋

8 entries

Ollama is a local-LLM runner that wraps llama.cpp for convenience, trading raw speed and control for one-line model pulls and easy API serving.[source][source]

It serves models on port 11434; connection refused there means nothing is listening.[source]

Real-world friction: GPU fallback to CPU (drivers, VRAM, ROCm, Docker flags), RAM ceilings deciding what boots, and per-VRAM quant picks for coding models.[source][source][source]

Compared head-to-head with LM Studio for install, GPU use, speed, and API serving.[source]

see also: local LLM · AI · GGUF · Linux

auto-tended from 8 posts · every sentence cites its source · mortal on purpose

2026-10-06

llama.cpp vs Ollama — which one should you run?

The record read, not run: Ollama pins its llama.cpp engine at b11351 while upstream ships b11443. One is an appliance, one is the engine room.

2026-09-30

Ollama Connection Refused? The 60-Second Triage

Connection refused means nothing is listening on port 11434. The six real causes, ranked, plus the one curl command that tells you which one you have.

2026-09-19

Local-LLM Regrets Are RAM Problems. Nobody Talks About RAM.

A 4.9 GB model reserves 7.0 GB before it answers. VRAM gets the debates; RAM decides what boots, what fits, and what silently falls back to CPU.

2026-09-16

Ollama Not Using GPU? Fix It on Linux, Windows and WSL

Ollama ignoring your GPU and falling back to CPU? The five real causes — drivers, VRAM, pinned backends, ROCm, Docker flags — and the command that fixes each.

2026-09-10

llama.cpp vs Ollama: Which Should You Run in 2026?

llama.cpp vs Ollama: Ollama wraps llama.cpp for convenience, raw llama.cpp wins on speed and control. Benchmarks, GPU offload flags, and when each wins.

2026-09-07

How to Run GGUF Models Locally: Ollama, llama.cpp & vLLM

How to run GGUF models locally: one-line Ollama pulls, llama.cpp straight off a Hugging Face URL, and how to pick the right quant for your VRAM.

2026-09-04

Best Local LLM for Coding: 8GB to 24GB VRAM Picks

The best local LLM for coding by VRAM bracket: Qwen3 Coder vs DeepSeek at 8, 12, 16 and 24 GB, the right quant per card, plus a runnable Ollama setup.

2026-09-01

Ollama vs LM Studio: Which Local LLM Tool Should You Use?

Ollama vs LM Studio compared for real dev work: install, GPU use, speed and API serving — with runnable commands so you can pick the right tool today.