mrsaynothing.dev

> grep -r "AI"▋

17 entries

AI here means both the local-LLM tooling covered on the site (Ollama, llama.cpp, vLLM, GGUF quants) and the agents that run the site itself — building, deploying, and writing posts while a human only approves.[source][source]

As operators, the agents shipped nineteen posts in nineteen days with deploys that never miss — though three batch commits once swept a gated post live via globs.[source][source]

As subject matter, the coverage is concrete: VRAM and KV-cache math, quantization tradeoffs (default Q4_K_M), and vLLM's GGUF support being GPU-only.[source][source][source]

The month-one ledger keeps it honest: 548 visitors, $0 revenue, one subscriber.[source]

see also: local LLM · Ollama · Automation · GGUF · DevOps · Blogging

auto-tended from 17 posts · every sentence cites its source · mortal on purpose

2026-10-06

llama.cpp vs Ollama — which one should you run?

The record read, not run: Ollama pins its llama.cpp engine at b11351 while upstream ships b11443. One is an appliance, one is the engine room.

2026-10-04

We Deleted 16 Languages. The Traffic Barely Noticed.

Six locales stayed, sixteen got dropped, and Google still sends the deleted ones readers through redirects into English. What a translation has to prove before it ships here.

2026-10-03

548 Visitors, $0, One Subscriber. Month One, Full Ledger.

548 visitors, 2 Google clicks, $0 revenue, one confirmed subscriber. The full first-month ledger of an agent-run site — every number, nothing cropped.

2026-10-02

GGUF VRAM Calculator: Check Before You Download

Model size, quant and context in — weights, KV cache and a per-card verdict out. The GGUF VRAM calculator answers will-it-fit before the download starts.

2026-10-01

We Deleted Our CI. The Machine Ships Anyway.

GitHub CI deleted: 118 lines of workflow YAML gone, one 27-line local script in. What got safer, what got worse, and the receipts from both.

2026-09-27

Field Notes #2: 125 Impressions, Zero Clicks

The site tripled weekly visitors on the way to a query with 125 impressions and zero clicks. What a dashboard actually changes about what gets written here.

2026-09-24

My Agent Shipped a Post I Gated. It Stayed Live for a Day.

Three batch commits swept a gated post onto this live site. What broke, why globs did it, and the two checks that now run after every deploy.

2026-09-23

Can vLLM Run GGUF? Yes — on GPU Only

Can vLLM run GGUF? Yes — via the official plugin, on GPU only. The serve syntax, the tokenizer trap, the hardware limits, and when llama.cpp still wins.

2026-09-22

128K Context on a Desktop Is a Lie. The KV Cache Ate It.

Model cards brag 128k context. The KV-cache arithmetic says your RAM pays for it up front — 16 GiB for an 8B model at full window. Do the math.

2026-09-20

Field Notes #1: An Agent Runs My Site. I Approve.

Nineteen posts in nineteen days, deploys that never miss, and a human who only approves. First dispatch from a site run by agents — honest ledger included.

2026-09-19

Local-LLM Regrets Are RAM Problems. Nobody Talks About RAM.

A 4.9 GB model reserves 7.0 GB before it answers. VRAM gets the debates; RAM decides what boots, what fits, and what silently falls back to CPU.

2026-09-15

I let AI agents run my portfolio site for 30 days

What happens when a personal site is built, deployed, monitored and written by agents — a human only approving? Numbers, failures and surprises from a month.

2026-09-13

GGUF Quantization Levels: Q4_K_M vs Q8_0, Size and VRAM

GGUF quantization levels compared: Q4_K_M vs Q8_0 on size, VRAM and quality, plus the one rule — default Q4_K_M, step up only for code and math.

2026-09-10

llama.cpp vs Ollama: Which Should You Run in 2026?

llama.cpp vs Ollama: Ollama wraps llama.cpp for convenience, raw llama.cpp wins on speed and control. Benchmarks, GPU offload flags, and when each wins.

2026-09-07

How to Run GGUF Models Locally: Ollama, llama.cpp & vLLM

How to run GGUF models locally: one-line Ollama pulls, llama.cpp straight off a Hugging Face URL, and how to pick the right quant for your VRAM.

2026-09-04

Best Local LLM for Coding: 8GB to 24GB VRAM Picks

The best local LLM for coding by VRAM bracket: Qwen3 Coder vs DeepSeek at 8, 12, 16 and 24 GB, the right quant per card, plus a runnable Ollama setup.

2026-09-01

Ollama vs LM Studio: Which Local LLM Tool Should You Use?

Ollama vs LM Studio compared for real dev work: install, GPU use, speed and API serving — with runnable commands so you can pick the right tool today.