Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

llama.cpp: a model server on almost anything

llama-server is llama.cpp’s own server: one program, one model file, on a Mac, on Linux with or without a GPU. LM Studio and Ollama are built on the same engine; this is the bare one.

What we ran (10 October 2026): llama.cpp 0.4.0 from Homebrew on a Mac with an M3 Max and 128 GB, serving Qwen3.5 9B at four bits (a 5.6 GB file), with Otōto 2026.19. ototo init’s tool-call test answered in 1.4 seconds. Two questions about ripgrep’s source were answered rightly, in 16 and 79 seconds. Nine billion parameters is a third of what we measure Otōto with, and two questions are a trial, not a measurement.

1. Install

brew install llama.cpp

on macOS or Linux; llama.cpp’s README has the other ways.

2. Fetch a model

A model is one .gguf file. The one we ran:

mkdir -p ~/models && cd ~/models
curl -L -C - -o Qwen3.5-9B-Q4_K_M.gguf \
  https://huggingface.co/lmstudio-community/Qwen3.5-9B-GGUF/resolve/main/Qwen3.5-9B-Q4_K_M.gguf

-C - carries on where a broken download stopped. Q4_K_M is four bits a weight, the usual balance of size and quality; the file should fit in your machine’s memory with several gigabytes to spare.

3. Start the server

llama-server -m ~/models/Qwen3.5-9B-Q4_K_M.gguf -c 65536 --host 127.0.0.1 --port 8080
  • -c 65536 is the context, shared by the server’s slots. Left out, the model’s own default is used, which for many is too short; at 32,768 ototo doctor warns that a long question outgrows it, and the second question above took 230 seconds and 16 turns where it took 79 and 9 with 65,536.
  • Tool calls need llama.cpp’s template engine (--jinja), which is on unless you turn it off.
  • --port 8080 is where ototo init looks. Say it: a newer llama.cpp may choose another port by default.
  • --host 127.0.0.1 keeps the server to this machine.

4. Point Otōto at it

ototo init

finds a server on port 8080. For another address: ototo init --base-url http://<host>:<port>/v1.

5. Check

ototo doctor

Its Model lines should say serves the model, a context of 65,536 tokens, calls tools, and that the server answers Otōto’s first turn with a tool call.