Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Ollama: the shortest way to a first model

Ollama fetches a model by name and serves it, with nothing to configure but the context.

What we ran (10 October 2026): Ollama 0.40.2 from Homebrew on a Mac with an M3 Max and 128 GB, serving qwen3.5:9b from its library (7.6 GB) with a context of 65,536 tokens, with Otōto 2026.19. ototo init found it by itself, and its tool-call test passed, in 9 seconds while the model loaded and 1.4 once it had. Two questions about ripgrep’s source were answered rightly, in 86 and 124 seconds. The first question we ever asked it took far longer, with Ollama unloading the model part of the way through; asked again it took the 86 seconds.

1. Install

brew install ollama

or Ollama’s installer for macOS, Linux and Windows.

2. Fetch a model

ollama pull qwen3.5:9b

The model must be one that calls tools: Ollama’s library marks them, and ollama show qwen3.5:9b lists tools among its capabilities.

3. Start the server, with enough context

OLLAMA_CONTEXT_LENGTH=65536 ollama serve

The context is the thing to get right. Ollama chooses a default by the machine’s memory: by its documentation as little as 4,096 tokens on a small machine, which is too short for a question about code. OLLAMA_CONTEXT_LENGTH sets it for the server; if Ollama runs as an app or a service, set it there and restart it. Check with:

ollama ps

whose CONTEXT column should say 65536 once a model is loaded. It listens on 127.0.0.1:11434.

4. Point Otōto at it

ototo init

finds Ollama on port 11434 and uses the model it serves. With several models pulled, name one: ototo init --base-url http://127.0.0.1:11434/v1 --model qwen3.5:9b.

5. Check

ototo doctor

It should say serves the model and calls tools. It says nothing of the context, since Ollama’s API does not report it: ollama ps is the check for that.