Ollama: the shortest way to a first model
Ollama fetches a model by name and serves it, with nothing to configure but the context.
What we ran (10 October 2026): Ollama 0.40.2 from Homebrew on a Mac with an M3 Max and 128 GB, serving
qwen3.5:9b from its library (7.6 GB) with a context of 65,536 tokens, with Otōto 2026.19. ototo init found it
by itself, and its tool-call test passed, in 9 seconds while the model loaded and 1.4 once it had. Two questions
about ripgrep’s source were answered rightly, in 86 and 124 seconds. The first question we ever asked it took far
longer, with Ollama unloading the model part of the way through; asked again it took the 86 seconds.
1. Install
brew install ollama
or Ollama’s installer for macOS, Linux and Windows.
2. Fetch a model
ollama pull qwen3.5:9b
The model must be one that calls tools: Ollama’s library marks them, and ollama show qwen3.5:9b lists tools
among its capabilities.
3. Start the server, with enough context
OLLAMA_CONTEXT_LENGTH=65536 ollama serve
The context is the thing to get right. Ollama chooses a default by the machine’s memory: by its
documentation as little as 4,096 tokens on a small machine, which is too short for a question about code. OLLAMA_CONTEXT_LENGTH sets it for the server; if Ollama runs as an app or a service, set it there and
restart it. Check with:
ollama ps
whose CONTEXT column should say 65536 once a model is loaded. It listens on 127.0.0.1:11434.
4. Point Otōto at it
ototo init
finds Ollama on port 11434 and uses the model it serves. With several models pulled, name one:
ototo init --base-url http://127.0.0.1:11434/v1 --model qwen3.5:9b.
5. Check
ototo doctor
It should say serves the model and calls tools. It says nothing of the context, since Ollama’s API does not
report it: ollama ps is the check for that.