Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Models worth using

A recommendation here is a measurement: a model, on a named server and machine, asked our set of questions about real open-source code, each with a checked answer. Where we have only tried a model, on two questions, it says so. The list is short because it is only what we have run.

ModelDownloadMeasured onRightTime a question
Qwen3.8 27B, four bitsfetched by vLLMvLLM, on a GPU server19 of 204 to 45 seconds
Bonsai 2 27B, two bits8.6 GBMLX (mlx-vlm), on a Mac with an M3 Max and 128 GB20 of 20about four minutes
Qwen3.5 9B, four bits5.6 GBMLX (mlx-lm), on the same Mac17 of 1925 seconds at the median, 8 to 198

The 9B was measured on 10 October 2026 with Otōto 2026.19; the two 27B models earlier, on earlier releases, when the set had twenty questions. One has since been retired, so the set is nineteen now.

Which one

  • A GPU server a team shares: Qwen3.8 27B on vLLM. It is what Otōto’s own benchmarks are run with: vLLM’s page.
  • A laptop or a desktop: Qwen3.5 9B. A third of the size, most of the answers, and quick: on a Mac, MLX or llama.cpp; Ollama and LM Studio serve it too. With a context of 65,536 tokens Ollama reported it using 10 to 15 GB of memory, so a 16 GB machine is tight and 24 GB or more is comfortable. Both answers it got wrong were ones Otōto’s own check flagged (“claims not backed by the cited code”), which is what that check is for; it also flagged three that were right.
  • A Mac with memory to spare, and patience: Bonsai 2 27B. The best score, at minutes a question.
  • Nothing smaller, yet. Gemma 4 E4B (3.4 GB, on LiteRT-LM) calls tools as it should and got both of the two questions we tried wrong.

What a model needs

  • Tool calls, through the server it runs on: the model asks for a search or a read. ototo init and ototo doctor test it.
  • A long context: 32,000 tokens at the least, 64,000 to be comfortable.
  • A size your machine holds: the file, plus the context, in memory.

Any model that does the first two can be tried: ototo init --base-url <server>/v1 --model <name>, then ask it something about your code. Otōto says at the foot of each answer how many turns and how long it took, and flags what the cited code does not back.