Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

A model server for Otōto: quick starts

Otōto’s ask, locate, callers and edit hand a question to a small model that you run. This is how to get one running, a page a server. Otōto works without one too (search, read, outline, changes, history and replace_all need no model), so you can install first and come back here.

What Otōto needs of a server

  • OpenAI’s chat completions API (/v1/chat/completions), which every server below speaks.
  • Tool calls through it: the model asks for a search or a read, Otōto runs it and sends the result back. A server that answers in prose where a tool call was wanted cannot be used. ototo init and ototo doctor test this with a real request, and say so in a line.
  • Room to read: 32,000 tokens of context at the least, 64,000 to be comfortable. A question’s conversation is sent again at every turn and grows with what the model reads. Several servers default to far less, and each page says how to raise it.

Which server

ServerForQuick startTried with Otōto
vLLMA GPU server a team sharesvllm.mdYes: what we run and measure with
llama.cppAlmost anythingllama-cpp.mdYes: a 9B model answered our two trial questions rightly
LiteRT-LMVery small machineslitert-lm.mdYes: it works, and its models are too small to answer well
MLXA Mac with Apple siliconmlx.mdYes: a 9B model answered our two trial questions rightly, the quickest on a Mac
OllamaThe shortest way to a first modelollama.mdYes: a 9B model answered our two trial questions rightly
LM StudioA desktop, with an app to manage modelslm-studio.mdYes: a 9B model answered our two trial questions rightly
mistral.rsA Mac, or Linux with or without a GPUmistral-rs.mdYes: right answers from a 9B model, and slow ones

Any other server that speaks the same API with tool calls should work: ototo init --base-url <its address>/v1 tests it and tells you. Which model to serve: Models worth using.

The shape of every quick start

  1. Install the server.
  2. Fetch a model that fits the machine’s memory.
  3. Start the server, with tool calls on and enough context.
  4. Point Otōto at it: ototo init finds a server on this machine’s usual ports (8000, 8080, 8081, 11434, 1234); for any other address, ototo init --base-url http://<host>:<port>/v1. It tests the server, says what it will change in your agent’s settings, and asks.
  5. Check: ototo doctor. Its Model lines say the server serves the model, how much context it gives it, and that it calls tools. Then start a new session of your agent and ask it about your code.

Each page says what we ran it with: the server’s version, the model, the machine, and how long a question took. We write a page only for what we have run.

If the code must not leave the machine

It does not: Otōto sends what the model reads to the server you name and nowhere else. A server on another machine should be reached over HTTPS; ototo doctor warns when it is plain HTTP. A server on your own machine should listen on 127.0.0.1 only, and some listen on every network interface unless told: the pages say which.