A model server for Otōto: quick starts
Otōto’s ask, locate, callers and edit hand a question to a small model that you run. This is how to get one
running, a page a server. Otōto works without one too (search, read, outline, changes, history and
replace_all need no model), so you can install first and come back here.
What Otōto needs of a server
- OpenAI’s chat completions API (
/v1/chat/completions), which every server below speaks. - Tool calls through it: the model asks for a search or a read, Otōto runs it and sends the result back. A
server that answers in prose where a tool call was wanted cannot be used.
ototo initandototo doctortest this with a real request, and say so in a line. - Room to read: 32,000 tokens of context at the least, 64,000 to be comfortable. A question’s conversation is sent again at every turn and grows with what the model reads. Several servers default to far less, and each page says how to raise it.
Which server
| Server | For | Quick start | Tried with Otōto |
|---|---|---|---|
| vLLM | A GPU server a team shares | vllm.md | Yes: what we run and measure with |
| llama.cpp | Almost anything | llama-cpp.md | Yes: a 9B model answered our two trial questions rightly |
| LiteRT-LM | Very small machines | litert-lm.md | Yes: it works, and its models are too small to answer well |
| MLX | A Mac with Apple silicon | mlx.md | Yes: a 9B model answered our two trial questions rightly, the quickest on a Mac |
| Ollama | The shortest way to a first model | ollama.md | Yes: a 9B model answered our two trial questions rightly |
| LM Studio | A desktop, with an app to manage models | lm-studio.md | Yes: a 9B model answered our two trial questions rightly |
| mistral.rs | A Mac, or Linux with or without a GPU | mistral-rs.md | Yes: right answers from a 9B model, and slow ones |
Any other server that speaks the same API with tool calls should work: ototo init --base-url <its address>/v1
tests it and tells you. Which model to serve: Models worth using.
The shape of every quick start
- Install the server.
- Fetch a model that fits the machine’s memory.
- Start the server, with tool calls on and enough context.
- Point Otōto at it:
ototo initfinds a server on this machine’s usual ports (8000, 8080, 8081, 11434, 1234); for any other address,ototo init --base-url http://<host>:<port>/v1. It tests the server, says what it will change in your agent’s settings, and asks. - Check:
ototo doctor. Its Model lines say the server serves the model, how much context it gives it, and that it calls tools. Then start a new session of your agent and ask it about your code.
Each page says what we ran it with: the server’s version, the model, the machine, and how long a question took. We write a page only for what we have run.
If the code must not leave the machine
It does not: Otōto sends what the model reads to the server you name and nowhere else. A server on another machine
should be reached over HTTPS; ototo doctor warns when it is plain HTTP. A server on your own machine should
listen on 127.0.0.1 only, and some listen on every network interface unless told: the pages say which.