Otōto’s documentation
Otōto (弟, “little brother”) is an MCP server for coding agents. Your agent hands it a question about your code; a
small model on a server you choose reads the code and brings back a short answer, with the lines that back it
checked. Tools that need no model (search, read, outline, changes, history, replace_all) answer at
once.
curl -fsSL https://ototo.sh | sh
- Quick start: install it, give it a small model or none, and ask the first question.
- How it works: what happens inside an
ask, with diagrams, how the answer’s citations are checked, and what leaves the machine. - The tools: each one, what it takes and what it is for.
- A model server: quick starts for vLLM, llama.cpp, MLX, Ollama, LM Studio, mistral.rs and LiteRT-LM, each as we ran it, and the models worth using.
- Your coding agent: Claude Code and OpenCode, which
ototo initsets up; Pi; and any other MCP client. - The plugins: the nine that come with it, and how to add and update them. Writing one, for a file format or a service Otōto does not read yet.
- Metrics: what Otōto reports over OpenTelemetry, and what never leaves the machine.
- For organisations: one settings file for a fleet, model servers that are the only ones code may go to, your own plugins and update channel, and a managed rollout.
Every setting, the dashboard, and what we measured of each part are still in the README, and each release’s changes in the changelog. They move here as these pages grow.