Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Otōto’s documentation

Otōto (弟, “little brother”) is an MCP server for coding agents. Your agent hands it a question about your code; a small model on a server you choose reads the code and brings back a short answer, with the lines that back it checked. Tools that need no model (search, read, outline, changes, history, replace_all) answer at once.

curl -fsSL https://ototo.sh | sh
  • Quick start: install it, give it a small model or none, and ask the first question.
  • How it works: what happens inside an ask, with diagrams, how the answer’s citations are checked, and what leaves the machine.
  • The tools: each one, what it takes and what it is for.
  • A model server: quick starts for vLLM, llama.cpp, MLX, Ollama, LM Studio, mistral.rs and LiteRT-LM, each as we ran it, and the models worth using.
  • Your coding agent: Claude Code and OpenCode, which ototo init sets up; Pi; and any other MCP client.
  • The plugins: the nine that come with it, and how to add and update them. Writing one, for a file format or a service Otōto does not read yet.
  • Metrics: what Otōto reports over OpenTelemetry, and what never leaves the machine.
  • For organisations: one settings file for a fleet, model servers that are the only ones code may go to, your own plugins and update channel, and a managed rollout.

Every setting, the dashboard, and what we measured of each part are still in the README, and each release’s changes in the changelog. They move here as these pages grow.