Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Otōto’s documentation

Otōto (弟, “little brother”) is an MCP server for coding agents. Your agent hands it a question about your code; a small model on a server you choose reads the code and brings back a short answer, with the lines that back it checked. Tools that need no model (search, read, outline, changes, history, replace_all) answer at once.

curl -fsSL https://ototo.sh | sh
  • Quick start: install it, give it a small model or none, and ask the first question.
  • How it works: what happens inside an ask, with diagrams, how the answer’s citations are checked, and what leaves the machine.
  • The tools: each one, what it takes and what it is for.
  • A model server: quick starts for vLLM, llama.cpp, MLX, Ollama, LM Studio, mistral.rs and LiteRT-LM, each as we ran it, and the models worth using.
  • Your coding agent: Claude Code and OpenCode, which ototo init sets up; Pi; and any other MCP client.
  • The plugins: the nine that come with it, and how to add and update them. Writing one, for a file format or a service Otōto does not read yet.
  • Settings: every one, with its default: the model servers, how far a question may go, what your agent is offered.
  • The dashboard: ototo ui, the live tail, the run log behind them, and a report to send.
  • Metrics: what Otōto reports over OpenTelemetry, and what never leaves the machine.
  • For organisations: one settings file for a fleet, model servers that are the only ones code may go to, your own plugins and update channel, and a managed rollout.
  • The hub: a collector, Prometheus and Grafana to run, for usage and savings by person and team and the fleet’s versions, as a compose file, a Helm chart or kustomize manifests.

What we measured of each part (the claim checks, the ranked search, each plugin) is in the README, with what someone building from the source needs; each release’s changes are in the changelog.

Quick start

Install Otōto, give it a small model if you have one, and ask your agent a question about your code.

1. Install

curl -fsSL https://ototo.sh | sh

On macOS with Apple silicon, or Linux on x86_64 or arm64. It fetches the newest release, checks it against Otōto’s release key (which is written into the script), puts ototo in ~/.local/bin with its plugins, and runs ototo init, the next step. To read the script first: ototo.dev/installer. To check and download and then stop: curl -fsSL https://ototo.sh | sh -s -- --dry-run.

You need a coding agent. ototo init sets up Claude Code (2.1 or later) and OpenCode; Pi and any other MCP client take one command by hand.

2. A small model, or none yet

ototo init looks for a model server on this machine, on the usual ports of vLLM, llama.cpp, MLX, Ollama and LM Studio. It sends the model one request with a tool, to check it can call tools, then lists what it will change and asks before it writes.

You haveDo
A model server on this machinenothing: ototo init finds it
One elsewhere (a team’s GPU server)ototo init --base-url http://<host>:<port>/v1
None yetcarry on: Otōto sets up the tools that need no model (search, read, outline, changes, history, replace_all). Run ototo init again once a server runs, and ask, locate, callers and edit join them

A model server: quick starts has one for each server we have run, and the models worth using: one for a team’s GPU server, and one a laptop holds.

3. Check

ototo doctor

checks each part and says what to do about anything wrong: the settings, each model server (reachable, serving the model, calling tools), your agent’s setup, and whether the plugins load.

4. Ask

Start a new session of your agent in a repository (one that is already running keeps the tools it started with), and ask about the code as you would anyway:

How are cached tiles evicted, and what decides the limit?

The agent hands the question to ask and answers from what comes back: a short explanation, with path:line citations that Otōto has checked against the files. For what it already knows it wants, it calls read, outline and search in place of opening whole files.

Every tool also runs from a shell, which is the quickest way to see what your agent sees:

ototo ask "which class validates a new pet?"
ototo outline src
ototo read src/cache/Store.kt#evict

5. Watch it work

ototo tail

follows each delegated call as it runs: the question, every turn of the small model with the tool it called, the claim check, and how it ended. ototo ui shows the same as a dashboard, at http://127.0.0.1:7777.

Later

  • Updating. ototo update fetches the newest release, checks it against the release key and installs it; ototo plugins update does the same for the plugins. Neither runs by itself.
  • Switching models. ototo init --base-url <url> again; your other settings are kept.
  • Removing it. The installer ends by printing the command that undoes everything it did (sh …/uninstall.sh). Each file it changed was backed up first, as <file>.before-ototo-<time>.

Where next

How it works

A coding agent finds its way round a codebase by opening files and running searches, and every file it opens, every result it scans and every dead end it follows stays in its context for the rest of the session. Each later turn carries all of it again, and has more to sift.

Otōto keeps that exploring out of the agent’s context. The agent hands over a whole question; a small model, on a server you choose, searches and reads in a context of its own, which is thrown away after the call; and the agent gets back a short answer, with the lines that back it checked against the files. For what the agent already knows it wants to see, a set of exact tools answers at once, with no model at all.

The parts

  • Your coding agent starts ototo serve in the repository it is working in, and talks to it over MCP (the Model Context Protocol, on standard input and output). Claude Code, OpenCode, Pi and others.
  • Otōto is one static binary, written in Rust. It reads the repository, changes it only when an edit is asked for, and ototo serve opens no port.
  • A model server runs the small model: vLLM, llama.cpp, MLX, Ollama, LM Studio and others, over the OpenAI-compatible API with tool calls. It can be on the same machine or a team’s GPU server. With none, the direct tools still work (quick starts).
  • Plugins read kinds of file Otōto does not know itself, each as its own tool would evaluate it (the plugins).

Two kinds of tool

DirectDelegated
Toolsread, outline, search, changes, history, replace_allask, locate, callers, edit
Answered byOtōto, from the files and from gitthe small model, exploring with Otōto’s tools
Takesmillisecondsseconds to minutes
Forwhat the agent already knows it wants: these lines, this declaration, this stringan open question: how something works, where it is decided, who uses it

The tools has each of them. The delegated ones are where the saving is; the direct ones are what the agent uses in place of grep, cat and whole-file reads, and they return less: a declaration by name, not the file it is in.

One ask, step by step

  1. The agent calls ask with a whole question: “how is auth wired into the router?”. Up to five independent questions go in one call, and are answered side by side.
  2. Otōto opens a fresh context with the small model: its own prompt, the question, and read-only tools. Nothing of the agent’s conversation goes with it.
  3. The small model works in turns. It calls find (the repository’s declarations, ranked against a few words, for when it does not know the name), search (exact text or a regex), outline and read. Otōto runs each call on your files, through the same checks as for the agent: nothing outside the repository, and never a secrets file. A read gives it at most 250 lines, since its context is the one that fills. The turns end at max_turns, or when the question has sent the model its budget of input tokens (400,000 unless set), after which the next turn is its last; an answer cut short that way says so.
  4. It answers, and each path:line citation in the answer carries a claim: what that line shows.
  5. Otōto checks every claim against the cited code, with no model. The identifiers and the values a claim names must be in the cited line’s enclosing declaration; a file it names must be the cited file; “X is defined here” must cite a declaration; a count of callers must match the call sites cited. A claim that fails is listed under “Not backed by the cited code (check before relying on it)”, with where the missing name does appear. Passing means “about the right code”, not “true”: with verify = true, the claims go back to the small model once more, in a fresh request with the code each cites and none of the reasoning, to catch a claim about the right code that says the wrong thing.
  6. The agent gets the answer: a few thousand characters, the citations checked, and the code of the main ones attached (the enclosing function, or the line and those around it), so that it need not read them again. The small model’s context is thrown away.

locate and callers are the same loop with a narrower question and a stricter check: a locate answer must cite a declaration, and a callers answer’s count must match the call sites it cites, each checked to be that symbol and not a namesake.

edit hands over a small, mechanical change. The small model proposes it and Otōto returns a diff; nothing is written unless the call says apply. With a check_cmd in the settings (cargo check, npx tsc --noEmit), the model can try its change: the pending edits are written, the command runs, and every file is put back afterwards.

Reading code as it is declared

Otōto parses source (tree-sitter does most of it), so a file is its declarations, not only its lines: outline lists them with their line ranges, read path#Name returns one, read path:137 the function round a line, and changes says which declarations a diff touched. It reads Java, Kotlin, TypeScript and JavaScript, Rust, Python, Go, C#, shell and SQL this way, and outlines Terraform, YAML, JSON, TOML, XML, properties, CSV, Dockerfiles, Makefiles and Markdown by their own names: a key, a section, a target, a column.

The index for find is built in memory the first time it is used, and a file is parsed again only when it changes. Nothing is written into the repository, and there is no daemon and no database.

Plugins

Some files mean more than their text. A GitLab CI job is its extends chain, its includes and its defaults merged; a Maven dependency’s version may be set three parents away. A plugin reads one such kind of file and gives Otōto its outline and its reads, so that read .gitlab-ci.yml#deploy:prod returns the job as GitLab runs it, each line with the file it came from.

A plugin is a WebAssembly component, run in a sandbox with no network, no filesystem and no environment. It reaches the repository only by asking Otōto, through the same path checks as everything else; each call has an instruction budget and a memory limit; and it loads only if it is signed by a key you trust. One that fails is skipped, and Otōto answers as it would without it. The plugins lists them, and Writing a plugin builds one from nothing.

What leaves the machine

  • To the model servers you configure, and nowhere else: the question, and the parts of the repository the small model reads to answer it. If those servers are yours, the code stays with you; if you list a hosted model as a fallback, it goes there when the others do not answer, and the reply says so.
  • Counts, if you set a collector: with otlp_endpoint, how many calls, tokens and seconds, to a collector of yours (Metrics). Never the questions, the code or the answers.
  • A forge, if you grant a forge plugin: read-only requests to the hosts it declared, for the forge tool’s questions about a merge request or a pipeline. Nothing until the grant.
  • The download channel, when you run ototo update or one of ototo plugins available, add and update: never by itself.
  • Nothing to us. Otōto sends no usage data, crash reports or identifiers to its makers.

SECURITY.md has all of it: what Otōto reads, writes and runs, and what it never reads (environment files, private keys, credentials).

When no model server answers

A delegated call fails at once and says which servers were tried and which tools need none, so the agent goes on with search, read and outline. With several servers listed, each request goes to the first that answers, and one that failed waits at the back of the queue for a minute. ototo doctor checks each of them: reachable, serving the model, calling tools.

Watching it

ototo tail follows every delegated call as it runs: the question, each turn of the small model with the tool it called, the claim check and how it ended. ototo ui shows the same as a dashboard, call by call. Both read what Otōto logs on your machine (~/.ototo/runs.jsonl and the live trace beside it).

12:39:17  ask  start   Which class validates a new pet, and what does it check?
12:39:22  ask  turn 1  find({"query":"validate new pet"}) · 1.6k tok · 4.7s
12:39:24  ask  turn 2  read({"addresses":["src/main/java/…/PetValidator.java"]}) · 3.8k tok · 7.1s
12:39:40  ask  turn 3  finish({"answer":"The class is PetValidator …"}) · 7.3k tok · 22.9s
12:39:40  ask  end     23.0s · 3 turns · 7.3k tok · ok

The tools

What your agent gets from Otōto over MCP, and what each is for. Every one also runs from a shell, as ototo <tool>, in the current directory.

ToolNeeds the small modelIn a line
askyesA whole question about the code, answered with checked citations
locateyesWhere something is defined, by name or by description
callersyesThe call sites of a symbol, counted and checked not to be a namesake
edityesA small, mechanical change, returned as a diff
readnoCode by address: lines, a declaration, a key, an entry of a JAR
outlinenoThe shape of a file, a directory or an archive, before reading it
searchnoExact text or a regex across the repository
changesnoWhat a checkout, a commit or a range changed, by declaration
historynoThe commits behind some lines
replace_allnoFind-and-replace across the repository, as a diff first
forgenoA merge request, a pipeline, a failed job’s log; with a forge plugin granted

The four that need the small model are offered only when one is set up; the rest work without (Quick start). How it works says what happens inside a delegated call.

ask

Hand over a whole question about what the code does: “how is auth wired into the router?”, “where is the retry count decided, and under what conditions?”. The small model explores and answers briefly, and every path:line it cites is checked against the files, with the code of the main ones attached. Seconds to minutes.

  • Several at once. Up to five independent questions go in one call (questions), and are answered side by side.
  • One thing per question. A long question in several parts is cut off more often than its parts asked apart.
  • It finds and explains; it does not decide. Ask where a value is set and what reads it, not what the fix should be: that is the agent’s to work out from the answer.
  • Read the footer. An answer that lists claims “Not backed by the cited code”, or says it was cut off, is the one to check before relying on it. The rest need not be read again.
ototo ask "which class validates a new pet?" --also "what page size lists owners?"

locate and callers

locate finds where something is defined or implemented, by its name or by a description (“the retry logic in the HTTP client”), and must cite a declaration. callers finds the usages of a function, method, type or field, each checked to be that symbol and not another of the same name, and returns the count with the citations; hint names the defining class or file when the name alone is ambiguous.

edit

Delegates a small, mechanical change: add a parameter and update the callers, fix the imports. It returns a unified diff and writes nothing unless the call says apply. With a check_cmd in the settings (a build or a type check), the small model tries its change against it before answering, and the verdict is in the reply. For a plain rename, replace_all is the tool: exact, instant, and no model.

read

Code by address, many addresses in one call:

AddressReturns
path:120-160those lines
path:137the whole function round that line
path#Name, path#Class.methodone declaration
Name, Class.methodthe declaration, found in the repository
patha small file whole, a large one as its outline
docs/plan.md#Goal, package.json#scripts.builda section of a document, a key of a config file
lib/core.jar!org.demo.Shelfa class in an archive, as its declarations; !META-INF/MANIFEST.MF any other entry

Up to 2,000 lines a call; a longer range is cut with a note saying where to continue. With a plugin for the file’s kind, path#name returns the thing as its tool evaluates it: a CI job with its extends merged, a dependency with its effective version (the plugins).

ototo read write_exceeded_line crates/printer/src/standard.rs:47

outline

The shape of code before reading it. A file gives every declaration with its line range; a directory or a glob (src/**/*Controller.java) each file’s length and top-level declarations; an archive its entries. It is how an agent chooses what to read, in place of ls and opening files.

It reads Java, Kotlin, TypeScript and JavaScript, Rust, Python, Go, C#, shell and SQL as their declarations, and Terraform, YAML, JSON, TOML, XML, properties, CSV, Dockerfiles, Makefiles and Markdown by their own names: a resource, a key, a column, a stage, a target, a heading.

ototo outline src/ui/control

Exact text or a regex across the repository: path:line:text for each matching line, 100 unless limit says otherwise, then the total, and where the rest are. fixed takes the text as it is, word whole words, ignore_case, and glob narrows it (src/**/*.ts, several comma-separated, !dir/** to leave one out). Ignored, binary and secrets files are skipped. An empty search says why: no file matches the glob, or none holds the text.

It is for a known string, identifier or message. For code you can only describe, ask or locate.

changes and history

changes is what the checkout changed since its base (the merge-base with the default branch), committed or not: each file with its added and removed lines and the declarations touched (changed Orders.place (12-48), added Orders.cancel). Or one commit, or a range. history is the commits behind path:10-40, path:25, path#Name or path, newest first, one line each. Between them they stand in for git diff, show, log and blame, and say it in declarations.

replace_all

Find-and-replace across the repository in one step: literal text, or a regex with ${1} groups; word, ignore_case and glob as for search. It returns the count in each file and a unified diff, and writes only with apply.

ototo replace max_columns_preview max_cols_preview --word

forge

The repository’s forge, through a plugin: with no arguments, the checked-out branch’s open merge request and its latest pipeline; change for one by number; pipeline for its jobs by stage; job for a failed job’s log, cut to what failed. It is offered only once a forge plugin is installed and granted (ototo plugins grant gitlab), since it reaches outside the machine.

Where a call works

Every tool takes root, to work in another git worktree of the repository or in a directory added to the session. Absolute paths inside one of those do the same without it.

Beside the tools

  • ototo digest -- <command> runs a build or a test command and prints the failures and the summary, not every line of progress; the exit code is the command’s.
  • ototo find "a few words" is the small model’s own first tool, on the command line: the repository’s declarations ranked against the words, for when the name is not known.
  • Routines (experimental, off unless routines = true): a job described once in a Markdown file, which the small model does with its read-only tools; the agent then gets a routine tool.

A shorter menu

Every tool’s definition is carried in the agent’s context on every turn. tools = "ask,read,search" in ~/.config/ototo/config.toml (or OTOTO_TOOLS) offers only those, and Otōto’s instructions then name only them.

A model server for Otōto: quick starts

Otōto’s ask, locate, callers and edit hand a question to a small model that you run. This is how to get one running, a page a server. Otōto works without one too (search, read, outline, changes, history and replace_all need no model), so you can install first and come back here.

What Otōto needs of a server

  • OpenAI’s chat completions API (/v1/chat/completions), which every server below speaks.
  • Tool calls through it: the model asks for a search or a read, Otōto runs it and sends the result back. A server that answers in prose where a tool call was wanted cannot be used. ototo init and ototo doctor test this with a real request, and say so in a line.
  • Room to read: 32,000 tokens of context at the least, 64,000 to be comfortable. A question’s conversation is sent again at every turn and grows with what the model reads. Several servers default to far less, and each page says how to raise it.

Which server

ServerForQuick startTried with Otōto
vLLMA GPU server a team sharesvllm.mdYes: what we run and measure with
llama.cppAlmost anythingllama-cpp.mdYes: a 9B model answered our two trial questions rightly
LiteRT-LMVery small machineslitert-lm.mdYes: it works, and its models are too small to answer well
MLXA Mac with Apple siliconmlx.mdYes: a 9B model answered our two trial questions rightly, the quickest on a Mac
OllamaThe shortest way to a first modelollama.mdYes: a 9B model answered our two trial questions rightly
LM StudioA desktop, with an app to manage modelslm-studio.mdYes: a 9B model answered our two trial questions rightly
mistral.rsA Mac, or Linux with or without a GPUmistral-rs.mdYes: right answers from a 9B model, and slow ones

Any other server that speaks the same API with tool calls should work: ototo init --base-url <its address>/v1 tests it and tells you. Which model to serve: Models worth using.

The shape of every quick start

  1. Install the server.
  2. Fetch a model that fits the machine’s memory.
  3. Start the server, with tool calls on and enough context.
  4. Point Otōto at it: ototo init finds a server on this machine’s usual ports (8000, 8080, 8081, 11434, 1234); for any other address, ototo init --base-url http://<host>:<port>/v1. It tests the server, says what it will change in your agent’s settings, and asks.
  5. Check: ototo doctor. Its Model lines say the server serves the model, how much context it gives it, and that it calls tools. Then start a new session of your agent and ask it about your code.

Each page says what we ran it with: the server’s version, the model, the machine, and how long a question took. We write a page only for what we have run.

If the code must not leave the machine

It does not: Otōto sends what the model reads to the server you name and nowhere else. A server on another machine should be reached over HTTPS; ototo doctor warns when it is plain HTTP. A server on your own machine should listen on 127.0.0.1 only, and some listen on every network interface unless told: the pages say which.

vLLM: a model server a team shares

vLLM serves one model to many people from a machine with a GPU. It is what we run Otōto’s small model on, and what its measurements are made with. Use it when a team shares a server; for one laptop, another page is shorter.

What we ran (10 October 2026): Qwen3.8 27B at four bits on vLLM, on a small server without a datacentre GPU, with Otōto 2026.19. ototo init’s tool-call test answered in 1.9 seconds; the server gave the model a context of 128,000 tokens; two questions about ripgrep’s source were answered rightly in 46 and 61 seconds. On our twenty-question suite this setup answers 19 rightly, in 4 to 45 seconds a question.

1. Install

As vLLM’s own quick start says: a Python package on Linux, or its container image. We run the image, pinned to a version and not latest: a restart otherwise changes vLLM under you.

2. and 3. Fetch a model and start the server

vLLM fetches the model from Hugging Face when it starts:

vllm serve Qwen/Qwen3.8-27B-FP8 \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen3 \
  --max-model-len 65536 --enable-prefix-caching
  • --enable-auto-tool-choice --tool-call-parser qwen3_coder are what make tool calls work: the parser turns the model’s own way of writing a tool call into the API’s tool_calls. Without them every answer is prose. Another model wants another parser: vLLM’s list.
  • --max-model-len 65536: at least 64,000 tokens of context.
  • --enable-prefix-caching: each turn’s prompt starts with the last one’s, so only the new part is computed.

It listens on port 8000. dist/vllm/README.md has more of what we learned running it: a chat template for coding clients, with a script that checks it; a Kubernetes example; and what to do when the server’s prompt cache goes bad.

4. Point Otōto at it

On each developer’s machine:

ototo init --base-url https://<your server>/v1

Put TLS in front of the server (a reverse proxy such as Caddy or nginx) and give its https:// address: over plain HTTP the code the model reads can be read on the network between, and ototo doctor says so.

For a whole team, ototo managed writes the settings once for every user, and can make the organisation’s servers the only ones code may go to: see For organisations.

5. Check

ototo doctor

The Model lines should say serves the model, a context of 64,000 tokens or more, calls tools, and that the server answers Otōto’s first turn with a tool call from its prompt cache and fresh. The last is the check for a prompt cache gone bad.

llama.cpp: a model server on almost anything

llama-server is llama.cpp’s own server: one program, one model file, on a Mac, on Linux with or without a GPU. LM Studio and Ollama are built on the same engine; this is the bare one.

What we ran (10 October 2026): llama.cpp 0.4.0 from Homebrew on a Mac with an M3 Max and 128 GB, serving Qwen3.5 9B at four bits (a 5.6 GB file), with Otōto 2026.19. ototo init’s tool-call test answered in 1.4 seconds. Two questions about ripgrep’s source were answered rightly, in 16 and 79 seconds. Nine billion parameters is a third of what we measure Otōto with, and two questions are a trial, not a measurement.

1. Install

brew install llama.cpp

on macOS or Linux; llama.cpp’s README has the other ways.

2. Fetch a model

A model is one .gguf file. The one we ran:

mkdir -p ~/models && cd ~/models
curl -L -C - -o Qwen3.5-9B-Q4_K_M.gguf \
  https://huggingface.co/lmstudio-community/Qwen3.5-9B-GGUF/resolve/main/Qwen3.5-9B-Q4_K_M.gguf

-C - carries on where a broken download stopped. Q4_K_M is four bits a weight, the usual balance of size and quality; the file should fit in your machine’s memory with several gigabytes to spare.

3. Start the server

llama-server -m ~/models/Qwen3.5-9B-Q4_K_M.gguf -c 65536 --host 127.0.0.1 --port 8080
  • -c 65536 is the context, shared by the server’s slots. Left out, the model’s own default is used, which for many is too short; at 32,768 ototo doctor warns that a long question outgrows it, and the second question above took 230 seconds and 16 turns where it took 79 and 9 with 65,536.
  • Tool calls need llama.cpp’s template engine (--jinja), which is on unless you turn it off.
  • --port 8080 is where ototo init looks. Say it: a newer llama.cpp may choose another port by default.
  • --host 127.0.0.1 keeps the server to this machine.

4. Point Otōto at it

ototo init

finds a server on port 8080. For another address: ototo init --base-url http://<host>:<port>/v1.

5. Check

ototo doctor

Its Model lines should say serves the model, a context of 65,536 tokens, calls tools, and that the server answers Otōto’s first turn with a tool call.

MLX: a model server on a Mac

MLX is Apple’s framework for running models on Apple silicon, and mlx-lm serves one over the API Otōto speaks. On a Mac it was the quickest of the servers we tried.

What we ran (10 October 2026): mlx-lm 0.32.0 on a Mac with an M3 Max and 128 GB, serving Qwen3.5 9B at four bits (5.6 GB), with Otōto 2026.19. ototo init’s tool-call test answered in 2.8 seconds. Two questions about ripgrep’s source were answered rightly, in 27 and 57 seconds.

1. Install

uv venv ~/.venvs/mlx --python 3.12
uv pip install --python ~/.venvs/mlx/bin/python mlx-lm

2. Fetch a model

A model is a folder of files in MLX’s own format. The one we ran:

hf download lmstudio-community/Qwen3.5-9B-MLX-4bit --local-dir ~/models/Qwen3.5-9B-MLX-4bit

(hf is Hugging Face’s command-line tool. Over a poor link it stalled for us, and we fetched the same files one by one with curl -L -C -.)

3. Start the server

~/.venvs/mlx/bin/python -m mlx_lm server --model ~/models/Qwen3.5-9B-MLX-4bit --host 127.0.0.1 --port 8081

Port 8081 is one ototo init looks on.

4. Point Otōto at it

ototo init --base-url http://127.0.0.1:8081/v1 --model ~/models/Qwen3.5-9B-MLX-4bit

Give --model the folder you started the server with: the server answers as the model it was started with under that name.

5. Check

ototo doctor

It should say serves the model and calls tools.

If it says the server did not answer with a model list, while questions work: the server’s list of models reads your Hugging Face cache, and fails when the cache holds files it cannot read. Ours did: the cache is on an exFAT drive, where macOS leaves a ._ file of its own beside each file, and with those deleted from the cache the list answered again. Otōto 2026.19’s doctor calls the server down meanwhile; later releases ask the server for the model instead, and say the list failed as a warning. ototo init with --model, as above, does not need the list.

A larger model

The install guide’s own setup for a Mac is Bonsai 2 27B on mlx-vlm: the model our measurements are of on a Mac, at 20 of 20 on our twenty questions, and at several minutes a question on this machine. Its commands are in dist/INSTALL.md, “Pick a model”.

Ollama: the shortest way to a first model

Ollama fetches a model by name and serves it, with nothing to configure but the context.

What we ran (10 October 2026): Ollama 0.40.2 from Homebrew on a Mac with an M3 Max and 128 GB, serving qwen3.5:9b from its library (7.6 GB) with a context of 65,536 tokens, with Otōto 2026.19. ototo init found it by itself, and its tool-call test passed, in 9 seconds while the model loaded and 1.4 once it had. Two questions about ripgrep’s source were answered rightly, in 86 and 124 seconds. The first question we ever asked it took far longer, with Ollama unloading the model part of the way through; asked again it took the 86 seconds.

1. Install

brew install ollama

or Ollama’s installer for macOS, Linux and Windows.

2. Fetch a model

ollama pull qwen3.5:9b

The model must be one that calls tools: Ollama’s library marks them, and ollama show qwen3.5:9b lists tools among its capabilities.

3. Start the server, with enough context

OLLAMA_CONTEXT_LENGTH=65536 ollama serve

The context is the thing to get right. Ollama chooses a default by the machine’s memory: by its documentation as little as 4,096 tokens on a small machine, which is too short for a question about code. OLLAMA_CONTEXT_LENGTH sets it for the server; if Ollama runs as an app or a service, set it there and restart it. Check with:

ollama ps

whose CONTEXT column should say 65536 once a model is loaded. It listens on 127.0.0.1:11434.

4. Point Otōto at it

ototo init

finds Ollama on port 11434 and uses the model it serves. With several models pulled, name one: ototo init --base-url http://127.0.0.1:11434/v1 --model qwen3.5:9b.

5. Check

ototo doctor

It should say serves the model and calls tools. It says nothing of the context, since Ollama’s API does not report it: ollama ps is the check for that.

LM Studio: a desktop app that also serves

LM Studio is an app for downloading and running models, with a server and a command-line tool, lms, beside it. Use it when you would rather manage models in a window than in a terminal.

What we ran (10 October 2026): LM Studio 0.4.26 on a Mac with an M3 Max and 128 GB, serving Qwen3.5 9B at four bits (a 5.6 GB GGUF file) with a context of 65,536 tokens, with Otōto 2026.19. ototo init’s tool-call test answered in 2.5 seconds and ototo doctor found nothing wrong. Two questions about ripgrep’s source were answered rightly, in 105 and 191 seconds.

1. Install

brew install --cask lm-studio

on a Mac, or the app from lmstudio.ai for macOS, Linux and Windows. lms comes with it. The first lms command starts LM Studio in the background, and can fail once while it starts (“Invalid passkey”): run it again.

2. Fetch a model

In the app, or from a file you already have, which is what we did:

lms import --copy --user-repo lmstudio-community/Qwen3.5-9B-GGUF ~/models/Qwen3.5-9B-Q4_K_M.gguf

(llama.cpp’s page has the command that fetches that file.) lms ls lists the models LM Studio has, by the names it gives them.

3. Load it with enough context, and start the server

lms load qwen3.5-9b --context-length 65536
lms server start --port 1234

lms ps should show the model with 65536 under CONTEXT.

4. Point Otōto at it

ototo init --base-url http://127.0.0.1:1234/v1 --model qwen3.5-9b

Name the model: LM Studio’s server lists every model it has downloaded, loaded or not, and Otōto should use the one you loaded with the long context.

5. Check

ototo doctor

It should say serves the model and calls tools. It says nothing of the context, since this server does not report it: lms ps is the check for that.

mistral.rs: a model server in one Rust binary

mistral.rs serves a model straight from its Hugging Face repository, quantising it as it loads, on a Mac (Metal) or on Linux with or without a GPU.

What we ran (10 October 2026): mistral.rs 0.9.4, the prebuilt Metal binary of its GitHub release, on a Mac with an M3 Max and 128 GB, serving Qwen3.5 9B quantised to four bits as it loaded, with Otōto 2026.19. ototo init’s tool-call test answered in 1.5 seconds and ototo doctor found nothing wrong. Two questions about ripgrep’s source were answered rightly, in 305 and 664 seconds. That is slow: the same model at four bits on llama.cpp, on the same machine, took 16 and 79 seconds. We have not found out why.

1. Install

mistral.rs’s installer fetches the prebuilt binary for your machine:

curl -fsSL https://mistralrs.dev/install.sh | sh

It puts the binary in ~/.mistralrs, links it into ~/.local/bin, and adds a line to your shell’s startup files. We read the script and ran the same release’s binary without it. To build from source instead, a Mac needs Xcode’s Metal Toolchain first (xcodebuild -downloadComponent MetalToolchain): without it the build stops at “Compiling metal -> air failed”.

2. and 3. Fetch a model and start the server

The server fetches the model by its Hugging Face name when it starts:

mistralrs serve -m Qwen/Qwen3.5-9B --isq 4 --max-model-len 65536 --host 127.0.0.1 -p 1234
  • -m names the repository. This is the model’s full weights (some 18 GB for this one), not a ready-quantised file.
  • --isq 4 quantises it to four bits as it loads, which took about a minute here.
  • --max-model-len 65536 is the context.
  • --host 127.0.0.1 keeps the server to this machine. Port 1234 is also LM Studio’s: run one of the two there.

4. Point Otōto at it

ototo init

looks for a server on port 1234 among others. For another address, or to be sure which server it means: ototo init --base-url http://127.0.0.1:1234/v1.

5. Check

ototo doctor

It should say serves the model and calls tools. It says nothing of the context: this server does not report one.

LiteRT-LM: very small models, on very small machines

LiteRT-LM is Google’s runtime for models small enough for a phone or a thin laptop. Its command-line tool has an OpenAI-compatible server, and Otōto’s tool calls work through it. The models it runs are too small to answer questions about code well: use it to see Otōto’s small-model tools working on a machine that can run nothing larger, not for answers to rely on.

What we ran (10 October 2026): LiteRT-LM 0.17.1 on a Mac with an M3 Max, serving Gemma 4 E4B (3.4 GB), with Otōto 2026.19. ototo init’s tool-call test passed, and ototo doctor found nothing wrong. Two questions about ripgrep’s source came back in 110 and 38 seconds, both wrong: the first named the right file and the wrong function. On the same questions a 9-billion-parameter model on llama.cpp was right both times.

1. and 2. Install it and import a model

As LiteRT-LM’s documentation says; litert-lm list then shows the models it has. We ran one imported earlier and did not follow those steps afresh.

3. Start the server

litert-lm serve --host 127.0.0.1 --port 9379

Say --host 127.0.0.1: without it the server listens on every network interface of the machine.

4. Point Otōto at it

ototo init --base-url http://127.0.0.1:9379/v1 --model gemma-4-E4B-it-hf

The server serves every model it has imported, so name the one you mean: litert-lm list has the names. Otōto 2026.19 and earlier do not look on port 9379 by themselves, so the address is named too; later releases find it, and take the first model when none is named.

5. Check

ototo doctor

It says serves the model and calls tools. It says nothing of the context, since this server does not report one: we do not know how long a question it can hold.

Models worth using

A recommendation here is a measurement: a model, on a named server and machine, asked our set of questions about real open-source code, each with a checked answer. Where we have only tried a model, on two questions, it says so. The list is short because it is only what we have run.

ModelDownloadMeasured onRightTime a question
Qwen3.8 27B, four bitsfetched by vLLMvLLM, on a GPU server19 of 204 to 45 seconds
Bonsai 2 27B, two bits8.6 GBMLX (mlx-vlm), on a Mac with an M3 Max and 128 GB20 of 20about four minutes
Qwen3.5 9B, four bits5.6 GBMLX (mlx-lm), on the same Mac17 of 1925 seconds at the median, 8 to 198

The 9B was measured on 10 October 2026 with Otōto 2026.19; the two 27B models earlier, on earlier releases, when the set had twenty questions. One has since been retired, so the set is nineteen now.

Which one

  • A GPU server a team shares: Qwen3.8 27B on vLLM. It is what Otōto’s own benchmarks are run with: vLLM’s page.
  • A laptop or a desktop: Qwen3.5 9B. A third of the size, most of the answers, and quick: on a Mac, MLX or llama.cpp; Ollama and LM Studio serve it too. With a context of 65,536 tokens Ollama reported it using 10 to 15 GB of memory, so a 16 GB machine is tight and 24 GB or more is comfortable. Both answers it got wrong were ones Otōto’s own check flagged (“claims not backed by the cited code”), which is what that check is for; it also flagged three that were right.
  • A Mac with memory to spare, and patience: Bonsai 2 27B. The best score, at minutes a question.
  • Nothing smaller, yet. Gemma 4 E4B (3.4 GB, on LiteRT-LM) calls tools as it should and got both of the two questions we tried wrong.

What a model needs

  • Tool calls, through the server it runs on: the model asks for a search or a read. ototo init and ototo doctor test it.
  • A long context: 32,000 tokens at the least, 64,000 to be comfortable.
  • A size your machine holds: the file, plus the context, in memory.

Any model that does the first two can be tried: ototo init --base-url <server>/v1 --model <name>, then ask it something about your code. Otōto says at the foot of each answer how many turns and how long it took, and flags what the cited code does not back.

Otōto in your coding agent

Otōto is an MCP server: ototo serve speaks the Model Context Protocol over standard input and output, and any agent that can start a command as an MCP server can use it.

AgentSet up byPage
Claude Codeototo init: the registration, Otōto’s block in CLAUDE.md, and its permissions and hooks in settings.jsonclaude-code.md
OpenCodeototo init: the server in opencode.json, with time for long questions, its instructions, and OpenCode’s own search offopencode.md
Piby hand, one commandpi.md
Another MCP clientby handbelow

Another MCP client

Register this command as a local (stdio) MCP server, under the name ototo:

~/.local/bin/ototo serve

Three things ototo init does for Claude Code and OpenCode are then yours to do:

  • Time. ask, locate, callers and edit wait for a small model, and can take minutes. Give the server’s calls a long timeout if your agent has one to set (Otōto gives OpenCode thirty minutes).
  • Instructions. An agent uses new tools well when it is told when to: for any question about the code, first call ask with the whole question; read with read and outline; search with search. Otōto’s block for Claude Code, dist/CLAUDE.ototo.md, is the text to adapt.
  • The model server. Otōto’s own settings are in ~/.config/ototo/config.toml (dist/config.toml.example): the model server’s address and model. ototo init finds a model server, tests it and writes that file. In 2026.19 and earlier it runs only where Claude Code or OpenCode is installed; without them, copy the example and fill it in, and check with ototo doctor. Later releases set up Otōto’s own side wherever they are run, and say what is left to register.

Claude Code

ototo init sets Otōto up in Claude Code, and the installer runs it for you. This page says what it changes, so that you can read it before you say yes, do it by hand, or take it out again.

What we ran (11 October 2026): Claude Code 2.1.296 with Otōto 2026.19 on a Mac, set up by ototo init: claude mcp get ototo says Connected, and ototo doctor passes its checks for Claude Code. It is also how Otōto is written, and the figures on ototo.dev were measured with Claude Code.

Claude Code 2.1 or later.

Set up

ototo init

finds a model server, tests it, lists what it will change and asks before it writes. --dry-run only lists, --yes does not ask, and it can be run again at any time: it is also the step after an upgrade. Each file it changes is backed up first, as <file>.before-ototo-<time>.

Then start a new Claude Code session: one that is already running keeps the tools it started with.

What it changes

The registration. ototo as an MCP server for all your projects. By hand, that is:

claude mcp add ototo --scope user -- ~/.local/bin/ototo serve

It serves the directory Claude Code starts it in. -e OTOTO_…=… on the registration sets a setting for it, and wins over ~/.config/ototo/config.toml; a project’s .mcp.json can register it with settings of its own.

~/.claude/CLAUDE.md. A marked block that tells Claude when to use Otōto: for a question about the code, first ask with the whole question; read with read and outline; search with search. The text is dist/CLAUDE.ototo.md. A CLAUDE.md that already tells Claude about Otōto in your own words is left as it is. This block is where the savings come from: the same advice in the MCP server’s own instructions alone was weaker.

~/.claude/settings.json. Three things:

KeyWhat is addedWhy
permissions.allowmcp__ototoOtōto’s tools run without a prompt each time
permissions.denyGrep, Glob, Bash(grep:*), Bash(rg:*), Bash(find:*), Bash(git grep:*)Claude’s own search is off, so a question goes to ask and exact text to search. It is how we run it, and how it was measured
hooksfour, below
HookWhat it does
UserPromptSubmitSays with each prompt what the block says, briefly
SubagentStartSays the same to the agents Claude starts, which see neither CLAUDE.md nor the prompt hook, only the brief Claude writes for them
PreToolUse on Bashototo hook bash refuses a command that only prints code (sed -n '10,40p', head, cat, and unzip -l or javap on an archive) on a file Otōto can read, and names the ototo read to make. Pipes, redirections, other commands and files outside the repository run as before; about 10 ms
PreToolUse on Readototo hook read refuses nothing. When Claude reads a file one of the plugins reads (a pipeline, a compose file), it adds a note naming the ototo read that gives the file as its tool evaluates it

ototo init --keep-search leaves Claude’s own search on: no permissions.deny, and no hook on Bash. The rest is the same.

With no small model

ototo init on a machine with no model server sets up the tools that need none: search, read, outline, changes, history and replace_all. The block and the hooks then speak of those alone. Run ototo init again once a server runs (quick starts), and ask, locate, callers and edit join them.

Check

claude mcp get ototo

should say Status: ✔ Connected. ototo doctor checks the three parts above, with the model server and the plugins, and says what to do about anything wrong.

Worth knowing

  • A long question. An ask can take minutes. If Claude Code gives up on one, raise MCP_TOOL_TIMEOUT (milliseconds) in its environment.
  • Git worktrees. Otōto serves the checkout the session started in and that repository’s other worktrees: a call works in one when its paths point inside it, or when it passes root.
  • Directories added to the session (/add-dir, claude --add-dir) are served the same way. A home directory or / is too broad, and is not.
  • Agents. Claude’s own agents get Otōto’s tools too, and the SubagentStart hook tells them so. When you brief one yourself, say it as well: outline a file before reading it, read path#Name for the parts that matter.

Taking it out

sh uninstall.sh, from the package, undoes all of it and removes the binary. By parts:

claude mcp remove ototo -s user
ototo settings remove ~/.claude/settings.json

and delete the block between <!-- ototo:begin and <!-- ototo:end --> in ~/.claude/CLAUDE.md. Everything else in those files stays as it was.

For everyone on a machine

ototo managed writes the registration, the permissions and the hooks as Claude Code’s managed settings (managed-mcp.json, managed-settings.json), which apply to every user and sit above anything a user or a project sets: see For organisations.

OpenCode

ototo init sets Otōto up in OpenCode when OpenCode is on the machine (opencode on the PATH, or its config directory), beside Claude Code or on its own. This page says what it changes.

What we ran (11 October 2026): OpenCode 1.18.30 with Otōto 2026.19 on a Mac, set up by ototo init: opencode mcp list shows ototo connected, and ototo doctor passes its checks for OpenCode. The figures on ototo.dev were measured with Claude Code, and we have none of our own for OpenCode. A tester runs it with one model on both sides, Qwen as OpenCode’s model and as Otōto’s small model, and finds the main context stays small and on the task through long sessions.

Set up

ototo init

finds a model server, tests it, lists what it will change and asks before it writes (--dry-run only lists). Each file it changes is backed up first, as <file>.before-ototo-<time>. Then start a new OpenCode session.

What it changes

OpenCode’s global config, ~/.config/opencode/opencode.json (under $XDG_CONFIG_HOME when that is set), gains three things. Everything else in it stays as it was.

{
  "mcp": {
    "ototo": { "type": "local", "command": ["/home/you/.local/bin/ototo", "serve"], "enabled": true, "timeout": 1800000 }
  },
  "instructions": ["/home/you/.config/ototo/opencode.md"],
  "permission": {
    "grep": "deny",
    "glob": "deny",
    "bash": { "grep *": "deny", "rg *": "deny", "find *": "deny", "git grep *": "deny" }
  }
}
  • mcp.ototo: the server, with thirty minutes for a call. OpenCode otherwise gives up on a tool call after about a minute, and an ask can take longer.
  • instructions: Otōto’s own file, ~/.config/ototo/opencode.md, added to the list. It says when to use each tool, by the names OpenCode gives them (ototo_ask, ototo_read, ototo_search…). An AGENTS.md of yours, and OpenCode’s fallback to ~/.claude/CLAUDE.md, stay as they are.
  • permission: OpenCode’s own grep and glob are denied, and grep, rg, find and git grep through bash, so that a question goes to ototo_ask and exact text to ototo_search. A single rule you had for bash ("bash": "ask") is kept as the rule for every other command.

ototo init --keep-search leaves OpenCode’s own search on: no permission entries. The rest is the same.

A config with comments (opencode.jsonc) is not rewritten: ototo init prints what to add, for you to put in by hand.

With no small model

ototo init on a machine with no model server sets up the tools that need none: ototo_search, ototo_read, ototo_outline, ototo_changes, ototo_history and ototo_replace_all, and the instructions speak of those alone. Run ototo init again once a server runs (quick starts).

Check

opencode mcp list

should show ototo connected. ototo doctor checks the registration, the instructions and the permissions, with the model server and the plugins.

Taking it out

ototo opencode remove

takes Otōto’s three entries out of opencode.json, backing it up first, and deletes opencode.md. sh uninstall.sh, from the package, runs it with the rest. A config with comments is left alone, and it says so.

Pi

Pi takes MCP servers since its 0.99, so Otōto is one command to add.

What we ran (10 October 2026): Pi 1.1.0 with Otōto 2026.19 on a Mac: the registration below, and pi mcp list showing Otōto connected with its tools. We have not yet run a session of Pi’s through it, so this page says how to connect the two and no more.

Pi 1.1.0 needs Node 22 or later; on Node 20 it stops at once with a syntax error.

Register Otōto

pi mcp add ototo -- ~/.local/bin/ototo serve

writes Otōto into ~/.pi/agent/mcp.json, for every project (-l writes .pi/mcp.json in this one instead).

Check

pi mcp list

should say ototo: connected, and list its tools: read, outline, search, changes, history, replace_all, and with a small model set up locate, callers, ask and edit.

What is still yours to do

ototo init sets Otōto up for Claude Code and OpenCode, not yet for Pi. See “Another MCP client”: the model server’s settings, and telling Pi when to use Otōto’s tools.

The plugins

Some files mean more than their text: a CI job is its includes and its extends chain merged, a dependency’s version may be set in a parent three files away. A plugin reads one such kind of file and answers for it through Otōto’s own tools, outline and read, so that what comes back is the thing as its tool would evaluate it, each line with the file and line it came from.

Nine come with Otōto. The installer puts them in ~/.config/ototo/plugins; each has its own version and is published on its own, between Otōto’s releases.

PluginReadsAsk it
gitlab-ciGitLab CI pipelinesread .gitlab-ci.yml#deploy:prod
composedocker compose filesread compose.yaml#web
mavenMaven POMsread pom.xml#spring-core
spring-configSpring’s application* and bootstrap* filesread application.yml#spring.datasource.url
terraformTerraform and OpenTofu modulesread main.tf#aws_db_instance.main@prod
kubeKustomize overlays and Helm valuesread overlays/prod/kustomization.yaml#Deployment/api
codeownersCODEOWNERSread .github/CODEOWNERS#src/api/server.rs
archiveJARs, WARs, wheels and other ZIPs; compiled Java classesread lib/core.jar!org.demo.Shelf
gitlabGitLab itself: merge requests and pipelinesthe forge tool

ototo.dev/plugins has the same list as the download channel has it today, with each one’s version.

Managing them

ototo plugins

lists the ones installed: on or off, version, who signed each, the files it reads, and why one does not load.

CommandDoes
ototo plugins available [tag]what the channel offers (java, infrastructure, gitlab…), and how each stands to the one here
ototo plugins add <plugin>installs one from the channel
ototo plugins updatebrings the installed ones up to the channel’s; never back a version, and a plugin of your own is left alone
ototo plugins remove <plugin>deletes one
ototo plugins help <plugin>which of this repository’s files it reads, and its own notes
ototo plugins grant <plugin>lets a forge plugin reach the hosts it declares (revoke takes that back)

Nothing is trusted for having been downloaded. The channel’s index must be signed by the release key built into Otōto; a plugin’s file must have the checksum the index gives; and the plugin must then carry its own signature, by a key in your ~/.config/ototo/allowed_signers. The channel is asked only when you run one of these commands.

How it works says what a plugin can and cannot reach; Writing a plugin builds one from an empty directory.

gitlab-ci

Pipelines outlined a line per job and template, with stages, include, variables and default. read .gitlab-ci.yml#deploy:prod (or read deploy:prod from anywhere, #deploy:prod.script for one key) returns the job as GitLab runs it: local includes followed, the extends chain deep-merged, default and global variables inherited unless inherit says not, anchors and !reference expanded, and every key marked with the file, lines and job or template it came from. Project, remote, template and component includes are named, not followed.

compose

Compose files and their overlays outlined a line per service. read compose.yaml#web returns the service as docker compose runs it: includes, the file and its override (or an overlay’s base under it) merged as compose merges them, extends resolved, anchors expanded, and ${VAR} filled in from the project’s .env, with each line’s source. Variables come from .env only: a plugin sees no environment.

maven

POMs outlined as coordinates, parent, modules, properties, dependency management, dependencies, build plugins and profiles. read pom.xml#spring-core returns the dependency with its effective version and scope as Maven resolves them: parents found by relativePath or by coordinates, properties merged and ${…} filled in, a version not declared taken from dependency management in Maven’s order, imported BOMs included. A parent or BOM outside the repository is named, and no version is made up. Also #effective, #properties, #parent, #modules.

With dependency_caches = true in Otōto’s settings, a parent or a BOM that is not in the repository is read where Maven keeps what it has downloaded, ~/.m2/repository, and the dependency’s JAR is named when it is there. From version 0.2 of the plugin, and an Otōto later than 2026.19.

spring-config

application* and bootstrap* .properties, .yml and .yaml, read with the files Spring loads beside them. read application.yml#spring.datasource.url gives the value with no profile active and for each profile that changes it, each with its file and line and what it overrides; #@prod gives every property as that profile sees it. Placeholders are filled in; a key no file sets shows as coming from the environment, never guessed.

terraform

Modules read with their variables’ values as Terraform gives them: the default, the tfvars files in its order, then an environment’s (@prod). read main.tf#aws_db_instance.main@prod gives the block as written, each expression that uses a variable or a local marked with what it works out to, and where every value came from. read var.region gives a variable’s declaration, value, every file that sets it and every use. No function is called, sensitive variables are never shown, and state, workspaces and remote modules are out of sight.

kube

Kustomize and Helm, from the repository’s files alone: no kustomize, kubectl or helm runs. read overlays/prod/kustomization.yaml#Deployment/api builds the resource in kustomize’s order (resources, components, generators, patches, namespace, prefixes, labels, replicas, images), each line with the file that set it and what changed it. read values.yaml#image.tag gives a chart’s value across its values files and the template lines that use it. Remote bases and generator plugins are named, not applied; templates are not rendered.

codeowners

Any CODEOWNERS, outlined as its rules. read .github/CODEOWNERS#src/api/server.rs answers who owns that path: the owners, the rule that decides with its line, and the earlier rules it overrides; a GitLab file with sections is answered section by section. #@backend lists the lines that name an owner. It says whether it read the patterns as GitHub or as GitLab does, and when that cannot be told and the two differ.

archive

Inside a JAR, WAR, wheel or any other ZIP without unpacking it, and a compiled Java class as its declarations, with no javap: outline lib/core.jar lists it, read lib/core.jar!org.demo.Shelf gives a class with its generics, parameter names and constants, and !META-INF/MANIFEST.MF any other entry; a.war!WEB-INF/lib/b.jar!… reaches inside. It is lent only the bytes of the files it opens, and writes and runs nothing. With dependency_caches = true, it reads a dependency’s JAR where the build tool keeps it.

gitlab

A forge plugin: it answers the forge tool’s questions about the repository on gitlab.com, or a GitLab of your own added in its settings. The checked-out branch’s open merge request and latest pipeline, a change by number, a pipeline’s jobs by stage, a failed job’s log cut to what failed. It reaches nothing until you grant it:

ototo plugins grant gitlab

It may only GET, over https, from the hosts it declares, and it never sees a token: Otōto adds the one you set for that host. Without a token it sees what is public; job logs need one. ototo plugins help gitlab has its settings.

Writing a plugin

A plugin teaches Otōto a kind of file it does not know: a build tool’s model, a CI system’s pipelines, a team’s own format. Otōto then outlines that file and reads one part of it by name, the way it does for code, and so does the small model when it answers a question. This page takes one from an empty directory to a loaded, signed plugin.

The plugin it follows is src/plugins/codeowners: who owns a path, by the CODEOWNERS rule that decides it. It is about 300 lines and their tests, and CI builds and tests it with the other plugins, so what this page says of it stays true.

What a plugin is, and what it cannot do

A plugin is a WebAssembly component, built against src/wit/plugin.wit and run in a sandbox (wasmtime). It has no network, no filesystem and no environment: it reaches only what Otōto lends its kind, and each call it makes goes through the same checks as Otōto’s own tools. A call has a fuel budget and the instance a memory limit; a plugin that traps, runs out or fails is logged and skipped, and Otōto answers as it would without it. One .wasm runs on every OS and CPU.

There are three kinds. Pick by what you need to be lent:

KindForIt is lent
A file plugin (world plugin)A kind of text file: a pipeline, a POM, a CODEOWNERSThe repository’s files as text: read-file, list-files, grep-files. When the user’s dependency_caches setting lends them, also a POM in a dependency cache: list-files("dependency-cache:<its path in the cache>") answers with where it is, and read-file reads it there
An archive plugin (archive-plugin)Files that are not text: an archive’s listing, a compiled classThe bytes of the files its own globs name, a part at a time, and no others
A forge plugin (forge-plugin)A forge’s merge requests and pipelinesHTTPS GET to the hosts it declares, once the user grants it; no files at all

Most plugins are file plugins, and the rest of this page is about one. src/plugins/archive and src/plugins/gitlab are the other two kinds to read.

What a file plugin gives

Four functions (interface view in plugin.wit):

  • describe: its name, its version, and the files it reads, as globs. Otōto asks it about no others.
  • outline(path, text): the file’s declarations, each with its lines, or nothing to say “this file is not my kind”, in which case Otōto reads it as before. A glob like **/*.yml matches many files that are not yours.
  • read(path, symbol): what ototo read path#symbol returns, or nothing to leave it to Otōto, which then shows the lines of the outline’s entry of that name.
  • find(symbol): what ototo read symbol returns with no file named, or nothing.

A declaration (entry) is a name, a line range, a depth, and the one line the outline shows for it.

A first plugin

You need Rust and the WebAssembly target (rustup target add wasm32-wasip2).

cargo new --lib owners && cd owners
mkdir wit && curl -fsSL https://gitlab.com/handmadedigital/projects/ototo/-/raw/main/src/wit/plugin.wit -o wit/plugin.wit

Cargo.toml: a cdylib is the component, the rlib lets cargo test run your code natively. The kit is what Otōto’s own plugins share: the declaration’s shape, the repository as a plugin sees it, an in-memory repository for tests, and YAML parsed with each value’s lines.

[lib]
crate-type = ["cdylib", "rlib"]

[dependencies]
wit-bindgen = "0.62"
ototo-plugin-kit = { git = "https://gitlab.com/handmadedigital/projects/ototo.git" }

[profile.release]
opt-level = "s"
lto = true
strip = true

src/lib.rs is the glue, and the same in every file plugin: copy src/plugins/codeowners/src/lib.rs and change the names (and the path to the interface, path: "wit"). It generates the bindings, implements the kit’s Repo over what Otōto lends, and hands each of the four functions to your own module. It is compiled only for WebAssembly, so that everything else builds and tests as ordinary Rust.

Your own module is plain functions over text and a &dyn Repo (src/plugins/codeowners/src/owners.rs):

#![allow(unused)]
fn main() {
pub fn outline(path: &str, text: &str) -> Option<Vec<Entry>>      // None: not my kind of file
pub fn read(repo: &dyn Repo, path: &str, symbol: &str) -> Option<String>   // None: leave it to Otōto
}

Build it, and try it without changing your settings: a directory of your own for plugins, and unsigned ones allowed for these commands only.

cargo build --release --target wasm32-wasip2
mkdir -p /tmp/try && cp target/wasm32-wasip2/release/owners.wasm /tmp/try/
export OTOTO_PLUGINS=/tmp/try OTOTO_ALLOW_UNSIGNED_PLUGINS=true
ototo plugins                                    # owners  0.1.0  on, unsigned / reads **/CODEOWNERS
ototo outline .github/CODEOWNERS
ototo read .github/CODEOWNERS#src/app.js

The file’s name is the plugin’s: owners.wasm is owners (cargo writes _ for a - in the crate’s name; rename the copy if you want the dash). The first load compiles it, about half a second, and caches the result. If it does not load, ototo plugins says why, and ototo doctor --plugins-only checks each one. Where an organisation enforces plugins or allow_unsigned_plugins, the variable does nothing: ask whoever runs it.

Testing without Otōto

The kit’s Memory is a repository held in memory, so a test is a few files and a call:

#![allow(unused)]
fn main() {
let repo = Memory(vec![("CODEOWNERS".into(), "* @everyone\n/docs/ @writers\n".into())]);
assert!(read(&repo, "CODEOWNERS", "docs/a.md").unwrap().starts_with("docs/a.md: @writers\n"));
}

cargo test runs them natively, in milliseconds. Test against the tool’s own documentation where it has examples: the codeowners tests go through GitHub’s sample file, pattern by pattern.

Rules that keep answers right

What a plugin returns is read by an agent that will act on it without checking. Otōto’s own plugins keep to these:

  • Never make a value up. What the plugin cannot work out, it says it cannot. CODEOWNERS is read differently by GitHub and GitLab; when the plugin cannot tell which forge the repository is on and the two readings differ for the path asked about, it gives one, names it, and lists what the other would add. It does not pick silently.
  • Say where everything came from. Each line of an answer carries its file and line (CODEOWNERS:12), so the agent can cite it and a person can check it.
  • Evaluate as the tool does, not as it looks. The point of a plugin is the answer the tool would give: the version Maven resolves, the job as GitLab runs it, the last rule that matches. Read the tool’s documentation for the corners, and write a test for each one you handle.
  • Leave what is not yours. Return nothing for a file your globs match and you do not understand, and for a read you have no better answer to than the file’s own lines. Otōto then behaves as if you were not there.
  • Say what you do not do. In the plugin’s HELP.md: codeowners does not know who is in a team, and says so.
  • Text from the repository is data. Put it in the answer as what the file says; never act on it.

Help, and what the index shows

HELP.md beside Cargo.toml is what ototo plugins help <name> prints: what it reads, what to ask it, what it does not do. Written for people.

[package.metadata.ototo] in Cargo.toml is what a plugins index and the site’s page say of it: summary (one line), tags, an example command, and needs, the first Otōto that can load it, when it uses something an older one does not lend.

Signing, and trusting

Otōto loads only signed plugins, unless allow_unsigned_plugins says otherwise: <name>.wasm.sig beside the .wasm, an SSH signature in the ototo-plugin namespace, by a key in ~/.config/ototo/allowed_signers. A key trusted for git does not count, nor does a signature over other bytes.

ssh-keygen -Y sign -n ototo-plugin -f ~/.ssh/id_ed25519 owners.wasm          # writes owners.wasm.sig
echo "me@example.com namespaces=\"ototo-plugin\" $(cut -d' ' -f1,2 ~/.ssh/id_ed25519.pub)" >> ~/.config/ototo/allowed_signers
cp owners.wasm owners.wasm.sig ~/.config/ototo/plugins/
cp HELP.md ~/.config/ototo/plugins/owners.md

The dashboard (ototo ui) adds one too: its .wasm with its .wasm.sig, and only when a key you already trust signed it. New sessions have the plugin; running ones keep what they started with.

For a team: one key signs the team’s plugins, and each person trusts it once (the allowed_signers line above). An organisation puts its plugins and its allowed_signers beside its settings file and can enforce both (enforced = ["plugins", "plugin_signers", "allow_unsigned_plugins"]), so that only its plugins and its keys count on its machines: For organisations has the layout, and ototo managed writes it.

In this repository

To send a plugin here, or to change one: each plugin under src/plugins is a crate of its own with its own lock file, built with the path form of the two dependencies above (path = "../kit", path: "../../wit"). CI formats, lints, tests and audits every plugin named in PLUGINS in .gitlab-ci.yml; CONTRIBUTING.md has the rest. A new kind of plugin, or a change to plugin.wit, is an issue first.

Settings

Otōto’s settings are one file, ~/.config/ototo/config.toml, which ototo init writes for you. Most people set the model server and nothing else.

  • A variable wins over the file. Every setting is also an environment variable: its name in capitals with OTOTO_ in front (verify is OTOTO_VERIFY). OTOTO_CONFIG points at another file.
  • A session reads the file when it starts. A change reaches the sessions of your agent started after it. ototo serve says on stderr which file it read, and names any key it does not know.
  • An organisation’s file sits under yours, and can enforce some settings over it: For organisations.
  • ototo doctor checks the file and each model server; the dashboard shows what the file says now.

dist/config.toml.example is a file to start from.

The model server

base_url = "http://<host>:<port>/v1"
model = "<the name the server gives the model>"
SettingDefaultWhat it is
base_urlhttp://127.0.0.1:8080/v1the server’s OpenAI-compatible address
modelthe model’s name there: curl -s <base_url>/models shows it
api_keya key, where the server wants one
endpointsseveral servers, in the order to try them (below)

The server needs OpenAI-style tool calls, 32,000 tokens of context and 4,000 of output; thinking is switched off in each request. A model server: quick starts has the flags for each server we have run.

Several servers

List several and each request goes to the first that answers. A spot GPU that has been reclaimed, a workstation that is switched off, or a gateway’s error page hands over in seconds, and a server that failed waits at the back of the queue for a minute, so that the next requests do not wait on it. A hosted model makes a last resort: paid for by the token, but always there.

[[endpoints]]
name = "gpu"
base_url = "http://<gpu host>/v1"
model = "Qwen/Qwen3.8-27B-FP8"

[[endpoints]]
name = "workstation"
base_url = "http://<workstation>:8080/v1"
model = "RedHatAI/Qwen3.8-27B-INT4"

[[endpoints]]
provider = "anthropic"          # https://api.anthropic.com unless base_url says otherwise
model = "claude-haiku-4-5"      # key from ANTHROPIC_API_KEY, or api_key / api_key_env

The other settings stay at the top of the file and apply to every server. A reply that another server answered says so in its footer (via haiku (gpu, workstation down) · $0.031), and the run log records which answered, why the others failed and what a paid model cost: list prices for the Claude models, or price_input and price_output on the endpoint, in US dollars per million tokens. OTOTO_BASE_URL in the environment still pins one server.

How far a question may go

SettingDefaultWhat it is
max_turns8 (ototo init writes 36)the small model’s turns for one question
max_input_tokens400000what one question may send the model over all its turns; past it the next turn is its last. 0 for no such budget
max_output_tokens4096the most the model may write in one turn
temperature0.2
thinkingfalsethe model’s own reasoning before each turn
tool_choiceautosee below

An answer that ran out of turns or tokens says so, and is the small model’s best by then. A long question in several parts reaches the budget most often, and does better asked apart.

tool_choice. With auto, no tool_choice is sent: the model may answer in prose, and is nudged back to its tools. required is for servers with constrained decoding (vLLM): every turn must call a tool, with arguments held to the tool’s schema. That rules out mangled parameter names, but with Qwen3.8 on vLLM it also made the model repeat its last call in place of changing course, a search it could not narrow five turns running, where auto answered the same question in eight. Leave it on auto, and try required only if a model keeps mangling its arguments. A server that refuses tool_choice (llama.cpp) has it dropped after the first refusal.

Checking the answers

SettingDefaultWhat it is
verifyoff, unless ototo init turns it on for your modela second pass by the small model over each ask and locate answer’s claims, in a fresh request with the code each cites. One more request an answer: 2 to 3 s on an idle GPU server. Leave it off on a slow or busy one
excerpts2how many of an answer’s citations get their code attached, under “Key code”: the enclosing function up to 25 lines, else the cited line and three each side. 0 turns it off
findonthe small model’s ranked search over declarations. off removes it

The checks that need no model, of every name and value a claim states against the code it cites, are always on: How it works.

What your agent is offered

SettingDefaultWhat it is
toolsall of themthe tools on the menu, comma-separated: "ask,read,search". Every tool’s definition is carried in your agent’s context on each turn, so a shorter menu costs less; Otōto’s instructions then name only those
hide_model_toolsoffwhile no model server answers, take ask, locate, callers and edit off the menu, and put them back when one does
routinesoffoffer the routine tool (experimental)

hide_model_tools. Without it, the tools that need a model stay listed while none answers, and fail at once, saying what works meanwhile. With it, the session that finds every server gone tells your agent its list has changed, asks once a minute (a request of a few tokens) whether a server is back, and leaves a note for the prompt hook. Two things to know first. Run ototo init after setting it, since the prompt hook must then ask Otōto what to say (ototo doctor says when that is still to do), and a process of Otōto’s starts with each prompt. And a list that changes in the middle of a session costs Claude Code its prompt cache, once each way. Tried with Claude Code; not yet with OpenCode.

How much a read returns

SettingDefaultWhat it is
read_max_lines2000the most lines one read call returns
read_target_lines1000the most for one address in it
read_whole_file1000a file up to this long comes back whole when read by its path; a longer one as its outline
read_whole_documentas read_whole_filethe same for a Markdown document: lower, a long one comes back as its sections, to read by name (docs/guide.md#install)

A longer range is cut, with a note saying where to continue. The small model’s own reads stay at 250 lines a call, whatever these say: its context is the one that fills.

Edits

SettingDefaultWhat it is
check_cmda build or a type check for edit: cargo check, ./gradlew compileJava, npx tsc --noEmit
check_timeout300seconds it may take

With a check_cmd, the small model gets a check tool. Its pending edits are written, the command runs in the repository, and every file is put back afterwards; one that someone else changed meanwhile is left alone, and reported. The verdict is in the reply’s footer. It is the one thing Otōto runs, and Otōto takes it only from its settings and its environment, never from the files of a repository it reads. (A project’s .mcp.json, which Claude Code asks you to approve, can set OTOTO_CHECK_CMD in the registration’s environment.)

Outside the repository

SettingDefaultWhat it is
dependency_cachesofftrue for Maven’s and Gradle’s caches (~/.m2/repository, ~/.gradle/caches), or a list of directories. Otōto then reads the archives there, and Maven’s POMs, by absolute path: read <jar>!<class>. Nothing else, no listing, no search, never a write
read_secretsofflet Otōto read files it takes for secrets by their names: environment files, private keys, credentials files, Terraform state

Both are off because each widens what Otōto reads. dependency_caches is the one case where it reads outside the directories a session was given; a cache that is a home directory or a whole disk is ignored.

Plugins

SettingDefaultWhat it is
plugins~/.config/ototo/pluginswhere plugins load from: directories, :-separated
plugin_signers~/.config/ototo/allowed_signersthe keys whose signature lets a plugin load, in git’s allowed-signers format
allow_unsigned_pluginsoffload unsigned plugins: for writing one of your own
plugin_grantsthe forge plugins that may reach their hosts: ototo plugins grant gitlab writes it
plugin_settingsa plugin’s own: for gitlab, which token goes to which host, project or group

The plugins has the commands; ototo plugins help gitlab the forge plugin’s settings.

The run log and the dashboard

SettingDefaultWhat it is
run_log~/.ototo/runs.jsonlwhere each call is logged, for the dashboard. off turns it off
run_log_max_mb20the size at which the log is rotated
run_log_keep3how many rotated files are kept
traceonthe live trace beside the log, which ototo tail and the dashboard’s “Running now” follow

The log holds questions, the small model’s steps and replies, and stays on your machine.

Reporting

otlp_endpoint, otlp_headers, otlp_attributes and otlp_interval send counts, never questions, code or answers, to an OpenTelemetry collector of yours: Metrics.

Updates

SettingDefaultWhat it is
update_urlhttps://ototo.sh/betathe download channel ototo update and ototo plugins ask, and only when you run them
update_tokena bearer token, for a private channel

ototo update installs a newer release from the channel, or, when the channel does not answer, the newest package for this machine in ~/Downloads. Either way it checks the release’s checksums against their signature by Otōto’s release key, which is built into the binary, and the package against its checksum, then runs the package’s install.sh, which keeps your settings. Only a newer release is installed, so a channel cannot move anyone back. --check only says whether there is one; --dry-run downloads and checks.

For an organisation

enforced and min_version belong in the organisation’s file, with the settings it wants everyone to have: For organisations.

The dashboard, and watching it work

Otōto logs every call on your machine, and three things read that log: a dashboard, a live tail, and a report to send when something goes wrong. None of it leaves the machine unless you send it.

The dashboard

ototo ui

serves a page at http://127.0.0.1:7777 (--port for another), and prints the link to open. It reads the log file, not a running server, so it covers every session and every repository that logged to that file, and needs nothing running but itself.

For a time range, a repository, a tool or a model, it shows:

  • Totals: delegated and direct calls, the small model’s tokens in and out, what went back to your agent, and the model’s time (median and 90th percentile).
  • Reliability: answers that finished and were not cut off, with what cut the others off (out of turns, out of tokens, repeating its calls, answered without finishing; more budget helps only the first two), citations dropped as unverifiable, build checks passed, and errors. No grade is given under twenty answers.
  • Activity over time, delegated against direct, and latency of each delegated call over time, which shows when the model server was busy.
  • By tool, by repository and by model: calls, errors, times, turns, tokens and reply size. Point Otōto at another model and it gets a row of its own to compare.
  • What the small model did: its steps across all turns: its own tools, tools it made up (it has seen other agents’ Grep, Read and Bash in training), turns it answered in prose, and calls it garbled, with the share of turns that went on those. A row opens to recent examples.
  • Recent calls, each opening to the question, the small model’s steps turn by turn, and the exact reply your agent got. Running now shows the calls in progress.
  • Model servers: a probe of each (latency, the models it serves), in the order they are tried.
  • Settings: the settings file’s path and what it says now, read again each time the page opens.
  • Plugins: each one’s version, the files it reads, who signed it, and whether it loads. One can be switched off or on, removed, or added from its .wasm and signature, and only if a key you already trust signed it; trusting another key is never done from the page.

Light or dark, or following the system.

Who can open it

The log holds questions, code and replies, so the dashboard keeps to this machine:

  • It listens on 127.0.0.1 only, and answers only requests addressed to 127.0.0.1 or localhost, so a web page whose name is pointed at 127.0.0.1 cannot read it through your browser.
  • Its data needs a key, kept in ~/.ototo/ui.key where only you can read it, so another user of a shared machine gets nothing from the port. The link ototo ui prints carries the key after a #, which browsers never send to a server; ototo ui --link prints it again, for a dashboard a service started.
  • Changes to plugins are accepted only from the page itself.

From another machine, tunnel to it, and open the printed link:

ssh -L 7777:127.0.0.1:7777 <host>

A live tail

ototo tail

follows every delegated call as it runs, from every session on the machine: its question, each turn of the small model (the tool it called, tokens so far, time), the claim check, and how it ended.

12:39:17  21572.1    ask      start   Which class validates a new pet, and what does it check?
12:39:22  21572.1    ask      turn 1  find({"query":"validate new pet"}) · 1.6k tok · 4.7s
12:39:24  21572.1    ask      turn 2  read({"addresses":["src/main/java/…/PetValidator.java"]}) · 3.8k tok · 7.1s
12:39:40  21572.1    ask      turn 3  finish({"answer":"The class is PetValidator …"}) · 7.3k tok · 22.9s
12:39:40  21572.1    ask      end     23.0s · 3 turns · 7.3k tok · ok

Your agent is told the same while a delegated call runs, when it asks for progress (Claude Code does): a notice when it starts, then every few seconds the turns so far and the latest step (turn 5 · read src/lib.rs). When no turn ends for 30 seconds, as when the model server is busy, it says so (waiting for the model's reply to turn 6, 4 min). That also keeps Claude Code from ending a long call that is still working.

From a shell, ototo ask, locate, callers and edit show the call on the terminal’s last line while it runs, and clear it before the answer is printed:

(•) ask · gpu · 4 s · turn 2 · read src/cache/Store.kt#evict

The run log

~/.ototo/runs.jsonl: one line a call, from your agent’s sessions and from the command line alike. A line has the kind of call, the repository, the session, the input, the small model’s steps, the reply (its first 6,000 characters), turns, tool calls, tokens, timings, citations kept and dropped, a check’s verdict, and whether an edit was applied. It also names the Otōto that made the call and, for a delegated one, a short hash of the prompts and tool descriptions the small model was given, so that two calls with the same hash had the same instructions.

The log is rotated at 20 MB, three old files kept, and the dashboard reads those too. run_log = "off" turns it off, run_log moves it, and trace = "off" stops the live trace beside it: Settings.

A report to send

ototo report

writes one file for when something goes wrong: the version and the machine, what ototo doctor finds, the settings with their credentials taken out (keys, headers, passwords in URLs), the log’s counts per tool for the last week (--days), and its latest errors’ messages (--errors, 20 unless set). Not the questions, the code read or the answers: --with-inputs adds each error’s input. The file is readable by you only, and nothing is sent: it is yours to read and to send.

Across a team

The dashboard is one machine’s. For usage across many, Otōto sends counts, never questions, code or answers, to an OpenTelemetry collector of yours: Metrics and For organisations.

Metrics: reporting over OpenTelemetry

Otōto can send its counts to an OpenTelemetry collector, beside the ones Claude Code sends of itself (claude_code.token.usage, claude_code.cost.usage). Together they show what handing questions to a small model saves, measured rather than estimated: calls, the small model’s tokens and time, and Claude’s cost, by person and by team.

Only counts leave the machine: never a question, code, or an answer. Reporting is off unless you set where it goes.

Turn it on

In ~/.config/ototo/config.toml (or the organisation’s file, for everyone on a machine):

otlp_endpoint = "http://<collector>:4318"         # the collector's OTLP/HTTP port
otlp_headers = "Authorization=Bearer <token>"     # optional, comma-separated
otlp_attributes = "team=platform"                # optional, added to every series
# otlp_interval = 60                              # seconds between sends; also sent when the process exits

ototo doctor says whether the collector accepts what is sent, and which e-mail address goes with it.

What is sent

MetricUnitAttributes
ototo.calls1tool, repo, model, endpoint, outcome (ok, flagged, unfinished, error)
ototo.local_tokenstokenstool, model, endpoint, type (input, output)
ototo.turns1tool, model
ototo.durations (histogram)tool, delegated
ototo.fallbacks1failed, answered_by
ototo.cost.usageUSDmodel, endpoint
ototo.claims1tool, verdict (backed, unbacked, unchecked, contradicted)
ototo.reply.charscharstool
ototo.servergaugeversion, OS, CPU, plugins, where its settings come from, whether Claude Code’s setup is managed, its endpoints
mcp.server.operation.durations (histogram)mcp.method.name (tools/call), gen_ai.tool.name, error.type (tool_error)
gen_ai.client.token.usage{token} (histogram)gen_ai.operation.name (chat), gen_ai.provider.name, gen_ai.request.model, gen_ai.token.type, ototo.endpoint

The last two are OpenTelemetry’s own names (its MCP and GenAI conventions), sent beside Otōto’s so that a dashboard built on those conventions covers Otōto with other MCP servers and model clients. ototo.server is what each running ototo serve says of itself, so a dashboard can show the servers and developers running now, and the versions in use.

Values are cumulative from each process’s start, sent as OTLP/HTTP in JSON. Through a collector’s Prometheus exporter the names gain unit suffixes, as Claude Code’s do: ototo_calls_total, ototo_local_tokens_total, ototo_cost_usage_USD_total beside claude_code_cost_usage_USD_total.

Who and where

Every series carries service.name (ototo), service.namespace, service.version, service.instance.id (one a process), host.name, user.name and user.email.

  • user.email is the address of the Claude account Claude Code is signed in with, which Claude Code puts on its own metrics: the two then compare person by person with nothing set per person. user.email=<address> in otlp_attributes sends another, and user.email= (empty) sends none.
  • The standard OTEL_RESOURCE_ATTRIBUTES adds attributes (deployment.environment.name=prod, team=…); otlp_attributes wins over it. Either can set service.namespace; neither can change service.name.

Claude Code’s own metrics, to the same collector

Claude Code reports its tokens and cost when told where, by its own settings (CLAUDE_CODE_ENABLE_TELEMETRY, OTEL_METRICS_EXPORTER, OTEL_EXPORTER_OTLP_ENDPOINT). For a team, ototo managed --collector <url> --team <name> writes those into Claude Code’s managed settings with the same team label as Otōto’s: see For organisations.

With no collector of your own, the hub is one to run, with Prometheus and Grafana and two dashboards over both sets of metrics.

For organisations

Everything else in these pages works a developer at a time. This is for rolling Otōto out to many machines, and it is all in the open-source Otōto: there is no other edition.

One settings file for everyone on a machine

/etc/ototo/config.toml (/Library/Application Support/Ototo/config.toml on macOS; OTOTO_SYSTEM_CONFIG moves it) sits under each user’s own ~/.config/ototo/config.toml. Its settings apply where a user’s file and environment say nothing, and the ones it lists in enforced apply whatever they say.

enforced = ["otlp_endpoint", "otlp_attributes", "plugin_signers", "allow_unsigned_plugins", "endpoints"]
otlp_endpoint = "http://collector.corp:4318"
otlp_attributes = "team=platform"
plugin_signers = "allowed_signers"   # relative paths are beside this file

[[endpoints]]
name = "gpu"
base_url = "https://gpu.corp/v1"
model = "Qwen/Qwen3.8-27B"

The file should belong to root and be writable by no one else, or anyone could change what it enforces. ototo doctor checks that, shows what is enforced, and says when a user’s own setting is overridden.

Code goes only to your model servers

With endpoints enforced, the organisation’s model servers are the only ones: a user cannot add another, the Claude Haiku fallback included, and ototo init will not set one. The server itself is vLLM’s page.

Your plugins, and whose signatures count

Beside the settings file, plugins/ holds the organisation’s plugins and allowed_signers the keys it trusts to sign them. Both are used alongside a user’s own, unless plugins or plugin_signers is enforced: then only the organisation’s are. Writing a plugin says how one is made and signed.

Rolling it out

ototo managed writes what to push to every machine with whatever you push settings with (MDM, Ansible, a golden image):

ototo managed --out ototo-managed --collector http://collector.corp:4318 --team platform \
  --base-url https://gpu.corp/v1 --model Qwen/Qwen3.8-27B
  • Claude Code’s managed-settings.json: Otōto’s permissions and hooks for every user, above anything a user or a project sets; with --collector, Claude Code’s own telemetry to the same collector with the same team label.
  • managed-mcp.json: the ototo server, registered for every user.
  • The organisation’s config.toml and allowed_signers.
  • A README of where each goes on macOS and on Linux.

--install writes them into this machine’s system locations instead, merged into managed settings already there and backed up. With --base-url the organisation’s servers become the only ones; --allow-user-endpoints lets users add their own. On a managed machine ototo init adds nothing a user would duplicate, and ototo doctor says the setup comes from the organisation.

Somewhere for the counts to go

--collector wants an OpenTelemetry collector. If you have none, the hub is one to run: a collector, Prometheus and Grafana with two dashboards, usage and savings by person and team, and which Otōto servers are running, on what version. It comes as a compose file, a Helm chart and kustomize manifests, in the source’s deploy/.

Updates from your own channel, and the oldest version you accept

  • update_url points ototo update at a mirror of the download channel that you host, in place of https://ototo.sh/beta. A release is checked against Otōto’s release key wherever it was fetched from.
  • min_version = "2026.18" is the oldest Otōto the organisation accepts: ototo doctor fails on an older one and says ototo update. Nothing is asked of anyone for it and nothing stops working: it is the check a rollout’s script or a fleet’s dashboard reads. ototo doctor also says which sessions are still running the Otōto that was there before the last update.

What is reported, and to whom

Point otlp_endpoint at the collector you already run: Metrics lists every count, and that nothing else leaves the machine. Each running server also reports what it is (ototo.server), so a dashboard can show who is running which version.

For a security review

SECURITY.md covers what runs, what it reads and writes, what leaves the machine, the plugin sandbox and signing, and releases. Every package carries sbom.cdx.json, a CycloneDX bill of materials of every dependency with its licence. Releases are built in the open pipeline, signed, and can be rebuilt to the same bytes: dist/reproduce.sh.

The hub: usage across a team

The dashboard is one machine’s. The hub is where many machines’ counts meet: an OpenTelemetry Collector that Otōto and Claude Code both send to, Prometheus keeping what arrives, and Grafana with two dashboards. Only counts reach it, never questions, code or answers (Metrics).

It is three upstream images and their configuration, in deploy/ of the source. Nothing of Otōto’s runs in the hub, and nothing in Otōto needs one: any OpenTelemetry collector you already run will do, and the two dashboards can be loaded into a Grafana of your own.

What we run: the kustomize manifests, on a single k3s node, which is where our own machines report. The compose file and the Helm chart are the same configuration, and on 11 October 2026 both render clean (docker compose config, helm lint, helm template); we have not run those two for this page. Collector 0.161.0, Prometheus 3.15.0, Grafana 13.2.2.

What it shows

Otōto and Claude Code, Grafana’s home page, for a time range and a team:

  • At a glance: what Claude Code cost and how many tokens it used, Otōto’s calls, the small model’s tokens, what a paid fallback cost, and the calls that need a second look.
  • Claude Code: cost by model, tokens by type, Claude’s tokens by MCP server, and cost by team.
  • Otōto: calls by tool, the small model’s tokens by server, how long a delegated call takes (median and 90th percentile, by tool), outcomes, the claims in answers and how many were not backed, fallbacks, the tokens handed back to Claude, and the turns a delegated call takes.
  • People: usage person by person, and calls by person and repository.

Otōto fleet:

  • Right now: the servers that reported in the last half hour, the developers and machines behind them, the versions in use, how many are managed by an organisation, and who was seen this week.
  • What is running: by version, team and operating system, with the plugins loaded, where each server’s settings come from, and the model servers they use.
  • Over time, and a table of every server.

Run it

Clone the source for the files:

git clone https://gitlab.com/handmadedigital/projects/ototo.git && cd ototo
WhereCommand
One machine, with Docker or Podmandocker compose -f deploy/hub/compose.yaml up -d
A cluster, with Helmhelm install hub deploy/helm/ototo-hub --namespace ototo-hub --create-namespace
A single-node k3s, as we run itkubectl apply -k deploy/k3s
ServicePortWhat
collector4318 (HTTP), 4317 (gRPC)OTLP in, from Otōto and Claude Code
grafana3000the dashboards. First login admin / admin, and Grafana then asks for a new password
prometheus9090queries and raw series

Who can reach them differs, and matters, since Prometheus has no login:

  • Compose publishes the collector and Grafana on the machine, and Prometheus on 127.0.0.1 only.
  • The chart keeps all three inside the cluster (ClusterIP) until you say otherwise: --set collector.service.type=LoadBalancer, or an ingress of your own in front of 4318. Grafana is reached with kubectl -n ototo-hub port-forward svc/grafana 3000:3000. One release a namespace: the services have fixed names, which the configuration refers to. values.yaml has the images, Prometheus’s retention, the volumes’ sizes and storage class, and each service’s type and resources.
  • The kustomize manifests give all three a LoadBalancer, which on k3s binds the ports on the node itself: for a network you trust.

Prometheus keeps 30 days, 4 GB at most. Grafana’s volume holds its users and settings.

Point machines at it

Otōto, in ~/.config/ototo/config.toml:

otlp_endpoint = "http://<the hub>:4318"
otlp_attributes = "team=platform"

Claude Code, in the env block of ~/.claude/settings.json:

"env": {
  "CLAUDE_CODE_ENABLE_TELEMETRY": "1",
  "OTEL_METRICS_EXPORTER": "otlp",
  "OTEL_EXPORTER_OTLP_PROTOCOL": "http/protobuf",
  "OTEL_EXPORTER_OTLP_ENDPOINT": "http://<the hub>:4318",
  "OTEL_RESOURCE_ATTRIBUTES": "team=platform"
}

For many machines, ototo managed --collector http://<the hub>:4318 --team platform writes both, as managed settings to push to each (For organisations).

Use the same team in both: the dashboard’s Team filter covers both sets of metrics. Claude Code puts the address of its Claude account (user.email) on every metric, and Otōto sends the same address, read from the account Claude Code is signed in with, so the By person table matches the two with no setting per person. user.email=<address> in otlp_attributes sends another, and user.email= none.

On a Mac, give the hub an IP address or a DNS name, not a .local one: where the host has no IPv6 address, macOS waits five seconds for one on every lookup of a .local name, which is as long as Otōto waits to connect. A line in /etc/hosts also cures it.

What arrives

Prometheus nameFromLabels worth knowing
claude_code_cost_usage_USD_totalClaude Codemodel, user_email, team, session_id
claude_code_token_usage_tokens_totalClaude Codetype (input, output, cacheRead, cacheCreation), model, mcp_server_name
claude_code_session_count_total, claude_code_active_time_seconds_totalClaude Code
ototo_calls_totalOtōtotool, outcome, repo, model, endpoint
ototo_local_tokens_totalOtōtotype, model, endpoint
ototo_duration_seconds_bucketOtōtotool, delegated
ototo_cost_usage_USD_total, ototo_fallbacks_total, ototo_claims_total, ototo_turns_total, ototo_reply_chars_totalOtōto

Resource attributes become labels too: job (ototo or claude-code), host_name and user_name (Otōto), user_email, user_account_uuid and organization_id (Claude Code), and anything set with team=…. Metrics has every metric Otōto sends, with its attributes.

Worth knowing before you rely on it

  • People are identifiable. E-mail addresses and account ids are kept, so that usage can be compared person by person. To report without them, add the attributes/people processor sketched in the collector’s configuration (deploy/k3s/collector/config.yaml): it deletes the address and hashes the account ids.
  • Nothing here is encrypted or behind a login, but Grafana. The collector takes plain HTTP, and Prometheus answers anyone who can reach it. On anything but a network you trust, keep the services inside the cluster and put an ingress with TLS, and whatever sign-in you use, in front of the collector and Grafana.
  • Every process counts from zero. Otōto runs a process a session, and Claude Code sends deltas, which the collector turns into running totals, so series come and go. The collector serves OpenMetrics, which carries each counter’s start time; Prometheus records a 0 there (created-timestamp-zero-ingestion), so that a process’s first call is counted, and the dashboards count exactly over a range, not by extrapolating (promql-extended-range-selectors). Both are set in the files here; a Prometheus of your own needs them too.
  • Events are not stored. Claude Code’s log events, if you turn them on, reach the collector’s log only. Add Loki to keep them.
  • The dashboards are files. An edit made in Grafana is lost on restart: export the JSON and replace deploy/k3s/grafana/ototo.json or fleet.json, then sh deploy/sync.sh to give the chart its copy.