How it works
A coding agent finds its way round a codebase by opening files and running searches, and every file it opens, every result it scans and every dead end it follows stays in its context for the rest of the session. Each later turn carries all of it again, and has more to sift.
Otōto keeps that exploring out of the agent’s context. The agent hands over a whole question; a small model, on a server you choose, searches and reads in a context of its own, which is thrown away after the call; and the agent gets back a short answer, with the lines that back it checked against the files. For what the agent already knows it wants to see, a set of exact tools answers at once, with no model at all.
The parts
- Your coding agent starts
ototo servein the repository it is working in, and talks to it over MCP (the Model Context Protocol, on standard input and output). Claude Code, OpenCode, Pi and others. - Otōto is one static binary, written in Rust. It reads the repository, changes it only when an edit is asked
for, and
ototo serveopens no port. - A model server runs the small model: vLLM, llama.cpp, MLX, Ollama, LM Studio and others, over the OpenAI-compatible API with tool calls. It can be on the same machine or a team’s GPU server. With none, the direct tools still work (quick starts).
- Plugins read kinds of file Otōto does not know itself, each as its own tool would evaluate it (the plugins).
Two kinds of tool
| Direct | Delegated | |
|---|---|---|
| Tools | read, outline, search, changes, history, replace_all | ask, locate, callers, edit |
| Answered by | Otōto, from the files and from git | the small model, exploring with Otōto’s tools |
| Takes | milliseconds | seconds to minutes |
| For | what the agent already knows it wants: these lines, this declaration, this string | an open question: how something works, where it is decided, who uses it |
The tools has each of them. The delegated ones are where the saving is; the direct ones are what the
agent uses in place of grep, cat and whole-file reads, and they return less: a declaration by name, not the
file it is in.
One ask, step by step
- The agent calls
askwith a whole question: “how is auth wired into the router?”. Up to five independent questions go in one call, and are answered side by side. - Otōto opens a fresh context with the small model: its own prompt, the question, and read-only tools. Nothing of the agent’s conversation goes with it.
- The small model works in turns. It calls
find(the repository’s declarations, ranked against a few words, for when it does not know the name),search(exact text or a regex),outlineandread. Otōto runs each call on your files, through the same checks as for the agent: nothing outside the repository, and never a secrets file. A read gives it at most 250 lines, since its context is the one that fills. The turns end atmax_turns, or when the question has sent the model its budget of input tokens (400,000 unless set), after which the next turn is its last; an answer cut short that way says so. - It answers, and each
path:linecitation in the answer carries a claim: what that line shows. - Otōto checks every claim against the cited code, with no model. The identifiers and the values a claim
names must be in the cited line’s enclosing declaration; a file it names must be the cited file; “X is defined
here” must cite a declaration; a count of callers must match the call sites cited. A claim that fails is listed
under “Not backed by the cited code (check before relying on it)”, with where the missing name does appear.
Passing means “about the right code”, not “true”: with
verify = true, the claims go back to the small model once more, in a fresh request with the code each cites and none of the reasoning, to catch a claim about the right code that says the wrong thing. - The agent gets the answer: a few thousand characters, the citations checked, and the code of the main ones attached (the enclosing function, or the line and those around it), so that it need not read them again. The small model’s context is thrown away.
locate and callers are the same loop with a narrower question and a stricter check: a locate answer must
cite a declaration, and a callers answer’s count must match the call sites it cites, each checked to be that
symbol and not a namesake.
edit hands over a small, mechanical change. The small model proposes it and Otōto returns a diff; nothing is
written unless the call says apply. With a check_cmd in the settings (cargo check, npx tsc --noEmit), the
model can try its change: the pending edits are written, the command runs, and every file is put back afterwards.
Reading code as it is declared
Otōto parses source (tree-sitter does most of it), so a file is its declarations, not only its lines: outline lists them with
their line ranges, read path#Name returns one, read path:137 the function round a line, and changes says which
declarations a diff touched. It reads Java, Kotlin, TypeScript and JavaScript, Rust, Python, Go, C#, shell and SQL
this way, and outlines Terraform, YAML, JSON, TOML, XML, properties, CSV, Dockerfiles, Makefiles and Markdown by
their own names: a key, a section, a target, a column.
The index for find is built in memory the first time it is used, and a file is parsed again only when it changes.
Nothing is written into the repository, and there is no daemon and no database.
Plugins
Some files mean more than their text. A GitLab CI job is its extends chain, its includes and its defaults merged;
a Maven dependency’s version may be set three parents away. A plugin reads one such kind of file and gives Otōto
its outline and its reads, so that read .gitlab-ci.yml#deploy:prod returns the job as GitLab runs it, each line
with the file it came from.
A plugin is a WebAssembly component, run in a sandbox with no network, no filesystem and no environment. It reaches the repository only by asking Otōto, through the same path checks as everything else; each call has an instruction budget and a memory limit; and it loads only if it is signed by a key you trust. One that fails is skipped, and Otōto answers as it would without it. The plugins lists them, and Writing a plugin builds one from nothing.
What leaves the machine
- To the model servers you configure, and nowhere else: the question, and the parts of the repository the small model reads to answer it. If those servers are yours, the code stays with you; if you list a hosted model as a fallback, it goes there when the others do not answer, and the reply says so.
- Counts, if you set a collector: with
otlp_endpoint, how many calls, tokens and seconds, to a collector of yours (Metrics). Never the questions, the code or the answers. - A forge, if you grant a forge plugin: read-only requests to the hosts it declared, for the
forgetool’s questions about a merge request or a pipeline. Nothing until the grant. - The download channel, when you run
ototo updateor one ofototo plugins available,addandupdate: never by itself. - Nothing to us. Otōto sends no usage data, crash reports or identifiers to its makers.
SECURITY.md has all of it: what Otōto reads, writes and runs, and what it never reads (environment files, private keys, credentials).
When no model server answers
A delegated call fails at once and says which servers were tried and which tools need none, so the agent goes on
with search, read and outline. With several servers listed, each request goes to the first that answers, and
one that failed waits at the back of the queue for a minute. ototo doctor checks each of them: reachable, serving
the model, calling tools.
Watching it
ototo tail follows every delegated call as it runs: the question, each turn of the small model with the tool it
called, the claim check and how it ended. ototo ui shows the same as a dashboard, call by call. Both read what
Otōto logs on your machine (~/.ototo/runs.jsonl and the live trace beside it).
12:39:17 ask start Which class validates a new pet, and what does it check?
12:39:22 ask turn 1 find({"query":"validate new pet"}) · 1.6k tok · 4.7s
12:39:24 ask turn 2 read({"addresses":["src/main/java/…/PetValidator.java"]}) · 3.8k tok · 7.1s
12:39:40 ask turn 3 finish({"answer":"The class is PetValidator …"}) · 7.3k tok · 22.9s
12:39:40 ask end 23.0s · 3 turns · 7.3k tok · ok