Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Settings

Otōto’s settings are one file, ~/.config/ototo/config.toml, which ototo init writes for you. Most people set the model server and nothing else.

  • A variable wins over the file. Every setting is also an environment variable: its name in capitals with OTOTO_ in front (verify is OTOTO_VERIFY). OTOTO_CONFIG points at another file.
  • A session reads the file when it starts. A change reaches the sessions of your agent started after it. ototo serve says on stderr which file it read, and names any key it does not know.
  • An organisation’s file sits under yours, and can enforce some settings over it: For organisations.
  • ototo doctor checks the file and each model server; the dashboard shows what the file says now.

dist/config.toml.example is a file to start from.

The model server

base_url = "http://<host>:<port>/v1"
model = "<the name the server gives the model>"
SettingDefaultWhat it is
base_urlhttp://127.0.0.1:8080/v1the server’s OpenAI-compatible address
modelthe model’s name there: curl -s <base_url>/models shows it
api_keya key, where the server wants one
endpointsseveral servers, in the order to try them (below)

The server needs OpenAI-style tool calls, 32,000 tokens of context and 4,000 of output; thinking is switched off in each request. A model server: quick starts has the flags for each server we have run.

Several servers

List several and each request goes to the first that answers. A spot GPU that has been reclaimed, a workstation that is switched off, or a gateway’s error page hands over in seconds, and a server that failed waits at the back of the queue for a minute, so that the next requests do not wait on it. A hosted model makes a last resort: paid for by the token, but always there.

[[endpoints]]
name = "gpu"
base_url = "http://<gpu host>/v1"
model = "Qwen/Qwen3.8-27B-FP8"

[[endpoints]]
name = "workstation"
base_url = "http://<workstation>:8080/v1"
model = "RedHatAI/Qwen3.8-27B-INT4"

[[endpoints]]
provider = "anthropic"          # https://api.anthropic.com unless base_url says otherwise
model = "claude-haiku-4-5"      # key from ANTHROPIC_API_KEY, or api_key / api_key_env

The other settings stay at the top of the file and apply to every server. A reply that another server answered says so in its footer (via haiku (gpu, workstation down) · $0.031), and the run log records which answered, why the others failed and what a paid model cost: list prices for the Claude models, or price_input and price_output on the endpoint, in US dollars per million tokens. OTOTO_BASE_URL in the environment still pins one server.

How far a question may go

SettingDefaultWhat it is
max_turns8 (ototo init writes 36)the small model’s turns for one question
max_input_tokens400000what one question may send the model over all its turns; past it the next turn is its last. 0 for no such budget
max_output_tokens4096the most the model may write in one turn
temperature0.2
thinkingfalsethe model’s own reasoning before each turn
tool_choiceautosee below

An answer that ran out of turns or tokens says so, and is the small model’s best by then. A long question in several parts reaches the budget most often, and does better asked apart.

tool_choice. With auto, no tool_choice is sent: the model may answer in prose, and is nudged back to its tools. required is for servers with constrained decoding (vLLM): every turn must call a tool, with arguments held to the tool’s schema. That rules out mangled parameter names, but with Qwen3.8 on vLLM it also made the model repeat its last call in place of changing course, a search it could not narrow five turns running, where auto answered the same question in eight. Leave it on auto, and try required only if a model keeps mangling its arguments. A server that refuses tool_choice (llama.cpp) has it dropped after the first refusal.

Checking the answers

SettingDefaultWhat it is
verifyoff, unless ototo init turns it on for your modela second pass by the small model over each ask and locate answer’s claims, in a fresh request with the code each cites. One more request an answer: 2 to 3 s on an idle GPU server. Leave it off on a slow or busy one
excerpts2how many of an answer’s citations get their code attached, under “Key code”: the enclosing function up to 25 lines, else the cited line and three each side. 0 turns it off
findonthe small model’s ranked search over declarations. off removes it

The checks that need no model, of every name and value a claim states against the code it cites, are always on: How it works.

What your agent is offered

SettingDefaultWhat it is
toolsall of themthe tools on the menu, comma-separated: "ask,read,search". Every tool’s definition is carried in your agent’s context on each turn, so a shorter menu costs less; Otōto’s instructions then name only those
hide_model_toolsoffwhile no model server answers, take ask, locate, callers and edit off the menu, and put them back when one does
routinesoffoffer the routine tool (experimental)

hide_model_tools. Without it, the tools that need a model stay listed while none answers, and fail at once, saying what works meanwhile. With it, the session that finds every server gone tells your agent its list has changed, asks once a minute (a request of a few tokens) whether a server is back, and leaves a note for the prompt hook. Two things to know first. Run ototo init after setting it, since the prompt hook must then ask Otōto what to say (ototo doctor says when that is still to do), and a process of Otōto’s starts with each prompt. And a list that changes in the middle of a session costs Claude Code its prompt cache, once each way. Tried with Claude Code; not yet with OpenCode.

How much a read returns

SettingDefaultWhat it is
read_max_lines2000the most lines one read call returns
read_target_lines1000the most for one address in it
read_whole_file1000a file up to this long comes back whole when read by its path; a longer one as its outline
read_whole_documentas read_whole_filethe same for a Markdown document: lower, a long one comes back as its sections, to read by name (docs/guide.md#install)

A longer range is cut, with a note saying where to continue. The small model’s own reads stay at 250 lines a call, whatever these say: its context is the one that fills.

Edits

SettingDefaultWhat it is
check_cmda build or a type check for edit: cargo check, ./gradlew compileJava, npx tsc --noEmit
check_timeout300seconds it may take

With a check_cmd, the small model gets a check tool. Its pending edits are written, the command runs in the repository, and every file is put back afterwards; one that someone else changed meanwhile is left alone, and reported. The verdict is in the reply’s footer. It is the one thing Otōto runs, and Otōto takes it only from its settings and its environment, never from the files of a repository it reads. (A project’s .mcp.json, which Claude Code asks you to approve, can set OTOTO_CHECK_CMD in the registration’s environment.)

Outside the repository

SettingDefaultWhat it is
dependency_cachesofftrue for Maven’s and Gradle’s caches (~/.m2/repository, ~/.gradle/caches), or a list of directories. Otōto then reads the archives there, and Maven’s POMs, by absolute path: read <jar>!<class>. Nothing else, no listing, no search, never a write
read_secretsofflet Otōto read files it takes for secrets by their names: environment files, private keys, credentials files, Terraform state

Both are off because each widens what Otōto reads. dependency_caches is the one case where it reads outside the directories a session was given; a cache that is a home directory or a whole disk is ignored.

Plugins

SettingDefaultWhat it is
plugins~/.config/ototo/pluginswhere plugins load from: directories, :-separated
plugin_signers~/.config/ototo/allowed_signersthe keys whose signature lets a plugin load, in git’s allowed-signers format
allow_unsigned_pluginsoffload unsigned plugins: for writing one of your own
plugin_grantsthe forge plugins that may reach their hosts: ototo plugins grant gitlab writes it
plugin_settingsa plugin’s own: for gitlab, which token goes to which host, project or group

The plugins has the commands; ototo plugins help gitlab the forge plugin’s settings.

The run log and the dashboard

SettingDefaultWhat it is
run_log~/.ototo/runs.jsonlwhere each call is logged, for the dashboard. off turns it off
run_log_max_mb20the size at which the log is rotated
run_log_keep3how many rotated files are kept
traceonthe live trace beside the log, which ototo tail and the dashboard’s “Running now” follow

The log holds questions, the small model’s steps and replies, and stays on your machine.

Reporting

otlp_endpoint, otlp_headers, otlp_attributes and otlp_interval send counts, never questions, code or answers, to an OpenTelemetry collector of yours: Metrics.

Updates

SettingDefaultWhat it is
update_urlhttps://ototo.sh/betathe download channel ototo update and ototo plugins ask, and only when you run them
update_tokena bearer token, for a private channel

ototo update installs a newer release from the channel, or, when the channel does not answer, the newest package for this machine in ~/Downloads. Either way it checks the release’s checksums against their signature by Otōto’s release key, which is built into the binary, and the package against its checksum, then runs the package’s install.sh, which keeps your settings. Only a newer release is installed, so a channel cannot move anyone back. --check only says whether there is one; --dry-run downloads and checks.

For an organisation

enforced and min_version belong in the organisation’s file, with the settings it wants everyone to have: For organisations.