Settings
Otōto’s settings are one file, ~/.config/ototo/config.toml, which ototo init writes for you. Most people set the
model server and nothing else.
- A variable wins over the file. Every setting is also an environment variable: its name in capitals with
OTOTO_in front (verifyisOTOTO_VERIFY).OTOTO_CONFIGpoints at another file. - A session reads the file when it starts. A change reaches the sessions of your agent started after it.
ototo servesays on stderr which file it read, and names any key it does not know. - An organisation’s file sits under yours, and can enforce some settings over it: For organisations.
ototo doctorchecks the file and each model server; the dashboard shows what the file says now.
dist/config.toml.example
is a file to start from.
The model server
base_url = "http://<host>:<port>/v1"
model = "<the name the server gives the model>"
| Setting | Default | What it is |
|---|---|---|
base_url | http://127.0.0.1:8080/v1 | the server’s OpenAI-compatible address |
model | the model’s name there: curl -s <base_url>/models shows it | |
api_key | a key, where the server wants one | |
endpoints | several servers, in the order to try them (below) |
The server needs OpenAI-style tool calls, 32,000 tokens of context and 4,000 of output; thinking is switched off in each request. A model server: quick starts has the flags for each server we have run.
Several servers
List several and each request goes to the first that answers. A spot GPU that has been reclaimed, a workstation that is switched off, or a gateway’s error page hands over in seconds, and a server that failed waits at the back of the queue for a minute, so that the next requests do not wait on it. A hosted model makes a last resort: paid for by the token, but always there.
[[endpoints]]
name = "gpu"
base_url = "http://<gpu host>/v1"
model = "Qwen/Qwen3.8-27B-FP8"
[[endpoints]]
name = "workstation"
base_url = "http://<workstation>:8080/v1"
model = "RedHatAI/Qwen3.8-27B-INT4"
[[endpoints]]
provider = "anthropic" # https://api.anthropic.com unless base_url says otherwise
model = "claude-haiku-4-5" # key from ANTHROPIC_API_KEY, or api_key / api_key_env
The other settings stay at the top of the file and apply to every server. A reply that another server answered says
so in its footer (via haiku (gpu, workstation down) · $0.031), and the run log records which answered, why the
others failed and what a paid model cost: list prices for the Claude models, or price_input and price_output
on the endpoint, in US dollars per million tokens. OTOTO_BASE_URL in the environment still pins one server.
How far a question may go
| Setting | Default | What it is |
|---|---|---|
max_turns | 8 (ototo init writes 36) | the small model’s turns for one question |
max_input_tokens | 400000 | what one question may send the model over all its turns; past it the next turn is its last. 0 for no such budget |
max_output_tokens | 4096 | the most the model may write in one turn |
temperature | 0.2 | |
thinking | false | the model’s own reasoning before each turn |
tool_choice | auto | see below |
An answer that ran out of turns or tokens says so, and is the small model’s best by then. A long question in several parts reaches the budget most often, and does better asked apart.
tool_choice. With auto, no tool_choice is sent: the model may answer in prose, and is nudged back to its
tools. required is for servers with constrained decoding (vLLM): every turn must call a tool, with arguments held
to the tool’s schema. That rules out mangled parameter names, but with Qwen3.8 on vLLM it also made the model
repeat its last call in place of changing course, a search it could not narrow five turns running, where auto
answered the same question in eight. Leave it on auto, and try required only if a model keeps mangling its
arguments. A server that refuses tool_choice (llama.cpp) has it dropped after the first refusal.
Checking the answers
| Setting | Default | What it is |
|---|---|---|
verify | off, unless ototo init turns it on for your model | a second pass by the small model over each ask and locate answer’s claims, in a fresh request with the code each cites. One more request an answer: 2 to 3 s on an idle GPU server. Leave it off on a slow or busy one |
excerpts | 2 | how many of an answer’s citations get their code attached, under “Key code”: the enclosing function up to 25 lines, else the cited line and three each side. 0 turns it off |
find | on | the small model’s ranked search over declarations. off removes it |
The checks that need no model, of every name and value a claim states against the code it cites, are always on: How it works.
What your agent is offered
| Setting | Default | What it is |
|---|---|---|
tools | all of them | the tools on the menu, comma-separated: "ask,read,search". Every tool’s definition is carried in your agent’s context on each turn, so a shorter menu costs less; Otōto’s instructions then name only those |
hide_model_tools | off | while no model server answers, take ask, locate, callers and edit off the menu, and put them back when one does |
routines | off | offer the routine tool (experimental) |
hide_model_tools. Without it, the tools that need a model stay listed while none answers, and fail at once,
saying what works meanwhile. With it, the session that finds every server gone tells your agent its list has
changed, asks once a minute (a request of a few tokens) whether a server is back, and leaves a note for the prompt
hook. Two things to know first. Run ototo init after setting it, since the prompt hook must then ask Otōto what to
say (ototo doctor says when that is still to do), and a process of Otōto’s starts with each prompt. And a list
that changes in the middle of a session costs Claude Code its prompt cache, once each way. Tried with Claude Code;
not yet with OpenCode.
How much a read returns
| Setting | Default | What it is |
|---|---|---|
read_max_lines | 2000 | the most lines one read call returns |
read_target_lines | 1000 | the most for one address in it |
read_whole_file | 1000 | a file up to this long comes back whole when read by its path; a longer one as its outline |
read_whole_document | as read_whole_file | the same for a Markdown document: lower, a long one comes back as its sections, to read by name (docs/guide.md#install) |
A longer range is cut, with a note saying where to continue. The small model’s own reads stay at 250 lines a call, whatever these say: its context is the one that fills.
Edits
| Setting | Default | What it is |
|---|---|---|
check_cmd | a build or a type check for edit: cargo check, ./gradlew compileJava, npx tsc --noEmit | |
check_timeout | 300 | seconds it may take |
With a check_cmd, the small model gets a check tool. Its pending edits are written, the command runs in the
repository, and every file is put back afterwards; one that someone else changed meanwhile is left alone, and
reported. The verdict is in the reply’s footer. It is the one thing Otōto runs, and Otōto takes it only from its
settings and its environment, never from the files of a repository it reads. (A project’s .mcp.json, which
Claude Code asks you to approve, can set OTOTO_CHECK_CMD in the registration’s environment.)
Outside the repository
| Setting | Default | What it is |
|---|---|---|
dependency_caches | off | true for Maven’s and Gradle’s caches (~/.m2/repository, ~/.gradle/caches), or a list of directories. Otōto then reads the archives there, and Maven’s POMs, by absolute path: read <jar>!<class>. Nothing else, no listing, no search, never a write |
read_secrets | off | let Otōto read files it takes for secrets by their names: environment files, private keys, credentials files, Terraform state |
Both are off because each widens what Otōto reads. dependency_caches is the one case where it reads outside the
directories a session was given; a cache that is a home directory or a whole disk is ignored.
Plugins
| Setting | Default | What it is |
|---|---|---|
plugins | ~/.config/ototo/plugins | where plugins load from: directories, :-separated |
plugin_signers | ~/.config/ototo/allowed_signers | the keys whose signature lets a plugin load, in git’s allowed-signers format |
allow_unsigned_plugins | off | load unsigned plugins: for writing one of your own |
plugin_grants | the forge plugins that may reach their hosts: ototo plugins grant gitlab writes it | |
plugin_settings | a plugin’s own: for gitlab, which token goes to which host, project or group |
The plugins has the commands; ototo plugins help gitlab the forge plugin’s settings.
The run log and the dashboard
| Setting | Default | What it is |
|---|---|---|
run_log | ~/.ototo/runs.jsonl | where each call is logged, for the dashboard. off turns it off |
run_log_max_mb | 20 | the size at which the log is rotated |
run_log_keep | 3 | how many rotated files are kept |
trace | on | the live trace beside the log, which ototo tail and the dashboard’s “Running now” follow |
The log holds questions, the small model’s steps and replies, and stays on your machine.
Reporting
otlp_endpoint, otlp_headers, otlp_attributes and otlp_interval send counts, never questions, code or
answers, to an OpenTelemetry collector of yours: Metrics.
Updates
| Setting | Default | What it is |
|---|---|---|
update_url | https://ototo.sh/beta | the download channel ototo update and ototo plugins ask, and only when you run them |
update_token | a bearer token, for a private channel |
ototo update installs a newer release from the channel, or, when the channel does not answer, the newest package
for this machine in ~/Downloads. Either way it checks the release’s checksums against their signature by Otōto’s
release key, which is built into the binary, and the package against its checksum, then runs the package’s
install.sh, which keeps your settings. Only a newer release is installed, so a channel cannot move anyone back.
--check only says whether there is one; --dry-run downloads and checks.
For an organisation
enforced and min_version belong in the organisation’s file, with the settings it wants everyone to have:
For organisations.