Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

LiteRT-LM: very small models, on very small machines

LiteRT-LM is Google’s runtime for models small enough for a phone or a thin laptop. Its command-line tool has an OpenAI-compatible server, and Otōto’s tool calls work through it. The models it runs are too small to answer questions about code well: use it to see Otōto’s small-model tools working on a machine that can run nothing larger, not for answers to rely on.

What we ran (10 October 2026): LiteRT-LM 0.17.1 on a Mac with an M3 Max, serving Gemma 4 E4B (3.4 GB), with Otōto 2026.19. ototo init’s tool-call test passed, and ototo doctor found nothing wrong. Two questions about ripgrep’s source came back in 110 and 38 seconds, both wrong: the first named the right file and the wrong function. On the same questions a 9-billion-parameter model on llama.cpp was right both times.

1. and 2. Install it and import a model

As LiteRT-LM’s documentation says; litert-lm list then shows the models it has. We ran one imported earlier and did not follow those steps afresh.

3. Start the server

litert-lm serve --host 127.0.0.1 --port 9379

Say --host 127.0.0.1: without it the server listens on every network interface of the machine.

4. Point Otōto at it

ototo init --base-url http://127.0.0.1:9379/v1 --model gemma-4-E4B-it-hf

The server serves every model it has imported, so name the one you mean: litert-lm list has the names. Otōto 2026.19 and earlier do not look on port 9379 by themselves, so the address is named too; later releases find it, and take the first model when none is named.

5. Check

ototo doctor

It says serves the model and calls tools. It says nothing of the context, since this server does not report one: we do not know how long a question it can hold.