vLLM: a model server a team shares
vLLM serves one model to many people from a machine with a GPU. It is what we run Otōto’s small model on, and what its measurements are made with. Use it when a team shares a server; for one laptop, another page is shorter.
What we ran (10 October 2026): Qwen3.8 27B at four bits on vLLM, on a small server without a datacentre GPU,
with Otōto 2026.19. ototo init’s tool-call test answered in 1.9 seconds; the server gave the model a context of
128,000 tokens; two questions about ripgrep’s source were answered rightly in 46 and 61 seconds. On our
twenty-question suite this setup answers 19 rightly, in 4 to 45 seconds a question.
1. Install
As vLLM’s own quick start says: a Python
package on Linux, or its container image. We run the image, pinned to a version and not latest: a restart
otherwise changes vLLM under you.
2. and 3. Fetch a model and start the server
vLLM fetches the model from Hugging Face when it starts:
vllm serve Qwen/Qwen3.8-27B-FP8 \
--enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen3 \
--max-model-len 65536 --enable-prefix-caching
--enable-auto-tool-choice --tool-call-parser qwen3_coderare what make tool calls work: the parser turns the model’s own way of writing a tool call into the API’stool_calls. Without them every answer is prose. Another model wants another parser: vLLM’s list.--max-model-len 65536: at least 64,000 tokens of context.--enable-prefix-caching: each turn’s prompt starts with the last one’s, so only the new part is computed.
It listens on port 8000. dist/vllm/README.md has more of what we learned running
it: a chat template for coding clients, with a script that checks it; a Kubernetes example; and what to do when the
server’s prompt cache goes bad.
4. Point Otōto at it
On each developer’s machine:
ototo init --base-url https://<your server>/v1
Put TLS in front of the server (a reverse proxy such as Caddy or nginx) and give its https:// address: over plain
HTTP the code the model reads can be read on the network between, and ototo doctor says so.
For a whole team, ototo managed writes the settings once for every user, and can make the organisation’s servers
the only ones code may go to: see For organisations.
5. Check
ototo doctor
The Model lines should say serves the model, a context of 64,000 tokens or more, calls tools, and that the
server answers Otōto’s first turn with a tool call from its prompt cache and fresh. The last is the check for a
prompt cache gone bad.