Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

vLLM: a model server a team shares

vLLM serves one model to many people from a machine with a GPU. It is what we run Otōto’s small model on, and what its measurements are made with. Use it when a team shares a server; for one laptop, another page is shorter.

What we ran (10 October 2026): Qwen3.8 27B at four bits on vLLM, on a small server without a datacentre GPU, with Otōto 2026.19. ototo init’s tool-call test answered in 1.9 seconds; the server gave the model a context of 128,000 tokens; two questions about ripgrep’s source were answered rightly in 46 and 61 seconds. On our twenty-question suite this setup answers 19 rightly, in 4 to 45 seconds a question.

1. Install

As vLLM’s own quick start says: a Python package on Linux, or its container image. We run the image, pinned to a version and not latest: a restart otherwise changes vLLM under you.

2. and 3. Fetch a model and start the server

vLLM fetches the model from Hugging Face when it starts:

vllm serve Qwen/Qwen3.8-27B-FP8 \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen3 \
  --max-model-len 65536 --enable-prefix-caching
  • --enable-auto-tool-choice --tool-call-parser qwen3_coder are what make tool calls work: the parser turns the model’s own way of writing a tool call into the API’s tool_calls. Without them every answer is prose. Another model wants another parser: vLLM’s list.
  • --max-model-len 65536: at least 64,000 tokens of context.
  • --enable-prefix-caching: each turn’s prompt starts with the last one’s, so only the new part is computed.

It listens on port 8000. dist/vllm/README.md has more of what we learned running it: a chat template for coding clients, with a script that checks it; a Kubernetes example; and what to do when the server’s prompt cache goes bad.

4. Point Otōto at it

On each developer’s machine:

ototo init --base-url https://<your server>/v1

Put TLS in front of the server (a reverse proxy such as Caddy or nginx) and give its https:// address: over plain HTTP the code the model reads can be read on the network between, and ototo doctor says so.

For a whole team, ototo managed writes the settings once for every user, and can make the organisation’s servers the only ones code may go to: see For organisations.

5. Check

ototo doctor

The Model lines should say serves the model, a context of 64,000 tokens or more, calls tools, and that the server answers Otōto’s first turn with a tool call from its prompt cache and fresh. The last is the check for a prompt cache gone bad.