Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

mistral.rs: a model server in one Rust binary

mistral.rs serves a model straight from its Hugging Face repository, quantising it as it loads, on a Mac (Metal) or on Linux with or without a GPU.

What we ran (10 October 2026): mistral.rs 0.9.4, the prebuilt Metal binary of its GitHub release, on a Mac with an M3 Max and 128 GB, serving Qwen3.5 9B quantised to four bits as it loaded, with Otōto 2026.19. ototo init’s tool-call test answered in 1.5 seconds and ototo doctor found nothing wrong. Two questions about ripgrep’s source were answered rightly, in 305 and 664 seconds. That is slow: the same model at four bits on llama.cpp, on the same machine, took 16 and 79 seconds. We have not found out why.

1. Install

mistral.rs’s installer fetches the prebuilt binary for your machine:

curl -fsSL https://mistralrs.dev/install.sh | sh

It puts the binary in ~/.mistralrs, links it into ~/.local/bin, and adds a line to your shell’s startup files. We read the script and ran the same release’s binary without it. To build from source instead, a Mac needs Xcode’s Metal Toolchain first (xcodebuild -downloadComponent MetalToolchain): without it the build stops at “Compiling metal -> air failed”.

2. and 3. Fetch a model and start the server

The server fetches the model by its Hugging Face name when it starts:

mistralrs serve -m Qwen/Qwen3.5-9B --isq 4 --max-model-len 65536 --host 127.0.0.1 -p 1234
  • -m names the repository. This is the model’s full weights (some 18 GB for this one), not a ready-quantised file.
  • --isq 4 quantises it to four bits as it loads, which took about a minute here.
  • --max-model-len 65536 is the context.
  • --host 127.0.0.1 keeps the server to this machine. Port 1234 is also LM Studio’s: run one of the two there.

4. Point Otōto at it

ototo init

looks for a server on port 1234 among others. For another address, or to be sure which server it means: ototo init --base-url http://127.0.0.1:1234/v1.

5. Check

ototo doctor

It should say serves the model and calls tools. It says nothing of the context: this server does not report one.