mistral.rs: a model server in one Rust binary
mistral.rs serves a model straight from its Hugging Face repository, quantising it as it loads, on a Mac (Metal) or on Linux with or without a GPU.
What we ran (10 October 2026): mistral.rs 0.9.4, the prebuilt Metal binary of its GitHub release, on a Mac with
an M3 Max and 128 GB, serving Qwen3.5 9B quantised to four bits as it loaded, with Otōto 2026.19. ototo init’s
tool-call test answered in 1.5 seconds and ototo doctor found nothing wrong. Two questions about ripgrep’s source
were answered rightly, in 305 and 664 seconds. That is slow: the same model at four bits on llama.cpp, on the
same machine, took 16 and 79 seconds. We have not found out why.
1. Install
mistral.rs’s installer fetches the prebuilt binary for your machine:
curl -fsSL https://mistralrs.dev/install.sh | sh
It puts the binary in ~/.mistralrs, links it into ~/.local/bin, and adds a line to your shell’s startup files.
We read the script and ran the same release’s binary without it. To build from source instead, a Mac needs
Xcode’s Metal Toolchain first (xcodebuild -downloadComponent MetalToolchain): without it the build stops at
“Compiling metal -> air failed”.
2. and 3. Fetch a model and start the server
The server fetches the model by its Hugging Face name when it starts:
mistralrs serve -m Qwen/Qwen3.5-9B --isq 4 --max-model-len 65536 --host 127.0.0.1 -p 1234
-mnames the repository. This is the model’s full weights (some 18 GB for this one), not a ready-quantised file.--isq 4quantises it to four bits as it loads, which took about a minute here.--max-model-len 65536is the context.--host 127.0.0.1keeps the server to this machine. Port 1234 is also LM Studio’s: run one of the two there.
4. Point Otōto at it
ototo init
looks for a server on port 1234 among others. For another address, or to be sure which server it means:
ototo init --base-url http://127.0.0.1:1234/v1.
5. Check
ototo doctor
It should say serves the model and calls tools. It says nothing of the context: this server does not report
one.