MLX: a model server on a Mac
MLX is Apple’s framework for running models on Apple silicon, and mlx-lm serves one over the API Otōto speaks.
On a Mac it was the quickest of the servers we tried.
What we ran (10 October 2026): mlx-lm 0.32.0 on a Mac with an M3 Max and 128 GB, serving Qwen3.5 9B at four
bits (5.6 GB), with Otōto 2026.19. ototo init’s tool-call test answered in 2.8 seconds. Two questions about
ripgrep’s source were answered rightly, in 27 and 57 seconds.
1. Install
uv venv ~/.venvs/mlx --python 3.12
uv pip install --python ~/.venvs/mlx/bin/python mlx-lm
2. Fetch a model
A model is a folder of files in MLX’s own format. The one we ran:
hf download lmstudio-community/Qwen3.5-9B-MLX-4bit --local-dir ~/models/Qwen3.5-9B-MLX-4bit
(hf is Hugging Face’s command-line tool. Over a poor link it stalled for us, and we fetched the same files one by
one with curl -L -C -.)
3. Start the server
~/.venvs/mlx/bin/python -m mlx_lm server --model ~/models/Qwen3.5-9B-MLX-4bit --host 127.0.0.1 --port 8081
Port 8081 is one ototo init looks on.
4. Point Otōto at it
ototo init --base-url http://127.0.0.1:8081/v1 --model ~/models/Qwen3.5-9B-MLX-4bit
Give --model the folder you started the server with: the server answers as the model it was started with under
that name.
5. Check
ototo doctor
It should say serves the model and calls tools.
If it says the server did not answer with a model list, while questions work: the server’s list of models
reads your Hugging Face cache, and fails when the cache holds files it cannot read. Ours did: the cache is on an
exFAT drive, where macOS leaves a ._ file of its own beside each file, and with those deleted from the cache the
list answered again. Otōto 2026.19’s doctor calls the server down meanwhile; later releases ask the server for the
model instead, and say the list failed as a warning. ototo init with --model, as above, does not need the
list.
A larger model
The install guide’s own setup for a Mac is Bonsai 2 27B on mlx-vlm: the model our measurements are of on a Mac,
at 20 of 20 on our twenty questions, and at several minutes a question on this machine. Its commands are in
dist/INSTALL.md, “Pick a model”.