Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

MLX: a model server on a Mac

MLX is Apple’s framework for running models on Apple silicon, and mlx-lm serves one over the API Otōto speaks. On a Mac it was the quickest of the servers we tried.

What we ran (10 October 2026): mlx-lm 0.32.0 on a Mac with an M3 Max and 128 GB, serving Qwen3.5 9B at four bits (5.6 GB), with Otōto 2026.19. ototo init’s tool-call test answered in 2.8 seconds. Two questions about ripgrep’s source were answered rightly, in 27 and 57 seconds.

1. Install

uv venv ~/.venvs/mlx --python 3.12
uv pip install --python ~/.venvs/mlx/bin/python mlx-lm

2. Fetch a model

A model is a folder of files in MLX’s own format. The one we ran:

hf download lmstudio-community/Qwen3.5-9B-MLX-4bit --local-dir ~/models/Qwen3.5-9B-MLX-4bit

(hf is Hugging Face’s command-line tool. Over a poor link it stalled for us, and we fetched the same files one by one with curl -L -C -.)

3. Start the server

~/.venvs/mlx/bin/python -m mlx_lm server --model ~/models/Qwen3.5-9B-MLX-4bit --host 127.0.0.1 --port 8081

Port 8081 is one ototo init looks on.

4. Point Otōto at it

ototo init --base-url http://127.0.0.1:8081/v1 --model ~/models/Qwen3.5-9B-MLX-4bit

Give --model the folder you started the server with: the server answers as the model it was started with under that name.

5. Check

ototo doctor

It should say serves the model and calls tools.

If it says the server did not answer with a model list, while questions work: the server’s list of models reads your Hugging Face cache, and fails when the cache holds files it cannot read. Ours did: the cache is on an exFAT drive, where macOS leaves a ._ file of its own beside each file, and with those deleted from the cache the list answered again. Otōto 2026.19’s doctor calls the server down meanwhile; later releases ask the server for the model instead, and say the list failed as a warning. ototo init with --model, as above, does not need the list.

A larger model

The install guide’s own setup for a Mac is Bonsai 2 27B on mlx-vlm: the model our measurements are of on a Mac, at 20 of 20 on our twenty questions, and at several minutes a question on this machine. Its commands are in dist/INSTALL.md, “Pick a model”.