LM Studio: a desktop app that also serves
LM Studio is an app for downloading and running models, with a server and a command-line tool, lms, beside it.
Use it when you would rather manage models in a window than in a terminal.
What we ran (10 October 2026): LM Studio 0.4.26 on a Mac with an M3 Max and 128 GB, serving Qwen3.5 9B at four
bits (a 5.6 GB GGUF file) with a context of 65,536 tokens, with Otōto 2026.19. ototo init’s tool-call test
answered in 2.5 seconds and ototo doctor found nothing wrong. Two questions about ripgrep’s source were answered
rightly, in 105 and 191 seconds.
1. Install
brew install --cask lm-studio
on a Mac, or the app from lmstudio.ai for macOS, Linux and Windows. lms comes with it.
The first lms command starts LM Studio in the background, and can fail once while it starts (“Invalid
passkey”): run it again.
2. Fetch a model
In the app, or from a file you already have, which is what we did:
lms import --copy --user-repo lmstudio-community/Qwen3.5-9B-GGUF ~/models/Qwen3.5-9B-Q4_K_M.gguf
(llama.cpp’s page has the command that fetches that file.) lms ls lists the models LM Studio
has, by the names it gives them.
3. Load it with enough context, and start the server
lms load qwen3.5-9b --context-length 65536
lms server start --port 1234
lms ps should show the model with 65536 under CONTEXT.
4. Point Otōto at it
ototo init --base-url http://127.0.0.1:1234/v1 --model qwen3.5-9b
Name the model: LM Studio’s server lists every model it has downloaded, loaded or not, and Otōto should use the one you loaded with the long context.
5. Check
ototo doctor
It should say serves the model and calls tools. It says nothing of the context, since this server does not
report it: lms ps is the check for that.