Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

LM Studio: a desktop app that also serves

LM Studio is an app for downloading and running models, with a server and a command-line tool, lms, beside it. Use it when you would rather manage models in a window than in a terminal.

What we ran (10 October 2026): LM Studio 0.4.26 on a Mac with an M3 Max and 128 GB, serving Qwen3.5 9B at four bits (a 5.6 GB GGUF file) with a context of 65,536 tokens, with Otōto 2026.19. ototo init’s tool-call test answered in 2.5 seconds and ototo doctor found nothing wrong. Two questions about ripgrep’s source were answered rightly, in 105 and 191 seconds.

1. Install

brew install --cask lm-studio

on a Mac, or the app from lmstudio.ai for macOS, Linux and Windows. lms comes with it. The first lms command starts LM Studio in the background, and can fail once while it starts (“Invalid passkey”): run it again.

2. Fetch a model

In the app, or from a file you already have, which is what we did:

lms import --copy --user-repo lmstudio-community/Qwen3.5-9B-GGUF ~/models/Qwen3.5-9B-Q4_K_M.gguf

(llama.cpp’s page has the command that fetches that file.) lms ls lists the models LM Studio has, by the names it gives them.

3. Load it with enough context, and start the server

lms load qwen3.5-9b --context-length 65536
lms server start --port 1234

lms ps should show the model with 65536 under CONTEXT.

4. Point Otōto at it

ototo init --base-url http://127.0.0.1:1234/v1 --model qwen3.5-9b

Name the model: LM Studio’s server lists every model it has downloaded, loaded or not, and Otōto should use the one you loaded with the long context.

5. Check

ototo doctor

It should say serves the model and calls tools. It says nothing of the context, since this server does not report it: lms ps is the check for that.