mlx-lm: Generating Tokens From Python, Not the CLI
9 min read · updated August 11, 2026
mlx_lm.generate on the command line is a thin wrapper over four functions. Calling them directly gets you streaming, per-request sampling, a reusable prompt cache and the timing figures as numbers rather than as printed text.
Install and load
pip install mlx-lm, on a native arm64 Python 3.10 or newer, on macOS 14.0 or newer. Then:
from mlx_lm import load
model, tokenizer = load("mlx-community/Qwen3-14B-4bit")The argument is a Hub repository id or a local directory. A Hub id is downloaded into ~/.cache/huggingface on first use and read from there afterwards; if you want the download to be explicit and the load to be offline, convert or fetch first and pass a path.
load returns the model and a TokenizerWrapper, not a raw Hugging Face tokenizer. The wrapper adds the detokenizer the streaming loop needs, and passes attribute access through, so tokenizer.apply_chat_template and tokenizer.eos_token_id work as you expect.
The chat template is not optional
An instruction-tuned model was trained on a specific set of role markers and will behave noticeably worse without them — rambling past the answer, answering as though continuing your text, ignoring a system prompt. The template lives in the tokenizer, so applying it is one call:
from mlx_lm import generate
messages = [
{"role": "system", "content": "Answer in one sentence."},
{"role": "user", "content": "Why is decode slower than prefill?"},
]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True)
text = generate(model, tokenizer, prompt=prompt, max_tokens=256, verbose=True)add_generation_prompt=True appends the opening marker for the assistant turn, which is what tells the model it is its go. Without it the model frequently starts a new user turn instead.
Some newer families take extra template arguments — a reasoning-mode toggle is the common one — and those go through as keyword arguments to apply_chat_template. If output arrives wrapped in unexpected thinking tags, that is the switch to look for, and the template rendered as a plain string will show you what the model is actually being handed.
Streaming, and what each chunk carries
generate is a loop over stream_generate that accumulates the text, so anything you want other than the final string means using the generator directly:
from mlx_lm import stream_generate
for response in stream_generate(model, tokenizer, prompt, max_tokens=512):
print(response.text, end="", flush=True)
print()
print("prompt:", response.prompt_tokens, "tokens at",
round(response.prompt_tps, 1), "tok/s")
print("generation:", response.generation_tokens, "tokens at",
round(response.generation_tps, 1), "tok/s")
print("peak memory:", round(response.peak_memory, 2), "GB")
print("finish reason:", response.finish_reason)Each yielded GenerationResponse carries text (the newly detokenized segment, which may be empty mid-multibyte-character), token, logprobs, from_draft, prompt_tokens, prompt_tps, generation_tokens, generation_tps, peak_memory and finish_reason. The last is null until the run ends and then "length" or "stop" — the same distinction between hitting your token budget and the model choosing to stop that every hosted API exposes, and the one you want to branch on before parsing the output as JSON.
generation_tps is the number to quote when anyone asks how fast a model runs on your machine. It is measured on the run that just happened, on your hardware, at your context length, which is more than can be said for any figure derived from a spec sheet — including the ceiling this cluster derives from memory bandwidth, which is an upper bound and not a prediction.
Sampling is an object now
There is no temp keyword on generate. Sampling is a callable built by make_sampler and passed in, which is what lets the same model serve two requests with different settings without reloading:
from mlx_lm.sample_utils import make_sampler
sampler = make_sampler(temp=0.7, top_p=0.9, min_p=0.0, top_k=0)
text = generate(model, tokenizer, prompt=prompt,
max_tokens=512, sampler=sampler)make_sampler takes temp, top_p, min_p, min_tokens_to_keep, top_k and the three XTC parameters. A temperature of exactly 0 short-circuits the whole chain to an argmax, so passing a top-p alongside temp=0 does nothing at all — a real source of confusion when someone sets both and sees no effect from either. The general behaviour of these knobs is in sampling parameters.
Anything not listed on generate or stream_generate is forwarded to generate_step, which is where max_kv_size, prompt_cache, prefill_step_size, kv_bits, kv_group_size and logits_processors live. That forwarding is why the documented signature looks so short, and why a mistyped keyword is silently accepted rather than raising — it is passed down the chain until something ignores it.
One default worth overriding on purpose: max_tokens is 256. A model that appears to stop mid-sentence in your first script has usually not stopped at all, and finish_reason of "length" on the last response tells you so directly.
The KV cache is yours to manage
Two arguments matter on a memory-constrained machine, and both are passed straight through.
max_kv_sizecaps the cache, evicting the oldest entries once the conversation exceeds it. It bounds memory at the cost of the model forgetting the start of a long chat, which is the right trade for a machine that would otherwise start swapping and the wrong one for a document-summarisation prompt.kv_bitsquantizes the cache itself, withkv_group_sizedefaulting to 64 andquantized_kv_startsetting the token offset after which quantization kicks in. On a long context the cache can exceed the weights, so this is a larger lever than it sounds — the arithmetic is worked through in how much RAM a 70B needs on a Mac.
A cache can also be built once and reused. prompt_cache takes the structure returned by make_prompt_cache, so a long fixed system prompt is prefilled once and every subsequent request starts from it:
from mlx_lm.models.cache import make_prompt_cache
cache = make_prompt_cache(model)
# first call pays prefill for the shared prefix and keeps it
generate(model, tokenizer, prompt=shared_prefix + first_question,
prompt_cache=cache, max_tokens=256)
# subsequent calls reuse it; only the new tokens are prefilled
generate(model, tokenizer, prompt=second_question,
prompt_cache=cache, max_tokens=256)This is the same idea as a hosted provider’s prompt caching, with two differences worth internalising. It costs memory rather than money, at the per-token rate the architecture dictates — so a long shared prefix is gigabytes you are choosing to hold. And the cache object is mutable and stateful: it accumulates every token that passes through it, so reusing one cache across two unrelated conversations leaks the first into the second. One cache per conversation, dropped when the conversation ends. The general mechanism is covered in the KV cache.
Two smaller controls round out the loop. logits_processors takes a list of callables applied to the logits before sampling, which is the hook for banning tokens or forcing a grammar without touching the sampler. And extra_eos_token — a flag on the CLI, a keyword on the library path — adds stop tokens beyond the ones in the tokenizer configuration, which is the fix for a model that keeps generating past its own turn because its converted configuration lost an end-of-turn marker.