Skip to content

Mistral's Instruct Prompt Template, and What Raw Completion Mode Skips

9 min read · updated August 11, 2026

A Mistral instruct model was fine-tuned on text with a specific shape: instructions wrapped in [INST] and [/INST], a beginning-of-sequence token at the front, an end-of-sequence token after each answer. Send it text in a different shape and you are asking for a continuation of something it has not seen much of.

The template

For the Mistral 7B instruct series — the release that established the format — the single-turn layout documented on the model card is:

<s>[INST] What is the capital of France? [/INST]

The model then generates the answer and emits </s> when it is done. Four details in that one line are load-bearing:

  • <s> is a token, not two characters. The tokeniser has a dedicated beginning-of-sequence id for it. If you type the literal characters into a string and tokenise the result, you get several ordinary text tokens instead of the one control token, and the model sees something it has never seen. Most tokeniser APIs add BOS automatically; passing it in the string as well gives you two.
  • [INST] and [/INST] are plain text in the original 7B format. They tokenise as ordinary bracket and word pieces. That is why they appear literally in the string rather than as angle-bracket control tokens.
  • The spaces matter. The card’s canonical form has a space after [INST] and before [/INST]. Tokenisers are whitespace-sensitive: " What" and "What" are different tokens, so an inconsistent space shifts every token that follows.
  • </s> closes the answer, not the instruction. It goes after the assistant’s reply, which is what makes multi-turn layout work.

Multi-turn layout

A conversation is the same pattern repeated, with each completed exchange terminated by the end-of-sequence token, and exactly one BOS at the very front of the whole sequence:

<s>[INST] First question [/INST] First answer</s>[INST] Second question [/INST]

Note what is not there. There is no BOS before the second [INST]. There is no role label for the assistant — the answer simply follows [/INST] as bare text. And the original 7B format had no separate slot for a system prompt at all; the convention was to prepend the system text inside the first [INST] block. Later revisions and later models change this, which is the subject of how Mistral handles the system role.

You should not be assembling this by hand. The tokeniser ships the template and can apply it for you, which removes the entire class of whitespace and BOS bugs:

from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-Instruct-v0.3")

messages = [
    {"role": "user", "content": "First question"},
    {"role": "assistant", "content": "First answer"},
    {"role": "user", "content": "Second question"},
]

# add_generation_prompt=True leaves the sequence ready for the model to answer.
ids = tok.apply_chat_template(messages, add_generation_prompt=True)
print(tok.decode(ids))       # inspect exactly what the model will see

That last line is the single most useful debugging step available here. Decode the ids and read them. Almost every “the model is ignoring my instructions” report from someone running weights locally is visible in that output as a doubled BOS, a missing [/INST], or a system message that quietly went nowhere.

What raw completion mode skips

Raw mode means sending your text to the model with no template applied — a bare string, tokenised as-is. The model does what a language model does: continues it. Nothing about that is broken, but three things you were getting for free disappear.

  • The instruction boundary. Without [INST], the model has no marker saying “the preceding is a request and the following is your answer”. Ask “What is the capital of France?” raw and a plausible continuation is another question, or a numbered list of similar quiz items, because that is what follows a question in a lot of text.
  • The instruction-tuned behaviour. Refusals, the helpful-assistant register, the tendency to answer concisely and stop — all of that was trained in on templated data and is only reliably evoked in the shape it was trained in.
  • The stop signal. Templated generation ends at the end-of-sequence token. Raw generation has nothing obvious to stop at, so it runs to max_tokens and you get a wall of continuation.

Raw mode is right for a genuine completion task — finishing a code block, continuing a document, computing perplexity on a corpus, or probing base-model behaviour without the instruct layer in the way. It is the wrong tool for asking a question, and reaching for it because the template was inconvenient is how people conclude a model is worse than it is.

Newer models moved the goalposts

The [INST] format above is the Mistral 7B instruct format, and it is what a search for this question overwhelmingly returns. It is not universal across Mistral’s catalogue. Later models — the ones using the Tekken tokeniser, and anything supporting tool calling — carry their own template with dedicated control tokens for tool definitions, tool calls and tool results, because a tool call cannot be expressed in a format that has only two brackets.

The reliable move is therefore never to copy a template from a page like this one into your code. Take it from the tokeniser of the exact checkpoint you are running, which is where the ground truth lives and which is why apply_chat_template exists. The version differences are the subject of Mistral’s tokeniser versions.

When this matters to you

If you call Mistral’s hosted API with a messages array, none of this is your problem. The server applies the correct template for the model you named, and it will be the right one even after the alias moves to a new snapshot with a different format. That is most of the value of a chat completions endpoint.

There is one hosted case where it does resurface: an endpoint that takes raw text rather than messages, offered for exactly the completion tasks described above. Sending a chat-shaped question to that endpoint gets you raw-mode behaviour from an instruct-tuned model, which is the confusing middle case — the model is capable of answering properly and is not being asked in the shape that evokes it. If a model that behaves well through messages behaves oddly through a different endpoint, check which one you are calling before you change anything about the prompt.

It becomes your problem the moment you self-host — with vLLM, llama.cpp, Ollama or a bare transformers loop — or the moment you fine-tune. The self-hosting failure has a characteristic signature worth recognising. A model that produces reasonable but slightly off answers, that ignores a system prompt, that rambles past where it should stop, or that occasionally emits the literal string [/INST] in its own output, is almost always being fed a malformed template rather than being a bad model. That last symptom is the giveaway: the model is continuing a conversation transcript rather than participating in one, because your string put it in a position where the next plausible token was another turn marker.

Training data must be formatted in the template the base model expects, and a fine-tune performed on differently-templated data teaches the model a format that then has to be reproduced exactly at inference time. A subtle mismatch there degrades quality in a way that looks like a bad dataset and is not.