Skip to content

Quantizing Your Own Model With llama-quantize

11 min read · updated August 11, 2026

Most people download a quantized GGUF somebody else made. Producing your own matters in three cases: the model is a fine-tune nobody has converted, you need a quant level nobody uploaded, or you need to know exactly what was done to the weights.

Why produce your own

A GGUF on a model hub is the output of somebody else’s decisions: which type, whether an importance matrix was used, what calibration data fed it, whether the token embedding and output tensor were left at higher precision. Those choices change quality measurably and are not always recorded. llama.cpp writes provenance into the file’s metadata when an importance matrix is used — keys including quantize.imatrix.file, quantize.imatrix.dataset, quantize.imatrix.entries_count and quantize.imatrix.chunks_count — so you can inspect a file you did not build, but only if whoever built it used the tool properly.

Model weights carry licences and some are gated. Converting a checkpoint you have legitimate access to is fine; obtaining gated weights around the gate is not, and redistributing a quantized derivative may be restricted even where using it locally is not. Check the licence on the original repository before you publish anything you build here.

Step one: a full-precision GGUF

llama-quantize does not read safetensors. The conversion is a Python script in the llama.cpp repository:

python convert_hf_to_gguf.py ./models/my-finetune \
  --outfile my-finetune-f16.gguf \
  --outtype f16

--outtype takes f32, f16, bf16 and q8_0 among others. Use f16 or bf16 as the intermediate: quantizing from an already-quantized file is possible with --allow-requantize, and the tool’s own help warns that it “can severely reduce quality compared to quantizing from 16bit or 32bit”. The reason is that quantization error compounds — the second pass is fitting a grid to values that are already snapped to a coarser grid, and the two grids do not align.

Conversion is also where architectures fail. The script supports the architectures llama.cpp knows; a brand-new one will not convert until support lands, and the error names the unrecognised architecture string from the config.

Choosing a type, with llama.cpp’s own numbers

llama-quantize --help prints a table of types with a size and a perplexity delta for each, measured against Llama-3-8B, and those figures live in tools/quantize/quantize.cpp. They are llama.cpp’s measurements, not anyone else’s, and they are the right thing to reason from:

Q2_K     2.96G   +3.5199 ppl @ Llama-3-8B
Q3_K_S   3.41G   +1.6321 ppl
Q3_K_M   3.74G   +0.6569 ppl
Q4_0     4.34G   +0.4685 ppl
Q4_K_S   4.37G   +0.2689 ppl
Q4_K_M   4.58G   +0.1754 ppl
Q5_K_M   5.33G   +0.0569 ppl
Q6_K     6.14G   +0.0217 ppl
Q8_0     7.96G   +0.0026 ppl

Read the shape rather than the digits. Between Q8_0 and Q4_K_M you give up 3.4 GB and 0.17 perplexity. Between Q4_K_M and Q3_K_M you give up 0.8 GB and gain 0.48 — nearly three times the damage for a quarter of the saving. Below Q3 the curve turns sharply: Q2_K costs +3.52, twenty times the Q4_K_M delta, for 1.6 GB. That is why Q4_K_M is the usual recommendation and why it is not arbitrary; it sits just before the knee.

The _S, _M and _L suffixes are mixtures rather than uniform bit widths. A k-quant keeps some tensors — attention projections, the output layer — at higher precision than the bulk, because those are where quantization error propagates furthest. --pure disables the mixture and quantizes everything to one type, which is useful for experiments and worse for quality. The IQ types use a different, non-linear codebook and reach lower bits per weight at more compute cost per token; they are also the types that benefit most from an importance matrix.

The importance matrix, and what it is not

An importance matrix records, per tensor, which weights carried the most activation magnitude over a calibration corpus. The quantizer uses it to spend its error budget where it matters least. You produce one with the llama-imatrix tool over a text file:

llama-imatrix -m my-finetune-f16.gguf \
  -f calibration.txt \
  -o my-finetune.imatrix

Two things it is not. It is not fine-tuning — no weight is changed by the calibration data, only the choice of how to round it. And it is not free of the calibration set: a matrix built on English prose will preserve English prose preferentially, which matters if the model is used for code or another language. Use calibration text that resembles your workload, and be aware that this makes the resulting file a slightly specialised artefact rather than a general one.

The quantize command

  1. Build the tool with the rest of llama.cpp: cmake -B build && cmake --build build --config Release.
  2. Run a dry pass first. --dry-run reports what would be produced without writing gigabytes, which catches an unrecognised type name before you wait.
  3. Quantize. The positional arguments are input, output, type, and optionally a thread count:
    llama-quantize --imatrix my-finetune.imatrix \
      my-finetune-f16.gguf \
      my-finetune-q4_k_m.gguf \
      Q4_K_M 8
  4. If you want the output tensor kept at higher precision — often worth it on small models, where the output layer is a large share of the parameters — add --leave-output-tensor, or set types individually with --output-tensor-type and --token-embedding-type.
  5. Load it and confirm the type is what you asked for. The loader prints the file type in its metadata block, and the size on disk should be close to the table above scaled by your parameter count.

Verifying you did not ruin it

A quantized model that loads and produces fluent text can still be badly damaged, because fluency is the last thing to go. The check that catches it is perplexity on a fixed text, run identically against the f16 file and the quantized one, and compared as a delta rather than as an absolute — the number depends on the text, so only the difference between two runs on the same text means anything. That is the methodology behind the table above, and the perplexity tool page covers running it.

Perplexity is a blunt instrument, though. It averages over every token and so understates damage to rare capabilities — structured output, tool-call formatting, non-English text and arithmetic degrade faster than a perplexity delta suggests, because they depend on precise behaviour in a small number of positions. If the model has a job, test the job as well. And if the file is going into a llama-server deployment, test it through the same chat template it will be served with.