Skip to content

“Failed Building Wheel for llama-cpp-python”

11 min read · updated August 11, 2026

The two lines everybody pastes into search are the last two lines pip prints, and neither of them contains the cause. The cause is in the CMake output twenty lines above, and once you know which line to look for this becomes a five-minute fix.

The strings, and which one matters

ERROR: Failed building wheel for llama-cpp-python
Failed to build llama-cpp-python
ERROR: Could not build wheels for llama-cpp-python, which is
required to install pyproject.toml-based projects

Those are pip’s summary. Scroll up to the block that begins *** scikit-build-core or CMake Error, and find one of these instead. Each names a different missing piece:

  • No CMAKE_C_COMPILER could be found. or No CMAKE_CXX_COMPILER could be found. — CMake ran but found no compiler.
  • Failed to find nmake, or a list of Visual Studio generators CMake tried and rejected — Windows, no MSVC build tools.
  • error: Microsoft Visual C++ 14.0 or greater is required. — same cause, different messenger.
  • Could not find CMAKE_ROOT or a missing cmake executable — CMake itself is not installed.

The two most-cited reports in the upstream trackers are llama-cpp-python issue 1664 and issue 789, and both resolve to a toolchain that was not there rather than to anything about the package.

Why pip is compiling anything

llama-cpp-python is a binding around llama.cpp, which is C++. Installing it means building llama.cpp. pip only avoids that when a prebuilt wheel exists for your exact combination of platform, CPU architecture, Python version and ABI — and the moment one of those is new, no wheel matches and pip falls back to building from the source distribution.

That is why this error tends to arrive the week after a Python upgrade on a machine where the same command worked for months. Nothing about your toolchain changed; the wheel that was covering for its absence stopped matching. Check first:

python -c "import sys, platform; print(sys.version, platform.machine())"
pip install --only-binary :all: llama-cpp-python

If the second command reports that no matching distribution was found, you have confirmed there is no wheel for you and a build is unavoidable — unless you use the project’s own wheel index, which is the last section here.

The cause, by platform

Windows. This is where the great majority of reports come from. You need the Visual Studio Build Tools with the “Desktop development with C++” workload, which is what supplies cl.exe and nmake, plus CMake. Installing Visual Studio Code does not install a compiler; installing Python does not either. After installing the build tools, open a fresh terminal — the PATH changes do not reach a shell that was already running, and a surprising share of “I installed it and it still fails” reports are exactly that.

Linux. You need a C and C++ compiler and CMake: build-essential and cmake on Debian and Ubuntu, gcc-c++, make and cmake on Fedora and RHEL. Slim container images are the usual culprit here, because python:3.12-slim ships no compiler at all.

macOS. xcode-select --install supplies the command line tools. A second, separate failure on macOS is a missing OpenMP runtime rather than a missing compiler, which produces a different message at import time rather than at build time — see the libomp page if the build succeeds and the import does not.

When the CUDA build is what fails

A GPU build needs everything above plus the CUDA toolkit, and nvcc must be on PATH — the driver alone is not enough, because you are compiling device code rather than merely running it. On Windows the CUDA installer also has to have registered its MSBuild integration with the Visual Studio version you have; installing CUDA before Visual Studio leaves that integration missing and produces a CMake failure that mentions neither product clearly.

The build is requested through an environment variable:

# Linux / macOS
CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python --no-cache-dir

# Windows PowerShell
$env:CMAKE_ARGS = "-DGGML_CUDA=on"
pip install llama-cpp-python --no-cache-dir
That flag was renamed. Guides written before the llama.cpp consolidation of its build options say -DLLAMA_CUBLAS=on, and passing the old name to a current version silently produces a CPU-only build rather than an error — you get a successful install and no GPU. If a GPU build appears to succeed and the model still runs on CPU, the flag name is the first thing to check. The Metal equivalent moved the same way, from -DLLAMA_METAL=on to -DGGML_METAL=on.

Always add --no-cache-dir when changing build flags. pip caches the built wheel, and without it your second attempt reinstalls the first attempt’s result and you conclude the flag did nothing.

Not building at all

The project publishes prebuilt wheels for common CUDA and Metal configurations at its own index, which removes the toolchain from the problem entirely. This is the right first move on a machine you do not want to turn into a build host:

pip install llama-cpp-python \
  --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cu121

Replace the trailing path segment with the CUDA version you have, or with metal on Apple Silicon. If no wheel exists for your Python version there either, the honest options are to install an older Python in a virtual environment, install the toolchain, or use a different runtime: Ollama ships precompiled binaries and needs no Python build step, and llama.cpp itself publishes release binaries for the server.

Two smaller things that turn a fixed build back into a broken one. First, pip’s build isolation creates a fresh environment for the build, so a package you installed to help — a specific CMake, a specific setuptools — may not be visible to it; --no-build-isolation is the escape hatch, at the cost of having to supply the build requirements yourself. Second, the build reads environment variables at the moment pip runs, so a CMAKE_ARGS exported in one terminal has no effect on a build started in another, and none at all on a build started by an installer script that clears the environment.

Finally, verify that you got what you asked for rather than assuming it. A CPU build and a CUDA build install to the same place under the same name, and the only difference visible from Python is what the library reports when it loads a model — look for the offload line naming your device and a non-zero layer count. If you asked for CUDA and see every layer on the CPU, the build succeeded and the flag did not, which sends you back to the rename described in the note above rather than to the toolchain.

from llama_cpp import Llama
llm = Llama(model_path="model.gguf", n_gpu_layers=-1, verbose=True)
# look for a line reporting layers assigned to the GPU;
# 0 layers offloaded means you have a CPU-only build.
The wheel index path scheme, the CUDA versions it covers and the flag names above all change with upstream releases. Read the llama-cpp-python README rather than trusting a command copied from anywhere, including here, if it is more than a couple of releases old.