|
| 1 | +# GPU Backends & Build Options |
| 2 | + |
| 3 | +`llama-cpp-python` builds the bundled `llama.cpp` from source by default, which produces a |
| 4 | +CPU-only build. To use a GPU (or a tuned CPU BLAS backend) you either pass the matching |
| 5 | +`llama.cpp` cmake option through `CMAKE_ARGS`, or install a pre-built wheel for your |
| 6 | +platform. |
| 7 | + |
| 8 | +This page consolidates the build flags, the available pre-built wheels, and how to verify |
| 9 | +that GPU offload is actually active. For the CLI/`pip` invocation syntax (and the |
| 10 | +`-C cmake.args=...` alternative) see the [Getting Started](../index.md) page. |
| 11 | + |
| 12 | +## CMAKE_ARGS mapping table |
| 13 | + |
| 14 | +All `llama.cpp` cmake build options can be set via the `CMAKE_ARGS` environment variable |
| 15 | +before installing, e.g.: |
| 16 | + |
| 17 | +```bash |
| 18 | +CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python |
| 19 | +``` |
| 20 | + |
| 21 | +| Backend | `CMAKE_ARGS` | Requirements | |
| 22 | +| --- | --- | --- | |
| 23 | +| CPU (default) | *(none)* | Any supported platform | |
| 24 | +| OpenBLAS (CPU) | `-DGGML_BLAS=ON -DGGML_BLAS_VENDOR=OpenBLAS` | OpenBLAS installed | |
| 25 | +| CUDA | `-DGGML_CUDA=on` | NVIDIA GPU + CUDA toolkit | |
| 26 | +| Metal | `-DGGML_METAL=on` | macOS 11.0+ on Apple Silicon / supported GPU | |
| 27 | +| HIP (ROCm) | `-DGGML_HIP=on` | AMD GPU + ROCm/HIP toolkit | |
| 28 | +| Vulkan | `-DGGML_VULKAN=on` | Vulkan-capable GPU + SDK | |
| 29 | +| SYCL | `-DGGML_SYCL=on -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx` | Intel oneAPI (`source /opt/intel/oneapi/setvars.sh` first) | |
| 30 | +| RPC | `-DGGML_RPC=on` | Distributed inference over RPC | |
| 31 | + |
| 32 | +!!! note |
| 33 | + SYCL builds require the Intel oneAPI environment to be sourced in the same shell before |
| 34 | + installing: |
| 35 | + |
| 36 | + ```bash |
| 37 | + source /opt/intel/oneapi/setvars.sh |
| 38 | + CMAKE_ARGS="-DGGML_SYCL=on -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx" pip install llama-cpp-python |
| 39 | + ``` |
| 40 | + |
| 41 | +## Pre-built wheels |
| 42 | + |
| 43 | +For some backends you can skip the source build entirely by adding the matching |
| 44 | +`--extra-index-url`. Pre-built wheels are currently published for Python 3.10, 3.11 and |
| 45 | +3.12. |
| 46 | + |
| 47 | +```bash |
| 48 | +pip install llama-cpp-python \ |
| 49 | + --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/<tag> |
| 50 | +``` |
| 51 | + |
| 52 | +| Backend | `<tag>` | Notes | |
| 53 | +| --- | --- | --- | |
| 54 | +| CUDA 11.8 | `cu118` | Compute capability 6.0–8.9 | |
| 55 | +| CUDA 12.1 | `cu121` | Compute capability 6.0+ | |
| 56 | +| CUDA 12.2 | `cu122` | Compute capability 6.0+ | |
| 57 | +| CUDA 12.3 | `cu123` | Compute capability 6.0+ | |
| 58 | +| CUDA 12.4 | `cu124` | Compute capability 6.0+ | |
| 59 | +| CUDA 12.5 | `cu125` | Compute capability 6.0+ | |
| 60 | +| CUDA 13.0 | `cu130` | Compute capability 7.5+ | |
| 61 | +| CUDA 13.2 | `cu132` | Compute capability 7.5+ | |
| 62 | +| Metal | `metal` | macOS 11.0+ | |
| 63 | +| ROCm (Linux) | `rocm72` | AMD GPU on Linux | |
| 64 | +| HIP Radeon (Windows) | `hip-radeon` | AMD Radeon on Windows | |
| 65 | +| Vulkan | `vulkan` | Linux or Windows | |
| 66 | + |
| 67 | +For example, to install the CUDA 12.1 wheel: |
| 68 | + |
| 69 | +```bash |
| 70 | +pip install llama-cpp-python \ |
| 71 | + --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cu121 |
| 72 | +``` |
| 73 | + |
| 74 | +!!! warning "Match the wheel to your driver" |
| 75 | + A CUDA wheel must match a CUDA toolkit/driver that supports your GPU's compute |
| 76 | + capability. CUDA 11.8 wheels cover compute capability 6.0–8.9; CUDA 12 wheels cover |
| 77 | + 6.0 and newer; CUDA 13 wheels require 7.5 and newer. If no wheel matches your setup, |
| 78 | + fall back to a source build with the relevant `CMAKE_ARGS` above. |
| 79 | + |
| 80 | +## Verifying GPU offload is active |
| 81 | + |
| 82 | +A build can silently fall back to CPU (for example, if the GPU backend wasn't compiled in, |
| 83 | +or the wheel didn't match your toolkit). Two quick checks: |
| 84 | + |
| 85 | +**1. Does this build support GPU offload at all?** |
| 86 | + |
| 87 | +```python |
| 88 | +import llama_cpp |
| 89 | + |
| 90 | +print("GPU offload supported:", llama_cpp.llama_supports_gpu_offload()) |
| 91 | +print(llama_cpp.llama_print_system_info().decode()) |
| 92 | +``` |
| 93 | + |
| 94 | +If `llama_supports_gpu_offload()` returns `False`, the installed build is CPU-only — |
| 95 | +re-install with the appropriate backend above. |
| 96 | + |
| 97 | +**2. Are layers actually placed on the GPU?** |
| 98 | + |
| 99 | +Load a model with `n_gpu_layers=-1` (offload all layers) and `verbose=True`. `llama.cpp` |
| 100 | +logs how many layers were assigned to the GPU and the per-device memory usage: |
| 101 | + |
| 102 | +```python |
| 103 | +from llama_cpp import Llama |
| 104 | + |
| 105 | +llm = Llama(model_path="model.gguf", n_gpu_layers=-1, verbose=True) |
| 106 | +``` |
| 107 | + |
| 108 | +Look for log lines such as `offloaded N/N layers to GPU` and a non-CPU buffer in the model |
| 109 | +load summary. If you see `0` layers offloaded despite requesting `-1`, the backend isn't |
| 110 | +being used. |
| 111 | + |
| 112 | +## Platform-specific notes |
| 113 | + |
| 114 | +- **macOS (Metal):** see the [macOS (Metal)](macos.md) guide for detailed Apple Silicon |
| 115 | + build/runtime instructions. |
| 116 | +- **Windows:** if cmake can't find `nmake` or `CMAKE_C_COMPILER`, install |
| 117 | + [w64devkit](https://github.com/ggml-org/llama.cpp) and point `CMAKE_ARGS` at its |
| 118 | + compilers: |
| 119 | + |
| 120 | + ```ps |
| 121 | + $env:CMAKE_GENERATOR = "MinGW Makefiles" |
| 122 | + $env:CMAKE_ARGS = "-DGGML_OPENBLAS=on -DCMAKE_C_COMPILER=C:/w64devkit/bin/gcc.exe -DCMAKE_CXX_COMPILER=C:/w64devkit/bin/g++.exe" |
| 123 | + ``` |
0 commit comments