Skip to content

Commit f9ea95c

Browse files
committed
docs(install): add consolidated GPU backends & build options guide
1 parent 3bda091 commit f9ea95c

2 files changed

Lines changed: 125 additions & 1 deletion

File tree

‎docs/install/gpu.md‎

Lines changed: 123 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,123 @@
1+
# GPU Backends & Build Options
2+
3+
`llama-cpp-python` builds the bundled `llama.cpp` from source by default, which produces a
4+
CPU-only build. To use a GPU (or a tuned CPU BLAS backend) you either pass the matching
5+
`llama.cpp` cmake option through `CMAKE_ARGS`, or install a pre-built wheel for your
6+
platform.
7+
8+
This page consolidates the build flags, the available pre-built wheels, and how to verify
9+
that GPU offload is actually active. For the CLI/`pip` invocation syntax (and the
10+
`-C cmake.args=...` alternative) see the [Getting Started](../index.md) page.
11+
12+
## CMAKE_ARGS mapping table
13+
14+
All `llama.cpp` cmake build options can be set via the `CMAKE_ARGS` environment variable
15+
before installing, e.g.:
16+
17+
```bash
18+
CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python
19+
```
20+
21+
| Backend | `CMAKE_ARGS` | Requirements |
22+
| --- | --- | --- |
23+
| CPU (default) | *(none)* | Any supported platform |
24+
| OpenBLAS (CPU) | `-DGGML_BLAS=ON -DGGML_BLAS_VENDOR=OpenBLAS` | OpenBLAS installed |
25+
| CUDA | `-DGGML_CUDA=on` | NVIDIA GPU + CUDA toolkit |
26+
| Metal | `-DGGML_METAL=on` | macOS 11.0+ on Apple Silicon / supported GPU |
27+
| HIP (ROCm) | `-DGGML_HIP=on` | AMD GPU + ROCm/HIP toolkit |
28+
| Vulkan | `-DGGML_VULKAN=on` | Vulkan-capable GPU + SDK |
29+
| SYCL | `-DGGML_SYCL=on -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx` | Intel oneAPI (`source /opt/intel/oneapi/setvars.sh` first) |
30+
| RPC | `-DGGML_RPC=on` | Distributed inference over RPC |
31+
32+
!!! note
33+
SYCL builds require the Intel oneAPI environment to be sourced in the same shell before
34+
installing:
35+
36+
```bash
37+
source /opt/intel/oneapi/setvars.sh
38+
CMAKE_ARGS="-DGGML_SYCL=on -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx" pip install llama-cpp-python
39+
```
40+
41+
## Pre-built wheels
42+
43+
For some backends you can skip the source build entirely by adding the matching
44+
`--extra-index-url`. Pre-built wheels are currently published for Python 3.10, 3.11 and
45+
3.12.
46+
47+
```bash
48+
pip install llama-cpp-python \
49+
--extra-index-url https://abetlen.github.io/llama-cpp-python/whl/<tag>
50+
```
51+
52+
| Backend | `<tag>` | Notes |
53+
| --- | --- | --- |
54+
| CUDA 11.8 | `cu118` | Compute capability 6.0–8.9 |
55+
| CUDA 12.1 | `cu121` | Compute capability 6.0+ |
56+
| CUDA 12.2 | `cu122` | Compute capability 6.0+ |
57+
| CUDA 12.3 | `cu123` | Compute capability 6.0+ |
58+
| CUDA 12.4 | `cu124` | Compute capability 6.0+ |
59+
| CUDA 12.5 | `cu125` | Compute capability 6.0+ |
60+
| CUDA 13.0 | `cu130` | Compute capability 7.5+ |
61+
| CUDA 13.2 | `cu132` | Compute capability 7.5+ |
62+
| Metal | `metal` | macOS 11.0+ |
63+
| ROCm (Linux) | `rocm72` | AMD GPU on Linux |
64+
| HIP Radeon (Windows) | `hip-radeon` | AMD Radeon on Windows |
65+
| Vulkan | `vulkan` | Linux or Windows |
66+
67+
For example, to install the CUDA 12.1 wheel:
68+
69+
```bash
70+
pip install llama-cpp-python \
71+
--extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cu121
72+
```
73+
74+
!!! warning "Match the wheel to your driver"
75+
A CUDA wheel must match a CUDA toolkit/driver that supports your GPU's compute
76+
capability. CUDA 11.8 wheels cover compute capability 6.0–8.9; CUDA 12 wheels cover
77+
6.0 and newer; CUDA 13 wheels require 7.5 and newer. If no wheel matches your setup,
78+
fall back to a source build with the relevant `CMAKE_ARGS` above.
79+
80+
## Verifying GPU offload is active
81+
82+
A build can silently fall back to CPU (for example, if the GPU backend wasn't compiled in,
83+
or the wheel didn't match your toolkit). Two quick checks:
84+
85+
**1. Does this build support GPU offload at all?**
86+
87+
```python
88+
import llama_cpp
89+
90+
print("GPU offload supported:", llama_cpp.llama_supports_gpu_offload())
91+
print(llama_cpp.llama_print_system_info().decode())
92+
```
93+
94+
If `llama_supports_gpu_offload()` returns `False`, the installed build is CPU-only —
95+
re-install with the appropriate backend above.
96+
97+
**2. Are layers actually placed on the GPU?**
98+
99+
Load a model with `n_gpu_layers=-1` (offload all layers) and `verbose=True`. `llama.cpp`
100+
logs how many layers were assigned to the GPU and the per-device memory usage:
101+
102+
```python
103+
from llama_cpp import Llama
104+
105+
llm = Llama(model_path="model.gguf", n_gpu_layers=-1, verbose=True)
106+
```
107+
108+
Look for log lines such as `offloaded N/N layers to GPU` and a non-CPU buffer in the model
109+
load summary. If you see `0` layers offloaded despite requesting `-1`, the backend isn't
110+
being used.
111+
112+
## Platform-specific notes
113+
114+
- **macOS (Metal):** see the [macOS (Metal)](macos.md) guide for detailed Apple Silicon
115+
build/runtime instructions.
116+
- **Windows:** if cmake can't find `nmake` or `CMAKE_C_COMPILER`, install
117+
[w64devkit](https://github.com/ggml-org/llama.cpp) and point `CMAKE_ARGS` at its
118+
compilers:
119+
120+
```ps
121+
$env:CMAKE_GENERATOR = "MinGW Makefiles"
122+
$env:CMAKE_ARGS = "-DGGML_OPENBLAS=on -DCMAKE_C_COMPILER=C:/w64devkit/bin/gcc.exe -DCMAKE_CXX_COMPILER=C:/w64devkit/bin/g++.exe"
123+
```

‎mkdocs.yml‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -47,6 +47,7 @@ watch:
4747
nav:
4848
- "Getting Started": "index.md"
4949
- "Installation Guides":
50+
- "GPU Backends & Build Options": "install/gpu.md"
5051
- "macOS (Metal)": "install/macos.md"
5152
- "API Reference": "api-reference.md"
5253
- "OpenAI Compatible Web Server": "server.md"
@@ -71,4 +72,4 @@ markdown_extensions:
7172
- pymdownx.tabbed:
7273
alternate_style: true
7374
- pymdownx.tilde
74-
- tables
75+
- tables

0 commit comments

Comments
 (0)