Skip to content

Commit 94fd771

Browse files
committed
fix(ci): shrink CUDA wheel fatbins
1 parent d6f46a5 commit 94fd771

1 file changed

Lines changed: 4 additions & 3 deletions

File tree

‎.github/workflows/build-wheels-cuda.yaml‎

Lines changed: 4 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -153,9 +153,10 @@ jobs:
153153
}
154154
$cudaTagVersion = $nvccVersion.Replace('.','')
155155
$env:VERBOSE = '1'
156-
# Keep a portable SM set, including sm_70, instead of CMake's `all`,
157-
# which now pulls in future targets the hosted-runner toolchains cannot assemble.
158-
$env:CMAKE_ARGS = "-DGGML_CUDA_FORCE_MMQ=ON -DGGML_CUDA=on -DCMAKE_CUDA_ARCHITECTURES=70;75;80;86;89;90 -DCMAKE_CUDA_FLAGS=--allow-unsupported-compiler $env:CMAKE_ARGS"
156+
# Build real cubins for the supported GPUs, including sm_70, and keep
157+
# one forward-compatible PTX target instead of embedding PTX for every
158+
# SM. This keeps the wheel under GitHub's 2 GiB release-asset limit.
159+
$env:CMAKE_ARGS = "-DGGML_CUDA_FORCE_MMQ=ON -DGGML_CUDA=on -DCMAKE_CUDA_ARCHITECTURES=70-real;75-real;80-real;86-real;89-real;90-real;90-virtual -DCMAKE_CUDA_FLAGS=--allow-unsupported-compiler $env:CMAKE_ARGS"
159160
# if ($env:AVXVER -eq 'AVX') {
160161
$env:CMAKE_ARGS = $env:CMAKE_ARGS + ' -DGGML_AVX2=off -DGGML_FMA=off -DGGML_F16C=off'
161162
# }

0 commit comments

Comments
 (0)