This document records the July 2026 work that moved CuMetal's CUDA source and binary-shim paths away from implicit CPU emulation and established positive, numerically checked Apple GPU execution.
Covered CUDA kernels now execute as Metal compute commands on Apple Silicon. CPU kernel emulation and host helper fallbacks are disabled by default and can only be enabled explicitly for diagnostics.
The following end-to-end results were verified on an Apple M4 Pro:
- A standalone
.cuvector-add program compiled through CuMetal and produced the expected values on the Apple GPU. - NVIDIA's unmodified
cuda-samplesvectorAddsource compiled through the in-tree CUDA toolchain shims, printedTest PASSED, and emitted completedgeneric_ptxApple-GPU provenance. - llm.c's GPT-2 FP32 conformance workload passed logits, loss, tensor, and overall numerical checks with CPU emulation disabled.
- llama.cpp's unmodified GGML CUDA backend, linked against
libcumetal, loaded SmolLM2-135M-Instruct-Q4_K_M with one GPU-offloaded layer and greedily completedThe capital of France isasThe capital of France is Paris.The verified run generated at 5.8 tokens/s.
This is proof for the covered paths, not a claim of general CUDA compatibility. Higher llama.cpp offload counts and arbitrary models still encounter unsupported GGML kernels.
The legacy llm.c CPU implementation is disabled unless
CUMETAL_ENABLE_LLMC_CPU_EMULATION=1 is set and emits a warning when enabled.
GGML's k_compute_batched_ptrs path is not CPU kernel emulation: it is an exact
runtime ABI helper that synchronizes its input stream and constructs native
Metal-address tables for batched cuBLAS calls.
The strict conformance scripts reject provenance from cpu_fallback and stub
sources, so a CPU result cannot be mistaken for a GPU pass.
llama.cpp allocates a large CUDA arena and passes adjacent regions of the same underlying Metal buffer through several CUDA streams. The initial asynchronous registration path produced timing-dependent stale reads: isolated RMS, conversion, dequantization, and GEMM probes were exact, while the live model alternated between exact and badly corrupted RMS results.
CuMetal now keeps those launches asynchronous while preserving visibility with
public Metal synchronization APIs. Each stream owns an MTLSharedEvent; every
command records one signal value, and commands touching a buffer wait on its
previous queue access. Identical dependencies across several arguments are
coalesced. Because aliases resolve to the same tracked Buffer, adjacent CUDA
suballocations receive the same dependency chain, including transitions between
typed kernels and MPS/cuBLAS commands.
CUMETAL_SYNC_REGISTERED_LAUNCH=1 restores the former host-side wait for
diagnosis and A/B performance comparisons. The focused cross-queue regression
is functional_metal_backend_cross_queue_fence.
CUMETAL_SYNC_EACH_LAUNCH=1 remains a broader diagnostic switch that also
synchronizes direct launches.
Set CUMETAL_TRACE_GPU=1 to print one CUMETAL_PROVENANCE record per completed
Metal dispatch. Records identify:
- the Metal device;
- the kernel name;
- lowering source (
generic_ptx,specialized_msl, ormetallib); - cache status;
- grid and block dimensions;
- launch success;
- completed GPU duration when available.
GPU conformance requires device=apple_gpu and launch_success=true. The gates
reject CPU fallback and stub provenance.
The PTX-to-Metal path gained:
cvta.to.globalsupport;mul.wide.s32andmul.wide.u32support;- floating-point opcode register typing fixes;
- a fast negative filter for unsupported large GGML kernel families;
- explicit classification of approximate/passthrough templates so they are refused by default instead of silently producing wrong values;
- exact specialized MSL for the covered GGML output path:
rms_norm_f32, including strided 3D input and mul/add broadcasting;k_bin_bcastfloat add and multiply;- float-to-half and half-to-float
convert_unary; - Q8_0-to-f16 block dequantization.
The RMS lowering uses one Metal threadgroup per CUDA row, a fixed 32-lane SIMD
width, simd_sum subtotals, and a small threadgroup reduction. The previous
global-thread mapping wrote beyond the destination arena.
Registration cache keys include a schema/version tag, so behavior-changing lowering updates invalidate stale generated Metal sources. Approximate registered kernels are unconditionally refused; no environment switch can enable known-wrong output.
All-FP16 cublasGemmEx with FP16 compute now lowers directly to an
MPSMatrixMultiplication over the tracked Metal buffers. Other mixed-type
combinations retain the FP32 conversion path. The implementation waits for
producer work before the library substitution and for the Metal/MPS result
before host-side conversion or reuse. Tests cover the exact SmolLM2 output-head
dimensions:
- Q8_0 dequantization of 28,311,552 weights;
- f32-to-f16 conversion of a 576-element activation;
- a
49152 x 1 x 576mixed GEMM; - f16-to-f32 conversion of 49,152 logits.
This removes the previous per-token CPU expansion of 28,311,552 FP16 weights. On Apple M4 Pro, five warm NGL=1 runs measured a 0.57 s median for the one-token gate and 0.61 s for the full 16-token coherence gate; generation improved from 8.1 to 279.2 tokens/s median.
CUMETAL_CUBLAS_CPU_REFERENCE=1 is an opt-in diagnostic oracle used to separate
GEMM errors from surrounding kernel/scheduling errors. It is not enabled by
default and is not used by GPU conformance.
CUMETAL_VALIDATE_GGML_RMS=1 synchronizes live GGML RMS launches and compares
their exact bound buffers, strides, dimensions, and broadcast metadata against a
CPU oracle. It was used to expose the alternating exact/corrupt results in the
asynchronous registration path. It is diagnostic-only and disabled by default.
The Metal backend records asynchronous command completion and reports positive GPU duration provenance. macOS build/test helpers strip stale provenance metadata and apply ad-hoc signatures to generated binaries where needed.
CuMetal supplies narrow ptxas and fatbinary command-line shims for Clang and
upstream CUDA projects. They accept the invocation shapes used by the verified
samples and preserve embedded PTX for runtime registration. They do not emulate
NVIDIA SASS generation.
The strict CUDA project harnesses:
- require a real Apple GPU provenance record;
- require a numerical pass marker or exact output comparison;
- reject CPU fallbacks, approximate stubs, and silent skips;
- keep unsupported larger projects as explicit exit-77 skips.
Configure and build the normal binary-shim path:
cmake -B build -DCMAKE_BUILD_TYPE=Debug
cmake --build build -j8Run the strict source and GGML operator gates:
ctest --test-dir build -R \
'functional_cuda_source_gpu_vector_add|conformance_cuda_samples_vectoradd_gpu|functional_cuda_projects_ggml_output_ops' \
--output-on-failureRun the verified llama.cpp smoke test:
CUMETAL_LLAMA_NGL=1 \
bash tests/conformance/run_llama_cpp_cumetal.shThe llama.cpp harness forces --simple-io so the generated response is written
to the pipe used by the coherence gate rather than only to an interactive
terminal. Because CuMetal provenance and llama.cpp token output share the
capture, the checker parses it as a byte stream and removes provenance records
before matching the expected response. A focused regression covers invalid
UTF-8, token fragments split by provenance, missing GPU provenance, forbidden
fallback/stub provenance, incoherent output, and nonzero llama exit status.
The 2026-07-20 NGL=1 recheck on Apple M4 Pro produced
The capital of France is Paris. at 8.4 tokens/s with completed
specialized_msl Apple-GPU provenance and no fallback or approximate kernels.
Validate that the source-first build does not depend on the binary shim:
cmake -B build-nosshim -DCMAKE_BUILD_TYPE=Debug \
-DCUMETAL_ENABLE_BINARY_SHIM=OFF
cmake --build build-nosshim -j8
ctest --test-dir build-nosshim --output-on-failureThe July 20 completion audit selected 182 non-benchmark/non-external-model tests in the default build and 171 in the binary-shim-off build. Both runs completed with zero failures. Environment-dependent AIR/Xcode, standalone CUDA, and binary-shim-only cases reported explicit skips instead of false passes.
- The verified SmolLM2 result passes NGL=1 through NGL=99, including saturated full offload. This does not imply that arbitrary models or all GGML kernels are supported.
- Registered launches use shared-event resource fencing across Metal command
queues;
CUMETAL_SYNC_REGISTERED_LAUNCH=1remains available for diagnostics. - Fatbinary registration indexes entry signatures lazily; a kernel still pays its PTX lowering and Metal compilation cost on first use.
- Approximate template bodies still exist for development but are refused by default.
- The generic PTX lowering surface remains a documented subset; successful specialized MSL kernels do not imply arbitrary PTX support.
See also known-gaps.md, testing.md, and status.md.