Summary
llvm.amdgcn.raw.ptr.buffer.store is modeled as an argmem-only write through
its AS8 resource operand. However, the effective memory location of the access
also depends on the voffset and soffset operands, while the access size is
determined by the stored type.
For this intrinsic, generic call-memory-location modeling does not preserve
that information precisely enough for DSE to prove that two calls completely
overwrite the same location, or for LICM to sink a repeated loop store.
Under the raw-descriptor restrictions required by AS7, expressing the same
access as an ordinary AS7 store enables the expected generic optimizations:
- DSE removes the earlier store.
- LICM sinks the loop store to the loop exit.
-O3 replaces the loop with a guarded closed-form calculation and one final
store.
This was initially observed in a controlled Triton compilation. Both paths
ultimately select buffer_store_dword on gfx942, so the code-quality
difference arises before instruction selection rather than from choosing a
different final store opcode.
Standalone LLVM reproducer
The pair functions provide the minimal DSE case. The loop functions demonstrate
the corresponding LICM and -O3 behavior.
target datalayout = "e-m:e-p:64:64-p1:64:64-p2:32:32-p3:32:32-p4:64:64-p5:32:32-p6:32:32-p7:160:256:256:32-p8:128:128:128:48-p9:192:256:256:32-i64:64-v16:16-v24:32-v32:32-v48:64-v96:128-v192:256-v256:256-v512:512-v1024:1024-v2048:2048-n32:64-S32-A5-G1-ni:7:8:9"
target triple = "amdgcn-amd-amdhsa"
define void @buffer_pair(
ptr addrspace(8) %rsrc,
i32 %offset,
i32 %first,
i32 %last) {
entry:
call void @llvm.amdgcn.raw.ptr.buffer.store.i32(
i32 %first, ptr addrspace(8) %rsrc,
i32 %offset, i32 0, i32 0)
call void @llvm.amdgcn.raw.ptr.buffer.store.i32(
i32 %last, ptr addrspace(8) %rsrc,
i32 %offset, i32 0, i32 0)
ret void
}
define void @fat_pointer_pair(
ptr addrspace(8) %rsrc,
i32 %offset,
i32 %first,
i32 %last) {
entry:
%fat.base = addrspacecast ptr addrspace(8) %rsrc to ptr addrspace(7)
%out = getelementptr i8, ptr addrspace(7) %fat.base, i32 %offset
store i32 %first, ptr addrspace(7) %out, align 1
store i32 %last, ptr addrspace(7) %out, align 1
ret void
}
define void @buffer_loop(
ptr addrspace(8) %rsrc,
i32 %offset,
i32 %iterations) {
entry:
%has.iterations = icmp sgt i32 %iterations, 0
br i1 %has.iterations, label %loop, label %exit
loop:
%step = phi i32 [ 0, %entry ], [ %next, %loop ]
call void @llvm.amdgcn.raw.ptr.buffer.store.i32(
i32 %step, ptr addrspace(8) %rsrc,
i32 %offset, i32 0, i32 0)
%next = add nuw nsw i32 %step, 1
%continue = icmp slt i32 %next, %iterations
br i1 %continue, label %loop, label %exit
exit:
ret void
}
define void @fat_pointer_loop(
ptr addrspace(8) %rsrc,
i32 %offset,
i32 %iterations) {
entry:
%fat.base = addrspacecast ptr addrspace(8) %rsrc to ptr addrspace(7)
%out = getelementptr i8, ptr addrspace(7) %fat.base, i32 %offset
%has.iterations = icmp sgt i32 %iterations, 0
br i1 %has.iterations, label %loop, label %exit
loop:
%step = phi i32 [ 0, %entry ], [ %next, %loop ]
store i32 %step, ptr addrspace(7) %out, align 1
%next = add nuw nsw i32 %step, 1
%continue = icmp slt i32 %next, %iterations
br i1 %continue, label %loop, label %exit
exit:
ret void
}
declare void @llvm.amdgcn.raw.ptr.buffer.store.i32(
i32,
ptr addrspace(8) writeonly captures(none),
i32,
i32,
i32 immarg) #0
attributes #0 = {
nocallback nofree nosync nounwind willreturn memory(argmem: write)
}
Commands:
opt -passes=dse -S repro.ll -o dse.ll
opt -passes='function(loop-simplify,lcssa,loop-mssa(licm))' \
-S repro.ll -o licm.ll
opt -O3 -mcpu=gfx942 -S repro.ll -o o3.ll
Observed results:
- DSE retains both raw intrinsic calls in
buffer_pair, but removes the first
AS7 store in fat_pointer_pair.
- LICM leaves the raw intrinsic inside
buffer_loop, but sinks the AS7 store in
fat_pointer_loop to the guarded loop exit.
- In this build,
-O3 unrolls the remaining raw-intrinsic loop into eight calls
in the main loop and one in the remainder loop. The AS7 loop is replaced with
a closed-form calculation and one guarded store.
The exact unroll count is not essential to the issue. The missed optimization
is that repeated non-volatile stores to the same effective buffer location are
not reduced to the final store.
Likely limitation
This does not appear to be a case where the intrinsic is completely invisible
to the memory infrastructure. The declaration identifies it as an argmem-only
write and marks the AS8 resource argument writeonly and captures(none).
The problem is that the effective location is composite:
(resource, voffset, soffset, access size)
Generic call-location handling can describe the resource pointer, but it has no
AMDGPU-specific representation of the two offset operands and the precise size
of this access. Consequently, two calls with identical resource and offset
operands are still not recognized as complete overwrites by DSE, and the loop
store is not sinkable by LICM.
Motivating Triton pipeline comparison
The original case was found by compiling the same noalias Triton kernel twice
for the HIP gfx942 target with num_warps=4 and LLVM -O3.
Both runs enable AMD buffer operations. The second run additionally enables
Triton's experimental buffer-fat-pointer store path. Stages 01 and 02 are
byte-identical; the first representation-level difference appears in stage 03,
before translation to LLVM IR and before LLVM -O3.
| Pipeline stage |
Raw-buffer path |
AS7 path |
| 03: LLVM dialect/ROCDL |
rocdl.raw.ptr.buffer.store |
AS8 to AS7 cast, GEP, ordinary llvm.store |
05: LLVM -O3 input |
llvm.amdgcn.raw.ptr.buffer.store |
AS7 ordinary store |
06: LLVM -O3 output |
Loop retained and unrolled; 9 static store calls in this build |
Loop eliminated; closed form and 1 store |
| Final AMDGCN |
9 static buffer_store_dword instructions |
1 static buffer_store_dword instruction |
Stage 03
The normal raw-buffer path forms a ROCDL raw buffer store:
%17 = rocdl.make.buffer.rsrc %output, %6, %4, %5 : <1> to <8>
%18 = llvm.insertelement %acc_23, %0[%10 : i32] : vector<1xi32>
%19 = llvm.bitcast %18 : vector<1xi32> to i32
rocdl.raw.ptr.buffer.store %19, %17, %value_13, %10, 0 : i32
The fat-pointer path retains an ordinary store:
%17 = rocdl.make.buffer.rsrc %output, %6, %4, %5 : <1> to <8>
%18 = llvm.insertelement %acc_23, %0[%10 : i32] : vector<1xi32>
%19 = llvm.bitcast %18 : vector<1xi32> to i32
%20 = llvm.addrspacecast %17 : !llvm.ptr<8> to !llvm.ptr<7>
%21 = llvm.getelementptr %20[%value_13]
: (!llvm.ptr<7>, i32) -> !llvm.ptr<7>, i8
llvm.store %19, %21 {alignment = 1 : i64} : i32, !llvm.ptr<7>
Stage 05: LLVM -O3 input
The stage-03 representations are carried through translation to the LLVM
optimizer input.
Raw-buffer path:
%rsrc = call ptr addrspace(8) @llvm.amdgcn.make.buffer.rsrc.p8.p1(...)
call void @llvm.amdgcn.raw.ptr.buffer.store.i32(
i32 %value, ptr addrspace(8) %rsrc,
i32 %byte_offset, i32 0, i32 0)
AS7 path:
%rsrc = call ptr addrspace(8) @llvm.amdgcn.make.buffer.rsrc.p8.p1(...)
%fat.base = addrspacecast ptr addrspace(8) %rsrc to ptr addrspace(7)
%out = getelementptr i8, ptr addrspace(7) %fat.base, i32 %byte_offset
store i32 %value, ptr addrspace(7) %out, align 1
Stage 06 and final AMDGCN
After -O3, the raw-buffer path still contains the loop. Subsequent unrolling
produces eight intrinsic calls in the main loop and one in the remainder loop in
this build. The AS7 path contains no loop and has one store after a closed-form
calculation.
Both paths eventually select the same store instruction form:
; Raw-buffer path
buffer_store_dword v2, v0, s[0:3], 0 offen
; AS7 path
buffer_store_dword v1, v0, s[0:3], 0 offen
Final static metrics:
| Metric |
Raw-buffer path |
AS7 path |
| Machine instructions |
70 |
33 |
| Branch instructions |
8 |
2 |
buffer_store_dword instructions |
9 |
1 |
LLVM codeLenInByte |
628 |
416 |
The VGPR/SGPR counts, occupancy, and spill count are unchanged in this case.
For positive I = iterations, the raw path dynamically issues one buffer store
per loop iteration for each executing wave, while the AS7 path issues one
guarded final store per wave. For I <= 0, neither path stores.
These are generated-code observations only; runtime performance has not been
measured on a GPU.
Scope and semantic assumptions
The demonstrated comparison is intentionally narrow:
- The descriptor is restricted to the raw form required by AS7: stride 0,
add_tid disabled, swizzling disabled, and extent measured in bytes.
- Every repeated access uses the same SSA resource,
voffset, and soffset.
- The access is a scalar
i32 store with soffset = 0.
- The auxiliary/cache-policy operand is zero, including a clear volatile bit.
- The AS7 stores use
align 1, matching the Triton-generated IR and avoiding
any additional base/offset alignment assumptions.
- The demonstrated byte offsets are in bounds. No conclusion is claimed about
transformations that change partially out-of-bounds behavior.
This does not attempt to cover structured or swizzled descriptors, explicit
cache modifiers, volatile accesses, atomics, async LDS operations, masked
accesses, or OOB sentinel behavior.
Expected result / questions
Is AS7 ordinary memory IR intended to be the mid-end representation for raw
buffer accesses that satisfy the buffer-fat-pointer descriptor restrictions,
with AMDGPULowerBufferFatPointersPass performing the late lowering before
instruction selection?
If llvm.amdgcn.raw.ptr.buffer.* intrinsics are also expected to pass through
the generic optimizer, is there an intended target-specific mechanism for
representing or comparing their effective locations:
(resource, voffset, soffset, access size)
so that DSE and LICM can recognize complete overwrites in cases such as this?
I am not assuming that extending generic MemoryLocation is necessarily the
right solution. Preserving AS7 ordinary memory operations longer in frontends,
or adding an AMDGPU-specific optimization for raw buffer intrinsics, may be more
appropriate.
Version
Triton version 3.8.0
LLVM version 24.0.0git
Git revision b010a18d2b648cab83c83967ff26b8fde11acdc6
Target: amdgcn-amd-amdhsa
CPU: gfx942
Optimized assertions build
Summary
llvm.amdgcn.raw.ptr.buffer.storeis modeled as an argmem-only write throughits AS8 resource operand. However, the effective memory location of the access
also depends on the
voffsetandsoffsetoperands, while the access size isdetermined by the stored type.
For this intrinsic, generic call-memory-location modeling does not preserve
that information precisely enough for DSE to prove that two calls completely
overwrite the same location, or for LICM to sink a repeated loop store.
Under the raw-descriptor restrictions required by AS7, expressing the same
access as an ordinary AS7 store enables the expected generic optimizations:
-O3replaces the loop with a guarded closed-form calculation and one finalstore.
This was initially observed in a controlled Triton compilation. Both paths
ultimately select
buffer_store_dwordongfx942, so the code-qualitydifference arises before instruction selection rather than from choosing a
different final store opcode.
Standalone LLVM reproducer
The pair functions provide the minimal DSE case. The loop functions demonstrate
the corresponding LICM and
-O3behavior.Commands:
opt -passes=dse -S repro.ll -o dse.ll opt -passes='function(loop-simplify,lcssa,loop-mssa(licm))' \ -S repro.ll -o licm.ll opt -O3 -mcpu=gfx942 -S repro.ll -o o3.llObserved results:
buffer_pair, but removes the firstAS7 store in
fat_pointer_pair.buffer_loop, but sinks the AS7 store infat_pointer_loopto the guarded loop exit.-O3unrolls the remaining raw-intrinsic loop into eight callsin the main loop and one in the remainder loop. The AS7 loop is replaced with
a closed-form calculation and one guarded store.
The exact unroll count is not essential to the issue. The missed optimization
is that repeated non-volatile stores to the same effective buffer location are
not reduced to the final store.
Likely limitation
This does not appear to be a case where the intrinsic is completely invisible
to the memory infrastructure. The declaration identifies it as an argmem-only
write and marks the AS8 resource argument
writeonlyandcaptures(none).The problem is that the effective location is composite:
Generic call-location handling can describe the resource pointer, but it has no
AMDGPU-specific representation of the two offset operands and the precise size
of this access. Consequently, two calls with identical resource and offset
operands are still not recognized as complete overwrites by DSE, and the loop
store is not sinkable by LICM.
Motivating Triton pipeline comparison
The original case was found by compiling the same
noaliasTriton kernel twicefor the HIP
gfx942target withnum_warps=4and LLVM-O3.Both runs enable AMD buffer operations. The second run additionally enables
Triton's experimental buffer-fat-pointer store path. Stages 01 and 02 are
byte-identical; the first representation-level difference appears in stage 03,
before translation to LLVM IR and before LLVM
-O3.rocdl.raw.ptr.buffer.storellvm.store-O3inputllvm.amdgcn.raw.ptr.buffer.storestore-O3outputbuffer_store_dwordinstructionsbuffer_store_dwordinstructionStage 03
The normal raw-buffer path forms a ROCDL raw buffer store:
The fat-pointer path retains an ordinary store:
Stage 05: LLVM
-O3inputThe stage-03 representations are carried through translation to the LLVM
optimizer input.
Raw-buffer path:
AS7 path:
Stage 06 and final AMDGCN
After
-O3, the raw-buffer path still contains the loop. Subsequent unrollingproduces eight intrinsic calls in the main loop and one in the remainder loop in
this build. The AS7 path contains no loop and has one store after a closed-form
calculation.
Both paths eventually select the same store instruction form:
Final static metrics:
buffer_store_dwordinstructionscodeLenInByteThe VGPR/SGPR counts, occupancy, and spill count are unchanged in this case.
For positive
I = iterations, the raw path dynamically issues one buffer storeper loop iteration for each executing wave, while the AS7 path issues one
guarded final store per wave. For
I <= 0, neither path stores.These are generated-code observations only; runtime performance has not been
measured on a GPU.
Scope and semantic assumptions
The demonstrated comparison is intentionally narrow:
add_tiddisabled, swizzling disabled, and extent measured in bytes.voffset, andsoffset.i32store withsoffset = 0.align 1, matching the Triton-generated IR and avoidingany additional base/offset alignment assumptions.
transformations that change partially out-of-bounds behavior.
This does not attempt to cover structured or swizzled descriptors, explicit
cache modifiers, volatile accesses, atomics, async LDS operations, masked
accesses, or OOB sentinel behavior.
Expected result / questions
Is AS7 ordinary memory IR intended to be the mid-end representation for raw
buffer accesses that satisfy the buffer-fat-pointer descriptor restrictions,
with
AMDGPULowerBufferFatPointersPassperforming the late lowering beforeinstruction selection?
If
llvm.amdgcn.raw.ptr.buffer.*intrinsics are also expected to pass throughthe generic optimizer, is there an intended target-specific mechanism for
representing or comparing their effective locations:
so that DSE and LICM can recognize complete overwrites in cases such as this?
I am not assuming that extending generic
MemoryLocationis necessarily theright solution. Preserving AS7 ordinary memory operations longer in frontends,
or adding an AMDGPU-specific optimization for raw buffer intrinsics, may be more
appropriate.
Version