AMDGPU: Fix true16 build_vector (0, x) pattern using a 16-bit shift operand
The real true16 pattern for (build_vector 0, VGPR_16:$x) fed the 16-bit
register directly to V_LSHLREV_B32, which takes a 32-bit operand. Widen
it with a REG_SEQUENCE first. This avoids redundant 16-bit moves in
SelectionDAG, and fixes a GlobalISel selection failure when the 16-bit
input is a G_TRUNC of a 32-bit value, as the shift's operand class
constrained the trunc result to vgpr_32.
I also don't know why this pattern is overcomplicating this. I would expect
true16 to literally translate build_vector to reg_sequence plus a materialize
of the 0.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
fix(X86): restore CCMP through boolean NOT
Demanded-bits simplification can turn XOR with 1 into NOT, hiding
SETCC operands from CCMP formation. Invert the condition under AND
when the other SETCC masks high bits and the NOT has one use.
Restore the original ccmp.ll checks for all four configurations.
Refs #230700
AMDGPU: Restore kernel tests in early-if-convert-cost.ll (#230764)
59911f438c47 converted these test kernels into functions using vgpr
function arguments, but this perturbed the code too much. This just
happened to run into a preexisting bug in later passes which appears
in expensive checks builds.
[AMDGPU] Don't select scalar loads for LDS/scratch (#227423)
A uniform load from LDS marked `!invariant.load` gets selected as
`s_load_dword`. SMEM can't access LDS, so it ends up reading global
memory at the LDS offset (address 0x0 + offset). On MI350 this faults.
```llvm
%v = load i32, ptr addrspace(3) %p, align 4, !invariant.load !0
```
```
s_mov_b32 s1, 0
s_load_dword s0, s[0:1], 0x0 ; LDS offset used as a global address
```
I hit this while marking LDS reads as `!invariant.load` from FlyDSL
(MLIR), to let LLVM rematerialize them instead of spilling.
Fix: only allow the SMRD load patterns for global, constant, and 32-bit
constant address spaces. GlobalISel already excluded LDS/scratch, so
[2 lines not shown]
AMDGPU: Restore kernel tests in early-if-convert-cost.ll
59911f438c47 converted these test kernels into functions using vgpr
function arguments, but this perturbed the code too much. This just
happened to run into a preexisting bug in later passes which appears
in expensive checks builds.