[mlir][linalg] Document and diagnose pack/unpack memref limits (#225773)
Scoped down from the [original
RFC](https://discourse.llvm.org/t/rfc-transformation-support-for-linalg-pack-linalg-unpack-on-memrefs/91832)
per discussion in #225650.
- Replace/add the `// TODO: Support Memref Pack/UnPackOp...` comment
across all sites with a comment stating the actual invariant, pointing
to #225650 for the reasoning.
- Document the invariant in the `Linalg_PackOp`/`Linalg_UnPackOp`
descriptions.
- Emit a dedicated diagnostic from `structured.pack`, `lower_pack`,
`lower_unpack`, and tiling when the target has memref operands, instead
of a generic/silent failure.
- Add test coverage for the new diagnostics, previously untested.
---------
Signed-off-by: Federico Bruzzone <federico.bruzzone.i at gmail.com>
[SlotIndexes] Add queries for stale indexes
An erased instruction leaves its index list entry in place, making the
index indistinguishable from a block boundary entry. Add
isBlockBoundaryIndex() and isStaleIndex() to tell the two apart, and
canonicalizeIndex() to resolve a stale index to the closest preceding
instruction's register slot, or the block start if none survives.
NFC. No caller yet. LiveDebugVariables is next.
[LiveDebugVariables] Repair stale SlotIndexes
The analysis keeps its indexes from before the first register allocator
until DBG_VALUEs are emitted, by which point passes in between have
erased some of the instructions they point at. Resolve them at the
start of each allocator run and before emitting.
SlotIndexes can then reclaim the entries of erased instructions without
sparing the ones held here, which would have made generated code depend
on -g. Emitted locations are unchanged, except that intervals resolving
to one position now emit a single DBG_VALUE rather than identical
consecutive ones.
[clang][AArch64] Fix range check for _Read/_WriteStatusReg (#227052)
PR #187290 added a range check of `[0x4000, 0x7fff]` to the
`_ReadStatusReg` and `_WriteStatusReg` builtins. Since `op0 == 2`
registers have bit 14 clear and encode below 0x4000, that check was
incorrect.
This change widens the range back to `[0x0, 0x7fff]` and adds a
regression test.
Fixes #226777
Assisted by: Claude Opus 5 (via VS Code).
[flang][cuda] Do not use cuf.alloc for derived-type function results (#227863)
Local variables of a derived type with device allocatable components
are allocated in managed memory with cuf.alloc/cuf.free. This also
applied to function results, but the storage of a derived-type function
result is replaced by the caller-provided buffer in the AbstractResult
pass. Allocating it with cuf.alloc is therefore incorrect, and the
matching cuf.free would release memory that is returned to the caller.
Skip function results in needCUDAAlloc so that they are lowered to a
regular fir.alloca. Explicit device/managed/shared/pinned attributes
on the symbol are still honored.
[mlir] Avoid rewriting unreachable blocks in the greedy driver
A rewrite can disconnect a block after the iteration's initial CFG
sweep. Queued operations can develop self-referential SSA uses in
unreachable code, causing crashes or repeated rewrites that prevent the
worklist pass from returning. Track reachability through rewriter
notifications and skip these operations until the next iteration removes
their blocks.
The reproducer in #221152 exposes this gap in the initial sweep added by
#153957 for #153732. #154038 similarly skips unreachable blocks in the
walk-based driver. Reachable graph-region self-cycles in #194824 and
#205064 remain separate; #207185 proposes a fold-specific fix for them.
A forwarding ReachabilityListener sits between the rewriter and the
worklist driver and records, per region, which blocks changed their
terminator and which blocks were removed. Queries happen between
rewrites, once per popped operation, and update a per-region cache of
reachable blocks and their successors incrementally: added edges extend
[41 lines not shown]
[flang][cuda] Keep length parameters when boxing in data transfer conversion (#227883)
When lowering cuf.data_transfer, CUFOpConversion creates descriptors for
non-descriptor operands in emboxSrc, emboxDst and asHLFIREntity. The
shape
was taken from the defining fir.declare/hlfir.declare, but the length
type
parameters were dropped. For a character entity with a non-constant
length
(e.g. an automatic `character(len=l) :: str(n)` assigned to a managed or
device array), this produced a fir.embox of !fir.char<1,?> without
typeparams, which later hit the `!lenParams.empty()` assertion in
EmboxCommonConversion::getCharacterByteSize during FIR to LLVM codegen.
Retrieve the type parameters from the declare alongside the shape and
pass
them when creating the box. Lengths already present in the type are
elided,
matching FirOpBuilder::createBox.
[clang] Avoid stack exhaustion in recursive `constexpr` calls (#201706)
Guard constexpr function-call evaluation with
`runWithSufficientStackSpace` so deeply recursive calls use a fresh
stack before exhausting the current one. This prevents crashes during
both constexpr analysis and constant folding during LLVM IR generation,
including recursive floating-point expressions.
Fixes #201418
Fixes #200673
Assisted by Codex.
[Clang] Initialize bypassed variables w/ trivial-auto-var-init (#181937)
When -ftrivial-auto-var-init=zero or -ftrivial-auto-var-init=pattern is
enabled, variables whose declarations are bypassed by goto or switch
statements were silently left uninitialized. This patch ensures they are
initialized, matching GCC 16's behavior.
The initialization is emitted at the jump source rather than the jump
target. This ensures correctness in loops: a goto whose source and
destination are both inside the variable's scope does not spuriously
reinitialize it, while a goto that actually bypasses the declaration
does. For computed gotos (where jump sources cannot be determined
statically), we fall back to initializing in the entry block.
The simplest example of the old behavior is:
```c
switch (x) {
int y;
case 1:
[15 lines not shown]
[LLVMABI] Classify vectors at their ABI size
The x86-64 classifier compared a vector's payload width against Clang type
sizes, so vectors with padding were misclassified. For example,
`struct { long double __attribute__((vector_size(16))) v; }` coerced to
`<2 x double>` instead of `<1 x x86_fp80>`.
getABISizeInBits() now counts an x87 element at its allocation size, and the
classifier uses it wherever Clang uses getTypeSize() for a vector. This also
applies to vectors with a non-power-of-two element count, bool vectors, and
vectors of sub-byte _BitInt. isIllegalVectorType now sends only __int128
vectors to memory, not _BitInt(128) ones, as Clang does. isSingleElementStruct
is shared, so AMDGPU picks up the fix too.
Assisted-by: Claude Code / Claude Opus 5.5
[LiveDebugVariables] Repair stale SlotIndexes
The analysis keeps its indexes from before the first register allocator
until DBG_VALUEs are emitted, by which point passes in between have
erased some of the instructions they point at. Resolve them at the
start of each allocator run and before emitting.
SlotIndexes can then reclaim the entries of erased instructions without
sparing the ones held here, which would have made generated code depend
on -g. Emitted locations are unchanged, except that intervals resolving
to one position now emit a single DBG_VALUE rather than identical
consecutive ones.
[SlotIndexes] Add queries for stale indexes
An erased instruction leaves its index list entry in place, making the
index indistinguishable from a block boundary entry. Add
isBlockBoundaryIndex() and isStaleIndex() to tell the two apart, and
canonicalizeIndex() to resolve a stale index to the closest preceding
instruction's register slot, or the block start if none survives.
NFC. No caller yet. LiveDebugVariables is next.
Fix variant context for lastprivate bounds
Use the owning directive's evaluation when collecting construct ancestors
for loop-control expressions. Lastprivate can re-evaluate bounds after
loop-body lowering leaves a body evaluation current, which otherwise adds
the loop construct before the loop-control context filter runs.
X86: Remove redundant SJLJ landing pad alignment in X86LFIRewritePass (#227295)
X86LFIRewritePass separately scanned for blocks holding a call site's landing pad label.
With SJLJ exception handling the dispatch block reaches those blocks through an indirect
jump, and they are no longer marked as EH pads by the time this pass runs, so they would
otherwise be missed.
EmitSjLjDispatchBlock puts those blocks in a jump table, so they are already aligned as jump
table targets. Dropping it removes the use of the TargetOptions exception model field from
this in preparation for its removal.
The new test checks the landing pad alignment, which was previously untested.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[Polly] Allocate packed arrays of matrix multiplication on heap (#226163)
Problem: The matrix multiplication optimization copies blocks of its
operands into the arrays Packed_A and Packed_B. Their sizes are derived
from the cache parameters rather than from the operands; with the
default parameters, Packed_B takes 4 MiB even for a 100x100 product.
They are allocated with alloca in the entry block of the function, so
two optimized multiplications in one function exceed the default 8 MiB
stack: two products of 100x100 int32 matrices declared as VLAs and
inlined into main crash with a segmentation fault.
Solution: Allocate a packed array on the heap, with malloc at the start
of the SCoP and free at its exit, only if it is larger than
-polly-pattern-matching-max-stack-array-size (1 MiB by default), and
keep smaller ones on the stack. -1 keeps all of them on the stack, 0
puts all of them on the heap. With the default parameters, Packed_B goes
to the heap and Packed_A (192 KiB) stays on the stack.
Default size depends only on the cache parameters. The default
[19 lines not shown]
[NVPTX] Add range attributes for cluster rank intrinsics to preserve signed non-negativity (#224727)
Adds `range` properties in `IntrinsicsNVVM.td` for:
* `llvm.nvvm.read.ptx.sreg.cluster.ctarank`
* `llvm.nvvm.read.ptx.sreg.cluster.nctarank`
The CTA rank is encoded in a single byte within a shared pointer, so the
max value of ctarank is 255 and the number of CTAs is one greater. This
enables InstCombine to apply signed-division optimizations (e.g. sdiv X,
16 -\> shift-based form) that require X \>= 0.
---------
Co-authored-by: Alex MacLean <amaclean at nvidia.com>
[Analysis] Delete SyntheticCountsUtils (#227741)
This isn't used anywhere, presumably getting dropped after we deleted
much of the rest of the synthetic profile infrastructure.
[libc] Handle PTHREAD_NULL in pthread_equal (#227856)
POSIX.1-2024 specifies that pthread_equal must accept PTHREAD_NULL for
either or both arguments. Previously, Thread::operator== dereferenced
attrib->tid, which crashed when passed a null or zero-initialized
pthread_t.
Compare Thread::attrib pointers directly in an inline constexpr
operator== in thread.h, matching how Thread::join and
pthread_getunique_np identify threads. Added PTHREAD_NULL test cases to
pthread_equal_test and updated file headers in the touched files.
Assisted-by: Automated tooling, human reviewed.
[LiveDebugVariables] Repair stale SlotIndexes
The analysis keeps its indexes from before the first register allocator
until DBG_VALUEs are emitted, by which point passes in between have
erased some of the instructions they point at. Resolve them at the
start of each allocator run and before emitting.
SlotIndexes can then reclaim the entries of erased instructions without
sparing the ones held here, which would have made generated code depend
on -g. Emitted locations are unchanged, except that intervals resolving
to one position now emit a single DBG_VALUE rather than identical
consecutive ones.
[SlotIndexes] Add queries for stale indexes
An erased instruction leaves its index list entry in place, making the
index indistinguishable from a block boundary entry. Add
isBlockBoundaryIndex() and isStaleIndex() to tell the two apart, and
canonicalizeIndex() to resolve a stale index to the closest preceding
instruction's register slot, or the block start if none survives.
NFC. No caller yet. LiveDebugVariables is next.
[flang][cuda] Count managed symbols per component chain in assignment check (#227825)
The check for a device to host transfer with more than one device object
on the right hand side compares the number of unique device symbols with
the number of managed or unified symbols, to allow the case where every
object is managed. The two counts were computed differently: device
symbols are counted once per component chain, while managed symbols were
counted individually. When both a derived type object and its component
are managed, the component chain contributed one device symbol but two
managed symbols, so the counts never matched and a valid assignment was
rejected:
```
type(t), managed, allocatable :: s(:) ! t has a managed component a
real, managed, allocatable :: b(:,:), c(:,:)
x = (b(i,j) - s(1)%a(i,j)) * c(i,j)
! error: More than one reference to a CUDA object on the right hand
! side of the assignment
```
[5 lines not shown]
[lldb] Fix adding shared ScriptInterpreter libraries to LLDB.framework (#223782)
These libraries were already being added to the framework in the build
tree but installation rules were still putting it in `/lib`. I pulled
the Python and Lua logic into a separate function and changed the
install logic when LLDB_BUILD_FRAMEWORK is enabled.
MIR: Serialize MachineBasicBlock::MaxBytesForAlignment (#227139)
Fix missing serialization of another field. The alignment was
already handled. The name is a bit verbose. Some places call it
"MaxSkip" which matches the name of the 2nd operand to the .p2align
directive this corresponds to.
Co-authored-by: Claude Sonnet 5 <noreply at anthropic.com>
[AMDGPU] Fold compares of S_CSELECT of two different constants (#227772)
If sX = S_CSELECT* A, B with A != B, then sX == A exactly when SCC was
set, so comparing sX with A or B recomputes SCC or its inverse:
s_cmp_eq_* sX, A => SCC s_cmp_lg_* sX, A => !SCC
s_cmp_eq_* sX, B => !SCC s_cmp_lg_* sX, B => SCC
Generalize optimizeCmpSelect to handle this for any pair of distinct
constants, instead of only S_CSELECT* (non-zero imm), 0 compared with 0.
Also try optimizeCmpSelect for S_CMP_EQ_U64.
This removes thousands of "s_cselect_b32 sN, 1, 0; s_cmp_lg_u32 sN, 1"
sequences from the AMDGPU codegen tests. They come from branches on an
inverted i1: the setcc against true is promoted to a compare with 1.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply at anthropic.com>