[MLIR][Bufferization] Fix IdentityLayoutMap allocation at function boundaries (#227253)
Resolves silent data corruption when passing subviews across function
boundaries under `LayoutMapOption::IdentityLayoutMap`.
### The Problem
When the `IdentityLayoutMap` option is specified for function boundary
bufferization, all function parameters are expected to have a fully
contiguous, zero-offset layout. However, if a caller passes a non-unit
stride or offset view (e.g. the result of `tensor.extract_slice`), the
bufferization pass incorrectly lowered this to a `memref.cast`.
Since `memref.cast` strips layout metadata but leaves the base pointer
unchanged, this caused silent wrong-value loads in the callee (reading
from offset 0 regardless of the actual dynamic offset).
### The Solution
This patch intercepts the operand materialization logic in
`FuncBufferizableOpInterfaceImpl.cpp` (specifically during `CallOp`
[19 lines not shown]
[mlir][ValueBounds] Skip analysis for identical slice components (#226894)
One-Shot Bufferize repeatedly compares subset slices whose offsets,
sizes, and strides often reuse the same SSA values and attributes. Avoid
constructing a ValueBounds constraint set when the two OpFoldResults are
already identical, while preserving the existing solver fallback for
distinct values.
[CIR][SYCL] Enable relocatable device code for SYCL (#226596)
Emit `sycl_external` functions with sycl-module-id, allow -fgpu-rdc
mangling, and embed offload objects in the host.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[lldb-dap][test] Let tests run under both stdio and server adapter modes (#227435)
Add create_debug_adapter(), which picks stdio or server mode based on
self.run_as_server. This allows tests to run under both modes when
toggling `LLDBDAP_RUN_AS_SERVER`, rather than being pinned to stdio.
[mlir][linalg] Document and diagnose pack/unpack memref limits (#225773)
Scoped down from the [original
RFC](https://discourse.llvm.org/t/rfc-transformation-support-for-linalg-pack-linalg-unpack-on-memrefs/91832)
per discussion in #225650.
- Replace/add the `// TODO: Support Memref Pack/UnPackOp...` comment
across all sites with a comment stating the actual invariant, pointing
to #225650 for the reasoning.
- Document the invariant in the `Linalg_PackOp`/`Linalg_UnPackOp`
descriptions.
- Emit a dedicated diagnostic from `structured.pack`, `lower_pack`,
`lower_unpack`, and tiling when the target has memref operands, instead
of a generic/silent failure.
- Add test coverage for the new diagnostics, previously untested.
---------
Signed-off-by: Federico Bruzzone <federico.bruzzone.i at gmail.com>
[SlotIndexes] Add queries for stale indexes
An erased instruction leaves its index list entry in place, making the
index indistinguishable from a block boundary entry. Add
isBlockBoundaryIndex() and isStaleIndex() to tell the two apart, and
canonicalizeIndex() to resolve a stale index to the closest preceding
instruction's register slot, or the block start if none survives.
NFC. No caller yet. LiveDebugVariables is next.
[LiveDebugVariables] Repair stale SlotIndexes
The analysis keeps its indexes from before the first register allocator
until DBG_VALUEs are emitted, by which point passes in between have
erased some of the instructions they point at. Resolve them at the
start of each allocator run and before emitting.
SlotIndexes can then reclaim the entries of erased instructions without
sparing the ones held here, which would have made generated code depend
on -g. Emitted locations are unchanged, except that intervals resolving
to one position now emit a single DBG_VALUE rather than identical
consecutive ones.
[clang][AArch64] Fix range check for _Read/_WriteStatusReg (#227052)
PR #187290 added a range check of `[0x4000, 0x7fff]` to the
`_ReadStatusReg` and `_WriteStatusReg` builtins. Since `op0 == 2`
registers have bit 14 clear and encode below 0x4000, that check was
incorrect.
This change widens the range back to `[0x0, 0x7fff]` and adds a
regression test.
Fixes #226777
Assisted by: Claude Opus 5 (via VS Code).
[flang][cuda] Do not use cuf.alloc for derived-type function results (#227863)
Local variables of a derived type with device allocatable components
are allocated in managed memory with cuf.alloc/cuf.free. This also
applied to function results, but the storage of a derived-type function
result is replaced by the caller-provided buffer in the AbstractResult
pass. Allocating it with cuf.alloc is therefore incorrect, and the
matching cuf.free would release memory that is returned to the caller.
Skip function results in needCUDAAlloc so that they are lowered to a
regular fir.alloca. Explicit device/managed/shared/pinned attributes
on the symbol are still honored.
[mlir] Avoid rewriting unreachable blocks in the greedy driver
A rewrite can disconnect a block after the iteration's initial CFG
sweep. Queued operations can develop self-referential SSA uses in
unreachable code, causing crashes or repeated rewrites that prevent the
worklist pass from returning. Track reachability through rewriter
notifications and skip these operations until the next iteration removes
their blocks.
The reproducer in #221152 exposes this gap in the initial sweep added by
#153957 for #153732. #154038 similarly skips unreachable blocks in the
walk-based driver. Reachable graph-region self-cycles in #194824 and
#205064 remain separate; #207185 proposes a fold-specific fix for them.
A forwarding ReachabilityListener sits between the rewriter and the
worklist driver and records, per region, which blocks changed their
terminator and which blocks were removed. Queries happen between
rewrites, once per popped operation, and update a per-region cache of
reachable blocks and their successors incrementally: added edges extend
[41 lines not shown]
[AMDGPU] Price scalar integer to fp casts by source width and sign
Scalar sources between a byte and 31 bits fell to the default cost of one
while the matching vector lanes were already priced, which skewed the
difference SLP weighs a bundle against. Such a source is extended before
the conversion, and what the extension takes depends on the width, on the
sign and on whether the subtarget has SDWA and 16 bit instructions.
Sources narrower than a byte are left alone, because their vector form is
not priced either.
[flang][cuda] Keep length parameters when boxing in data transfer conversion (#227883)
When lowering cuf.data_transfer, CUFOpConversion creates descriptors for
non-descriptor operands in emboxSrc, emboxDst and asHLFIREntity. The
shape
was taken from the defining fir.declare/hlfir.declare, but the length
type
parameters were dropped. For a character entity with a non-constant
length
(e.g. an automatic `character(len=l) :: str(n)` assigned to a managed or
device array), this produced a fir.embox of !fir.char<1,?> without
typeparams, which later hit the `!lenParams.empty()` assertion in
EmboxCommonConversion::getCharacterByteSize during FIR to LLVM codegen.
Retrieve the type parameters from the declare alongside the shape and
pass
them when creating the box. Lengths already present in the type are
elided,
matching FirOpBuilder::createBox.
[clang] Avoid stack exhaustion in recursive `constexpr` calls (#201706)
Guard constexpr function-call evaluation with
`runWithSufficientStackSpace` so deeply recursive calls use a fresh
stack before exhausting the current one. This prevents crashes during
both constexpr analysis and constant folding during LLVM IR generation,
including recursive floating-point expressions.
Fixes #201418
Fixes #200673
Assisted by Codex.
[Clang] Initialize bypassed variables w/ trivial-auto-var-init (#181937)
When -ftrivial-auto-var-init=zero or -ftrivial-auto-var-init=pattern is
enabled, variables whose declarations are bypassed by goto or switch
statements were silently left uninitialized. This patch ensures they are
initialized, matching GCC 16's behavior.
The initialization is emitted at the jump source rather than the jump
target. This ensures correctness in loops: a goto whose source and
destination are both inside the variable's scope does not spuriously
reinitialize it, while a goto that actually bypasses the declaration
does. For computed gotos (where jump sources cannot be determined
statically), we fall back to initializing in the entry block.
The simplest example of the old behavior is:
```c
switch (x) {
int y;
case 1:
[15 lines not shown]
[LLVMABI] Classify vectors at their ABI size
The x86-64 classifier compared a vector's payload width against Clang type
sizes, so vectors with padding were misclassified. For example,
`struct { long double __attribute__((vector_size(16))) v; }` coerced to
`<2 x double>` instead of `<1 x x86_fp80>`.
getABISizeInBits() now counts an x87 element at its allocation size, and the
classifier uses it wherever Clang uses getTypeSize() for a vector. This also
applies to vectors with a non-power-of-two element count, bool vectors, and
vectors of sub-byte _BitInt. isIllegalVectorType now sends only __int128
vectors to memory, not _BitInt(128) ones, as Clang does. isSingleElementStruct
is shared, so AMDGPU picks up the fix too.
Assisted-by: Claude Code / Claude Opus 5.5
[LiveDebugVariables] Repair stale SlotIndexes
The analysis keeps its indexes from before the first register allocator
until DBG_VALUEs are emitted, by which point passes in between have
erased some of the instructions they point at. Resolve them at the
start of each allocator run and before emitting.
SlotIndexes can then reclaim the entries of erased instructions without
sparing the ones held here, which would have made generated code depend
on -g. Emitted locations are unchanged, except that intervals resolving
to one position now emit a single DBG_VALUE rather than identical
consecutive ones.
[SlotIndexes] Add queries for stale indexes
An erased instruction leaves its index list entry in place, making the
index indistinguishable from a block boundary entry. Add
isBlockBoundaryIndex() and isStaleIndex() to tell the two apart, and
canonicalizeIndex() to resolve a stale index to the closest preceding
instruction's register slot, or the block start if none survives.
NFC. No caller yet. LiveDebugVariables is next.
Fix variant context for lastprivate bounds
Use the owning directive's evaluation when collecting construct ancestors
for loop-control expressions. Lastprivate can re-evaluate bounds after
loop-body lowering leaves a body evaluation current, which otherwise adds
the loop construct before the loop-control context filter runs.
X86: Remove redundant SJLJ landing pad alignment in X86LFIRewritePass (#227295)
X86LFIRewritePass separately scanned for blocks holding a call site's landing pad label.
With SJLJ exception handling the dispatch block reaches those blocks through an indirect
jump, and they are no longer marked as EH pads by the time this pass runs, so they would
otherwise be missed.
EmitSjLjDispatchBlock puts those blocks in a jump table, so they are already aligned as jump
table targets. Dropping it removes the use of the TargetOptions exception model field from
this in preparation for its removal.
The new test checks the landing pad alignment, which was previously untested.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[NFC][AMDGPU] Add more tests for int to fp casts of loaded odd width integers
Covers loads of i24, i40, i48 and i56 converted to fp, both the loads
that are split into narrower extending loads and the constant or
invariant ones that are widened to a scalar load.