[SLP] Add tests for affine loop address computations (NFC) (#226907)
Precommit tests for #226220: related affine address recurrences in a
loop whose index arithmetic SLP currently vectorizes and extracts per
lane for scalar loads and stores, which LSR would otherwise turn into
pointer increments.
`X86/strength-reducible-address.ll` has Haswell and skylake-avx512 RUN
lines. It covers strided loads, constant-offset GEPs, and a load/store
at the same address. The negative cases are loop-loaded offsets,
escaping addresses, and offsets computed in the preheader. The checks
are generated with current main.
[BasicAA] Compute minimal access extents lazily (#226931)
`aliasCheck` computes the minimal extent of an access before calling
`isObjectSmallerThan`, even though the latter immediately returns false
when the other pointer is not an identified object.
Move the extent computation behind the existing `isIdentifiedObject`
check. This preserves the result: whenever the extent can affect the
answer, the same value is computed and used as before; otherwise only
unused work is skipped.
[lldb][test] Fix undefined behaviour in multi process driver (#226949)
This code was using strcmp on function names. Function names can be null
(see SBFrame::GetFunctionName), and passing null to strcmp is UB.
[MergeFunctions] Add debug locations to redirected calls (#225625)
A call to a function without debug info does not need a debug location,
even in a function with debug info. If MergeFunctions redirects such a
call to an equivalent function that has debug info, the verifier fails:
```
inlinable function call in a function with debug info must have a !dbg location
```
Before redirecting such a call, give it a line 0 location in the scope
of the caller, as `DwarfEHPrepare` does for the call to the rewind
function. This covers both ways calls are redirected:
`replaceDirectCallers()`, and `replaceAllUsesWith()` for `unnamed_addr`
functions.
[SimplifyCFG] Drop UB-implying attributes when hoisting speculatively (#226746)
`hoistCommonCodeFromSuccessors()` can hoist identical instructions past
non-identical ones that it skips. If a skipped instruction may throw or
not return, the hoisted instruction is speculated, but it keeps its
UB-implying attributes and metadata, like `noundef`. Drop them in that
case, as SimplifyCFG already does when it speculates.
[SCEV] Rename SCEVFlags -> SCEVFlagsPair (NFC) (#226935)
Make it clear that it's a pair of no-wrap flags, and free up the
SCEVFlags name for renaming SCEVNoWrapFlags to it.
RegisterPressure: Detect dead physreg defs from LiveIntervals
This reverts the remainder of #222627, which was partially reverted by
on dead flags. This is a prerequisite to deleting LiveVariables.
When constructing the PressureDiff for an instruction during scheduling
DAG construction, dead defs were only recognized from the dead flag on the
operand. This implicitly relied on preprocessing done by LiveVariables to
fixup inconsistent dead flags with overlapping registers in other operands.
Dead flags have no verifier-enforced rules and are thus unreliable.
Before LiveVariables, consider this example:
dead $eax = MOV32r0 implicit-def dead $eflags, implicit-def $rax
; $rax is never used
$rax is never used, but only the $eax def is dead-flagged and the overlapping
implicit-def $rax is not. The shared $eax register units are covered by the
non-dead $rax def and so are counted as live defs. That shared unit is then
[26 lines not shown]
[AsmParser] Remove TBAA auto-upgrade in IR parser (#226914)
At this point, the struct path TBAA format is more than a decade old, so
remove support for upgrading from the old format in the textual IR
parser. Bitcode upgrade support is retained, as bitcode is backwards
compatible.
This avoids introduction of new tests using the old format. Existing
tests were migrated in https://github.com/llvm/llvm-project/pull/226051.
[X86] Don't narrow a foldable wide vector load to a scalar FP element
When only a single element of a wide vector load is demanded, a narrowing
combine can scalarize the load. On X86 this pessimizes cases where a
>128-bit load feeds a vector-producing operation (shuffle, permute, blend,
insertps): the wide load would otherwise fold a cheap 128-bit subvector
load into that instruction, but scalarizing forces a broadcast/scalar load
of one lane that must then be reinserted into a vector.
Teach shouldReduceLoadWidth to refuse narrowing a 256/512-bit load with a
vector-producing use all the way down to a scalar FP element. This is
restricted to loads wider than a vector register to preserve the
beneficial 128-bit-to-scalar movss narrowings.
This is NFC on its own; it prevents regressions once load scalarization is
enabled in SimplifyDemandedVectorElts.
Co-authored-by: Claude (Claude-Sonnet-5.0) <noreply at anthropic.com>
DAG: Fold extract_vector_elt of bitcasted scalar-rebuilt vector
When SimplifyDemandedVectorElts narrows a wide vector load to a scalar
load, it rebuilds the result with insert_vector_elt(undef, X, i), which
often becomes a build_vector. An extract_vector_elt through a bitcast of
that vector (e.g. extract i32, elt 2, of bitcast <2 x i64> to <4 x i32>)
was left as a full vector materialization plus a lane extract instead of
folding back to the originating scalar.
Fold extract_vector_elt (bitcast (build_vector/insert_vector_elt ... X)),
i to a trunc/srl of the wider scalar source operand X that contains the
extracted element. Matching against the insert_vector_elt/build_vector
source directly keeps the fold reachable at the point the extract is
naturally revisited, without re-queueing users.
Recovers scalar extracts across several targets (X86, AArch64, ARM,
PowerPC, AMDGPU), e.g. turning a movq+pshufd+movd or vmovddup+vextractps
sequence back into a single narrow load or GPR shift.
Co-authored-by: Claude (Claude-Sonnet-5.0) <noreply at anthropic.com>
DAG: Handle load in SimplifyDemandedVectorElts
This improves some AMDGPU cases and avoids future regressions.
The combiner likes to form shuffles for cases where an extract_vector_elt
would do perfectly well, and this recovers some of the regressions from
losing load narrowing.
AMDGPU, Arch64 and RISCV test changes look broadly better. Other targets have
some improvements, but mostly regressions. In particular X86 looks much
worse. I'm guessing this is because it's shouldReduceLoadWidth is wrong.
I mostly just regenerated the checks. I assume some set of them should
switch to use volatile loads to defeat the optimization.
Co-authored-by: Claude (Claude-Sonnet-5.0) <noreply at anthropic.com>
[X86] Shrink i32 unsigned compares to i8/i16 when high bits are zero (#197801)
Shrink i32 compares to 8 or 16 bit for non-constant operands when high
bits are zero, following the precedent of i64 -> i32 case.
Fixes the cases in #196827 where both sides of the compare are
non-constants, as this allows the `and(X, mask) `to be eliminated
downstream in DAGCombine.
I believe we can extend to constant operands as long as the length
changing prefix problem for i16 immediates is accounted for - however:
1. it's not sufficient to fix the remaining cases in the motivating
issue alone: when the rhs is constant, InstCombine clear the bits in the
mask that does not affect the comparison result (`icmp (and X, 255),
123` -> `icmp (and X, 252), 123`), so the mask does not fill the full
width of the truncated type, making the analysis trickier.
2. it leads to a lot of check line drifts, i'm not sure if it's worth
the effort to account for all of them, even if most don't look like
regressions.
AMDGPU: Make rewrite-vgpr-mfma-to-agpr-spill-multi-store.ll less allocator sensitive (#226913)
This test is sensitive to the exact split and spills which occur, and disappeared under
a future upstream improvement. Use basic RA with a fixed occupancy since it more
stably produces the spill pattern.
Also add a codegen reference test for the same kernel, so future codegen
improvements are visible.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[LiveDebugVariables] Repair stale SlotIndexes
The analysis keeps its indexes from before the first register allocator
until DBG_VALUEs are emitted, by which point passes in between have
erased some of the instructions they point at. Resolve them at the
start of each allocator run and before emitting.
SlotIndexes can then reclaim the entries of erased instructions without
sparing the ones held here, which would have made generated code depend
on -g. Emitted locations are unchanged, except that intervals resolving
to one position now emit a single DBG_VALUE rather than identical
consecutive ones.
[SlotIndexes] Add queries for stale indexes
An erased instruction leaves its index list entry in place, making the
index indistinguishable from a block boundary entry. Add
isBlockBoundaryIndex() and isStaleIndex() to tell the two apart, and
canonicalizeIndex() to resolve a stale index to the closest preceding
instruction's register slot, or the block start if none survives.
NFC. No caller yet. LiveDebugVariables is next.
[flang][Test] Cover the lowering of loops with a non-terminating body
A previous change leaves such a loop unstructured. Check what that produces:
the cycle survives as a block branching to itself, no fir.do_loop is emitted
for the loop control, and a loop that only needs a block of its own still
gets the structured form with its body in a region.
[flang] Lower loops whose branching is confined to their body structurally
Such a loop was classified separately by a previous change but still
lowered as unstructured, so its structured form was lost.
Lower it structurally instead, with its body folded into a region that can
hold the branching. The loop keeps its bounds on the op, so it remains
available to whatever transforms or parallelizes it. Only the body is
folded: the loop control statements are emitted as they are for any
structured loop, since a branch from outside may target either of them.
Loops an OpenACC or OpenMP directive owns are lowered the same way, so
they keep their form too.
[flang] Keep loops with a non-terminating body unstructured
A loop body that cannot run to completion must not be folded into an
scf.execute_region: the region carries no memory effects, so DCE deletes
it outright, dropping the non-termination and letting execution fall past
the loop. Branches survive that, being terminators.
An infinite DO was already rejected. Follow chains of unconditional GO TOs
as well and reject a body whose chain closes on itself, which is the same
bound the cf.br canonicalization applies to cyclic branches.
[flang] Honor the execute-region wrap flag when detecting loop internals
A loop whose branching is confined to its body is lowered with that body
in an scf.execute_region, since the CFG needs more than the single block
fir.do_loop's region admits. With the wrap disabled there is nowhere to
put those blocks, so skip the reclassification and leave the loop
unstructured.
[flang] Resolve an assigned GO TO's targets from the completed assign map
An assigned GO TO reaches any label ASSIGNed to its variable, and a label
list does not bound that: lowering allows a branch to any ASSIGNed label
whether or not the list names it. Branch analysis only sees the ASSIGNs
preceding the GO TO in program order, so the successors it records, and the
incoming branches derived from them, can be incomplete.
The symbol-to-labels map is complete once branch analysis has finished,
which is when the classification runs. Ask it for the full target set
instead of trusting the recorded successors, so a loop whose assigned GO TO
stays within its body is still recognised.
[flang] Detect loops whose branching is confined to their body
A DO loop is classified as either structured or unstructured, and a single
raw branch anywhere in its body forces the loop -- and every construct
enclosing it -- onto the unstructured path.
That is stronger than necessary. A loop keeps its structured control flow
as long as its branching neither leaves its body nor enters it from
outside. Classify such a loop separately from a fully unstructured one.
This only classifies: lowering is unchanged. PFT dumps mark the new
classification with '~', which is what the tests key on.