[MergeFunctions] Add debug locations to redirected calls (#225625)
A call to a function without debug info does not need a debug location,
even in a function with debug info. If MergeFunctions redirects such a
call to an equivalent function that has debug info, the verifier fails:
```
inlinable function call in a function with debug info must have a !dbg location
```
Before redirecting such a call, give it a line 0 location in the scope
of the caller, as `DwarfEHPrepare` does for the call to the rewind
function. This covers both ways calls are redirected:
`replaceDirectCallers()`, and `replaceAllUsesWith()` for `unnamed_addr`
functions.
[SimplifyCFG] Drop UB-implying attributes when hoisting speculatively (#226746)
`hoistCommonCodeFromSuccessors()` can hoist identical instructions past
non-identical ones that it skips. If a skipped instruction may throw or
not return, the hoisted instruction is speculated, but it keeps its
UB-implying attributes and metadata, like `noundef`. Drop them in that
case, as SimplifyCFG already does when it speculates.
[SCEV] Rename SCEVFlags -> SCEVFlagsPair (NFC) (#226935)
Make it clear that it's a pair of no-wrap flags, and free up the
SCEVFlags name for renaming SCEVNoWrapFlags to it.
RegisterPressure: Detect dead physreg defs from LiveIntervals
This reverts the remainder of #222627, which was partially reverted by
on dead flags. This is a prerequisite to deleting LiveVariables.
When constructing the PressureDiff for an instruction during scheduling
DAG construction, dead defs were only recognized from the dead flag on the
operand. This implicitly relied on preprocessing done by LiveVariables to
fixup inconsistent dead flags with overlapping registers in other operands.
Dead flags have no verifier-enforced rules and are thus unreliable.
Before LiveVariables, consider this example:
dead $eax = MOV32r0 implicit-def dead $eflags, implicit-def $rax
; $rax is never used
$rax is never used, but only the $eax def is dead-flagged and the overlapping
implicit-def $rax is not. The shared $eax register units are covered by the
non-dead $rax def and so are counted as live defs. That shared unit is then
[26 lines not shown]
[AsmParser] Remove TBAA auto-upgrade in IR parser (#226914)
At this point, the struct path TBAA format is more than a decade old, so
remove support for upgrading from the old format in the textual IR
parser. Bitcode upgrade support is retained, as bitcode is backwards
compatible.
This avoids introduction of new tests using the old format. Existing
tests were migrated in https://github.com/llvm/llvm-project/pull/226051.
[X86] Don't narrow a foldable wide vector load to a scalar FP element
When only a single element of a wide vector load is demanded, a narrowing
combine can scalarize the load. On X86 this pessimizes cases where a
>128-bit load feeds a vector-producing operation (shuffle, permute, blend,
insertps): the wide load would otherwise fold a cheap 128-bit subvector
load into that instruction, but scalarizing forces a broadcast/scalar load
of one lane that must then be reinserted into a vector.
Teach shouldReduceLoadWidth to refuse narrowing a 256/512-bit load with a
vector-producing use all the way down to a scalar FP element. This is
restricted to loads wider than a vector register to preserve the
beneficial 128-bit-to-scalar movss narrowings.
This is NFC on its own; it prevents regressions once load scalarization is
enabled in SimplifyDemandedVectorElts.
Co-authored-by: Claude (Claude-Sonnet-5.0) <noreply at anthropic.com>
DAG: Fold extract_vector_elt of bitcasted scalar-rebuilt vector
When SimplifyDemandedVectorElts narrows a wide vector load to a scalar
load, it rebuilds the result with insert_vector_elt(undef, X, i), which
often becomes a build_vector. An extract_vector_elt through a bitcast of
that vector (e.g. extract i32, elt 2, of bitcast <2 x i64> to <4 x i32>)
was left as a full vector materialization plus a lane extract instead of
folding back to the originating scalar.
Fold extract_vector_elt (bitcast (build_vector/insert_vector_elt ... X)),
i to a trunc/srl of the wider scalar source operand X that contains the
extracted element. Matching against the insert_vector_elt/build_vector
source directly keeps the fold reachable at the point the extract is
naturally revisited, without re-queueing users.
Recovers scalar extracts across several targets (X86, AArch64, ARM,
PowerPC, AMDGPU), e.g. turning a movq+pshufd+movd or vmovddup+vextractps
sequence back into a single narrow load or GPR shift.
Co-authored-by: Claude (Claude-Sonnet-5.0) <noreply at anthropic.com>
DAG: Handle load in SimplifyDemandedVectorElts
This improves some AMDGPU cases and avoids future regressions.
The combiner likes to form shuffles for cases where an extract_vector_elt
would do perfectly well, and this recovers some of the regressions from
losing load narrowing.
AMDGPU, Arch64 and RISCV test changes look broadly better. Other targets have
some improvements, but mostly regressions. In particular X86 looks much
worse. I'm guessing this is because it's shouldReduceLoadWidth is wrong.
I mostly just regenerated the checks. I assume some set of them should
switch to use volatile loads to defeat the optimization.
Co-authored-by: Claude (Claude-Sonnet-5.0) <noreply at anthropic.com>
[X86] Shrink i32 unsigned compares to i8/i16 when high bits are zero (#197801)
Shrink i32 compares to 8 or 16 bit for non-constant operands when high
bits are zero, following the precedent of i64 -> i32 case.
Fixes the cases in #196827 where both sides of the compare are
non-constants, as this allows the `and(X, mask) `to be eliminated
downstream in DAGCombine.
I believe we can extend to constant operands as long as the length
changing prefix problem for i16 immediates is accounted for - however:
1. it's not sufficient to fix the remaining cases in the motivating
issue alone: when the rhs is constant, InstCombine clear the bits in the
mask that does not affect the comparison result (`icmp (and X, 255),
123` -> `icmp (and X, 252), 123`), so the mask does not fill the full
width of the truncated type, making the analysis trickier.
2. it leads to a lot of check line drifts, i'm not sure if it's worth
the effort to account for all of them, even if most don't look like
regressions.
AMDGPU: Make rewrite-vgpr-mfma-to-agpr-spill-multi-store.ll less allocator sensitive (#226913)
This test is sensitive to the exact split and spills which occur, and disappeared under
a future upstream improvement. Use basic RA with a fixed occupancy since it more
stably produces the spill pattern.
Also add a codegen reference test for the same kernel, so future codegen
improvements are visible.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[LiveDebugVariables] Repair stale SlotIndexes
The analysis keeps its indexes from before the first register allocator
until DBG_VALUEs are emitted, by which point passes in between have
erased some of the instructions they point at. Resolve them at the
start of each allocator run and before emitting.
SlotIndexes can then reclaim the entries of erased instructions without
sparing the ones held here, which would have made generated code depend
on -g. Emitted locations are unchanged, except that intervals resolving
to one position now emit a single DBG_VALUE rather than identical
consecutive ones.
[SlotIndexes] Add queries for stale indexes
An erased instruction leaves its index list entry in place, making the
index indistinguishable from a block boundary entry. Add
isBlockBoundaryIndex() and isStaleIndex() to tell the two apart, and
canonicalizeIndex() to resolve a stale index to the closest preceding
instruction's register slot, or the block start if none survives.
NFC. No caller yet. LiveDebugVariables is next.
net/amnezia-kmod: Update 2.0.12 => 3.1.0
Changelog:
* if_amn: add AmneziaWG v3 protocol support
* Extend the FreeBSD kernel module with the protocol and configuration
features required for interoperability with AmneziaWG v3, while retaining
compatibility with existing AWG2 userspace.
* Implement AWG3 header protection using a small standalone ChaCha20-IETF
implementation. Use the random junk prefix as the 96-bit nonce and apply
header protection to handshake, cookie, and transport headers. Require
S1-S4 to provide enough nonce material whenever a header protection key
is configured. Add a known-answer ChaCha20 self-test to the module's
existing cryptographic self-test suite.
* Add support for the new AWG3 configuration parameters:
- HeaderProtectionKey
- ContentPaddingAddition
- RekeyAfterTime
- RekeyTimeout
- RejectAfterTime
[11 lines not shown]
[flang][Test] Cover the lowering of loops with a non-terminating body
A previous change leaves such a loop unstructured. Check what that produces:
the cycle survives as a block branching to itself, no fir.do_loop is emitted
for the loop control, and a loop that only needs a block of its own still
gets the structured form with its body in a region.
[flang] Lower loops whose branching is confined to their body structurally
Such a loop was classified separately by a previous change but still
lowered as unstructured, so its structured form was lost.
Lower it structurally instead, with its body folded into a region that can
hold the branching. The loop keeps its bounds on the op, so it remains
available to whatever transforms or parallelizes it. Only the body is
folded: the loop control statements are emitted as they are for any
structured loop, since a branch from outside may target either of them.
Loops an OpenACC or OpenMP directive owns are lowered the same way, so
they keep their form too.
[flang] Keep loops with a non-terminating body unstructured
A loop body that cannot run to completion must not be folded into an
scf.execute_region: the region carries no memory effects, so DCE deletes
it outright, dropping the non-termination and letting execution fall past
the loop. Branches survive that, being terminators.
An infinite DO was already rejected. Follow chains of unconditional GO TOs
as well and reject a body whose chain closes on itself, which is the same
bound the cf.br canonicalization applies to cyclic branches.
[flang] Honor the execute-region wrap flag when detecting loop internals
A loop whose branching is confined to its body is lowered with that body
in an scf.execute_region, since the CFG needs more than the single block
fir.do_loop's region admits. With the wrap disabled there is nowhere to
put those blocks, so skip the reclassification and leave the loop
unstructured.