CodeGen: Remove dead LiveVariables plumbing from MachineSink
MachineSink threaded a LiveVariables pointer through to
SplitCriticalEdge so the analysis would be updated. It never used
LiveVariables for any decision, and MachineSinking runs before
LiveVariables pass in every pipeline.
Co-authored-by: Claude (Claude-Opus-4.8)
AMDGPU: Use LiveIntervals in SIOptimizeVGPRLiveRange when available
LiveVariables has been long deprecated. Use LiveIntervals if available.
With the current pass structure, this will use LiveVariables.
Co-authored-by: Claude (Claude-Opus-4.8)
AMDGPU: Remove deprecated getArchAttr and ArchFeatures TableGen (#222462)
Everything should now use getFeatureBitset*
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
[libFuzzer] Modify stop-file.test to support remote devices (#222012)
Currently this test fails when running on remote devices because the
'rm' command only removes the file from the host, and the second part of
the test expects the file to have been removed. This patch changes the later
run command to point to a non-existent file in order to work around this.
rdar://186683123
[AArch64][NEON] Fold insert(zero, extract(X, 0), 0) -> X when X is known to zero lanes 1-N (#213940)
Add patterns for NEON across-lane reductions whose scalar result
is inserted into lane zero of a zero vector.
This is valid because these instructions place their result in lane zero
and
clear the remaining lanes.
Limit this change to matching input and result vector types. Mixed
vector
sizes and scalar conversions are not part of this PR.
Handle llvm.vector.reduce.add.v2i64 separately because it lowers to
ADDP.
NEON has no equivalent instruction for v2i64 signed or unsigned min/max
reductions, so they cannot use this fold. Double-precision
floating-point
reductions are already handled by pairwise instruction patterns.
[SLP][modularisation][NFC] Move BoUpSLP class declaration to SLPTree.h
Move the BoUpSLP class declaration and its nested types out of
SLPVectorizer.cpp into SLPVectorizer/SLPTree.h. This is a pure relocation:
method definitions stay in SLPVectorizer.cpp, and the prerequisite changes
(de-inlining the cl::opt users; seeding SLPTree.h with ReductionVectorPart
and MinScheduleRegionSize) landed earlier in the stack.
The header is included after DEBUG_TYPE is defined because BoUpSLP inline
methods use LLVM_DEBUG. The DenseMapInfo/GraphTraits specializations remain
in SLPVectorizer.cpp.
Part of the SLPVectorizer.cpp modularization effort:
https://discourse.llvm.org/t/modularizing-slpvectorizer-cpp/90922
[MLIR][Linalg] Remove Linalg named ops (#220916)
Removes the named ops from the Linalg dialect.
I have also updated the ElementwiseOp builder to simplify the default
case: kind + no affine map.
Depends on both unary and binary removal branches. #220905 #220912
This PR also has the final cleanup, which turned out to be simple
enough.
Ref:
https://discourse.llvm.org/t/rfc-update-semantics-of-linalg-named-operations-unary-binary-ternary/91531
Assisted by: Claude Opus
X86: Mark EFLAGS dead on MOV32r0 emitted outside SelectionDAG (#222464)
Currently these get set by LiveVariables after the fact, but
ideally we would not rely on that since it's long overdue for
deletion.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
[AMDGPU] Use new CSR cost calculation (#219220)
So far, AMDGPU backend relied on the legacy calculation for the cost of
the first use of a callee-save register. This legacy path scales the
entry block frequency by a fixed value, making a comparison with
alternatives, e.g., rematerialization, difficult due to different
scales. One example for that is the work in
https://github.com/llvm/llvm-project/pull/206756, where comparison of
rematerialization cost and first use of a CSR is hard to compare due to
the different scales.
Change to instead use the new calculation path with different target
hooks for better comparability.
As the new scale is different, the cost was changed. The new cost value
accounts for various factors impacting CSR cost.
This change (with different cost value) was originally part of
https://github.com/llvm/llvm-project/pull/202007. Another part of that
[5 lines not shown]
[libc++][pstl] Implementation of parallel std::minmax_element() based on parallel reduce (#221572)
This PR adds an implementation of parallel `std::minmax_element()` based
on parallel reduce.
The implementation formulates `minmax_element` as a reduction of
iterators over the input range.
The iterators are first transformed into a min-max iterator pair.
The reduction is provided in two forms: 2-element reduction that makes
an iterator pair pointing to a smaller and a greater value, and a
range-based reduction with init value.
The latter delegates the heavy-lifting to the serial implementation of
`std::minmax_element()`.
Part of #99938.
libclc: Update fmod implementations (#222369)
This was originally ported from rocm device libs in
93af966747b59d37c57312a0c0242151076c072b. Merge in more
recent changes. This should also approximately match the default
expansion in ExpandIRInsts
Co-authored-by: Claude <noreply at anthropic.com>
AMDGPU: Remove deprecated getArchAttr and ArchFeatures TableGen
Everything should now use getFeatureBitset*
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>