[SLP]Fix regressions after the fmul/fadd fusion cost changes
Revisit the standalone seeds skipped as dependent on the lanes of a
vectorized window, instead of dropping them until the next SLP run.
Do not credit the fusion of the root of a loop-carried fadd reduction:
the chain through the loop phi is bound by its latency, which the
vector code cuts.
Fixes regressions in RAJAPerf.
Assisted-by: Cursor
Reviewers: RKSimon, bababuck
Pull Request: https://github.com/llvm/llvm-project/pull/228720
[clang] Migrate away from PointerUnion::dyn_cast (NFC) (#228843)
Note that PointerUnion::dyn_cast has been soft deprecated in
PointerUnion.h:
// FIXME: Replace the uses of is(), get() and dyn_cast() with
// isa<T>, cast<T> and the llvm::dyn_cast<T>
Literal migration would result in dyn_cast_if_present (see the
definition of PointerUnion::dyn_cast), but this patch uses dyn_cast.
Note that the if-then-else chain ends with:
} else
llvm_unreachable("Unhandled CausingFact type");
which guarantees that CausingFact is nonnull.
Assisted-by: Antigravity
[CodeGen] Remove unused functions (NFC) (#228844)
LiveVariables::replaceKillInstruction: The last caller was removed on
October 1, 2026 in commit 8068182dd6fb1a8efae0291db1d2462d29025cd1.
TargetLoweringBase::getSupportedLibcallImpl: The last caller was
removed on October 24, 2025 in commit
26450db480761c0fedf319fc350178798ce7e547.
Assisted-by: Antigravity
PPC: Fold 64-bit zero-extending word load feeding extsw subregister
A gprc LWZ/LWZX feeding EXTSW_32_64 is rewritten into a sign-extending
LWA/LWAX load. Extend the same fold to the 64-bit zero-extending word
loads LWZ8/LWZX8 when the EXTSW_32_64 reads their sub_32 subregister,
producing a single LWA/LWAX instead of a redundant lwz+extsw pair.
Co-authored-by: Claude (Claude Opus 4.8, claude-opus-4-8) <noreply at anthropic.com>
PPC: Add MIR examples for missed extsw+word-load fold on subregister input
A gprc LWZ/LWZX feeding EXTSW_32_64 folds into a sign-extending LWA/LWAX
load. The equivalent 64-bit zero-extending word loads (LWZ8/LWZX8) whose
sub_32 feeds EXTSW_32_64 are not folded, leaving a redundant lwz+extsw
(or lwzx+extsw) pair. Add MIR examples documenting the missed fold.
Co-authored-by: Claude (Claude Opus 4.8, claude-opus-4-8) <noreply at anthropic.com>
PPC: Fix extsw elimination when the input reads a subregister
The EXTSW_32_64 sign-extend elimination previously assumed its input
was a full register value. It would then try using that value as the
source of the new (unnecessary) INSERT_SUBREG.
The new test would then hit this verifier error:
```
bb.0:
liveins: $x3
%0:g8rc = COPY killed $x3
%1:g8rc = RLDICL killed %0:g8rc, 0, 33
%3:g8rc = IMPLICIT_DEF
%2:g8rc = INSERT_SUBREG %3:g8rc(tied-def 0), %1:g8rc, %subreg.sub_32
$x3 = COPY killed %2:g8rc
BLR8 implicit $lr8, implicit $rm, implicit killed $x3
*** Bad machine code: INSERT_SUBREG expected inserted value to have equal or lesser size than the subreg it was inserted into ***
[8 lines not shown]
PPC: Fix EXTSW elimination promoting a subregister operand
promoteInstr32To64ForElimEXTSW copies operands from the 32-bit
instruction verbatim into its promoted 64-bit form. When an operand
reads the sub_32 subregister of a 64-bit register, the promoted
instruction (which takes a full register) ended up with an illegal
subregister use and failed the machine verifier.
Drop the sub_32 subregister and use the original full register, which
provides the low 32 bits the promoted instruction operates on. This
avoids verifier error regressions in a future change.
Co-Authored-By: Claude <noreply at anthropic.com> (Claude Opus 4.8, claude-opus-4-8)
CodeGen: Prefer getting the Triple from the Module
Continue replacing TargetMachine::getTargetTriple() with the module's
triple at sites where a Module is one hop away through an available
Function, GlobalValue or MachineModuleInfo.
Where the surrounding class already holds a Subtarget, use its triple
rather than routing through the Module.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[ADT] Use StorageT for temporary storage in DenseMapBase::grow (NFC) (#228802)
Without this patch, the non-relocatable path of DenseMapBase::grow
rehashes into a temporary DenseMapBase instance:
DenseMapBase Tmp(NumBuckets, ExactBucketCount{});
This patch switches to:
StorageT Tmp;
while converting initEmpty, initWithExactBucketCount, and moveFrom to
static helpers to operate directly on StorageT. This brings the
following benefits:
- grow() no longer calls the destructor for Tmp because StorageT does
not have one. Without this patch, we call Tmp.~DenseMapBase(), which
in turn calls destroyAll() even though the bucket array has no
elements to destruct. Even with inlining enabled, the host compiler
[11 lines not shown]
CodeGen: Prefer getting the Triple from the Module
Continue replacing TargetMachine::getTargetTriple() with the module's
triple at sites where a Module is one hop away through an available
Function, GlobalValue or MachineModuleInfo.
Where the surrounding class already holds a Subtarget, use its triple
rather than routing through the Module.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
CodeGen: Run LiveIntervals before PHIElimination and drop LiveVariables from it
Move LiveIntervals to run before PHIElimination in the optimized register
allocation pipeline, and make PHIElimination maintain LiveIntervals only.
This removes the last explicit use of LiveVariables. The actual analysis is no
longer used. There are implicit dependencies on the side effects of running the
analysis due to adjustments of dead flags, so further work is still needed to
complete the removal.
This perturbs register allocation in a number of tests. The same codegen result
can be achieved by not preserving the analysis and recomputing fresh. Greedy is
just sensitive to the exact slot index and value numbering with identical MIR.
Measured across every affected test the emitted instruction count goes from
145124 to 145196, +0.050%, with changes in both directions. The largest
regression is AArch64/phi.ll, where the GlobalISel output gains about 30
instructions and no longer matches the SelectionDAG output; the largest
improvements are ARM/fpclamptosat.ll and PowerPC/common-chain.ll.
[2 lines not shown]
[CodeGen] Declare command line options in TableGen
Move the cl::opts of TargetPassConfig.cpp and CodeGenPrepare.cpp into
CodeGenOptions.td, private to lib/CodeGen; other files will follow.
-enable-machine-outliner (cl::ValueOptional) and -regalloc
(RegisterPassParser) stay cl::opt.
CodeGenPrepare and its addressing-mode helpers hold
`const CodeGenOptions &Opts`; TargetPassConfig functions read
CodeGenOptions::Global. getCGPassBuilderOption() converts BoolOrDefault
members to the cl::boolOrDefault and std::optional<bool> fields of the
public CGPassBuilderOption. -basic-block-section-match-infer, which was
not cl::Hidden, is now listed by -help-hidden only.
Aided by Opus 5.5
[CodeGenPrepare] Prefix option names with cgp-
Rename the CodeGenPrepare options so that they start with cgp-, and turn
the -disable-* options into positive options that default to true:
```
-addr-sink-* -cgp-addr-sink-*
-cgpp-huge-func -cgp-huge-func
-disable-cgp-X -cgp-X=0
-disable-complex-addr-modes -cgp-complex-addr-modes=0
-disable-preheader-prot -cgp-preheader-prot=0
-enable-andcmp-sinking -cgp-andcmp-sinking
-force-split-store -cgp-force-split-store
-stress-cgp-X -cgp-stress-X
```
The section prefix options (-profile-guided-section-prefix, etc.) are
kept, as they are not specific to CodeGenPrepare.
Aided by Opus 5.5
[Option] Add BoolOrDefault for OptionalBoolField and DefaultOnOffField (#228835)
The members of OptionalBoolField and DefaultOnOffField are
std::optional<bool>, on which `if (X)` tests whether the option is
given, not its value. Make them a new 1-byte `enum class BoolOrDefault :
uint8_t { Default, True, False }`, which does not convert to bool, and
read them with `valueOr(X, Default)`.
Aided by Opus 5.5