AArch64/GlobalISel: Mark the LR def of the Mach-O TLS call dead (#227234)
The ordinary call path already does this, and the SelectionDAG emitter
marks unused implicit physreg defs dead for free. Without it the dead
flag needs to be reinferred later.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[AMDGPU] Price scalar integer to fp casts by source width and sign
Scalar sources between a byte and 31 bits fell to the default cost of one
while the matching vector lanes were already priced, which skewed the
difference SLP weighs a bundle against. Such a source is extended before
the conversion, and what the extension takes depends on the width, on the
sign and on whether the subtarget has SDWA and 16 bit instructions.
Sources narrower than a byte are left alone, because their vector form is
not priced either.
[AMDGPU] Price narrow integer to bfloat vector casts
A vector lane of 9 to 15 or 17 to 31 bits converted to bfloat got the
generic cost, which leaves out the rounding. Such a lane is converted to
f32 first like any other narrow lane, so price it as the f32 conversion of
the same source plus the rounding. Lanes of 8 and 16 bits keep their cost.
[AMDGPU] Model the cost of the expanded integer to/from floating point casts
No instruction converts to or from a 64 bit integer, and narrow vector
lanes are converted one at a time. Price these expansions by the FP64
rate, sdwa and 16 bit instruction support, and price i33 to i63 like
i64 and bf16 like f32 plus rounding.
Assisted-by: Claude Code Opus 5
[NFC][AMDGPU] Add cost tests for narrow integer to fp casts
Covers integer sources from a byte to 31 bits converted to half, float,
bfloat and double, as vector lanes and as scalars, over the subtarget
combinations that change the expansion. The existing cast tests get the
same subtarget coverage and the cases they were missing. The costs
recorded here are the ones the model reports today.
AMDGPU: Remove update-only LiveVariables maintenance from SILowerControlFlow
This was only maintained, never relied on. Part of staged LiveVariables
removal.
Co-authored-by: Claude (Claude-Opus-4.8)
NVPTX: Drop LiveVariables from the register allocation pipeline
The optimized RegAlloc pipeline ran LiveVariables only to satisfy PHIElimination
and TwoAddressInstruction, both of which no longer need it. Remove the
LiveVariables run (and, in the new pass manager, the UnreachableMachineBlockElim
that was there only as a LiveVariables prerequisite).
Co-authored-by: Claude (Claude-Opus-4.8)
CodeGen: Drop the LiveVariables parameter from convertToThreeAddress
This was used for analysis updates, but now the analysis is being
removed.
Co-authored-by: Claude (Claude-Opus-4.8)
CodeGen: Remove LiveVariables use from TwoAddressInstructionPass
Now that LiveIntervals is computed unconditionally before TwoAddressInstructions
in the pipeline, the pass no longer needs LiveVariables.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
AMDGPU: Use LiveIntervals instead of LiveVariables in SIOptimizeVGPRLiveRange
Drop the LiveVariables dependency and the hand-written VarInfo
maintenance; the pass already knew how to recompute the affected
intervals with LiveIntervals, so make that the only path.
The legacy pass manager cannot schedule a pass requiring both
LiveIntervals and LiveVariables here, since LiveVariables (and its
UnreachableMachineBlockElim dependency) invalidates the LiveIntervals
just computed for it. In the AMDGPU pipeline, anchor the pass after
MachineLoopInfo instead of PHIElimination so LiveIntervals is computed
before PHIElimination, which is required anyway: the pass needs SSA and
introduces new PHIs. Preserving SlotIndexes keeps the transitive
last-user chain intact.
Since LiveVariables is no longer maintained past this point, PHIElimination
now splits critical edges using LiveIntervals, which accounts for most of
the test churn: LiveIntervalCalc drops kill flags and adds dead flags.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
Mips: Mark the $gp setup copy dead when erasing the call's $gp use (#227216)
MipsOptimizePICCall drops the implicit $gp operand from a call when the
lazy binding stub for the callee has already run. The copy that set $gp
up for that call then has no reader left, but nothing flagged it, so the
MIR carried a live def until a later liveness recomputation cleaned it
up.
The pass already walks each block in order, so track the reaching
definition of $gp as it goes and mark it dead when the use is erased.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[libc++] Encode the standard version in the ABI tag (#218527)
This prevents ODR mismatch issues from biting us across standard
versions. The order of elements in the ABI tag is now: libc++ version,
hardening mode, assertion semantic, exceptions, standard version.
Fixes #218524
[AArch64] ISel support for nxv1i1 vector_reverse (#226939)
This ensures llvm.vector.reverse can be selected for nxv1i1 types. These intrinsics can be generated by LoopVectorizer for `VF = vscale x 1` and loops iterating in reverse order.
[AMDGPU] Use S_CMP to lower a copy of a lane mask to SCC (#221445)
SIFixSGPRCopies lowers a copy of a lane mask to SCC as an AND with EXEC,
whose destination register is created by the pass and is always dead.
The AND with EXEC is only needed because SCC has to be "any active lane
is set". When the lane mask already has 0 in the bits of all inactive
lanes, that is just "the mask is non-zero", which S_CMP computes without
needing a destination register. In some cases the S_CMP can be optimized
away later by SIInstrInfo::optimizeCompareInstr.
Co-authored-by: Claude Opus 5 (1M context) <noreply at anthropic.com>
[LifetimeSafety] Fix off-by-one crash in `lifetime_capture_by` argument indexing (#227231)
Fixes an off-by-one crash in the lifetime_capture_by attribute handling.
The CapturingArgIdx is already an index into the full Args array, but
the code was incorrectly creating a CallArgs array (dropping the first
argument for instance methods) and then indexing into it with
CapturingArgIdx, causing incorrect argument access and potential
out-of-bounds crashes.
Currently crashes at head: https://godbolt.org/z/7rYczK1z6
[GVN] Move `GVNPass` and `GVNLeaderMap` out of `GVN.h` and into 'GVN.cpp` (NFC)
Rename `GVNPass` to `GVNPassImpl`. Leave in the header only
the class for interfacing with the pass manager.
[clang] Also disable scalable vectorisation for vectorize(disable)
This ensures that #pragma clang loop vectorize(disable) also disables
scalable vectorisation. Only setting the width to 1 could potentially
enable LoopVectorizer to vectorize with VF = vscale x 1. Especially with
the -scalable-vectorization=preferred option.
X86: Fix missing ... separator between functions in mir test (#227228)
The update_mir_test_checks output is incomplete so the later functions
here were really manually checked.
[BFI] Use MapVector in combineWeightsByHashing. (#227027)
combineWeightsByHashing iterated over a DenseMap, which depends on the
hash. Use MapVector to get consistent results, independent of index
type/values.
This is mainly to ensure consistency between users that use different
index numbers, i.e. used for IR BFI and VPlan's use.
PR: https://github.com/llvm/llvm-project/pull/227027
[C++20] [Modules] Load friends for classes in ADL (#219094)
Close https://github.com/llvm/llvm-project/issues/218228
The root cause of the problem is the corresponding friend is not loaded
at the point of ADL.
This patch tries to fix this simply by loading the friends at the point
of ADL. Note that this may be best efficient if there are a lot of
friends. We just think it is rare. If it is really possible, we can
change the structure of friends from a list to a name lookup table.
(cherry picked from commit 01aedf3b34325ad2f74a77325a2da9a7e36ae3d5)