LLVM/project e0c49c0 — llvm/lib/Target/AArch64/GISel AArch64InstructionSelector.cpp, llvm/test/CodeGen/AArch64/GlobalISel select-tls-macho.ll

AArch64/GlobalISel: Mark the LR def of the Mach-O TLS call dead (#227234)

The ordinary call path already does this, and the SelectionDAG emitter
marks unused implicit physreg defs dead for free. Without it the dead
flag needs to be reinferred later.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+30-0llvm/test/CodeGen/AArch64/GlobalISel/select-tls-macho.ll
+1-0llvm/lib/Target/AArch64/GISel/AArch64InstructionSelector.cpp
+31-02 files

LLVM/project aa6d5e3 — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp

format
DeltaFile
+1-1llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+1-11 files

LLVM/project 10a25ec — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Analysis/CostModel/AMDGPU narrow-int-to-bfloat.ll narrow-int-to-fp.ll

[AMDGPU] Price scalar integer to fp casts by source width and sign

Scalar sources between a byte and 31 bits fell to the default cost of one
while the matching vector lanes were already priced, which skewed the
difference SLP weighs a bundle against. Such a source is extended before
the conversion, and what the extension takes depends on the width, on the
sign and on whether the subtarget has SDWA and 16 bit instructions.
Sources narrower than a byte are left alone, because their vector form is
not priced either.
DeltaFile
+259-259llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+45-45llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+29-9llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+333-3133 files

LLVM/project 605c023 — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Analysis/CostModel/AMDGPU narrow-int-to-bfloat.ll

[AMDGPU] Price narrow integer to bfloat vector casts

A vector lane of 9 to 15 or 17 to 31 bits converted to bfloat got the
generic cost, which leaves out the rounding. Such a lane is converted to
f32 first like any other narrow lane, so price it as the f32 conversion of
the same source plus the rounding. Lanes of 8 and 16 bits keep their cost.
DeltaFile
+45-45llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+11-7llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+56-522 files

LLVM/project 212b04c — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Analysis/CostModel/AMDGPU fptoui.ll fptosi.ll

fix 16-bit conv
DeltaFile
+136-136llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+69-31llvm/test/Analysis/CostModel/AMDGPU/expanded-int-fp-casts.ll
+30-14llvm/test/Analysis/CostModel/AMDGPU/fptoui.ll
+30-14llvm/test/Analysis/CostModel/AMDGPU/fptosi.ll
+41-3llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+306-1985 files

LLVM/project 67e50d4 — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Analysis/CostModel/AMDGPU fptoui.ll narrow-int-to-bfloat.ll

[AMDGPU] Model the cost of the expanded integer to/from floating point casts

No instruction converts to or from a 64 bit integer, and narrow vector
lanes are converted one at a time. Price these expansions by the FP64
rate, sdwa and 16 bit instruction support, and price i33 to i63 like
i64 and bf16 like f32 plus rounding.

Assisted-by: Claude Code Opus 5
DeltaFile
+312-215llvm/test/Analysis/CostModel/AMDGPU/cast.ll
+210-210llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+146-70llvm/test/Analysis/CostModel/AMDGPU/expanded-int-fp-casts.ll
+70-70llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+107-0llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+60-37llvm/test/Analysis/CostModel/AMDGPU/fptoui.ll
+905-6021 files not shown
+965-6397 files

LLVM/project 994f271 — llvm/test/Analysis/CostModel/AMDGPU expanded-int-fp-casts.ll narrow-int-to-bfloat.ll

apply review
DeltaFile
+545-483llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+189-189llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+15-15llvm/test/Analysis/CostModel/AMDGPU/expanded-int-fp-casts.ll
+749-6873 files

LLVM/project a9af90e — llvm/test/Analysis/CostModel/AMDGPU fptoui.ll fptosi.ll

[NFC][AMDGPU] Add cost tests for narrow integer to fp casts

Covers integer sources from a byte to 31 bits converted to half, float,
bfloat and double, as vector lanes and as scalars, over the subtarget
combinations that change the expansion. The existing cast tests get the
same subtarget coverage and the cases they were missing. The costs
recorded here are the ones the model reports today.
DeltaFile
+642-0llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+241-0llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+32-84llvm/test/Analysis/CostModel/AMDGPU/expanded-int-fp-casts.ll
+17-10llvm/test/Analysis/CostModel/AMDGPU/cast.ll
+11-6llvm/test/Analysis/CostModel/AMDGPU/fptoui.ll
+11-6llvm/test/Analysis/CostModel/AMDGPU/fptosi.ll
+954-1066 files

LLVM/project ffc1da1 — llvm/lib/Target/AMDGPU SILowerControlFlow.cpp, llvm/test/CodeGen/AMDGPU lower-control-flow-live-variables-update.xfail.mir lower-control-flow-live-variables-update.mir

AMDGPU: Remove update-only LiveVariables maintenance from SILowerControlFlow

This was only maintained, never relied on. Part of staged LiveVariables
removal.

Co-authored-by: Claude (Claude-Opus-4.8)
DeltaFile
+87-0llvm/test/CodeGen/AMDGPU/lower-control-flow-live-variables-update.mir
+5-71llvm/lib/Target/AMDGPU/SILowerControlFlow.cpp
+0-42llvm/test/CodeGen/AMDGPU/lower-control-flow-live-variables-update.xfail.mir
+92-1133 files

LLVM/project b4813ce — llvm/lib/Target/NVPTX NVPTXTargetMachine.cpp NVPTXCodeGenPassBuilder.cpp, llvm/test/CodeGen/NVPTX llc-pipeline-npm.ll

NVPTX: Drop LiveVariables from the register allocation pipeline

The optimized RegAlloc pipeline ran LiveVariables only to satisfy PHIElimination
and TwoAddressInstruction, both of which no longer need it. Remove the
LiveVariables run (and, in the new pass manager, the UnreachableMachineBlockElim
that was there only as a LiveVariables prerequisite).

Co-authored-by: Claude (Claude-Opus-4.8)
DeltaFile
+0-7llvm/lib/Target/NVPTX/NVPTXCodeGenPassBuilder.cpp
+0-2llvm/test/CodeGen/NVPTX/llc-pipeline-npm.ll
+0-1llvm/lib/Target/NVPTX/NVPTXTargetMachine.cpp
+0-103 files

LLVM/project 5388e53 — llvm/lib/Target/AMDGPU SIFoldOperands.cpp SIInstrInfo.cpp, llvm/lib/Target/RISCV RISCVInstrInfo.cpp

CodeGen: Drop the LiveVariables parameter from convertToThreeAddress

This was used for analysis updates, but now the analysis is being
removed.

Co-authored-by: Claude (Claude-Opus-4.8)
DeltaFile
+12-66llvm/lib/Target/X86/X86InstrInfo.cpp
+0-17llvm/lib/Target/AMDGPU/SIInstrInfo.cpp
+0-11llvm/lib/Target/RISCV/RISCVInstrInfo.cpp
+1-10llvm/lib/Target/SystemZ/SystemZInstrInfo.cpp
+2-4llvm/lib/Target/X86/X86InstrInfo.h
+2-2llvm/lib/Target/AMDGPU/SIFoldOperands.cpp
+17-1106 files not shown
+22-11712 files

LLVM/project 9a8a97c — llvm/lib/CodeGen TwoAddressInstructionPass.cpp, llvm/test/CodeGen/Hexagon two-addr-tied-subregs.mir

CodeGen: Remove LiveVariables use from TwoAddressInstructionPass

Now that LiveIntervals is computed unconditionally before TwoAddressInstructions
in the pipeline, the pass no longer needs LiveVariables.

Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
DeltaFile
+38-131llvm/lib/CodeGen/TwoAddressInstructionPass.cpp
+18-18llvm/test/CodeGen/X86/statepoint-vreg-unlimited-tied-opnds.ll
+10-8llvm/test/CodeGen/X86/two-address-subreg-to-reg-kill.mir
+8-8llvm/test/CodeGen/SystemZ/twoaddr-kill.mir
+7-7llvm/test/CodeGen/Hexagon/two-addr-tied-subregs.mir
+3-3llvm/test/CodeGen/X86/twoaddr-dbg-value.mir
+84-1751 files not shown
+86-1777 files

LLVM/project 27ce195 — llvm/lib/CodeGen TwoAddressInstructionPass.cpp

Unused variable
DeltaFile
+1-2llvm/lib/CodeGen/TwoAddressInstructionPass.cpp
+1-21 files

LLVM/project 3a6314d — llvm/test/CodeGen/AMDGPU amdgcn.bitcast.768bit.ll amdgcn.bitcast.832bit.ll

AMDGPU: Use LiveIntervals instead of LiveVariables in SIOptimizeVGPRLiveRange

Drop the LiveVariables dependency and the hand-written VarInfo
maintenance; the pass already knew how to recompute the affected
intervals with LiveIntervals, so make that the only path.

The legacy pass manager cannot schedule a pass requiring both
LiveIntervals and LiveVariables here, since LiveVariables (and its
UnreachableMachineBlockElim dependency) invalidates the LiveIntervals
just computed for it. In the AMDGPU pipeline, anchor the pass after
MachineLoopInfo instead of PHIElimination so LiveIntervals is computed
before PHIElimination, which is required anyway: the pass needs SSA and
introduces new PHIs. Preserving SlotIndexes keeps the transitive
last-user chain intact.

Since LiveVariables is no longer maintained past this point, PHIElimination
now splits critical edges using LiveIntervals, which accounts for most of
the test churn: LiveIntervalCalc drops kill flags and adds dead flags.

Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+37,859-37,709llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+3,855-3,814llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+3,291-3,302llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.960bit.ll
+3,094-3,085llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.896bit.ll
+2,730-2,686llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.832bit.ll
+1,773-1,731llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.768bit.ll
+52,602-52,32755 files not shown
+57,170-57,07961 files

LLVM/project d1851e9 — clang/test/Sema/LifetimeSafety safety.cpp

negative-test-capture-by
DeltaFile
+22-0clang/test/Sema/LifetimeSafety/safety.cpp
+22-01 files

LLVM/project 5c170b0 — llvm/lib/Target/Mips MipsOptimizePICCall.cpp, llvm/test/CodeGen/Mips optimize-pic-call-dead-gp.ll

Mips: Mark the $gp setup copy dead when erasing the call's $gp use (#227216)

MipsOptimizePICCall drops the implicit $gp operand from a call when the
lazy binding stub for the callee has already run. The copy that set $gp
up for that call then has no reader left, but nothing flagged it, so the
MIR carried a live def until a later liveness recomputation cleaned it
up.

The pass already walks each block in order, so track the reaching
definition of $gp as it goes and mark it dead when the use is erased.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+78-0llvm/test/CodeGen/Mips/optimize-pic-call-dead-gp.ll
+32-22llvm/lib/Target/Mips/MipsOptimizePICCall.cpp
+110-222 files

LLVM/project d5f6c30 — libcxx/docs Hardening.rst, libcxx/include/__configuration attributes.h

[libc++] Encode the standard version in the ABI tag (#218527)

This prevents ODR mismatch issues from biting us across standard
versions. The order of elements in the ABI tag is now: libc++ version,
hardening mode, assertion semantic, exceptions, standard version.

Fixes #218524
DeltaFile
+9-4libcxx/include/__configuration/attributes.h
+2-1libcxx/docs/Hardening.rst
+11-52 files

LLVM/project a64ab34 — llvm/lib/Transforms/Vectorize VPlanTransforms.cpp LoopVectorizationLegality.cpp, llvm/test/Transforms/LoopVectorize early_exit_store_legality.ll early_exit_legality.ll

Address comments
DeltaFile
+4-4llvm/test/Transforms/LoopVectorize/early_exit_legality.ll
+4-4llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp
+3-2llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
+2-2llvm/test/Transforms/LoopVectorize/early_exit_store_legality.ll
+2-2llvm/test/Transforms/LoopVectorize/X86/vectorization-remarks-missed.ll
+15-145 files

LLVM/project a4b0944 — llvm/lib/Target/AArch64 AArch64ISelLowering.h AArch64ISelLowering.cpp, llvm/test/CodeGen/AArch64 named-vector-shuffle-reverse-sve.ll

[AArch64] ISel support for nxv1i1 vector_reverse (#226939)

This ensures llvm.vector.reverse can be selected for nxv1i1 types. These intrinsics can be generated by LoopVectorizer for `VF = vscale x 1` and loops iterating in reverse order.
DeltaFile
+20-0llvm/test/CodeGen/AArch64/named-vector-shuffle-reverse-sve.ll
+15-0llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
+1-0llvm/lib/Target/AArch64/AArch64ISelLowering.h
+36-03 files

LLVM/project 5a24573 — llvm/test/CodeGen/Mips/GlobalISel/irtranslator extend_args.ll call.ll, llvm/test/CodeGen/Mips/GlobalISel/legalizer sitofp_and_uitofp.mir fptosi_and_fptoui.mir

Mips/GlobalISel: Mark call's implicit RA def dead (#227233)

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+10-10llvm/test/CodeGen/Mips/GlobalISel/irtranslator/float_args.ll
+8-8llvm/test/CodeGen/Mips/GlobalISel/legalizer/sitofp_and_uitofp.mir
+8-8llvm/test/CodeGen/Mips/GlobalISel/legalizer/fptosi_and_fptoui.mir
+8-8llvm/test/CodeGen/Mips/GlobalISel/legalizer/ceil_and_floor.mir
+8-8llvm/test/CodeGen/Mips/GlobalISel/irtranslator/call.ll
+6-6llvm/test/CodeGen/Mips/GlobalISel/irtranslator/extend_args.ll
+48-489 files not shown
+62-6115 files

LLVM/project 7cf8ecb — llvm/lib/Target/Mips MipsFastISel.cpp, llvm/test/CodeGen/Mips/Fast-ISel mul-dead-hilo.ll

Mips: Mark MUL's existing HI0/LO0 defs dead in FastISel (#227230)
DeltaFile
+19-0llvm/test/CodeGen/Mips/Fast-ISel/mul-dead-hilo.ll
+4-4llvm/lib/Target/Mips/MipsFastISel.cpp
+23-42 files

LLVM/project 74c86a7 — llvm/test/CodeGen/AMDGPU fix-sgpr-copies-scc-cmp.mir fcmp.f16.ll

[AMDGPU] Use S_CMP to lower a copy of a lane mask to SCC (#221445)

SIFixSGPRCopies lowers a copy of a lane mask to SCC as an AND with EXEC,
whose destination register is created by the pass and is always dead.
The AND with EXEC is only needed because SCC has to be "any active lane
is set". When the lane mask already has 0 in the bits of all inactive
lanes, that is just "the mask is non-zero", which S_CMP computes without
needing a destination register. In some cases the S_CMP can be optimized
away later by SIInstrInfo::optimizeCompareInstr.

Co-authored-by: Claude Opus 5 (1M context) <noreply at anthropic.com>
DeltaFile
+506-495llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+364-364llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+171-168llvm/test/CodeGen/AMDGPU/llvm.log10.ll
+171-168llvm/test/CodeGen/AMDGPU/llvm.log.ll
+129-129llvm/test/CodeGen/AMDGPU/fcmp.f16.ll
+224-0llvm/test/CodeGen/AMDGPU/fix-sgpr-copies-scc-cmp.mir
+1,565-1,32441 files not shown
+2,256-1,89047 files

LLVM/project 6ccc491 — clang/lib/Analysis/LifetimeSafety FactsGenerator.cpp, clang/test/Sema/LifetimeSafety safety.cpp

[LifetimeSafety] Fix off-by-one crash in `lifetime_capture_by` argument indexing (#227231)

Fixes an off-by-one crash in the lifetime_capture_by attribute handling.
The CapturingArgIdx is already an index into the full Args array, but
the code was incorrectly creating a CallArgs array (dropping the first
argument for instance methods) and then indexing into it with
CapturingArgIdx, causing incorrect argument access and potential
out-of-bounds crashes.

Currently crashes at head: https://godbolt.org/z/7rYczK1z6
DeltaFile
+23-0clang/test/Sema/LifetimeSafety/safety.cpp
+6-5clang/lib/Analysis/LifetimeSafety/FactsGenerator.cpp
+29-52 files

LLVM/project b9a4527 — llvm/include/llvm/Transforms/Scalar GVN.h, llvm/lib/Transforms/Scalar GVN.cpp

[GVN] Move `GVNPass` and `GVNLeaderMap` out of `GVN.h` and into 'GVN.cpp` (NFC)

Rename `GVNPass` to `GVNPassImpl`. Leave in the header only
the class for interfacing with the pass manager.
DeltaFile
+402-85llvm/lib/Transforms/Scalar/GVN.cpp
+0-335llvm/include/llvm/Transforms/Scalar/GVN.h
+402-4202 files

LLVM/project c4cc8fa — clang/lib/CodeGen CGLoopInfo.cpp, clang/test/CodeGenCXX pragma-loop-safety.cpp pragma-loop.cpp

[clang] Also disable scalable vectorisation for vectorize(disable)

This ensures that #pragma clang loop vectorize(disable) also disables
scalable vectorisation. Only setting the width to 1 could potentially
enable LoopVectorizer to vectorize with VF = vscale x 1. Especially with
the -scalable-vectorization=preferred option.
DeltaFile
+7-7clang/test/CodeGenCXX/pragma-loop-predicate.cpp
+3-3clang/test/CodeGenCXX/pragma-loop.cpp
+3-2clang/lib/CodeGen/CGLoopInfo.cpp
+2-1clang/test/CodeGenCXX/pragma-loop-safety.cpp
+15-134 files

LLVM/project fa3d170 — llvm/lib/Transforms/IPO FunctionAttrs.cpp, llvm/test/Transforms/FunctionAttrs noalias.ll

[FunctionAttrs] Infer noalias through a null check. (#226956)

Only comparing against null does not captures provenance and should not
impact whether a pointer is noalias or not.

This enables noalias inference in a number of cases for malloc-like
functions:
https://github.com/dtcxzyw/llvm-opt-benchmark-nightly/pull/1459.

PR: https://github.com/llvm/llvm-project/pull/226956
DeltaFile
+192-0llvm/test/Transforms/FunctionAttrs/noalias.ll
+5-1llvm/lib/Transforms/IPO/FunctionAttrs.cpp
+197-12 files

LLVM/project 724557a — llvm/test/CodeGen/X86 two-address-subreg-to-reg-kill.mir

X86: Fix missing ... separator between functions in mir test (#227228)

The update_mir_test_checks output is incomplete so the later functions
here were really manually checked.
DeltaFile
+36-14llvm/test/CodeGen/X86/two-address-subreg-to-reg-kill.mir
+36-141 files

LLVM/project f69661d — llvm/lib/Analysis BlockFrequencyInfoImpl.cpp, llvm/test/Analysis/BlockFrequencyInfo many-successors-order.ll

[BFI] Use MapVector in combineWeightsByHashing. (#227027)

combineWeightsByHashing iterated over a DenseMap, which depends on the
hash. Use MapVector to get consistent results, independent of index
type/values.

This is mainly to ensure consistency between users that use different
index numbers, i.e. used for IR BFI and VPlan's use.

PR: https://github.com/llvm/llvm-project/pull/227027
DeltaFile
+156-0llvm/test/Analysis/BlockFrequencyInfo/many-successors-order.ll
+3-6llvm/lib/Analysis/BlockFrequencyInfoImpl.cpp
+159-62 files

LLVM/project dfb1297 — clang/include/clang/AST DeclCXX.h, clang/lib/AST DeclFriend.cpp

[C++20] [Modules] Load friends for classes in ADL (#219094)

Close https://github.com/llvm/llvm-project/issues/218228

The root cause of the problem is the corresponding friend is not loaded
at the point of ADL.

This patch tries to fix this simply by loading the friends at the point
of ADL. Note that this may be best efficient if there are a lot of
friends. We just think it is rare. If it is really possible, we can
change the structure of friends from a list to a name lookup table.

(cherry picked from commit 01aedf3b34325ad2f74a77325a2da9a7e36ae3d5)
DeltaFile
+95-0clang/test/Modules/pr218228.cppm
+11-0clang/lib/Sema/SemaLookup.cpp
+9-0clang/lib/AST/DeclFriend.cpp
+4-0clang/include/clang/AST/DeclCXX.h
+119-04 files

LLVM/project 6abb6e7 — clang/include/clang/AST Redeclarable.h DeclCXX.h, clang/lib/Interpreter IncrementalParser.cpp

[clang-repl] Keep earlier declarations alive when an input fails (#218149)

Fixes #201844

(cherry picked from commit 37e2844bc4681c69f61cbb55376a29ae0445c18c)
DeltaFile
+127-54clang/lib/Interpreter/IncrementalParser.cpp
+102-0clang/test/Interpreter/failed-input-keeps-redecls.cpp
+1-0clang/include/clang/AST/Redeclarable.h
+1-0clang/include/clang/AST/DeclCXX.h
+231-544 files