LLVM/project 267acde — llvm/lib/Transforms/Vectorize VectorCombine.cpp, llvm/test/Transforms/VectorCombine deinterleave-interleave-tree.ll

[VectorCombine] Fold interleave and widen chained operations (#224005)

Generalize the existing single deinterleave-interleave pair fold by starting from the interleave
and walking backwards through its operands.
Starting from the interleave instead exposes the whole expression tree that produces its operands,
allowing the combine to discover multiple deinterleaves participating in the same reconstructed vector.
The walk follows supported element-wise operations and splats backwards until it reaches the originating deinterleaves. Once all interleave operands can be traced back consistently, the operations can be rebuilt
on the original wider vectors and the intermediate deinterleave-interleave operations removed.
This makes the fold handle patterns with multiple deinterleaved inputs.
DeltaFile
+506-0llvm/test/Transforms/VectorCombine/deinterleave-interleave-tree.ll
+219-210llvm/lib/Transforms/Vectorize/VectorCombine.cpp
+725-2102 files

LLVM/project 89681f1 — llvm/lib/IR Instructions.cpp, llvm/test/Transforms/InstCombine byte-cast-pairs.ll

[IR] Fold `bitcast` pairs through bytes with `ptrtoaddr`
DeltaFile
+18-4llvm/lib/IR/Instructions.cpp
+4-8llvm/test/Transforms/InstCombine/byte-cast-pairs.ll
+1-1llvm/test/Transforms/InstSimplify/byte-cast-pairs.ll
+23-133 files

LLVM/project 0c0e3a4 — llvm/include/llvm/Analysis TargetFolder.h, llvm/include/llvm/IR ConstantFolder.h ConstantFold.h

[ConstantFolding] Fold `bitinsert` and `bitextract`
DeltaFile
+28-56llvm/test/Transforms/InstSimplify/ConstProp/bitinsert-bitextract.ll
+72-0llvm/lib/IR/ConstantFold.cpp
+13-0llvm/include/llvm/IR/ConstantFold.h
+9-2llvm/include/llvm/IR/ConstantFolder.h
+9-0llvm/include/llvm/Analysis/TargetFolder.h
+5-0llvm/lib/Analysis/ConstantFolding.cpp
+136-586 files

LLVM/project b1cb400 — llvm/test/CodeGen/AMDGPU shrink-disjoint-or-scc.mir shrink-insts-scalar-bit-ops.mir

Merge the cases to an existing test file
DeltaFile
+0-35llvm/test/CodeGen/AMDGPU/shrink-disjoint-or-scc.mir
+35-0llvm/test/CodeGen/AMDGPU/shrink-insts-scalar-bit-ops.mir
+35-352 files

LLVM/project 339d95d — llvm/lib/Target/AMDGPU SIShrinkInstructions.cpp, llvm/test/CodeGen/AMDGPU s_or_b32_transformation.ll

[AMDGPU] Preserve live SCC when shrinking disjoint S_OR_B32
DeltaFile
+6-0llvm/lib/Target/AMDGPU/SIShrinkInstructions.cpp
+2-1llvm/test/CodeGen/AMDGPU/s_or_b32_transformation.ll
+8-12 files

LLVM/project 84dd4a1 — llvm/lib/Target/AMDGPU SIShrinkInstructions.cpp, llvm/test/CodeGen/AMDGPU s_or_b32_transformation.ll shrink-disjoint-or-scc.mir

Address review feedback:
1. Use allImplicitDefsAreDead() helper to check whether SCC is dead
2. Simplify .ll test case
3. Add a new .mir test case
DeltaFile
+35-0llvm/test/CodeGen/AMDGPU/shrink-disjoint-or-scc.mir
+6-8llvm/test/CodeGen/AMDGPU/s_or_b32_transformation.ll
+1-2llvm/lib/Target/AMDGPU/SIShrinkInstructions.cpp
+42-103 files

LLVM/project 7f415a6 — llvm/test/CodeGen/AMDGPU s_or_b32_transformation.ll

Precommit test case for s_or incorrectly transformed to s_addk when SCC is live
DeltaFile
+28-0llvm/test/CodeGen/AMDGPU/s_or_b32_transformation.ll
+28-01 files

LLVM/project 2c17d4d — bolt/lib/Passes LongJmp.cpp, bolt/test/AArch64 relax-branches-with-thunk-chain.s relax-cluster-size-too-large.s

[BOLT][AArch64] Remove thunk estimation and warn on margin exhaustion (#228069)

When performing branch range checks compare modeled hop distances with
MaxClusterSize inclusively without charging estimated thunk bytes
against it. This allows a thunk at an adjacent cluster boundary to reach
its target.

Cap MaxClusterSize at 126 MiB to reserve minimum alignment headroom.
Warn when thunk bytes plus text alignment exceed the remaining branch
safety margin.
DeltaFile
+30-63bolt/lib/Passes/LongJmp.cpp
+53-0bolt/test/AArch64/relax-adjacent-cluster-boundary.s
+8-16bolt/test/AArch64/relax-cross-fragment-calls.s
+24-0bolt/test/AArch64/relax-cluster-branch-margin-warning.s
+20-0bolt/test/AArch64/relax-cluster-size-too-large.s
+6-12bolt/test/AArch64/relax-branches-with-thunk-chain.s
+141-911 files not shown
+141-957 files

LLVM/project 7a14bed — llvm/test/Transforms/InstCombine byte-cast-pairs.ll

[IR] Pre-commit tests for folding `bitcast` pairs through bytes with `ptrtoaddr`
DeltaFile
+41-0llvm/test/Transforms/InstCombine/byte-cast-pairs.ll
+41-01 files

LLVM/project 09b9b83 — llvm/lib/Remarks YAMLRemarkParser.cpp YAMLRemarkParser.h, llvm/unittests/Remarks YAMLRemarksParsingTest.cpp

[Remarks] Fix use-after-free of block scalar values in the YAML remark parser

YAMLRemarkParser::parseStr returned BlockScalarNode::getValue() as is.
The YAML parser copies a block scalar's value into the document's node
allocator, and next() advances the document iterator before returning
the remark, which frees that allocator. So every argument value written
as a block scalar pointed at freed memory by the time the caller saw
it.

Copy block scalar values into an allocator owned by the parser.

Assisted-by: Claude
DeltaFile
+37-0llvm/unittests/Remarks/YAMLRemarksParsingTest.cpp
+3-1llvm/lib/Remarks/YAMLRemarkParser.cpp
+4-0llvm/lib/Remarks/YAMLRemarkParser.h
+44-13 files

LLVM/project a5633de — lldb/test/API/lang/cpp/breakpoint_in_member_func_w_non_primitive_params main.cpp

[lldb] Remove unused cstdio include in test (#229032)
DeltaFile
+0-1lldb/test/API/lang/cpp/breakpoint_in_member_func_w_non_primitive_params/main.cpp
+0-11 files

LLVM/project 3112686 — llvm/lib/Target/SPIRV SPIRVNonSemanticDebugHandler.h SPIRVNonSemanticDebugHandler.cpp, llvm/test/CodeGen/SPIRV/debug-info debug-function-declaration-composite-scope.ll

[SPIRV] Drop the per-kind NSDI type vectors.

Replace partitionTypes and the seven type lists with the type list from DebugInfoFinder.
DeltaFile
+44-89llvm/lib/Target/SPIRV/SPIRVNonSemanticDebugHandler.cpp
+1-17llvm/lib/Target/SPIRV/SPIRVNonSemanticDebugHandler.h
+4-4llvm/test/CodeGen/SPIRV/debug-info/debug-function-declaration-composite-scope.ll
+49-1103 files

LLVM/project b902a6c — llvm/test/Transforms/InstSimplify/ConstProp bitinsert-bitextract.ll

[ConstantFolding] Pre-commit tests for `bitinsert`/`bitextract` folding
DeltaFile
+380-0llvm/test/Transforms/InstSimplify/ConstProp/bitinsert-bitextract.ll
+380-01 files

LLVM/project a128514 — llvm/lib/Target/SPIRV SPIRVNonSemanticDebugHandler.h SPIRVNonSemanticDebugHandler.cpp, llvm/test/CodeGen/SPIRV/debug-info debug-function-declaration-composite-scope.ll

[SPIRV] Drop the per-kind NSDI type vectors.

Replace partitionTypes and the seven type lists with the type list from DebugInfoFinder.
DeltaFile
+47-89llvm/lib/Target/SPIRV/SPIRVNonSemanticDebugHandler.cpp
+1-17llvm/lib/Target/SPIRV/SPIRVNonSemanticDebugHandler.h
+4-4llvm/test/CodeGen/SPIRV/debug-info/debug-function-declaration-composite-scope.ll
+52-1103 files

LLVM/project 9683a9b — llvm/lib/Target/X86 X86ISelLowering.cpp, llvm/test/CodeGen/X86 atomic-cmp-fold.ll

[X86] Fix reversed condition in ADC/SBB fold of COND_A compares (#228693)

#161388 added a COND_A case to combineAddOrSubToADCOrSBB that reuses the
flags of `CMP X, Y` for ADC/SBB without swapping the operands, computing
X +/- (X <u Y) instead of X +/- (X >u Y).

It was unreachable until #221290 started emitting X86ISD::CMP instead of
X86ISD::SUB for compares of atomic loads. This patch removes it.

 `(a - b) - (a > b)` case #161388 targeted goes through the X86ISD::SUB path and is unaffected.

Fixes #228457

Assisted-by: Claude Opus 5.5
DeltaFile
+53-0llvm/test/CodeGen/X86/atomic-cmp-fold.ll
+0-10llvm/lib/Target/X86/X86ISelLowering.cpp
+53-102 files

LLVM/project fc7f642 — llvm/lib/IR Instructions.cpp, llvm/test/Transforms/InstCombine byte-cast-pairs.ll

[IR] Fold pointer-byte-integer `bitcast` pairs to `ptrtoaddr`
DeltaFile
+11-2llvm/lib/IR/Instructions.cpp
+2-4llvm/test/Transforms/InstCombine/byte-cast-pairs.ll
+1-1llvm/test/Transforms/InstSimplify/byte-cast-pairs.ll
+14-73 files

LLVM/project c6db326 — llvm/lib/CodeGen TargetLoweringObjectFileImpl.cpp TailDuplicator.cpp, llvm/lib/Target/AArch64 AArch64MCInstLower.cpp AArch64FrameLowering.cpp

CodeGen: Prefer getting the Triple from the Module

Continue replacing TargetMachine::getTargetTriple() with the module's
triple at sites where a Module is one hop away through an available
Function, GlobalValue or MachineModuleInfo.

Where the surrounding class already holds a Subtarget, use its triple
rather than routing through the Module.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+7-3llvm/lib/CodeGen/TailDuplicator.cpp
+4-4llvm/lib/CodeGen/TargetLoweringObjectFileImpl.cpp
+3-2llvm/lib/Target/AArch64/AArch64PointerAuth.cpp
+2-2llvm/lib/Target/AArch64/AArch64FrameLowering.cpp
+2-1llvm/lib/Target/AMDGPU/AMDGPUTargetObjectFile.cpp
+2-1llvm/lib/Target/AArch64/AArch64MCInstLower.cpp
+20-1313 files not shown
+34-2619 files

LLVM/project 70fb870 — llvm/lib/Target/X86 X86DynAllocaExpander.cpp, llvm/test/CodeGen/X86 dyn-alloca-expander-dead-eflags.ll

X86: Preserve the dead flag clobber when expanding dynamic allocas

The DYN_ALLOCA pseudos clobber EFLAGS, so propagate the pseudo's dead
flag to the stack adjustment they expand to.

Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+51-0llvm/test/CodeGen/X86/dyn-alloca-expander-dead-eflags.ll
+15-7llvm/lib/Target/X86/X86DynAllocaExpander.cpp
+66-72 files

LLVM/project aa7315f — llvm/test/CodeGen/VE/Scalar store_stk.ll stackframe_align.ll, llvm/test/CodeGen/VE/Vector store_stk_stvm.ll load_stk_ldvm.ll

VE: Use splitAt in expandExtendStackPseudo

Replace the manual block-splitting in expandExtendStackPseudo with
MachineBasicBlock::splitAt. Reduces boilerplate, but there's some
block renumbering churn in the output.

Co-authored-by: Claude (Claude-Opus-4.8)
DeltaFile
+57-57llvm/test/CodeGen/VE/Scalar/atomic_swap.ll
+57-57llvm/test/CodeGen/VE/Scalar/atomic_cmp_swap.ll
+48-48llvm/test/CodeGen/VE/Vector/store_stk_stvm.ll
+48-48llvm/test/CodeGen/VE/Vector/load_stk_ldvm.ll
+42-42llvm/test/CodeGen/VE/Scalar/store_stk.ll
+42-42llvm/test/CodeGen/VE/Scalar/stackframe_align.ll
+294-29461 files not shown
+759-76267 files

LLVM/project 330eb34 — llvm/test/Transforms/InstCombine byte-cast-pairs.ll

[IR] Pre-commit tests for folding `bitcast` pairs to `ptrtoaddr`
DeltaFile
+14-0llvm/test/Transforms/InstCombine/byte-cast-pairs.ll
+14-01 files

LLVM/project c73950c — llvm/lib/Target/Mips MipsISelLowering.cpp MipsFastISel.cpp, llvm/test/CodeGen/Mips/Fast-ISel mul-dead-hilo.ll

Mips: Stop setting kill flags on virtual registers before FinalizeISel

There is no point in maintaining these before register allocation anymore.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+3-3llvm/lib/Target/Mips/MipsISelLowering.cpp
+3-3llvm/lib/Target/Mips/MipsFastISel.cpp
+2-2llvm/test/CodeGen/Mips/Fast-ISel/mul-dead-hilo.ll
+8-83 files

LLVM/project 5b07ba7 — llvm/bindings/ocaml/llvm llvm_ocaml.c llvm.mli, llvm/include/llvm-c Core.h

[llvm-c][ocaml] Deprecate/replace constexpr GEP APIs (#228998)

In the C API, this deprecates LLVMConstGEP, LLVMConstInBoundsGEP and
LLVMConstGEPWithNoWrapFlags. The replacement APIs are LLVMConstPtrAdd
and LLVMConstPtrAddFromIndices. The latter exposes the DataLayout-aware
`getGetElementPtr()` API, and only exists for the sake of convenience.

Similarly, in the OCaml API, remove const_gep and const_inbounds_gep in
favor of const_ptradd and const_ptradd_from_indices, which mirror the C
APIs translated to OCaml.

Unlike the older APIs, the new ones work directly on GEPNoWrapFlags
instead of providing (an incomplete set of) separate APIs for different
flavors.

This completes https://github.com/llvm/llvm-project/pull/227601 by
extending the deprecation of constexpr GEP methods to the C API.
DeltaFile
+39-16llvm/include/llvm-c/Core.h
+21-10llvm/bindings/ocaml/llvm/llvm.mli
+28-0llvm/unittests/IR/ConstantsTest.cpp
+23-4llvm/test/Bindings/OCaml/core.ml
+14-12llvm/bindings/ocaml/llvm/llvm_ocaml.c
+21-0llvm/lib/IR/Core.cpp
+146-424 files not shown
+180-5410 files

LLVM/project 042302b — clang/test/CodeGen builtins-arm64.c, llvm/include/llvm/Support AArch64MemoryHints.h

[AArch64] Add CMH hints to store_with_hint intrinsic (#227326)

This patch extends __arm_atomic_store_with_hint intrinsic with new CMH
hints defined in
[ACLE](https://arm-software.github.io/acle/main/acle.html#atomic-store-with-hints-intrinsics)
DeltaFile
+403-1llvm/test/CodeGen/AArch64/Atomics/aarch64-atomic-store-hint.ll
+15-6llvm/lib/Target/AArch64/AArch64InstrAtomics.td
+12-6llvm/test/CodeGen/AArch64/Atomics/aarch64-relaxed-store-hint.ll
+16-1clang/test/CodeGen/builtins-arm64.c
+4-10llvm/lib/Target/AArch64/AArch64ISelDAGToDAG.cpp
+10-1llvm/include/llvm/Support/AArch64MemoryHints.h
+460-252 files not shown
+465-268 files

LLVM/project bb67802 — llvm/lib/IR Instructions.cpp, llvm/test/Transforms/InstCombine byte-cast-pairs.ll

[IR] Don't fold `bitcast` pairs through bytes into invalid casts
DeltaFile
+107-0llvm/test/Transforms/InstCombine/byte-cast-pairs.ll
+26-0llvm/test/Transforms/InstSimplify/byte-cast-pairs.ll
+20-1llvm/lib/IR/Instructions.cpp
+153-13 files

LLVM/project e1ab1b1 — llvm/lib/Target/X86 X86ISelLowering.cpp, llvm/test/CodeGen/X86 avx512fp16-combine-fmsubadd.ll avx512fp16-combine-fmsubadd-fadd.ll

[X86] Fold FMADDSUB into VFMULC for fp16 complex multiply (#227698)

We fold fmaddsub using isCFMulFromFMAddSub (formerly isCFMulFromFMSUBADD) into vfmulc.
DeltaFile
+54-0llvm/test/CodeGen/X86/avx512fp16-combine-fmsubadd-fadd.ll
+51-0llvm/test/CodeGen/X86/avx512fp16-combine-fmsubadd.ll
+15-11llvm/lib/Target/X86/X86ISelLowering.cpp
+120-113 files

LLVM/project 5b2879a — llvm/lib/Target/AMDGPU AMDGPUAtomicOptimizer.cpp, llvm/test/CodeGen/AMDGPU local-atomicrmw-fadd.ll atomic_optimizations_local_pointer.ll

[AMDGPU] Skip the iterative atomic scan on native LDS atomics (#227222)

The iterative scan runs a serial loop with one iteration per active lane. For an LDS atomic that the target executes natively, this costs more than the hardware serialization it replaces. Leave such atomics alone when the value is divergent.
DeltaFile
+733-5,435llvm/test/CodeGen/AMDGPU/atomic_optimizations_local_pointer.ll
+68-479llvm/test/CodeGen/AMDGPU/local-atomicrmw-fadd.ll
+7-0llvm/lib/Target/AMDGPU/AMDGPUAtomicOptimizer.cpp
+808-5,9143 files

LLVM/project 123ad4d — llvm/lib/IR Instructions.cpp, llvm/test/Transforms/InstCombine byte-cast-pairs.ll

[IR] Don't fold bitcast pairs through bytes into invalid casts

A pair of bitcasts through a byte can change whether a value is a
pointer, or its address space. Merging such a pair produced an invalid
bitcast and crashed InstCombine, InstSimplify, constant folding and the
IR parser. Give bitcast pairs their own case, and only fold them if
neither end is a pointer, or both are pointers in the same address
space.

Also keep inttoptr and ptrtoint/ptrtoaddr separate from a pointer-byte
bitcast. These pairs used to hit an assertion.

Fixes #207532.
DeltaFile
+107-0llvm/test/Transforms/InstCombine/byte-cast-pairs.ll
+26-0llvm/test/Transforms/InstSimplify/byte-cast-pairs.ll
+20-1llvm/lib/IR/Instructions.cpp
+153-13 files

LLVM/project 795e750 — llvm/lib/Target/PowerPC/GISel PPCInstructionSelector.cpp

PowerPC/GlobalISel: Stop setting kill flags on selected instructions (#229016)

There is no point in maintaining these before register allocation
anymore.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+18-18llvm/lib/Target/PowerPC/GISel/PPCInstructionSelector.cpp
+18-181 files

LLVM/project 73ad673 — llvm/lib/Target/VE VEISelLowering.cpp

VE: Stop setting kill flags on virtual registers before FinalizeISel (#229012)

There is no point in maintaining kill flags before register allocation
anymore.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+30-30llvm/lib/Target/VE/VEISelLowering.cpp
+30-301 files

LLVM/project 07cd029 — llvm/lib/Transforms/Vectorize VPlanTransforms.cpp

[LV] Make partial-reduction naming consistent (NFC) (#222377)

The partial-reduction code used "Chain" to refer to two different
concepts:

1. A `VPPartialReductionChain`, which is a collection of recipes that
   form a partial reduction.
2. A list of `VPPartialReductionChain` objects that forms a chained
   reduction.

This patch renames `VPPartialReductionChain` to
`PartialReductionDescriptor` and updates the terminology as follows: a
`Link` is a single `PartialReductionDescriptor`, a `Chain` is a list of
links, and `Chains` refers to a collection of chains.
DeltaFile
+60-60llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
+60-601 files