[VectorCombine] Fold interleave and widen chained operations (#224005)
Generalize the existing single deinterleave-interleave pair fold by starting from the interleave
and walking backwards through its operands.
Starting from the interleave instead exposes the whole expression tree that produces its operands,
allowing the combine to discover multiple deinterleaves participating in the same reconstructed vector.
The walk follows supported element-wise operations and splats backwards until it reaches the originating deinterleaves. Once all interleave operands can be traced back consistently, the operations can be rebuilt
on the original wider vectors and the intermediate deinterleave-interleave operations removed.
This makes the fold handle patterns with multiple deinterleaved inputs.
[BOLT][AArch64] Remove thunk estimation and warn on margin exhaustion (#228069)
When performing branch range checks compare modeled hop distances with
MaxClusterSize inclusively without charging estimated thunk bytes
against it. This allows a thunk at an adjacent cluster boundary to reach
its target.
Cap MaxClusterSize at 126 MiB to reserve minimum alignment headroom.
Warn when thunk bytes plus text alignment exceed the remaining branch
safety margin.
[Remarks] Fix use-after-free of block scalar values in the YAML remark parser
YAMLRemarkParser::parseStr returned BlockScalarNode::getValue() as is.
The YAML parser copies a block scalar's value into the document's node
allocator, and next() advances the document iterator before returning
the remark, which frees that allocator. So every argument value written
as a block scalar pointed at freed memory by the time the caller saw
it.
Copy block scalar values into an allocator owned by the parser.
Assisted-by: Claude
[X86] Fix reversed condition in ADC/SBB fold of COND_A compares (#228693)
#161388 added a COND_A case to combineAddOrSubToADCOrSBB that reuses the
flags of `CMP X, Y` for ADC/SBB without swapping the operands, computing
X +/- (X <u Y) instead of X +/- (X >u Y).
It was unreachable until #221290 started emitting X86ISD::CMP instead of
X86ISD::SUB for compares of atomic loads. This patch removes it.
`(a - b) - (a > b)` case #161388 targeted goes through the X86ISD::SUB path and is unaffected.
Fixes #228457
Assisted-by: Claude Opus 5.5
CodeGen: Prefer getting the Triple from the Module
Continue replacing TargetMachine::getTargetTriple() with the module's
triple at sites where a Module is one hop away through an available
Function, GlobalValue or MachineModuleInfo.
Where the surrounding class already holds a Subtarget, use its triple
rather than routing through the Module.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
X86: Preserve the dead flag clobber when expanding dynamic allocas
The DYN_ALLOCA pseudos clobber EFLAGS, so propagate the pseudo's dead
flag to the stack adjustment they expand to.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
VE: Use splitAt in expandExtendStackPseudo
Replace the manual block-splitting in expandExtendStackPseudo with
MachineBasicBlock::splitAt. Reduces boilerplate, but there's some
block renumbering churn in the output.
Co-authored-by: Claude (Claude-Opus-4.8)
Mips: Stop setting kill flags on virtual registers before FinalizeISel
There is no point in maintaining these before register allocation anymore.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[llvm-c][ocaml] Deprecate/replace constexpr GEP APIs (#228998)
In the C API, this deprecates LLVMConstGEP, LLVMConstInBoundsGEP and
LLVMConstGEPWithNoWrapFlags. The replacement APIs are LLVMConstPtrAdd
and LLVMConstPtrAddFromIndices. The latter exposes the DataLayout-aware
`getGetElementPtr()` API, and only exists for the sake of convenience.
Similarly, in the OCaml API, remove const_gep and const_inbounds_gep in
favor of const_ptradd and const_ptradd_from_indices, which mirror the C
APIs translated to OCaml.
Unlike the older APIs, the new ones work directly on GEPNoWrapFlags
instead of providing (an incomplete set of) separate APIs for different
flavors.
This completes https://github.com/llvm/llvm-project/pull/227601 by
extending the deprecation of constexpr GEP methods to the C API.
[X86] Fold FMADDSUB into VFMULC for fp16 complex multiply (#227698)
We fold fmaddsub using isCFMulFromFMAddSub (formerly isCFMulFromFMSUBADD) into vfmulc.
[AMDGPU] Skip the iterative atomic scan on native LDS atomics (#227222)
The iterative scan runs a serial loop with one iteration per active lane. For an LDS atomic that the target executes natively, this costs more than the hardware serialization it replaces. Leave such atomics alone when the value is divergent.
[IR] Don't fold bitcast pairs through bytes into invalid casts
A pair of bitcasts through a byte can change whether a value is a
pointer, or its address space. Merging such a pair produced an invalid
bitcast and crashed InstCombine, InstSimplify, constant folding and the
IR parser. Give bitcast pairs their own case, and only fold them if
neither end is a pointer, or both are pointers in the same address
space.
Also keep inttoptr and ptrtoint/ptrtoaddr separate from a pointer-byte
bitcast. These pairs used to hit an assertion.
Fixes #207532.
PowerPC/GlobalISel: Stop setting kill flags on selected instructions (#229016)
There is no point in maintaining these before register allocation
anymore.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
VE: Stop setting kill flags on virtual registers before FinalizeISel (#229012)
There is no point in maintaining kill flags before register allocation
anymore.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[LV] Make partial-reduction naming consistent (NFC) (#222377)
The partial-reduction code used "Chain" to refer to two different
concepts:
1. A `VPPartialReductionChain`, which is a collection of recipes that
form a partial reduction.
2. A list of `VPPartialReductionChain` objects that forms a chained
reduction.
This patch renames `VPPartialReductionChain` to
`PartialReductionDescriptor` and updates the terminology as follows: a
`Link` is a single `PartialReductionDescriptor`, a `Chain` is a list of
links, and `Chains` refers to a collection of chains.