LLVM/project f12785ellvm/lib/Target/AMDGPU SIInstrInfo.h SIInstrInfo.cpp

AMDGPU: Pass MachineRegisterInfo to getImmOrMaterializedImm

The only used the operand to reach the MachineRegisterInfo. Pass it
directly so it no longer depends on MachineOperand::getParent().

Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
DeltaFile
+8-8llvm/lib/Target/AMDGPU/SIFoldOperands.cpp
+3-3llvm/lib/Target/AMDGPU/SILoadStoreOptimizer.cpp
+2-2llvm/lib/Target/AMDGPU/SIInstrInfo.cpp
+2-1llvm/lib/Target/AMDGPU/SIInstrInfo.h
+15-144 files

LLVM/project dc3696bllvm/include/llvm/CodeGen TargetInstrInfo.h, llvm/lib/CodeGen MachineVerifier.cpp

CodeGen: Pass instruction and operand index to isPCRelRegisterOperandLegal

Replace the MachineOperand argument to the
TargetInstrInfo::isPCRelRegisterOperandLegal hook with the containing
instruction and operand index. The M68k implementation only used the operand
to recover its parent instruction and operand number, so this drops the
dependence on MachineOperand::getParent().

Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
DeltaFile
+8-10llvm/lib/Target/M68k/M68kInstrInfo.cpp
+5-3llvm/include/llvm/CodeGen/TargetInstrInfo.h
+2-1llvm/lib/Target/M68k/M68kInstrInfo.h
+1-1llvm/lib/CodeGen/MachineVerifier.cpp
+16-154 files

LLVM/project 9f89076clang/test/CodeGen/RISCV/rvv-intrinsics-autogenerated/zvabd/policy/non-overloaded vabdu.c, libcxx/test/std/language.support/support.limits/support.limits.general version.version.compile.pass.cpp

Merge branch 'main' into users/gandhi56/revert-combine-redundant-ballot
DeltaFile
+0-6,246llvm/test/CodeGen/AMDGPU/NextUseAnalysis/test_ers_nested_loops.mir
+0-4,877llvm/test/CodeGen/AMDGPU/NextUseAnalysis/test_ers_emit_restore_in_loop_preheader2.mir
+2,428-0llvm/test/CodeGen/SPIRV/extensions/SPV_EXT_long_vector/unmerge-crash-0.ll
+2,426-0llvm/test/CodeGen/SPIRV/extensions/SPV_EXT_long_vector/unmerge-crash-1.ll
+2,372-2libcxx/test/std/language.support/support.limits/support.limits.general/version.version.compile.pass.cpp
+1,449-52clang/test/CodeGen/RISCV/rvv-intrinsics-autogenerated/zvabd/policy/non-overloaded/vabdu.c
+8,675-11,1771,729 files not shown
+86,586-42,5911,735 files

LLVM/project fce0156llvm/lib/CodeGen/SelectionDAG TargetLowering.cpp, llvm/test/CodeGen/X86 intrinsic-cttz-elts.ll

[SelectionDAG] Preserve CTTZ_ELTS lane count during expansion (#217982)

Fixes #216649

`expandCttzElts` derives `VL` from its legalized auxiliary step vector.
When that helper vector is widened for target legality, its lane count
may differ from the logical lane count of the `CTTZ_ELTS` operand.

On X86 with AVX512F, the step vector for a semantic `<4 x i1>` mask is
widened from `v4i8` to `v16i8`. As a result, an all-zero
`llvm.experimental.cttz.elts` input returns 16 instead of the required
result of 4.

Preserve the `ElementCount` from the `CTTZ_ELTS` operand and use it when
materializing `VL`. Auxiliary step-vector legalization can still widen
its representation without changing the logical lane domain of the
operation.

Add AVX512F regression coverage for the affected `v4i1` and `v8i1`

    [3 lines not shown]
DeltaFile
+95-0llvm/test/CodeGen/X86/intrinsic-cttz-elts.ll
+5-5llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
+100-52 files

LLVM/project e86c19dclang/docs ReleaseNotes.md, clang/lib/Sema SemaCoroutine.cpp

[Clang] Fix -Wunused-parameter for implicit coroutine uses (#217518)

Clang's coroutine semantic analysis builds references to coroutine
parameters
while looking up a class-specific allocation function. When overload
resolution
falls back to a size-only `operator new`, these speculative references
currently
suppress `-Wunused-parameter`.

Preserve each parameter's referenced state while collecting placement
arguments, then mark the parameters referenced only when those arguments
are
included in the selected allocation call. Keep this distinction when
placement
arguments are replaced by `std::nothrow`.

Apply the same rule to promise initialization: when initialization using
the

    [16 lines not shown]
DeltaFile
+44-0clang/test/SemaCXX/warn-unused-parameters-coroutine.cpp
+30-3clang/lib/Sema/SemaCoroutine.cpp
+5-0clang/docs/ReleaseNotes.md
+79-33 files

LLVM/project b085518.github/workflows release-binaries.yml, clang/cmake/caches Release.cmake

Revert "Revert "workflows/release-binaries: Disable flang on Darwin (#164667)"" (#218978)

Reverts llvm/llvm-project#216667

This change was ported to the `release/23.x` branch in #217059, and when
we created the first release that included this change (3.1.0), the job
for the MacOS ARM binaries was killed when the job hit the 6 hour mark.
Previous 3.1.0-rc release did not include this change and all completed
well within the 6 hour time out.

To enable the job that builds the release binaries to complete within
the allotted time, I am reverting this change which will essentially
disable flang from building on Darwin.

In the future if we get faster builders, we can explore re-enabling
building flang.

(cherry picked from commit 3d771f477aa026ca41cb3c4e492d7a07fa852cbe)
DeltaFile
+8-2clang/cmake/caches/Release.cmake
+0-7.github/workflows/release-binaries.yml
+8-92 files

LLVM/project aedf7d6llvm/lib/Target/X86 X86ISelLowering.cpp

[X86] LowerAVXCONCAT_VECTORS - fix some clang-format messiness. NFC. (#218504)

Bad indentation and missing braces

Pulled out of #218381 to simplify functional diff

(cherry picked from commit a6eb1e1e75a2d29f8f458c6ee513ae5b030ff9a0)
DeltaFile
+11-12llvm/lib/Target/X86/X86ISelLowering.cpp
+11-121 files

LLVM/project 55a6a0bllvm/lib/Target/X86 X86ISelLowering.cpp, llvm/test/CodeGen/X86 vector-shuffle-combining-avx512f.ll

[X86] LowerAVXCONCAT_VECTORS - collect all subvector operands before ReplaceAllUsesWith call (#218381)

We can end up referencing a node replaced with ISD::DELETED_NODE

Fixes #218379

(cherry picked from commit 7bced7f78781d3bdc179e64c777f50ef46ab1375)
DeltaFile
+43-0llvm/test/CodeGen/X86/vector-shuffle-combining-avx512f.ll
+12-13llvm/lib/Target/X86/X86ISelLowering.cpp
+55-132 files

LLVM/project 4dea21ellvm/lib/Transforms/Vectorize SLPVectorizer.cpp, llvm/test/Transforms/SLPVectorizer/X86 reduction-vals-used-as-load-indices.ll

[SLP]Extend GEP pointer-chain cost to casts and non-root external uses

Cherry-pick of 084c5507ee0146a7506f3117868082162760b689 to release/23.x.
Fixes AArch64 regression introduced by 376311097a27a6eb99ec2614e7c39c97ce33172f.

Original Pull Request: #216520
Recommit after perf regression fixes: #217683
DeltaFile
+219-0llvm/test/Transforms/SLPVectorizer/X86/reduction-vals-used-as-load-indices.ll
+33-42llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+252-422 files

LLVM/project d81f5d2llvm/lib/Transforms/Vectorize SLPVectorizer.cpp, llvm/test/Transforms/SLPVectorizer/AArch64 long-non-power-of-2.ll abs-mul-buildvector-in-loop.ll

[SLP]Fix narrow-tree gate width and extract miscount

Measure the narrowness of in-loop trees by the actual vectorization
width instead of the feeder-load width, and skip vector-typed scalars
in the instruction count check to match the cost model.

Fixes #216715

Reviewers: efriedma-quic, bababuck, RKSimon, kartcq

Pull Request: https://github.com/llvm/llvm-project/pull/216797
DeltaFile
+111-0llvm/test/Transforms/SLPVectorizer/AArch64/abs-mul-buildvector-in-loop.ll
+17-2llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+8-6llvm/test/Transforms/SLPVectorizer/X86/multi-parent-instr-copyable-regular.ll
+5-5llvm/test/Transforms/SLPVectorizer/X86/deleted-instructions-clear.ll
+4-2llvm/test/Transforms/SLPVectorizer/AArch64/long-non-power-of-2.ll
+145-155 files

LLVM/project 490d4edllvm/lib/CodeGen/GlobalISel LegalizerHelper.cpp, llvm/test/CodeGen/AArch64 arm64ec-fp128-cmp.ll

[GlobalISel] Fix inverted libcall status check in createFCMPLibcall (#219242)

BuildLibcall tested `if (!Status)` on the LegalizeResult returned by
createLibcall. Since LegalizeResult is an enum with AlreadyLegal == 0,
that condition is false for both Legalized and UnableToLegalize, so a
failed libcall was never detected. The helper then built an ICMP against
the undefined result register and reported success, so the legalizer
never fell back.

Check `Status != Legalized` instead, so failures propagate and the
caller can bail out.

On arm64ec, GlobalISel call lowering is unimplemented, so this silently
dropped `fp128` compare libcalls at `-O0`. Add a test that checks the
expected `__eqtf2`, `__lttf2` and `__unordtf2` calls are emitted at both
`-O0` and `-O2`.

(cherry picked from commit ded76cf640b9b7c1425f9543c54da402bbda7ca5)
DeltaFile
+33-0llvm/test/CodeGen/AArch64/arm64ec-fp128-cmp.ll
+1-1llvm/lib/CodeGen/GlobalISel/LegalizerHelper.cpp
+34-12 files

LLVM/project f54a7b7llvm/lib/Target/LoongArch LoongArchFloat32InstrInfo.td, llvm/test/CodeGen/LoongArch pr215935.ll

[LoongArch] Fix selection of BRCOND with constant conditions (#216027)

LoongArch DAG instruction selection could fail to select
`LoongArchISD::BRCOND` when its condition was a constant integer. Add
patterns to lower constant zero and one conditions to `BEQZ` and `BNEZ`.

Co-authored-by: wanglei <wanglei at loongson.cn>
Fixes: https://github.com/llvm/llvm-project/issues/215935
(cherry picked from commit 17adf57f977f24685e535d649d2244610b104a13)
DeltaFile
+50-0llvm/test/CodeGen/LoongArch/pr215935.ll
+3-0llvm/lib/Target/LoongArch/LoongArchFloat32InstrInfo.td
+53-02 files

LLVM/project 910bfcdllvm/lib/Transforms/Vectorize VectorCombine.cpp, llvm/test/Transforms/VectorCombine/RISCV fold-vp-load.ll

[VectorCombine] Fix foldBitcastOfVPLoad reordering loads (#218336)

We were inserting the new vp.load where the bitcast was, which would
reorder loads. This should hopefully fix RISC-V buildbot failures that
were exposed after 93ac788df8ff

(cherry picked from commit 2eda652e5cd7d70de8dd4f33a7cece6b70465163)
DeltaFile
+14-0llvm/test/Transforms/VectorCombine/RISCV/fold-vp-load.ll
+1-0llvm/lib/Transforms/Vectorize/VectorCombine.cpp
+15-02 files

LLVM/project 5c70951llvm/lib/Target/RISCV RISCVTargetTransformInfo.cpp, llvm/test/Transforms/InstCombine/RISCV riscv-vsetvli-range.ll

[RISCV] Infer vl > 0 if AVL > 0 in range attributes (#219385)

A follow up from #218652, a non-zero AVL guarantees a non-zero vl per
the spec
DeltaFile
+15-2llvm/test/Transforms/InstCombine/RISCV/riscv-vsetvli-range.ll
+8-6llvm/lib/Target/RISCV/RISCVTargetTransformInfo.cpp
+23-82 files

LLVM/project ac7e309llvm/lib/Target/SPIRV SPIRVInstructionSelector.cpp, llvm/test/CodeGen/SPIRV/hlsl-intrinsics discard.ll

[SPIR-V] Erase all instructions after OpKill, not just the next one (#219022)
DeltaFile
+26-0llvm/test/CodeGen/SPIRV/hlsl-intrinsics/discard.ll
+6-4llvm/lib/Target/SPIRV/SPIRVInstructionSelector.cpp
+32-42 files

LLVM/project d5fc211lldb/include/lldb/Utility Locked.h, lldb/unittests/Utility LockedTest.cpp

[lldb] Add Guarded<T, Mutex> to Locked.h

LLDB's code base has many variables that have an associated mutex that
needs to be locked to safely access that variable from multiple
threads. However, this locking scheme is currently not enforced by the
compiler and code sometimes accesses these variables without aquiring
the respective mutex first.

This patch introduces a `Guarded` class that strictly enforces that
some memory is only accessed after the respective mutex has been
aquired. This class hands out `Locked` objects for every access which
guarentee that the mutex is held as long as the variable is in scope.
DeltaFile
+37-0lldb/unittests/Utility/LockedTest.cpp
+28-0lldb/include/lldb/Utility/Locked.h
+65-02 files

LLVM/project 1616a75llvm/lib/Target/X86 X86CompressEVEX.cpp

X86: Use use_instructions in CompressEVEX cross-block check (#219408)

The loop only checks the using instruction's parent block, so iterate
use_instructions() instead of the operands and checking their parents.

Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
DeltaFile
+2-2llvm/lib/Target/X86/X86CompressEVEX.cpp
+2-21 files

LLVM/project d1840bdmlir/lib/Dialect/LLVMIR/IR LLVMDialect.cpp, mlir/test/Dialect/LLVMIR func.mlir

[mlir][llvm] verify external llvm.func not have `function_entry_count` attr (#219231)

Fixes #218625.
DeltaFile
+7-0mlir/test/Dialect/LLVMIR/func.mlir
+4-0mlir/lib/Dialect/LLVMIR/IR/LLVMDialect.cpp
+11-02 files

LLVM/project 8e1e1adllvm/lib/Target/AMDGPU GCNSchedStrategy.cpp

AMDGPU: Use use_nodbg_instructions in GCNSchedStrategy MFMA check (#219403)
DeltaFile
+2-2llvm/lib/Target/AMDGPU/GCNSchedStrategy.cpp
+2-21 files

LLVM/project 8a87673llvm/lib/Target/Hexagon HexagonEarlyIfConv.cpp

Hexagon: Use use_instructions in EarlyIfConversion predicate check (#219405)

The loop only checks whether any user is a PHI, so iterate
use_instructions() instead of the operands to query the parents.

Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
DeltaFile
+2-2llvm/lib/Target/Hexagon/HexagonEarlyIfConv.cpp
+2-21 files

LLVM/project 1829cfallvm/include/llvm/CodeGen SelectionDAG.h, llvm/lib/CodeGen/SelectionDAG SelectionDAG.cpp

[SelectionDAG] Unique VT lists in a DenseSet instead of a FoldingSet (#219364)

Every getVTList call serializes the list's raw bits into a
FoldingSetNodeID and hashes it with xxh3, and every SDVTListNode carries
an interned copy of that profile. The list itself is the key.

Key VTLists on the ArrayRef the returned SDVTList already points at, and
let the two-, three- and four-type overloads share the ArrayRef one.

Aided by Opus 5
DeltaFile
+11-66llvm/lib/CodeGen/SelectionDAG/SelectionDAG.cpp
+15-37llvm/include/llvm/CodeGen/SelectionDAG.h
+26-1032 files

LLVM/project c9b0931lldb/packages/Python/lldbsuite/test decorators.py, lldb/packages/Python/lldbsuite/test/tools/lldb-server gdbremote_testcase.py

[lldb] Add requireSocketPermission decorator for tests that bind sockets (#219208)

We sometimes run tests in sandboxed environments that deny all calls to
`bind`. This breaks a few of our tests that e.g. use a mock GDB server
or any other functionality involving sockets.

This patch adds a requireSocketPermission decorator (and an equivalent
utility for unittests) that check whether we are allowed to call bind.
If we aren't allowed to call bind, we skip the few tests that need this
functionality.
DeltaFile
+25-1lldb/packages/Python/lldbsuite/test/decorators.py
+12-0lldb/unittests/SBTestingSupport/SBTestUtilities.cpp
+4-0lldb/unittests/API/SBProtocolServerTest.cpp
+2-1lldb/packages/Python/lldbsuite/test/tools/lldb-server/gdbremote_testcase.py
+3-0lldb/unittests/SBTestingSupport/SBTestUtilities.h
+2-0lldb/test/API/tools/lldb-dap/launch/TestDAP_launch_termination.py
+48-210 files not shown
+60-216 files

LLVM/project d59fa1cllvm/lib/Target/Hexagon CMakeLists.txt Hexagon.h, llvm/test/CodeGen/Hexagon align-global-arrays.ll

[Hexagon] Add GlobalArrayAlignment pass (#217850)

Add a module pass that raises the alignment of global integer arrays
(char, short, int), including multi-dimensional arrays, to an 8-byte
boundary. This gives their base address a wider alignment, which is
beneficial for the wide (double-word) loads and stores available on
Hexagon.

At -O1/-O2 the pass keeps byte and half-word arrays at their natural
alignment to reduce .rodata size; full 8-byte alignment is applied at
-O3. This size-reduction behavior can be disabled with
-hexagon-disable-align-opt-byte-half.

The pass is enabled by default and can be disabled with
-hexagon-disable-global-array-align.

Co-Authored by: Jyotsna Verma jverma at quicinc.com
DeltaFile
+125-0llvm/lib/Target/Hexagon/HexagonAlignGlobalArrays.cpp
+77-0llvm/test/CodeGen/Hexagon/align-global-arrays.ll
+6-0llvm/lib/Target/Hexagon/HexagonTargetMachine.cpp
+4-0llvm/lib/Target/Hexagon/Hexagon.h
+1-0llvm/lib/Target/Hexagon/CMakeLists.txt
+213-05 files

LLVM/project fc00f85llvm/lib/Target/PowerPC PPCTLSDynamicCall.cpp

PowerPC: Use use_instructions in TLSDynamicCall user collection

The loop only collects the using instructions, so iterate
use_instructions() instead of the operands to get their parents.

Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
DeltaFile
+2-2llvm/lib/Target/PowerPC/PPCTLSDynamicCall.cpp
+2-21 files

LLVM/project c8f86a9libcxx/include/__vector vector.h

[libc++] Use __copy_n in vector::__assign_with_size (#218370)

This was originally part of #214132. However, that patch has some
difficult to track down performance issue. I'm splitting this up to make
the search easier.
DeltaFile
+1-1libcxx/include/__vector/vector.h
+1-11 files

LLVM/project 7599928llvm/lib/Target/X86 X86CompressEVEX.cpp

X86: Use use_instructions in CompressEVEX cross-block check

The loop only checks the using instruction's parent block, so iterate
use_instructions() instead of the operands and checking their parents.

Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
DeltaFile
+2-2llvm/lib/Target/X86/X86CompressEVEX.cpp
+2-21 files

LLVM/project c1a0ce0mlir/include/mlir/Dialect/LLVMIR NVVMOps.td, mlir/lib/Dialect/LLVMIR/IR NVVMDialect.cpp

[MLIR][NVVM] Support tcgen05.mma{.block_scale}.decompress_b Ops (#218354)

This change adds support for `tcgen05.mma.decompress_b` and
`tcgen05.mma.block_scale.decompress_b` MLIR Ops.
DeltaFile
+307-0mlir/test/Target/LLVMIR/nvvm/tcgen05-mma-tensor-decompress-b.mlir
+307-0mlir/test/Target/LLVMIR/nvvm/tcgen05-mma-shared-decompress-b.mlir
+157-0mlir/test/Target/LLVMIR/nvvm/tcgen05-mma-block-scale-tensor-decompress-b.mlir
+157-0mlir/test/Target/LLVMIR/nvvm/tcgen05-mma-block-scale-shared-decompress-b.mlir
+144-0mlir/include/mlir/Dialect/LLVMIR/NVVMOps.td
+123-0mlir/lib/Dialect/LLVMIR/IR/NVVMDialect.cpp
+1,195-01 files not shown
+1,206-07 files

LLVM/project 6f7822allvm/lib/Target/Hexagon HexagonEarlyIfConv.cpp

Hexagon: Use use_instructions in EarlyIfConversion predicate check

The loop only checks whether any user is a PHI, so iterate
use_instructions() instead of the operands to query the parents.

Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
DeltaFile
+2-2llvm/lib/Target/Hexagon/HexagonEarlyIfConv.cpp
+2-21 files

LLVM/project cb05854llvm/lib/Target/AMDGPU GCNSchedStrategy.cpp

AMDGPU: Use use_nodbg_instructions in GCNSchedStrategy MFMA check

The loop only inspects the using instruction, so iterate
use_nodbg_instructions() instead of the operands and their parents.

Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
DeltaFile
+2-2llvm/lib/Target/AMDGPU/GCNSchedStrategy.cpp
+2-21 files

LLVM/project ed96f74llvm/lib/Target/RISCV RISCVInstrInfo.h RISCVVLOptimizer.cpp

RISCV: Pass MachineRegisterInfo to isVLKnownLE

isVLKnownLE and its getEffectiveImm helper used the VL operands to reach
the MachineRegisterInfo. Pass it directly so they no longer depend on
MachineOperand::getParent().

Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
DeltaFile
+8-8llvm/lib/Target/RISCV/RISCVVectorPeephole.cpp
+8-8llvm/lib/Target/RISCV/RISCVInstrInfo.cpp
+7-7llvm/lib/Target/RISCV/RISCVVLOptimizer.cpp
+2-1llvm/lib/Target/RISCV/RISCVInstrInfo.h
+25-244 files