LLVM/project fdd93c0 — llvm/lib/CodeGen TargetLoweringBase.cpp, llvm/lib/LTO ThinLTOCodeGenerator.cpp LTOBackend.cpp

CodeGen: Prefer getting the Triple from the Module (#228620)

Take the triple from the contextual module rather than TargetMachine 
when it's already readily available.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+4-4llvm/lib/CodeGen/TargetLoweringBase.cpp
+2-2llvm/lib/Target/SPIRV/SPIRVEmitIntrinsics.cpp
+1-1llvm/lib/Target/X86/X86ISelLowering.cpp
+1-1llvm/lib/Target/AMDGPU/AMDGPURemoveIncompatibleFunctions.cpp
+1-1llvm/lib/LTO/ThinLTOCodeGenerator.cpp
+1-1llvm/lib/LTO/LTOBackend.cpp
+10-101 files not shown
+11-117 files

LLVM/project 13b5de3 — llvm/lib/Target/AArch64 AArch64TargetTransformInfo.cpp

[AArch64] Stop count cost once force-unroll threshold is reached (NFC). (#228247)

The cost is only used to compare against Aarch64ForceUnrollThreshold.
Stop counting when we reach it, to avoid unnecessary cost queries.

Reduces compile-time on AArch64 by -0.08%.


https://llvm-compile-time-tracker.com/compare.php?from=5ffb13d6d1f87195bba8af13366a152a7de2bfe0&to=18f3cab804f98653c6b3a56b6ac02c9f6960b672&stat=instructions:u
DeltaFile
+3-0llvm/lib/Target/AArch64/AArch64TargetTransformInfo.cpp
+3-01 files

LLVM/project de20437 — clang/lib/AST ExprConstant.cpp, clang/lib/AST/ByteCode InterpBuiltin.cpp

[clang] Refactor CRC32 helper for x86 CRC builtins (#225774)

Move the reflected CRC32 calculation used by the x86 CRC builtins into
llvm::calculateReflectedCRC32() and reuse it from both the AST
interpreter and constant expression evaluator.

This removes duplicated CRC32 implementation logic from Clang. This is
planned to be used for implementing const folding CRC instructions for
x86 and AArch64 in the future in LLVM.

Based on suggestion in
https://github.com/llvm/llvm-project/pull/219452#discussion_r4070846498
DeltaFile
+20-0llvm/include/llvm/Support/CRC.h
+3-9clang/lib/AST/ExprConstant.cpp
+3-9clang/lib/AST/ByteCode/InterpBuiltin.cpp
+26-183 files

LLVM/project 5109478 — orc-rt/test/unit/support AllocActionTest.cpp

[orc-rt] Use the Error matchers in AllocActionTest (#228680)

Use the Error matchers introduced in 4c8a437d0487 to clean up error
checks in AllocActionTest.
DeltaFile
+15-21orc-rt/test/unit/support/AllocActionTest.cpp
+15-211 files

LLVM/project ac9ee31 — llvm/lib/CodeGen TargetLoweringObjectFileImpl.cpp TailDuplicator.cpp, llvm/lib/Target/AArch64 AArch64FrameLowering.cpp

CodeGen: Prefer getting the Triple from the Module

Continue replacing TargetMachine::getTargetTriple() with the module's
triple at sites where a Module is one hop away through an available
Function, GlobalValue or MachineModuleInfo.

Where the surrounding class already holds a Subtarget, use its triple
rather than routing through the Module.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+6-6llvm/lib/Target/ARM/ARMISelLowering.cpp
+6-3llvm/lib/CodeGen/TailDuplicator.cpp
+4-4llvm/lib/CodeGen/TargetLoweringObjectFileImpl.cpp
+2-2llvm/lib/Target/NVPTX/NVPTXISelDAGToDAG.cpp
+1-2llvm/lib/Target/AArch64/AArch64FrameLowering.cpp
+2-1llvm/lib/Target/AMDGPU/AMDGPUTargetObjectFile.cpp
+21-1814 files not shown
+35-3220 files

LLVM/project f71f2b8 — llvm/lib/CodeGen Analysis.cpp TargetLoweringBase.cpp, llvm/lib/LTO ThinLTOCodeGenerator.cpp LTOBackend.cpp

CodeGen: Prefer getting the Triple from the Module

Take the triple from the contextual module rather than TargetMachine
when it's already readily available.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+4-4llvm/lib/CodeGen/TargetLoweringBase.cpp
+2-2llvm/lib/Target/SPIRV/SPIRVEmitIntrinsics.cpp
+1-1llvm/lib/Target/X86/X86ISelLowering.cpp
+1-1llvm/lib/LTO/ThinLTOCodeGenerator.cpp
+1-1llvm/lib/LTO/LTOBackend.cpp
+1-1llvm/lib/CodeGen/Analysis.cpp
+10-101 files not shown
+11-117 files

LLVM/project 4aaffe8 — flang/test/Semantics/OpenMP requires06.f90 requires05.f90

[Flang][OpenMP] Reset REQUIRES directive with new program unit (#227829)

A REQUIRES directive with unified_address, unified_shared_memory, or
reverse_offload must appear lexically before any device construct or
device routine is scoped to a program unit.
DeltaFile
+31-0flang/test/Semantics/OpenMP/requires11.f90
+6-11flang/test/Semantics/OpenMP/requires08.f90
+6-11flang/test/Semantics/OpenMP/requires07.f90
+5-11flang/test/Semantics/OpenMP/requires03.f90
+3-5flang/test/Semantics/OpenMP/requires06.f90
+3-5flang/test/Semantics/OpenMP/requires05.f90
+54-432 files not shown
+61-488 files

LLVM/project 1e80bdc — llvm/lib/CodeGen/AsmPrinter AsmPrinter.cpp, llvm/lib/Target/AArch64 AArch64AsmPrinter.cpp

CodeGen: Prefer getting the Triple from the Module when convenient (#228429)

Take the triple from the contextual module rather than TargetMachine when it's 
already readily available.
DeltaFile
+17-16llvm/lib/CodeGen/AsmPrinter/AsmPrinter.cpp
+5-4llvm/lib/Target/X86/X86AsmPrinter.cpp
+4-4llvm/lib/Target/RISCV/RISCVAsmPrinter.cpp
+3-3llvm/lib/Target/ARM/ARMAsmPrinter.cpp
+3-3llvm/lib/Target/AArch64/AArch64AsmPrinter.cpp
+3-2llvm/lib/Target/Hexagon/HexagonAsmPrinter.cpp
+35-325 files not shown
+41-3811 files

LLVM/project 283cb39 — orc-rt/test/unit/bedrock/sys/darwin StandaloneMachOUnwindInfoRegistrarTest.cpp

[orc-rt] Use the Error matchers in StandaloneMachOUnwindInfoRegistrar… (#228679)

…Test

Use the Error matchers introduced in 4c8a437d0487 to clean up error
checks in StandaloneMachOUnwindInfoRegistrarTest.
DeltaFile
+51-42orc-rt/test/unit/bedrock/sys/darwin/StandaloneMachOUnwindInfoRegistrarTest.cpp
+51-421 files

LLVM/project edeaf21 — llvm/lib/Target/ARM ARMISelLowering.cpp, llvm/test/CodeGen/Thumb optional-def-dead-cpsr.ll

ARM: Preserve the dead flag when activating the optional CPSR def (#227655)

This did not preserve the original dead flag, so it would be recomputed later by 
LiveVariables or RegAllocFast.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+46-0llvm/test/CodeGen/Thumb/optional-def-dead-cpsr.ll
+1-0llvm/lib/Target/ARM/ARMISelLowering.cpp
+47-02 files

LLVM/project 8673f8d — offload/plugins-nextgen/amdgpu/dynamic_hsa hsa.h hsa.cpp, offload/plugins-nextgen/amdgpu/src rtl.cpp

[offload][AMDGPU] Add dynamic HSA version and symbol checks
DeltaFile
+39-2offload/plugins-nextgen/amdgpu/dynamic_hsa/hsa.cpp
+15-0offload/plugins-nextgen/amdgpu/src/rtl.cpp
+2-0offload/plugins-nextgen/amdgpu/dynamic_hsa/hsa.h
+56-23 files

LLVM/project 92ee379 — orc-rt/test/unit/bedrock/sps SimpleRemoteCATest.cpp

[orc-rt] Use the Error matchers in SimpleRemoteCATest (#228672)

Use the Error matchers introduced in 4c8a437d0487 to clean up error
checks in SimpleRemoteCATest.
DeltaFile
+16-16orc-rt/test/unit/bedrock/sps/SimpleRemoteCATest.cpp
+16-161 files

LLVM/project 7516fa2 —

[webkit.UncountedLambdaCapturesChecker] Treat call_once arguments as noescape (#224492)

call_once does not store the lambda in heap so consider all its
arguments as noescape.
DeltaFile
+0-00 files

LLVM/project 09648b6 —

[GlobalISel] [AArch64] Extend SelectShiftMask to handle ADD/SUB mod patterns (#225842)

AArch64 shift instructions only use the low log2(ShiftWidth) bits of the
shift amount, so shifting by X+N where N == 0 mod ShiftWidth is
equivalent to shifting by X alone.

This extends both SelectionDAG (SelectShiftMask) and GlobalISel
(selectShiftMask) implementations.

Part of the incremental work tracked in #224245 to replace
tryShiftAmountMod with ComplexPattern-based isel.

Depends on #223136.
DeltaFile
+0-00 files

LLVM/project 195c816 — llvm/lib/ExecutionEngine/Orc Core.cpp

ORC: Fix flaky OrcLazy tests (#228619)

I've seen this fail a few too many times so just let AI deal with it. I
don't
know anything about orc, but extending lifetime of lock_guard seems
plausible.

Notify lookupInitSymbols CV while holding the mutex

The init-symbol lookup completion callback decremented Count under
LookupMutex but called CV.notify_one() after releasing it. The waiting
thread could observe Count == 0, return from lookupInitSymbols, and
destroy the stack-allocated mutex and condition variable before the
callback signalled it. With concurrent compile threads the callback runs
on a dispatcher thread, so the late notify wrote into reused stack
memory, e.g. during endSession right after deinitialize.

This caused intermittent crashes in
ExecutionEngine/OrcLazy/multiple-compile-threads-basic.ll on macOS

    [6 lines not shown]
DeltaFile
+9-10llvm/lib/ExecutionEngine/Orc/Core.cpp
+9-101 files

LLVM/project 382e6ff — clang/lib/StaticAnalyzer/Checkers/WebKit RawPtrRefLambdaCapturesChecker.cpp, clang/test/Analysis/Checkers/WebKit uncounted-lambda-captures.cpp mock-types.h

[webkit.UncountedLambdaCapturesChecker] Treat call_once arguments as noescape (#224492)

call_once does not store the lambda in heap so consider all its
arguments as noescape.
DeltaFile
+10-3clang/lib/StaticAnalyzer/Checkers/WebKit/RawPtrRefLambdaCapturesChecker.cpp
+13-0clang/test/Analysis/Checkers/WebKit/mock-types.h
+9-0clang/test/Analysis/Checkers/WebKit/uncounted-lambda-captures.cpp
+32-33 files

LLVM/project 6f88885 — mlir/lib/Dialect/OpenMP/Transforms HostOpFiltering.cpp, mlir/test/Dialect/OpenMP host-op-filtering.mlir

[Flang][MLIR] Fix reset for OpenMP dyn_groupprivate clause (#228177)
DeltaFile
+26-0mlir/test/Dialect/OpenMP/host-op-filtering.mlir
+2-0mlir/lib/Dialect/OpenMP/Transforms/HostOpFiltering.cpp
+28-02 files

LLVM/project 6cd72b3 — clang/include/clang/Basic Diagnostic.h, clang/include/clang/Sema Sema.h

Reland "[clang] Don't add documentation comments to the AST if not requested (#206363)" (#221605)

The original PR caused a huge
[regression](https://llvm-compile-time-tracker.com/compare.php?from=aa1058e34a6127df91898829b0e60cbae3111cbf&to=e046dce4a4c80610b49d67bc02c85f86b1a6353d&stat=instructions:u)

The problem was that `areAllIgnored` ends up being called once per
declared entity, and each call walks all 26 diagnostics in
`-Wdocumentation` and `-Wdocumentation-pedantic`. So now the result is
cached and it's a performance
[improvement](https://llvm-compile-time-tracker.com/compare.php?from=5d063386f51b7d9925df7db5f05ec2f0a33f63c4&to=e8f7f96262a312739a990d1747a332ba6f1859b3&stat=instructions%3Au)
again.

The cache is keyed on the diagnostic state plus whether the location is
in a system header, because both change the answer. Keying it on the
state alone is wrong: a declaration coming from a system header would
cache "off" and silence the comments in the user's own code after it.
There is a test for that.

Also, Claude noticed that a similar thing is done in

    [7 lines not shown]
DeltaFile
+107-79clang/lib/Basic/DiagnosticIDs.cpp
+74-4clang/lib/Sema/Sema.cpp
+42-0clang/include/clang/Basic/Diagnostic.h
+38-0clang/test/Sema/warn-documentation-comment-retention.cpp
+28-0clang/include/clang/Sema/Sema.h
+28-0clang/test/AST/ast-dump-comment-retention.cpp
+317-8323 files not shown
+415-11329 files

LLVM/project 2c35201 — llvm/test/CodeGen/AArch64 and-mask-variable.ll shift-mod.ll

[GlobalISel] [AArch64] Extend SelectShiftMask to handle ADD/SUB mod patterns (#225842)

AArch64 shift instructions only use the low log2(ShiftWidth) bits of the
shift amount, so shifting by X+N where N == 0 mod ShiftWidth is
equivalent to shifting by X alone.

This extends both SelectionDAG (SelectShiftMask) and GlobalISel
(selectShiftMask) implementations.

Part of the incremental work tracked in #224245 to replace
tryShiftAmountMod with ComplexPattern-based isel.

Depends on #223136.
DeltaFile
+300-328llvm/test/CodeGen/AArch64/fcvt-i256.ll
+282-306llvm/test/CodeGen/AArch64/fsh.ll
+49-67llvm/test/CodeGen/AArch64/shift.ll
+26-28llvm/test/CodeGen/AArch64/funnel-shift.ll
+35-10llvm/test/CodeGen/AArch64/shift-mod.ll
+6-8llvm/test/CodeGen/AArch64/and-mask-variable.ll
+698-7472 files not shown
+725-7478 files

LLVM/project 6a61012 — llvm/lib/CodeGen/SelectionDAG DAGCombiner.cpp, llvm/test/CodeGen/AArch64 scalar_to_vector.ll

[AArch64][SelectionDAG] Avoid fold bitcast of scalar_to_vector to anyext in big-endian (#225442)

```
  int_vt (bitcast (vec_vt (scalar_to_vector elt_vt:x)))
    => int_vt (any_extend elt_vt:x)
```

This pattern not legal in big-endian, so disable in big-endian.

Fix https://github.com/llvm/llvm-project/issues/225436
DeltaFile
+29-0llvm/test/CodeGen/AArch64/scalar_to_vector.ll
+2-1llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp
+31-12 files

LLVM/project 0bdb393 — llvm/test/tools/llvm-objcopy/ELF section-index-unsupported.test unsupported-machine-specific-shndx.test

[llvm-objcopy,test] Reorganize reserved st_shndx tests (#228664)

Rename section-index-unsupported.test to reserved-shndx.test, fold
unsupported-machine-specific-shndx.test into it, and enhance it.
DeltaFile
+33-0llvm/test/tools/llvm-objcopy/ELF/reserved-shndx.test
+0-17llvm/test/tools/llvm-objcopy/ELF/unsupported-machine-specific-shndx.test
+0-15llvm/test/tools/llvm-objcopy/ELF/section-index-unsupported.test
+33-323 files

LLVM/project d97a5ad — llvm/include/llvm/ADT DenseMap.h

[ADT] Make protected members of DenseMapBase private (NFC) (#228651)

Neither DenseMap nor SmallDenseMap accesses these members after #227063.
This patch also moves getMemorySize into the main public section.

Assisted-by: Antigravity
DeltaFile
+9-11llvm/include/llvm/ADT/DenseMap.h
+9-111 files

LLVM/project 0fae4df — llvm/lib/Target/RISCV RISCVMergeBaseOffset.cpp, llvm/test/CodeGen/RISCV fold-mem-offset.ll

[RISCV] Make sure ADDI isn't a frame index in RISCVMergeBaseOffsetOpt::foldLargeOffset. (#228635)

Fixes #228053
DeltaFile
+35-0llvm/test/CodeGen/RISCV/fold-mem-offset.ll
+3-0llvm/lib/Target/RISCV/RISCVMergeBaseOffset.cpp
+38-02 files

LLVM/project b3e41fd — clang/docs HIPSupport.md, clang/lib/CodeGen CGException.cpp

[HIPStdPar] Report reachable C++ exceptions from GPU kernel (#228082)

HipStdPar removes all host functions which are unreachable but there can
be cases where some reachable function contains C++ exceptions it is
forced to device code and since GPU devices don't support C++ exceptions
it must be reported.

This change allows for any reachable C++ exception to be reported as
error.

Part of https://github.com/llvm/llvm-project/issues/221941
DeltaFile
+236-0llvm/test/Transforms/HipStdPar/unsupported-exceptions.ll
+149-0clang/test/CodeGenHipStdPar/unsupported-exceptions.cpp
+49-10llvm/lib/Transforms/HipStdPar/HipStdPar.cpp
+15-0clang/lib/CodeGen/CGException.cpp
+6-1clang/docs/HIPSupport.md
+455-115 files

LLVM/project 4146a3d — llvm/lib/Target/RISCV RISCVISelLowering.cpp

[RISCV] Convert mask to scalable vector in lowerVPREDUCE. (#228219)

Assisted-by: Claude
DeltaFile
+4-3llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+4-31 files

LLVM/project 39e3785 — clang/lib/CIR/CodeGen CIRGenBuiltin.cpp, clang/test/CIR/CodeGenBuiltins builtin-reduce-arithmetic-sve.c builtin-reduce-arithmetic.c

 [CIR] Lowering for __builtin_reduce_assoc_fadd (#226095)

Added lowering for __builtin_reduce_assoc_fadd with the use of the new
implementation of CIR FastMathFlags, following same lowering path as
Classic Codegen.

Added tests for the same.
DeltaFile
+29-10clang/lib/CIR/CodeGen/CIRGenBuiltin.cpp
+34-0clang/test/CIR/CodeGenBuiltins/builtin-reduce-arithmetic.c
+22-0clang/test/CIR/CodeGenBuiltins/builtin-reduce-arithmetic-sve.c
+85-103 files

LLVM/project 52b4562 — lldb/test/API/lang/cpp/decl-from-submodule TestDeclFromSubmodule.py

fixup! [LLDB] Mark some objective-C and ClangModule test as darwin only
DeltaFile
+2-2lldb/test/API/lang/cpp/decl-from-submodule/TestDeclFromSubmodule.py
+2-21 files

LLVM/project 076266f — llvm/lib/Transforms/Vectorize SLPVectorizer.cpp, llvm/test/Transforms/SLPVectorizer/AMDGPU alt-fmul-fadd-cost.ll

[SLP] Partially revert #224931 for alternate nodes

Do not pass scalar context to alternate node vector cost queries.

PR #224931 passes the scalar main/alternate op as the context instruction when
costing the vector ops of an alternate node. Those vector ops only feed the
lane-select shuffle, so use-based discounts derived from the scalar's users
do not apply. On AMDGPU this priced an [fmul | fadd] node's vector fmul as
fused (free) where the sole user of the scalar fmul is fadd/fsub, keeping
unprofitable subtrees vectorized. In a rocFFT kernel on gfx950 this raised
VGPR usage from 160 to 178 and caused a performance regression. Restore the
null context for these queries.
DeltaFile
+76-0llvm/test/Transforms/SLPVectorizer/AMDGPU/alt-fmul-fadd-cost.ll
+3-4llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+79-42 files

LLVM/project a7ab057 — llvm/lib/Transforms/Vectorize SLPVectorizer.cpp, llvm/test/Transforms/SLPVectorizer/AMDGPU alt-fmul-fadd-cost.ll

[SLP] Partially revert #224931 for alternate nodes

Do not pass scalar context to alternate node vector cost queries.

PR #224931 passes the scalar main/alternate op as the context instruction when
costing the vector ops of an alternate node. Those vector ops only feed the
lane-select shuffle, so use-based discounts derived from the scalar's users
do not apply. On AMDGPU this priced an [fmul | fadd] node's vector fmul as
fused (free) where the sole user of the scalar fmul is fadd/fsub, keeping
unprofitable subtrees vectorized. In a rocFFT kernel on gfx950 this raised
VGPR usage from 160 to 178 and caused a performance regression. Restore the
null context for these queries.
DeltaFile
+76-0llvm/test/Transforms/SLPVectorizer/AMDGPU/alt-fmul-fadd-cost.ll
+3-4llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+79-42 files

LLVM/project 0015ba5 — libc/test/UnitTest CMakeLists.txt ExecuteFunctionUnix.cpp, libc/test/src/math/smoke SubTest.h MulTest.h

[libc][test] Fix death test timeouts and dead-code elimination in math tests. (#227966)

- In `ExecuteFunctionUnix.cpp`, call `prctl(PR_SET_DUMPABLE, 0)` in
child processes to prevent external core dump handlers (such as apport)
from intercepting expected crashes during death tests, eliminating 10s
poll timeouts under parallel lit runs.
- Close `pipe_fds[0]` properly in `invoke_in_subprocess` to avoid
leaking file descriptors.
- In math smoke test templates (`AddTest.h`, `SubTest.h`, `MulTest.h`,
`DivTest.h`), assign test function results to `[[maybe_unused]] volatile
OutType res` to prevent GCC from dead-code eliminating floating point
operations in `test_inexact_results` when FMA optimization is disabled.

Assisted-by: Gemini
DeltaFile
+19-0utils/bazel/llvm-project-overlay/libc/BUILD.bazel
+13-0libc/test/UnitTest/ExecuteFunctionUnix.cpp
+6-0utils/bazel/llvm-project-overlay/libc/test/UnitTest/BUILD.bazel
+6-0libc/test/UnitTest/CMakeLists.txt
+1-1libc/test/src/math/smoke/SubTest.h
+1-1libc/test/src/math/smoke/MulTest.h
+46-22 files not shown
+48-48 files