LLVM/project 1b5f644 —

[libc++] Encode the standard version in the ABI tag (#218527)

This prevents ODR mismatch issues from biting us across standard
versions. The order of elements in the ABI tag is now: libc++ version,
hardening mode, assertion semantic, exceptions, standard version.

Fixes #218524
DeltaFile
+0-00 files

LLVM/project fc1884a — llvm/test/CodeGen/AMDGPU frem.ll

Update lit merge conflict
DeltaFile
+49-31llvm/test/CodeGen/AMDGPU/frem.ll
+49-311 files

LLVM/project 09cdef0 — libc/src/__support/math CMakeLists.txt cosf_double_eval.h, libc/test/src/math/exhaustive sinf_float_test.cpp

[libc][math] Reorganize sinf and cosf function selection (#226529)

Reorganises sinf and cosf similarly to #224735 such that:
src/__support/math/func.h selects implementation, and
src/__support/math/func_<type>_eval.h implements func with <type> as the
intermediate computational type.
DeltaFile
+18-174libc/src/__support/math/sinf.h
+182-0libc/src/__support/math/sinf_double_eval.h
+17-151libc/src/__support/math/cosf.h
+162-0libc/src/__support/math/cosf_double_eval.h
+56-4libc/src/__support/math/CMakeLists.txt
+0-47libc/test/src/math/exhaustive/sinf_float_test.cpp
+435-37611 files not shown
+621-55117 files

LLVM/project e15e5e6 — llvm/test/CodeGen/RISCV clmul.ll, llvm/test/CodeGen/RISCV/rvv clmulh-sdnode.ll clmul-sdnode.ll

[ISel] Improve `clmul` fallback implementation (#204802)

Generalize the approach from
https://github.com/llvm/llvm-project/pull/203727 to narrower and wider
integers.

We still need the fallback for when multiplication isn't available, and
it turns out that for some widths the fallback emits fewer instructions,
the naive fallback is still used for `i1`, `i3`, `i4` and `i9`. I've
also now enabled wider integers (`i128` and `i256` have uses in
cryptography).

Based on my local experiments, the Karatsuba approach (e.g. as in
https://github.com/rust-lang/rust/pull/152132#discussion_r2778609222) is
not actually better than zero extending the input and using
multiplication with holes on the wider type.

CC https://github.com/llvm/llvm-project/issues/203694
CC: @eisenwave
DeltaFile
+4,036-4,380llvm/test/CodeGen/RISCV/rvv/clmul-sdnode.ll
+2,544-2,872llvm/test/CodeGen/RISCV/rvv/clmulh-sdnode.ll
+1,020-1,412llvm/test/CodeGen/X86/clmul-vector-512.ll
+850-1,332llvm/test/CodeGen/X86/clmul-vector-256.ll
+772-1,269llvm/test/CodeGen/X86/clmul-vector.ll
+399-1,451llvm/test/CodeGen/RISCV/clmul.ll
+9,621-12,71614 files not shown
+11,877-17,47120 files

LLVM/project a537e6e — clang/test/CodeGen/RISCV rvp-intrinsics.c, llvm/test/CodeGen/AMDGPU load-constant-i8.ll load-local-i16.ll

Merge branch 'users/zGoldthorpe/wide-copies/precommit' into users/zGoldthorpe/wide-copies/materialise
DeltaFile
+42,441-42,426llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+7,527-7,623llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+4,022-8,362clang/test/CodeGen/RISCV/rvp-intrinsics.c
+2,603-2,582llvm/test/CodeGen/AMDGPU/load-constant-i1.ll
+1,833-1,839llvm/test/CodeGen/AMDGPU/load-local-i16.ll
+1,790-1,779llvm/test/CodeGen/AMDGPU/load-constant-i8.ll
+60,216-64,6111,268 files not shown
+125,107-116,5321,274 files

LLVM/project 15ec5f2 — llvm/test/CodeGen/AMDGPU bitcast-vector-extract.ll

[AMDGPU] Replace GCN-NOT with autogen checks in test (#227148)

The assembly has many v_mov_b32s, so minor scheduling changes can
trigger failure on the GCN-NOT. It seems the test is designed to show
CSE behavior, which still holds even if minor scheduling changes break
the NOT checks.
DeltaFile
+235-30llvm/test/CodeGen/AMDGPU/bitcast-vector-extract.ll
+235-301 files

LLVM/project b727f3d — clang/test/CodeGen/RISCV rvp-intrinsics.c, llvm/test/CodeGen/AMDGPU load-constant-i8.ll load-local-i16.ll

Merge branch 'main' into users/zGoldthorpe/wide-copies/precommit
DeltaFile
+42,441-42,426llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+7,527-7,623llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+4,022-8,362clang/test/CodeGen/RISCV/rvp-intrinsics.c
+2,603-2,582llvm/test/CodeGen/AMDGPU/load-constant-i1.ll
+1,833-1,839llvm/test/CodeGen/AMDGPU/load-local-i16.ll
+1,790-1,779llvm/test/CodeGen/AMDGPU/load-constant-i8.ll
+60,216-64,6111,268 files not shown
+125,100-116,5061,274 files

LLVM/project 34ba34d — llvm/lib/Transforms/InstCombine InstCombineCalls.cpp, llvm/test/Transforms/InstCombine bitreverse.ll

[InstCombine] Fix ProfCheck for optimizing bitreverse (#227189)
DeltaFile
+8-3llvm/test/Transforms/InstCombine/bitreverse.ll
+8-2llvm/lib/Transforms/InstCombine/InstCombineCalls.cpp
+0-1llvm/utils/profcheck-xfail.txt
+16-63 files

LLVM/project 4c79537 — llvm/include/llvm/Analysis AliasSetTracker.h, llvm/lib/Analysis AliasSetTracker.cpp

[LICM] Drop *only* per-iteration AA tags
DeltaFile
+48-27llvm/lib/Transforms/Scalar/LICM.cpp
+6-4llvm/test/Transforms/LICM/scalar-promote-aa-tags.ll
+4-3llvm/lib/Analysis/AliasSetTracker.cpp
+1-1llvm/include/llvm/Analysis/AliasSetTracker.h
+59-354 files

LLVM/project d9e7b3e — llvm/lib/Transforms/Scalar LICM.cpp

fix formatting
DeltaFile
+3-3llvm/lib/Transforms/Scalar/LICM.cpp
+3-31 files

LLVM/project 17352e5 — llvm/lib/Transforms/Scalar LICM.cpp

Collect per-iteration alias scopes in advance
DeltaFile
+37-50llvm/lib/Transforms/Scalar/LICM.cpp
+37-501 files

LLVM/project b614800 — llvm/lib/Transforms/Scalar LICM.cpp, llvm/test/Transforms/LICM scalar-promote-aa-tags.ll

[LICM] Drop per-iteration AA tags
DeltaFile
+46-1llvm/lib/Transforms/Scalar/LICM.cpp
+6-12llvm/test/Transforms/LICM/scalar-promote-aa-tags.ll
+52-132 files

LLVM/project 0df5a52 — llvm/test/Transforms/LICM scalar-promote-aa-tags.ll

Add precommit test for noalias scope declared outside loop
DeltaFile
+58-0llvm/test/Transforms/LICM/scalar-promote-aa-tags.ll
+58-01 files

LLVM/project 9c2a881 — llvm/test/Transforms/LICM scalar-promote-aa-tags.ll

[NFC][LICM] Precommit mishandled per-iteration scoped alias metadata
DeltaFile
+117-0llvm/test/Transforms/LICM/scalar-promote-aa-tags.ll
+117-01 files

LLVM/project 583eec6 — llvm/test/Transforms/LICM scalar-promote-aa-tags.ll

Add pre-commit test mixing scoped alias info and tbaa
DeltaFile
+52-0llvm/test/Transforms/LICM/scalar-promote-aa-tags.ll
+52-01 files

LLVM/project 64248e4 — offload/include PluginManager.h, offload/liboffload/src OffloadImpl.cpp

[offload][omp] Load plugins through liboffload
DeltaFile
+15-11offload/libompaccsupport/PluginManager.cpp
+6-0offload/liboffload/src/OffloadImpl.cpp
+2-1offload/include/PluginManager.h
+23-123 files

LLVM/project 0cf09fc — llvm/lib/CodeGen TwoAddressInstructionPass.cpp, llvm/test/CodeGen/X86 twoaddr-reschedule-copy-chain.mir

TwoAddressInstructions: Move the rescheduled copy chain back to front (#227289)

rescheduleMIBelowKill sinks an instruction below the kill of its tied
source, along with the run of copies that follows it. With LiveIntervals
the copies are spliced one at a time so handleMove sees a well-formed
block, but they were visited front to back and each inserted before the
previously moved one. This reversed them, and transiently moved a copy
below its use, asserting in handleMoveDown.

Walk them back to front instead, which also preserves the original
order.

Exposed by #225174, which made LiveIntervals available here by default.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+36-0llvm/test/CodeGen/X86/twoaddr-reschedule-copy-chain.mir
+3-2llvm/lib/CodeGen/TwoAddressInstructionPass.cpp
+39-22 files

LLVM/project 24ab874 — offload/liboffload/API Platform.td, offload/liboffload/src OffloadImpl.cpp

[OFFLOAD] add olIteratePlatforms
DeltaFile
+45-0offload/unittests/OffloadAPI/platform/olIteratePlatforms.cpp
+23-0offload/liboffload/API/Platform.td
+11-0offload/liboffload/src/OffloadImpl.cpp
+79-03 files

LLVM/project fb1a7d0 — llvm/test/Transforms/LoopVectorize pr59319-loop-access-info-invalidation.ll select-last-index-fp.ll, llvm/test/Transforms/LoopVectorize/ARM tail-folding-scalar-epilogue-fallback.ll

[LV] Regenerate some CHECK lines (#226468)
DeltaFile
+298-298llvm/test/Transforms/LoopVectorize/X86/induction-costs.ll
+170-170llvm/test/Transforms/LoopVectorize/optimal-epilog-vectorization-liveout.ll
+147-147llvm/test/Transforms/LoopVectorize/select-last-index-fp.ll
+109-109llvm/test/Transforms/PhaseOrdering/AArch64/hoist-runtime-checks.ll
+64-63llvm/test/Transforms/LoopVectorize/pr59319-loop-access-info-invalidation.ll
+32-31llvm/test/Transforms/LoopVectorize/ARM/tail-folding-scalar-epilogue-fallback.ll
+820-8186 files

LLVM/project 236e0ee — llvm/lib/Transforms/Scalar SROA.cpp, llvm/test/Transforms/SROA load-store-overlap.ll

[SROA] Remove splitSliceTails loop in presplitLoadsAndStores (#227061)

The asserts in this loop can trigger when we presplit an overlapping
load/store pair, and the extra splits it adds don't appear to have any
benefit: we only get them when we have an unsplittable slice that's
fully enclosed inside a splittable slice, and in the test I've managed
to create for this (no_move_enclosed_unsplittable_store, and there are
no existing tests for this) the final SROA output is the same (except
that the variables end up with different names).
DeltaFile
+198-0llvm/test/Transforms/SROA/load-store-overlap.ll
+0-22llvm/lib/Transforms/Scalar/SROA.cpp
+198-222 files

LLVM/project 13b63a5 — llvm/test/CodeGen/Hexagon/live-vars live-outs.ll

[Hexagon] Add missing -mtriple to live-outs test (#227079)

This test's RUN command did not specify a target triple, so lit could
execute it using the host target.

This causes the test to fail when LLVM is built with only the Hexagon
target enabled. Add -mtriple=hexagon so the test consistently runs as a
Hexagon test.
DeltaFile
+1-1llvm/test/CodeGen/Hexagon/live-vars/live-outs.ll
+1-11 files

LLVM/project 85406f7 — flang/test/Lower do-loop-infinite-body-cycle.f90

[flang][Test] Cover the lowering of loops with a non-terminating body

A previous change leaves such a loop unstructured. Check what that produces:
the cycle survives as a block branching to itself, no fir.do_loop is emitted
for the loop control, and a loop that only needs a block of its own still
gets the structured form with its body in a region.
DeltaFile
+50-28flang/test/Lower/do-loop-infinite-body-cycle.f90
+50-281 files

LLVM/project 88a09a6 — flang/lib/Lower Bridge.cpp, flang/test/Lower do-loop-branch-to-loop-header.f90 do_loop_unstructured.f90

[flang] Lower loops whose branching is confined to their body structurally

Such a loop was classified separately by a previous change but still
lowered as unstructured, so its structured form was lost.

Lower it structurally instead, with its body folded into a region that can
hold the branching. The loop keeps its bounds on the op, so it remains
available to whatever transforms or parallelizes it. Only the body is
folded: the loop control statements are emitted as they are for any
structured loop, since a branch from outside may target either of them.

Loops an OpenACC or OpenMP directive owns are lowered the same way, so
they keep their form too.
DeltaFile
+20-134flang/test/Lower/do_loop_unstructured.f90
+119-11flang/lib/Lower/Bridge.cpp
+73-0flang/test/Lower/OpenACC/acc-unstructured-internals.f90
+34-28flang/test/Lower/OpenMP/wsloop-unstructured-cycle.f90
+55-0flang/test/Lower/do-loop-branch-to-loop-header.f90
+50-0flang/test/Lower/OpenMP/metadirective-loop-unstructured.f90
+351-17310 files not shown
+453-24116 files

LLVM/project 7fb36ba — flang/include/flang/Lower PFTBuilder.h, flang/lib/Lower PFTBuilder.cpp

[flang] Detect loops whose branching is confined to their body (#225757)

A DO loop is classified as either structured or unstructured, and a
single raw branch anywhere in its body forces the loop -- and every
construct enclosing it -- onto the unstructured path.

That is stronger than necessary. A loop keeps its structured control
flow as long as its branching neither leaves its body nor enters it from
outside. Classify such a loop separately from a fully unstructured one.

This only classifies: lowering is unchanged. PFT dumps mark the new
classification with '~', which is what the tests key on.
DeltaFile
+305-34flang/lib/Lower/PFTBuilder.cpp
+170-0flang/test/Lower/do-loop-infinite-body-cycle.f90
+137-0flang/test/Lower/pre-fir-tree-unstructured-internals.f90
+51-6flang/include/flang/Lower/PFTBuilder.h
+47-0flang/test/Lower/pre-fir-tree-assigned-goto.f90
+43-0flang/test/Lower/pre-fir-tree-branch-into-body.f90
+753-401 files not shown
+762-467 files

LLVM/project 0b3eb21 — llvm/lib/Transforms/InstCombine InstCombineCalls.cpp, llvm/lib/Transforms/Utils SimplifyLibCalls.cpp

[InstCombine] Don't fold `sin(-x)` to `-sin(x)` if `denormals` may flush to `+0.0` (#227039)
DeltaFile
+122-0llvm/test/Transforms/InstCombine/cos-1.ll
+56-0llvm/test/Transforms/InstCombine/cos-sin-intrinsic.ll
+9-1llvm/lib/Transforms/Utils/SimplifyLibCalls.cpp
+9-1llvm/lib/Transforms/InstCombine/InstCombineCalls.cpp
+196-24 files

LLVM/project 14b0128 — mlir/lib/Dialect/OpenACC/Transforms ACCCGToGPU.cpp, mlir/test/Dialect/OpenACC acc-cg-to-gpu-predicate-region-reuse-barrier.mlir

[mlir][OpenACC] Match reuse barriers to the reused private scope (#225570)

A region that writes both gang- and worker-private slots was considered
to be a gang scope and a barrier was missed. This change keeps the two
store sets separate and pick the barrier from the scope that is actually
reused. Follow-up for
https://github.com/llvm/llvm-project/pull/224437#discussion_r4066418696.
DeltaFile
+60-37mlir/lib/Dialect/OpenACC/Transforms/ACCCGToGPU.cpp
+53-0mlir/test/Dialect/OpenACC/acc-cg-to-gpu-predicate-region-reuse-barrier.mlir
+113-372 files

LLVM/project 83cb885 — llvm/lib/Target/SPIRV SPIRVNonSemanticDebugHandler.h SPIRVNonSemanticDebugHandler.cpp, llvm/test/CodeGen/SPIRV/debug-info debug-type-pointer-to-composite-drop.ll debug-typedef-nested-drop.ll

Initial work.
DeltaFile
+316-278llvm/lib/Target/SPIRV/SPIRVNonSemanticDebugHandler.cpp
+162-79llvm/lib/Target/SPIRV/SPIRVNonSemanticDebugHandler.h
+14-12llvm/test/CodeGen/SPIRV/debug-info/debug-type-composite-nested-drop.ll
+10-15llvm/test/CodeGen/SPIRV/debug-info/debug-lexical-block-namespace-in-block.ll
+11-8llvm/test/CodeGen/SPIRV/debug-info/debug-typedef-nested-drop.ll
+9-9llvm/test/CodeGen/SPIRV/debug-info/debug-type-pointer-to-composite-drop.ll
+522-4016 files not shown
+563-42512 files

LLVM/project 2bf008c — clang/docs ReleaseNotes.md, clang/lib/CodeGen CGStmtOpenMP.cpp

[clang][OpenMP] Fix crashes on target regions inside namespace-scope lambdas and blocks (#226691)

Fixes #223397

A `target` region inside a lambda or block at namespace scope crashed
clang in two places. In Sema, `isOpenMPCapturedDecl` decides whether a
global must be captured by walking the function scope stack down to the
innermost OpenMP captured region, stopping at an ordinary function
scope. The capture initializers of a directive's outermost region are
built after all of its regions have been popped. Inside a function that
walk ends at the function's scope, but a namespace-scope lambda or block
has nothing underneath it, so the walk ran off the stack and asserted.
This happens for any global reference, such as `int &r = x; auto l = []
{ #pragma omp target r = 1; };`. The self-referential declaration in the
report is incidental. Once past Sema, CodeGen asserted too: it names the
outlined kernel after the region's parent function, and such a region
has none, even without a reference.

In Sema, running out of scopes now means the same as reaching a function

    [3 lines not shown]
DeltaFile
+67-0clang/test/OpenMP/target_global_ref_namespace_scope_lambda_codegen.cpp
+55-0clang/test/OpenMP/target_global_ref_namespace_scope_lambda.cpp
+5-3clang/lib/CodeGen/CGStmtOpenMP.cpp
+3-1clang/lib/Sema/SemaOpenMP.cpp
+1-0clang/docs/ReleaseNotes.md
+131-45 files

LLVM/project b7cc7ee — offload/test/offloading shared_lib_global_var.c

[offload][omp] Add NVIDIA LTO to unsupported list in shlib global var test (#227315)
DeltaFile
+1-0offload/test/offloading/shared_lib_global_var.c
+1-01 files

LLVM/project 7e19707 — mlir/include/mlir/Support InterfaceSupport.h, mlir/test/mlir-tblgen op-interface.td

[mlir][ODS] Copy constant generated interface model prototypes (#226323)

Generate constexpr constructors for interface models whose concept is a
table of callbacks. Construct exact generated models from a static
prototype; keep external and fallback models on their existing path.

In OpenMPDialect.cpp this saves about 0.22B compiler instructions and
10.5KB of text.

Assisted-by: Codex
DeltaFile
+25-3mlir/include/mlir/Support/InterfaceSupport.h
+6-1mlir/tools/mlir-tblgen/OpInterfacesGen.cpp
+4-1mlir/test/mlir-tblgen/op-interface.td
+35-53 files