LLVM/project 91b08a5 — llvm/include/llvm/CodeGen MachinePipeliner.h ScheduleDAGInstrs.h, llvm/lib/CodeGen ScheduleDAGInstrs.cpp

[MISched](NFC) Factor control dependency construction state and logic out of `buildSchedGraph` (#226533)

Continuing the groundwork for a proper implementation of #205689, factor
the control dependency construction state and logic out of
`ScheduleDAGInstrs` into a friend class. This will also enable us to
further break up `buildSchedGraph` to make, e.g., the barrier chain
handling easier to comprehend.
DeltaFile
+192-128llvm/lib/CodeGen/ScheduleDAGInstrs.cpp
+3-58llvm/include/llvm/CodeGen/ScheduleDAGInstrs.h
+1-0llvm/include/llvm/CodeGen/MachinePipeliner.h
+196-1863 files

LLVM/project 05a75e8 — llvm/include/llvm/ADT DenseMap.h

[ADT] Move DenseMapIterator above DenseMapBase (NFC) (#227208)

This patch moves the definition of DenseMapIterator right above
DenseMapBase so that all classes in DenseMap.h are defined in the
bottom-up order without forward declarations:

- DenseMapStorage
- SmallDenseMapStorage
- DenseMapIterator
- DenseMapBase
- DenseMap
- SmallDenseMap

Assisted-by: Antigravity
DeltaFile
+141-145llvm/include/llvm/ADT/DenseMap.h
+141-1451 files

LLVM/project 3272c28 — .github/workflows/test-suite aarch64.cmake

[GitHub] Fix test-suite.yml FFmpeg build with AArch64 (#227346)

On AArch64 there are assembly files that need CMAKE_ASM_COMPILER_TARGET
and CMAKE_ASM_FLAGS_INIT set
DeltaFile
+2-0.github/workflows/test-suite/aarch64.cmake
+2-01 files

LLVM/project 4840749 — llvm/docs GettingInvolved.md

docs: Remove Johannes' office hours
DeltaFile
+0-6llvm/docs/GettingInvolved.md
+0-61 files

LLVM/project e3c7196 — llvm/include/llvm/CodeGen TargetLowering.h, llvm/lib/Target/AMDGPU SIISelLowering.cpp

[AMDGPU] Don't apply gfx950 fetch-window loop align to the wrong block (#221821)

MachineBlockPlacement aligns the backedge destination after loop
rotation, which may not be the LoopInfo header. Query that block for the
32-byte request and the 4-byte pad cap so they stay consistent.
## Summary
- Pass the block being aligned into `getPrefLoopAlignment` so gfx950
fetch-window alignment uses the backedge destination, not only the
LoopInfo header.
- Unrotated 8-byte headers still get capped.
- Rotated 4-byte landing pads are no longer given uncapped 32-byte
alignment.
DeltaFile
+274-15llvm/test/CodeGen/AMDGPU/loop-header-align-gfx950.mir
+122-0llvm/test/CodeGen/AMDGPU/loop-header-align-gfx950.ll
+14-9llvm/lib/Target/AMDGPU/SIISelLowering.cpp
+8-2llvm/include/llvm/CodeGen/TargetLowering.h
+3-1llvm/lib/Target/X86/X86ISelLowering.h
+3-1llvm/lib/Target/PowerPC/PPCISelLowering.h
+424-285 files not shown
+436-3311 files

LLVM/project 8ce1ee4 — mlir/lib/Dialect/Arith/Utils Utils.cpp, mlir/unittests/Dialect CMakeLists.txt

[mlir][arith] Support equal-width float conversions (#225346)

Use arith.convertf in convertScalarToDtype when converting between
floating-point types that have the same bit width but different
semantics.

Add focused unit coverage for f16-to-bf16 and bf16-to-f16 conversions.

Assisted-by: Codex
DeltaFile
+50-0mlir/unittests/Dialect/Arith/ArithUtilsTest.cpp
+7-0mlir/unittests/Dialect/Arith/CMakeLists.txt
+3-2mlir/lib/Dialect/Arith/Utils/Utils.cpp
+1-0mlir/unittests/Dialect/CMakeLists.txt
+61-24 files

LLVM/project e90c043 — compiler-rt/lib/ubsan/offload ubsan_offload_hsa_interceptors.cpp

[compiler-rt] Remove dlsym interceptor for HSA UBSan (#227022)

Summary:
This existed to handle the OpenMP case that dynamically opened via
`dlsym`. However, https://github.com/llvm/llvm-project/pull/227020
removes the need for this by first checking the global space first.

The main motivation is that the `dlsym` interceptor layer was the most
janky part of this whole affair and was a major blocker to supporting
`-shared-libsan` in https://github.com/llvm/llvm-project/pull/226551.
DeltaFile
+0-74compiler-rt/lib/ubsan/offload/ubsan_offload_hsa_interceptors.cpp
+0-741 files

LLVM/project 046e55b — libcxx/test/std/containers/container.adaptors push_range_container_adaptors.h, libcxx/test/std/containers/container.adaptors/priority.queue/priqueue.members push_range.pass.cpp

[libc++] Add coverage for push_range on stack/queue with non-default underlying containers (#210730)

This increases the test coverage for stack/queue and actually caught an
issue in list::__invariants which was previously dead code.
DeltaFile
+7-7libcxx/test/std/containers/container.adaptors/push_range_container_adaptors.h
+5-0libcxx/test/std/containers/container.adaptors/stack/stack.defn/push_range.pass.cpp
+5-0libcxx/test/std/containers/container.adaptors/stack/stack.cons/from_range.pass.cpp
+2-2libcxx/test/std/containers/container.adaptors/priority.queue/priqueue.members/push_range.pass.cpp
+3-0libcxx/test/std/containers/container.adaptors/queue/queue.defn/push_range.pass.cpp
+3-0libcxx/test/std/containers/container.adaptors/queue/queue.cons/from_range.pass.cpp
+25-91 files not shown
+26-107 files

LLVM/project bbc6e8f — lldb/source/Plugins/Process/Linux NativeRegisterContextLinux_arm64.cpp

[lldb][AArch64][Linux] Add llvm_unreachable after some RegisterSetType switches (#223375)

The ones where you are supposed to case X: return Y;. We do enable the
not fully covered switch warning, so the unreachable just makes the
mistake more obvious.

I did not change GetInvalidationMask because this will be refactored by
#223373.
DeltaFile
+6-0lldb/source/Plugins/Process/Linux/NativeRegisterContextLinux_arm64.cpp
+6-01 files

LLVM/project 36422ef — llvm/lib/IR AutoUpgrade.cpp, llvm/unittests/Bitcode DataLayoutUpgradeTest.cpp

[llvm] Upgrade ARM data layouts that are missing Fi8 (#224639)

Such as the ones in
https://github.com/llvm/llvm-test-suite/tree/main/Bitcode/simd_ops,
which started failing to compile after
https://github.com/llvm/llvm-project/pull/224012
stopped us overriding the module data layout.

fatal error: error in backend: Can't create a MachineFunction using a
Module with a Target-incompatible DataLayout attached
  Target DataLayout: e-m:e-p:32:32-Fi8-i64:64-v128:64:128-a:0:32-n32-S64
  Module DataLayout: e-m:e-p:32:32-i64:64-v128:64:128-a:0:32-n32-S64

The old layout is p32:32-i64, the new is p32:32-Fi8-i64.

In this change I have added this case to the data layout upgrades. If
there's no p32:32, the layout is not changed, also if there is already a
"Fi" or "Fn" in the layout. These cases may not exist or be valid in
reality, but I need some fallback behaviour just in case.
DeltaFile
+25-0llvm/unittests/Bitcode/DataLayoutUpgradeTest.cpp
+9-0llvm/lib/IR/AutoUpgrade.cpp
+34-02 files

LLVM/project 2c9a8d4 — llvm/lib/Target/SystemZ SystemZXPLINKAsmPrinter.cpp, llvm/test/CodeGen/SystemZ zos-lower-constant.ll

[SystemZ][z/OS] Use a function descriptor for external functions in initializers (#226682)

The address of an external function in a static initializer was emitted
as a V-con, i.e. the entry point, while an XPLINK function pointer must
point to a function descriptor. A call through such a pointer loads the
environment and entry point from the machine code of the function.
Create a descriptor in the ADA for external functions as well, as is
already done for internal functions and for constructor/destructor lists
(`emitXXStructorList`).

`zos-lower-constant.ll` is updated: the pointer to the external function
`bar` now points to its descriptor in the ADA.

A minimal program (`int (*p)(void) = ext;` called from `main`) ended
with S0C6 at address `0707070707070707` on z/OS 3.1; with this change it
returns the expected value.

Tests: `llvm-lit test/CodeGen/SystemZ test/MC/SystemZ test/MC/GOFF`
passes (1303 passed, 19 unsupported).

    [5 lines not shown]
DeltaFile
+16-4llvm/lib/Target/SystemZ/SystemZXPLINKAsmPrinter.cpp
+15-3llvm/test/CodeGen/SystemZ/zos-lower-constant.ll
+31-72 files

LLVM/project b0fb935 — llvm/test/CodeGen/X86 vector-interleaved-load-i8-stride-4.ll vector-interleaved-store-i16-stride-8.ll

[X86] LowerStore - peek through oneuse bitcasts to see if vector is freely splittable. (#227282)

Peek through bitcasts when seeing if we can avoid an unnecessary
concat_vectors and store the subvectors directly - if the concat had
been worth it, combineConcatVectorOps would have moved the concat
further up.

Helps avoid more cases of unnecessary vzeroupper, 256-bit usage etc. in
particular on AVX1 targets (Sandybridge, Jaguar and Bulldozer all
benefit from this), but AVX2/512 targets as well.
DeltaFile
+864-863llvm/test/CodeGen/X86/vector-interleaved-load-i16-stride-3.ll
+698-688llvm/test/CodeGen/X86/vector-interleaved-load-i16-stride-4.ll
+422-424llvm/test/CodeGen/X86/vector-interleaved-store-i16-stride-6.ll
+292-291llvm/test/CodeGen/X86/frem.ll
+273-274llvm/test/CodeGen/X86/vector-interleaved-store-i16-stride-8.ll
+263-255llvm/test/CodeGen/X86/vector-interleaved-load-i8-stride-4.ll
+2,812-2,79517 files not shown
+3,551-3,55723 files

LLVM/project b690595 — clang/lib/CodeGen/TargetBuiltins RISCV.cpp, cross-project-tests/intrinsic-header-tests riscv_packed_simd.c

[RISCV][P-ext] Add packed multiply high parts accumulate intrinsics (#224261)

Add intrinsics, Clang builtins and SelectionDAG support for the RISC-V P
multiply high accumulate with byte/halfword index operations.
See
https://github.com/riscv/riscv-p-spec/blob/master/P-ext-intrinsics.adoc#packed-multiply-high-parts-accumulate.

The pmhacc.h.bXX forms operate on the whole register and select directly
on both RV32 and RV64.
The pmhacc.w.hXX forms select directly on RV64, on RV32 there is no
64-bit packed form, so the intrinsics split into a pair of scalar
mhacc.h0/mhacc.h1 accumulations, one per result word.
DeltaFile
+130-0llvm/test/CodeGen/RISCV/rvp-simd-64.ll
+98-0llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+84-0llvm/lib/Target/RISCV/RISCVInstrInfoP.td
+83-0cross-project-tests/intrinsic-header-tests/riscv_packed_simd.c
+45-0llvm/include/llvm/IR/IntrinsicsRISCV.td
+41-0clang/lib/CodeGen/TargetBuiltins/RISCV.cpp
+481-03 files not shown
+550-09 files

LLVM/project a2b6f77 — llvm/lib/Transforms/Vectorize SLPVectorizer.cpp, llvm/test/Transforms/SLPVectorizer/X86 runtime-alias-checks-alloca.ll

[SLP]Fix alloca handling in runtime alias check versioning (#227341)

Do not duplicate allocas when versioning a block: keep the leading
static allocas in the header block and reject blocks with any other
alloca. Duplicating them moved static allocas out of the entry block
and merged the clones through a PHI, which is invalid for lifetime
markers.

Fixes #227328
DeltaFile
+238-0llvm/test/Transforms/SLPVectorizer/X86/runtime-alias-checks-alloca.ll
+10-3llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+248-32 files

LLVM/project debb718 — offload/liboffload exports

restore export file
DeltaFile
+0-56offload/liboffload/exports
+0-561 files

LLVM/project e098def — offload/liboffload exports

remove two export lines
DeltaFile
+56-0offload/liboffload/exports
+56-01 files

LLVM/project d0f5882 — offload/include device.h PluginManager.h, offload/libompaccsupport device.cpp PluginManager.cpp

[offload][omp] Use olIterateCompatibleDevices for device init and image registration
DeltaFile
+116-100offload/libompaccsupport/PluginManager.cpp
+9-8offload/libompaccsupport/device.cpp
+0-15offload/plugins-nextgen/common/src/PluginInterface.cpp
+5-8offload/include/PluginManager.h
+0-6offload/plugins-nextgen/common/include/PluginInterface.h
+4-1offload/include/device.h
+134-1381 files not shown
+138-1387 files

LLVM/project 7dadbf7 — libcxx/test/benchmarks monotonic_buffer.bench.cpp

[libc++] Fix use-after-free in the monotonic_buffer benchmark (#227072)

We would release the monotonic_buffer_resource before the closing brace,
which runs the destructor of the list and accesses the nodes.
DeltaFile
+6-4libcxx/test/benchmarks/monotonic_buffer.bench.cpp
+6-41 files

LLVM/project 21b3f12 — offload/liboffload/API Program.td, offload/liboffload/src OffloadImpl.cpp

[OFFLOAD] Add olIterateCompatibleDevices API
DeltaFile
+73-0offload/unittests/OffloadAPI/program/olIterateCompatibleDevices.cpp
+25-2offload/liboffload/src/OffloadImpl.cpp
+15-0offload/liboffload/API/Program.td
+1-0offload/unittests/OffloadAPI/CMakeLists.txt
+114-24 files

LLVM/project 3b8bb27 — libcxx/test/benchmarks/containers/associative associative_container_benchmarks.h

[libc++] Use a union for uninitialized storage in associative container benchmarks (#227068)

This removes the need for reinterpret_cast when accessing the containers
constructed in the scratch space, and fixes the PMR constructor
benchmark passing the wrong pointer to DoNotOptimize.
DeltaFile
+29-24libcxx/test/benchmarks/containers/associative/associative_container_benchmarks.h
+29-241 files

LLVM/project 175d12d — llvm/test/CodeGen/X86 twoaddr-reschedule-copy-chain.mir

TwoAddressInstructions: Add another reschedule copy order test

Another test for the fix from #227289, which hit a different
assert condition.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+54-0llvm/test/CodeGen/X86/twoaddr-reschedule-copy-chain.mir
+54-01 files

LLVM/project 6183ad5 — offload/include PluginManager.h, offload/liboffload/src OffloadImpl.cpp

[offload][omp] Load plugins through liboffload
DeltaFile
+18-20offload/libompaccsupport/PluginManager.cpp
+6-0offload/liboffload/src/OffloadImpl.cpp
+2-1offload/include/PluginManager.h
+26-213 files

LLVM/project 3f4fd64 — offload/liboffload/API Platform.td, offload/liboffload/src OffloadImpl.cpp

[OFFLOAD] add olIteratePlatforms
DeltaFile
+44-0offload/unittests/OffloadAPI/platform/olIteratePlatforms.cpp
+23-0offload/liboffload/API/Platform.td
+10-0offload/liboffload/src/OffloadImpl.cpp
+77-03 files

LLVM/project b73d8ab — offload/include PluginManager.h, offload/liboffload/src OffloadImpl.cpp

[offload][omp] Load plugins through liboffload
DeltaFile
+18-20offload/libompaccsupport/PluginManager.cpp
+6-0offload/liboffload/src/OffloadImpl.cpp
+2-1offload/include/PluginManager.h
+26-213 files

LLVM/project 9796f2a — offload/liboffload/API Platform.td, offload/liboffload/src OffloadImpl.cpp

[OFFLOAD] add olIteratePlatforms
DeltaFile
+44-0offload/unittests/OffloadAPI/platform/olIteratePlatforms.cpp
+23-0offload/liboffload/API/Platform.td
+12-2offload/liboffload/src/OffloadImpl.cpp
+79-23 files

LLVM/project 584b74e — llvm/lib/Target/X86 X86ExpandPseudo.cpp X86ISelLowering.cpp, llvm/test/CodeGen/X86 expand-cmpxchg16b-dead-eflags.mir cmpxchg16b-dead-eflags.mir

X86: Preserve dead flags on EFLAGS when lowering CMPXCHG16B pseudos

Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+104-0llvm/test/CodeGen/X86/cmpxchg16b-dead-eflags.mir
+40-0llvm/test/CodeGen/X86/expand-cmpxchg16b-dead-eflags.mir
+3-0llvm/lib/Target/X86/X86ISelLowering.cpp
+1-0llvm/lib/Target/X86/X86ExpandPseudo.cpp
+148-04 files

LLVM/project 907a809 — libcxx/test/benchmarks/containers/associative associative_container_benchmarks.h

[libc++] Fix out-of-bounds read in the associative container query benchmarks (#227070)

The query benchmarks for associative containers would access the pool of
keys to use in the benchmark out-of-bounds.
DeltaFile
+1-1libcxx/test/benchmarks/containers/associative/associative_container_benchmarks.h
+1-11 files

LLVM/project f51e3f2 — flang/lib/Lower Bridge.cpp

[flang][NFC] Split the OpenACC construct lowering into two lanes

genFIR(OpenACCConstruct) decided twice, in three places, whether the
construct it lowers is structured, and reassigned the evaluation it works
from halfway through: before the descent that evaluation is the construct,
after it the loop the directive absorbs. Everything downstream had to know
which one it was holding.

Give each form its own function and leave genFIR to choose between them.
One lane allocates the exit selector, lowers the evaluations the construct
holds, and emits the jump table; the other reads the collapse clauses,
descends to the absorbed depth, and lowers what is inside it. The prologue
and epilogue are short enough to state in both rather than share.
DeltaFile
+149-101flang/lib/Lower/Bridge.cpp
+149-1011 files

LLVM/project d2d3041 — flang/lib/Lower PFTBuilder.cpp, flang/test/Lower/OpenACC acc-unstructured-combined-construct.f90 acc-directive-loop-bounds.f90

[flang] Let a directive keep the loop it owns when its body branches

A loop whose branching is confined to its body keeps its structured form,
but the construct holding it stayed unstructured. A directive does not
merely contain such a loop, it owns it, and its lowering reads the
construct's own classification to decide whether the loop op carries its
bounds. The directive was left with a bounds-free loop that nothing could
partition, and the loop it owns became a second one nested inside.

Reclassify a directive construct once the loops it holds no longer need it
to stay unstructured, and fold the body of the loop it takes over into a
region, which the DO lowering can no longer do for it.

A construct whose branching leaves it is untouched, as is one holding a
branch of its own.
DeltaFile
+8-124flang/test/Lower/OpenACC/Todo/acc-unstructured-loop-construct.f90
+120-0flang/test/Lower/OpenACC/acc-unstructured-loop-construct.f90
+66-0flang/test/Lower/OpenACC/acc-directive-loop-bounds.f90
+40-0flang/test/Lower/OpenACC/acc-unstructured-combined-construct.f90
+33-4flang/lib/Lower/PFTBuilder.cpp
+34-0flang/test/Lower/OpenMP/wsloop-directive-loop-bounds.f90
+301-1284 files not shown
+333-16810 files

LLVM/project 99ac64c — offload/liboffload/API Program.td, offload/liboffload/src OffloadImpl.cpp

[OFFLOAD] Add olIterateCompatibleDevices API
DeltaFile
+73-0offload/unittests/OffloadAPI/program/olIterateCompatibleDevices.cpp
+25-2offload/liboffload/src/OffloadImpl.cpp
+15-0offload/liboffload/API/Program.td
+1-2offload/unittests/OffloadAPI/platform/olIteratePlatforms.cpp
+1-0offload/unittests/OffloadAPI/CMakeLists.txt
+115-45 files