LLVM/project da3277f — llvm/docs AMDGPUUsage.rst, llvm/test/CodeGen/AMDGPU frem.ll

Rebase, address comments

Created using spr 1.3.7
DeltaFile
+1,609-1,040llvm/test/CodeGen/AMDGPU/frem.ll
+986-1,144llvm/test/CodeGen/AMDGPU/GlobalISel/sdiv.i64.ll
+914-1,072llvm/test/CodeGen/AMDGPU/GlobalISel/srem.i64.ll
+608-584llvm/test/CodeGen/NVPTX/arbitrary-fp-to-float.ll
+581-521llvm/test/CodeGen/X86/vector-compress.ll
+448-433llvm/docs/AMDGPUUsage.rst
+5,146-4,794652 files not shown
+24,603-14,582658 files

LLVM/project 2037f0e —

[SLP][NFC]Add an extra test with the reodering of the aggregate bitpacks, NFC



Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/229380
DeltaFile
+0-00 files

LLVM/project c4c287a —

[SLP][NFC]Add an extra test with the reodering of the aggregate bitpacks, NFC



Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/229380
DeltaFile
+0-00 files

LLVM/project 57f2572 — mlir/lib/Conversion/TosaToLinalg TosaToLinalg.cpp, mlir/test/Conversion/TosaToLinalg tosa-to-linalg-invalid.mlir tosa-to-linalg.mlir

[mlir][tosa] Lower ROW_GATHER to Linalg (#225417)

Lower ROW_GATHER to a linalg.generic that maps each expanded output row
back to its source index slot and consecutive row offset. Support
dynamic output dimensions and both i32 and i64 indices.

Assisted-by: Codex
DeltaFile
+71-0mlir/test/Conversion/TosaToLinalg/tosa-to-linalg.mlir
+70-0mlir/lib/Conversion/TosaToLinalg/TosaToLinalg.cpp
+9-0mlir/test/Conversion/TosaToLinalg/tosa-to-linalg-invalid.mlir
+150-03 files

LLVM/project 9a39608 — llvm/test/Transforms/SLPVectorizer/X86 bitpacked-aggregate.ll

[SLP][NFC]Add an extra test with the reodering of the aggregate bitpacks, NFC



Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/229380
DeltaFile
+54-0llvm/test/Transforms/SLPVectorizer/X86/bitpacked-aggregate.ll
+54-01 files

LLVM/project f30e753 — bolt/test high-segments.s, bolt/test/X86 high-segments.s

[BOLT][AArch64] Port X86 --custom-allocation-vma test to AArch64 (#229104)

**Before**: #136385 introduced the `--custom-allocation-vma` flag for
BOLT to be able to specify a suitable location to place rewritten
binary. This was accompanied by a corresponding test for `x86`.

**After**: This PR ports the existing test `high-segments.s` to AArch64,
improving test coverage.

Assisted-by: Codex
DeltaFile
+0-46bolt/test/X86/high-segments.s
+33-0bolt/test/high-segments.s
+33-462 files

LLVM/project 4e75021 — llvm/lib/CodeGen MachineOutliner.cpp, llvm/test/CodeGen/AArch64 machine-outliner-bundle.mir machine-outliner-operand-flags.mir

Process instructions inside bundles
DeltaFile
+47-3llvm/test/CodeGen/AArch64/machine-outliner-operand-flags.mir
+4-2llvm/lib/CodeGen/MachineOutliner.cpp
+1-1llvm/test/CodeGen/AArch64/machine-outliner-bundle.mir
+52-63 files

LLVM/project 1b825bd — llvm/test/CodeGen/AArch64 machine-outliner-operand-flags.mir

Fix comments
DeltaFile
+3-3llvm/test/CodeGen/AArch64/machine-outliner-operand-flags.mir
+3-31 files

LLVM/project 1c49151 — llvm/test/CodeGen/ARM machine-outliner-stack-fixup-thumb.mir

Update CodeGen/ARM/machine-outliner-stack-fixup-thumb.mir
DeltaFile
+1-1llvm/test/CodeGen/ARM/machine-outliner-stack-fixup-thumb.mir
+1-11 files

LLVM/project b6be3ae — llvm/lib/CodeGen MachineOutliner.cpp, llvm/test/CodeGen/AArch64 machine-outliner-operand-flags.mir

[AArch64][PAC] Reset `killed` operand flags in outlined functions

Presently, MachineOutliner does not take `killed` operand flags into
account when merging instruction sequences. While it sounds perfectly
reasonable not to inhibit merging of the instruction sequences that
only differ in `killed` flags (for N flags there is technically 2^N
valid ways to drop some subset of them), copying these flags from
an arbitrarily chosen representative instruction may result in
incorrect codegen of PAuth-related pseudo instructions on AArch64.

To keep `killed` flags conservatively correct as if `OUTLINED_FUNCTION`s
are virtually re-inserted at every call site, this patch takes the
simplest approach of resetting every `killed` flag inside the
outlined functions.
DeltaFile
+82-0llvm/test/CodeGen/AArch64/machine-outliner-operand-flags.mir
+1-0llvm/lib/CodeGen/MachineOutliner.cpp
+83-02 files

LLVM/project 22ddf4a — llvm/test/Transforms/SLPVectorizer/X86 bitpacked-aggregate.ll

[𝘀𝗽𝗿] initial version

Created using spr 1.3.7
DeltaFile
+54-0llvm/test/Transforms/SLPVectorizer/X86/bitpacked-aggregate.ll
+54-01 files

LLVM/project 0c53d0f — llvm/lib/Target/X86 X86ISelLowering.cpp, llvm/test/CodeGen/X86 avx512fp16-combine-fmsubadd.ll

[X86] Require contract on FMUL while folding FMADDSUB to VFMULC (#229309)

FMADDSUB/FMSUBADD has no FMF, so we check its only operand (third; FMUL) which still has it.
DeltaFile
+45-0llvm/test/CodeGen/X86/avx512fp16-combine-fmsubadd.ll
+1-1llvm/lib/Target/X86/X86ISelLowering.cpp
+46-12 files

LLVM/project 08483c5 — llvm/lib/Target/RISCV RISCVISelLowering.cpp

RISCV: Stop setting kill flags on virtual registers before FinalizeISel (#229013)

Kill flags have no remaining use before register allocation and are stripped by 
LiveIntervals.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+2-2llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+2-21 files

LLVM/project 4fadb2e — utils/bazel/llvm-project-overlay/mlir BUILD.bazel

[Bazel] Fixes fc90bd3 (#229344)

This fixes fc90bd3cfb0bab14a2bf7cbd33f5f1e904c7e133 (#228459).

Buildkite error link:
https://buildkite.com/llvm-project/upstream-bazel/builds?commit=fc90bd3cfb0bab14a2bf7cbd33f5f1e904c7e133

Co-authored-by: Google Bazel Bot <google-bazel-bot at google.com>
DeltaFile
+1-0utils/bazel/llvm-project-overlay/mlir/BUILD.bazel
+1-01 files

LLVM/project 94eb65e — clang/cmake/modules CMakeLists.txt ClangConfig.cmake.in, llvm/docs ReleaseNotes.md

[CMake] Make find_package(Clang) load MLIR when ClangIR uses a full MLIR (#228366)

With CLANG_ENABLE_CIR=ON, Clang's exported libraries link MLIR. When
MLIR is listed in LLVM_ENABLE_PROJECTS it installs its own package and
owns every target exported from mlir/. ClangTargets.cmake references
those targets without defining them, and ClangConfig.cmake never loaded
the MLIR package, so an external find_package(Clang) failed with:

```
  The following imported targets are referenced, but are missing:
  MLIRIR MLIRPass MLIRAnalysis ... MLIRSupport MLIRTransforms ...
```

unless the consumer happened to call find_package(MLIR) first.

When MLIR is enabled only implicitly as a ClangIR dependency
(LLVM_DEPENDENCY_ONLY_PROJECTS), no MLIR package exists and Clang
already promotes the reachable MLIR closure into its own exports, so
that mode worked. The two modes therefore need different plumbing but

    [131 lines not shown]
DeltaFile
+53-0clang/cmake/modules/ClangConfig.cmake.in
+36-0clang/cmake/modules/CMakeLists.txt
+18-4mlir/cmake/modules/CMakeLists.txt
+10-0llvm/docs/ReleaseNotes.md
+6-0mlir/cmake/modules/MLIRConfig.cmake.in
+123-45 files

LLVM/project 07e9854 — llvm/lib/Target/RISCV RISCVISelLowering.cpp

RISCV: Stop setting kill flags on virtual registers before FinalizeISel

Kill flags have no remaining use before register allocation and are
stripped by LiveIntervals.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+2-2llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+2-21 files

LLVM/project fab6b85 — llvm/lib/Target/AMDGPU AMDGPUCodeGenPrepare.cpp, llvm/test/CodeGen/AMDGPU amdgpu-codegenprepare-fmul-legacy.ll

[AMDGPU] Fold fmul with zero checks to fmul.legacy

Recognize f32 multiplication with zero checks in AMDGPUCodeGenPrepare,
including fabs/fneg wrappers.

Example:
```
  (x == 0 || y == 0) ? +0.0 : x * y
==>
  llvm.amdgcn.fmul.legacy(x, y)
```
DeltaFile
+21-107llvm/test/CodeGen/AMDGPU/amdgpu-codegenprepare-fmul-legacy.ll
+77-0llvm/lib/Target/AMDGPU/AMDGPUCodeGenPrepare.cpp
+98-1072 files

LLVM/project b6396d7 — llvm/test/CodeGen/AMDGPU amdgpu-codegenprepare-fmul-legacy.ll

[AMDGPU] Add fmul.legacy select fold tests. NFC

Add AMDGPUCodeGenPrepare tests for folding a select guarded by zero
checks into llvm.amdgcn.fmul.legacy, including negative tests and min
cases for future work.
DeltaFile
+1,030-0llvm/test/CodeGen/AMDGPU/amdgpu-codegenprepare-fmul-legacy.ll
+1,030-01 files

LLVM/project 92d7d31 — llvm/lib/Target/AArch64 AArch64Combine.td, llvm/lib/Target/AArch64/GISel AArch64PostLegalizerCombiner.cpp

[AArch64][SPIRV][GlobalISel] Migrate target-specific wip_match_opcode combines to MIR-pattern (#222903)

It replaces the `wip_match_opcode` with declarative MIR-pattern match
roots across the AArch64 and SPIRV target-specific GICombines.

Some important points to consider :

- **shuffle_vector_lowering:** these `G_SHUFFLE_VECTOR` rules are order-
sensitive. `fullrev` previously gained priority through a nested
`G_IMPLICIT_DEF` match; that predicate is relocated to C++ so all rules
are equal-priority and `fullrev` is ordered last, preserving the
original dispatch sequence and codegen.
- **SPIRV intrinsic roots:** the matrix/length/distance rules match and
apply on target intrinsics (`int_spv_*`, `int_matrix_*`), requiring the
`IntrinsicsSPIRV` enum in the combiner translation unit.
- **Deferred:** `vector_unmerge_lowering` and `unmerge_ext_to_unmerge`
remain on `wip_match_opcode`; their `G_UNMERGE_VALUES` variadic-def
roots are not yet expressible as MIR patterns.
DeltaFile
+247-410llvm/test/CodeGen/AArch64/neon-shuffle-vector-tbl.ll
+62-49llvm/lib/Target/AArch64/AArch64Combine.td
+0-63llvm/lib/Target/AArch64/GISel/AArch64PostLegalizerCombiner.cpp
+50-0llvm/test/CodeGen/AArch64/GlobalISel/postlegalizercombiner-mul-identity.mir
+0-33llvm/lib/Target/SPIRV/SPIRVCombinerHelper.cpp
+8-11llvm/lib/Target/SPIRV/SPIRVCombine.td
+367-5665 files not shown
+376-58811 files

LLVM/project 08652b2 — llvm/lib/Target/AArch64 AArch64TargetTransformInfo.h AArch64TargetTransformInfo.cpp, llvm/test/CodeGen/AArch64 zext-to-tbl.ll

[AArch64][LSR] Prefer pointer IVs for SVE accesses that need splitting (#228510)

Currently, LSR always prefers scaled accesses but that ends up doing
more work as only the first access can used the scaled register. Later
split accesses need to compute the base then another `mul vl` offset.

Assisted-by: Codex
DeltaFile
+187-3llvm/test/Transforms/LoopStrengthReduce/AArch64/vscale-fixups.ll
+38-0llvm/lib/Target/AArch64/AArch64TargetTransformInfo.cpp
+16-16llvm/test/CodeGen/AArch64/zext-to-tbl.ll
+5-0llvm/lib/Target/AArch64/AArch64TargetTransformInfo.h
+246-194 files

LLVM/project 54f957e — llvm/lib/Target/AArch64 AArch64InstrInfo.td, llvm/lib/Target/AArch64/AsmParser AArch64AsmParser.cpp

[AArch64][llvm] Armv9.8-A: Add support for FEAT_RLCS (release consistency scoping)

Add support for FEAT_RLCS (release consistency scoping) instructions:
  - SRLS
  - SLBND
DeltaFile
+33-0llvm/test/MC/AArch64/armv9.8a-rlcs.s
+27-0llvm/test/MC/AArch64/armv9.8a-rlcs-diagnostics.s
+3-1llvm/lib/Target/AArch64/AsmParser/AArch64AsmParser.cpp
+3-0llvm/lib/Target/AArch64/AArch64InstrInfo.td
+66-14 files

LLVM/project ddb6566 — llvm/lib/Target/AArch64 AArch64InstrInfo.td AArch64InstrFormats.td, llvm/test/MC/AArch64 directive-arch_extension-negative.s directive-arch-negative.s

[AArch64][llvm] Armv9.8-A: Add support for FEAT_LSC64B (single-copy atomic 64-byte load/store)

Add support for FEAT_LSC64B (single-copy atomic 64-byte load/store)
instructions:

  - LDA64B
  - STL64B
  - STL64BV
  - STL64BV0
DeltaFile
+79-0llvm/test/MC/AArch64/armv9.8a-lsc64b.s
+67-0llvm/test/MC/AArch64/armv9.8a-lsc64b-diagnostics.s
+10-8llvm/lib/Target/AArch64/AArch64InstrFormats.td
+13-4llvm/lib/Target/AArch64/AArch64InstrInfo.td
+6-0llvm/test/MC/AArch64/directive-arch_extension-negative.s
+6-0llvm/test/MC/AArch64/directive-arch-negative.s
+181-126 files not shown
+197-1312 files

LLVM/project f9ca99c — llvm/lib/Target/AArch64 AArch64InstrInfo.td AArch64InstrFormats.td, llvm/lib/Target/AArch64/AsmParser AArch64AsmParser.cpp

[AArch64][llvm] Armv9.8-A: Add support for FEAT_CFLT (Conditional Fault instructions)

Add support for FEAT_CFLT (Conditional Fault instructions), which
are optional from Armv9.7 onwards:
```
   cflteq, cfltne,
   cfltgt, cfltlt,
   cfltge, cfltle,
   cflthi, cfltlo,
   cflths, cfltls,
   cfltz,  cfltnz,
   tfltz,  tfltnz,
   flt.ne, flt.eq,
   flt.mi, flt.pl,
   flt.vs, flt.vc,
   flt.hi, flt.ls,
   flt.ge, flt.lt,
   flt.gt, flt.le,
   flt.al, flt.nv,

    [5 lines not shown]
DeltaFile
+329-0llvm/test/MC/AArch64/armv9.8a-cflt.s
+140-0llvm/lib/Target/AArch64/AArch64InstrFormats.td
+89-0llvm/test/MC/AArch64/armv9.8a-cflt-diagnostics.s
+49-0llvm/lib/Target/AArch64/AArch64InstrInfo.td
+29-11llvm/lib/Target/AArch64/AsmParser/AArch64AsmParser.cpp
+7-0llvm/lib/Target/AArch64/Disassembler/AArch64Disassembler.cpp
+643-117 files not shown
+670-1113 files

LLVM/project f0a0dc5 — clang/lib/Basic/Targets AArch64.cpp, clang/test/Driver arm-cortex-cpus-1.c aarch64-v98a.c

[ARM][AArch64] Introduce Armv9.8-A architecture version

This introduces the Armv9.8-A architecture version, including the
relevant command-line option for -march.

More details about the Armv9.8-A architecture version can be found at:
  * https://community.arm.com/arm-community-blogs/b/architectures-and-processors-blog/posts/arm-a-profile-architecture-developments-2026
  * https://support.arm.com/documentation/109697/2026_09/2026-Architecture-Extensions
  * https://support.arm.com/documentation/109697/2026_09/Future-Architecture-Technologies
  * https://support.arm.com/documentation/ddi0602/2026-09/
DeltaFile
+19-0clang/test/Driver/aarch64-v98a.c
+16-0clang/test/Driver/arm-cortex-cpus-1.c
+12-0llvm/lib/Target/ARM/ARMArchitectures.td
+11-0clang/lib/Basic/Targets/AArch64.cpp
+8-2llvm/unittests/TargetParser/TargetParserTest.cpp
+7-0llvm/lib/Target/AArch64/AArch64Features.td
+73-215 files not shown
+113-521 files

LLVM/project b4a2044 — flang/lib/Optimizer/OpenMP LowerWorkdistribute.cpp, flang/test/Transforms/OpenMP lower-workdistribute-fission-recompute-box.mlir

update
DeltaFile
+93-0flang/test/Transforms/OpenMP/lower-workdistribute-fission-recompute-box.mlir
+48-4flang/lib/Optimizer/OpenMP/LowerWorkdistribute.cpp
+141-42 files

LLVM/project ef18054 — llvm/lib/Analysis LoopAccessAnalysis.cpp, llvm/test/Analysis/LoopAccessAnalysis invariant-dep-same-ptr.ll non-affine-monotonic-bounds.ll

[LAA] Avoid duplicate runtime-check group members (#226829)

Found while working on #226816 and extracted into a separate PR for
easier review.

A read-modify-write pointer has separate read and write entries in the
dependence candidates. Both entries can map back to the same
runtime-check pointer, causing its index to be added to a checking group
twice.

Key the runtime-pointer map by MemAccessInfo, including the read/write
distinction, so each runtime pointer appears only once in a checking
group.

Duplicate members also consume the memory-check merge budget. With a low
merge threshold, the fix reduces two runtime checks to one. It also
restores the singleton-group invariant used by #226816.

Developed with assistance from OpenAI Codex.
DeltaFile
+58-0llvm/test/Analysis/LoopAccessAnalysis/runtime-checks-read-write-merge.ll
+6-10llvm/lib/Analysis/LoopAccessAnalysis.cpp
+0-9llvm/test/Analysis/LoopAccessAnalysis/pointer-phis.ll
+0-8llvm/test/Analysis/LoopAccessAnalysis/non-affine-monotonic-bounds.ll
+0-7llvm/test/Analysis/LoopAccessAnalysis/invariant-dep-same-ptr.ll
+0-4llvm/test/tools/UpdateTestChecks/update_analyze_test_checks/Inputs/loop-distribute.ll.expected
+64-385 files not shown
+64-5111 files

LLVM/project 23c05b5 — clang/test/Sema loadtime-comment-vars.cpp

[Clang][AIX] Test a dynamically initialized array for -mloadtime-comment-vars
DeltaFile
+8-3clang/test/Sema/loadtime-comment-vars.cpp
+8-31 files

LLVM/project 0f9d54e — mlir/test/Dialect/OpenMP invalid.mlir, mlir/test/IR traits.mlir

[MLIR][ODS] Require the value after oilist keyword (#229334)

The generated parser treated optional attribute inside an oilist clause
as optional even after its keyword was parsed

As a result, a keyword without a value was accepted, but the attribute
was simply dropped

Prerequisite for https://github.com/llvm/llvm-project/pull/228977
DeltaFile
+18-0mlir/test/IR/traits.mlir
+9-1mlir/test/Dialect/OpenMP/invalid.mlir
+9-0mlir/test/lib/Dialect/Test/TestOpsSyntax.td
+1-1mlir/tools/mlir-tblgen/OpFormatGen.cpp
+37-24 files

LLVM/project 164a15f — lldb/test/Shell/SymbolFile/NativePDB thread-locals.cpp

[lldb][NativePDB] Use host arch for thread locals test (#229340)

The test used to use `--arch=64`, but technically, it doesn't care about
the architecture, so it was removed. This intends to fix the
lldb-remote-linux-win buildbot:
<lab.llvm.org/buildbot/#/builders/197/builds/19717>.
DeltaFile
+1-1lldb/test/Shell/SymbolFile/NativePDB/thread-locals.cpp
+1-11 files

LLVM/project 2b68c68 — clang/test/CodeGenOpenCL atomics-unsafe-hw-remarks-gfx90a.cl, llvm/lib/Target/AMDGPU SIISelLowering.cpp

AMDGPU: Improve optimization remarks for atomic lowering

Improve the phrasing when the atomicrmw legalization emits the hardware
instruction. Replace the unhelpful "due to an unsafe request" wording in the
optimization remark with the actual reason the native instruction was legal
to use.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+113-43llvm/lib/Target/AMDGPU/SIISelLowering.cpp
+115-0llvm/test/CodeGen/AMDGPU/atomics-hw-remarks-gfx908.ll
+41-0llvm/test/CodeGen/AMDGPU/atomics-hw-remarks-scope.ll
+9-10llvm/test/CodeGen/AMDGPU/atomics-hw-remarks-gfx90a.ll
+3-3clang/test/CodeGenOpenCL/atomics-unsafe-hw-remarks-gfx90a.cl
+281-565 files