LLVM/project 37ecc8c — llvm/lib/Analysis ValueTracking.cpp, llvm/test/CodeGen/AMDGPU absdiff.ll

[ValueTracking] Compute known bits of and/or recurrences from start and step (#226164)

For simple phi recurrences of the form `%iv = %iv op %step` we currently
only derive trailing zero bits for and/or, in a case shared with
add/sub/mul. This moves and/or into their own case and propagates full
known bits:

* or: bits that are zero in both the start value and the step stay zero,
and bits that are one in the start value stay one.
* and: bits that are zero in the start value stay zero, and bits that
are one in both the start value and the step stay one.

The step only applies from the second iteration on, so every fact must
also hold for the start value alone. This subsumes the trailing-zeros
rule for these operations. The nsw handling of the add/sub/mul case
never applied to them.

The two AMDGPU tests are adjusted to use an opaque start value for their
recurrences, so that the improved known bits don't fold away the

    [4 lines not shown]
DeltaFile
+166-0llvm/test/Transforms/InstCombine/recurrence.ll
+22-2llvm/lib/Analysis/ValueTracking.cpp
+9-10llvm/test/CodeGen/AMDGPU/GlobalISel/is-safe-to-sink-bug.ll
+6-7llvm/test/CodeGen/AMDGPU/absdiff.ll
+203-194 files

LLVM/project c68c176 — llvm/lib/CodeGen MachineOutliner.cpp, llvm/test/CodeGen/AArch64 machine-outliner-bundle.mir machine-outliner-operand-flags.mir

[AArch64][PAC] Reset `killed` operand flags in outlined functions (#221041)

Presently, MachineOutliner does not take `killed` operand flags into
account when merging instruction sequences. While it sounds perfectly
reasonable not to inhibit merging of the instruction sequences that only
differ in `killed` flags (for N flags there is technically 2^N valid
ways to drop some subset of them), copying these flags from an
arbitrarily chosen representative instruction may result in incorrect
codegen of PAuth-related pseudo instructions on AArch64.

To keep `killed` flags conservatively correct as if `OUTLINED_FUNCTION`s
are virtually re-inserted at every call site, this patch takes the
simplest approach of resetting every `killed` flag inside the outlined
functions.
DeltaFile
+126-0llvm/test/CodeGen/AArch64/machine-outliner-operand-flags.mir
+5-2llvm/lib/CodeGen/MachineOutliner.cpp
+1-1llvm/test/CodeGen/ARM/machine-outliner-stack-fixup-thumb.mir
+1-1llvm/test/CodeGen/AArch64/machine-outliner-bundle.mir
+133-44 files

LLVM/project 0c28080 — flang/lib/Semantics check-omp-structure.cpp, flang/test/Semantics/OpenMP task-45.f90

[flang][OpenMP] Perform checks on features from different versions (#229116)

The check for DETACH and MERGEABLE used together was only done when the
OpenMP version was set to the versions that allow both clauses. The use
of these clauses is accepted with a warning in older versions as well,
so make sure to perform the checks for all versions.
DeltaFile
+14-0flang/test/Semantics/OpenMP/task-45.f90
+1-1flang/lib/Semantics/check-omp-structure.cpp
+15-12 files

LLVM/project 037b7c1 — clang/lib/Basic Diagnostic.cpp

[clang/wsm] Check for suppression section before computing presumed loc (#229158)

With --warning-suppression-mappings=, every
DiagnosticIDs::getDiagnosticSeverity() call for a diagnostic that isn't
ignored calls WarningsSpecialCaseList::isDiagSuppressed(), so it's
called fairly often.

It seems reasonable to assume that the warning suppression list has few
entries compared to all the diagnostics clang knows about. So checking
if a diag ID is in the list is a) fast and b) rejects most DiagIds.

So check if DiagId is in DiagToSection before calling getPresumedLoc, as
the latter is somewhat expensive.

For 60 random Chromium TUs (linux x64, -O2, with Chromium's suppression
mapping file) picked with probability proportional to their compile
time, sum over all TUs, mean of two runs:

    CPU time: 192.7 s => 191.4 s, -0.7% (runs differ by up to 0.7%)

    [2 lines not shown]
DeltaFile
+3-3clang/lib/Basic/Diagnostic.cpp
+3-31 files

LLVM/project df18ff3 — mlir/lib/Conversion/TosaToLinalg TosaToLinalg.cpp, mlir/test/Conversion/TosaToLinalg tosa-to-linalg-invalid.mlir tosa-to-linalg.mlir

Revert "[mlir][tosa] Lower ROW_GATHER to Linalg (#225417)"

This reverts commit 57f2572b545ae1083c663b6636bae61191ba8c02.
DeltaFile
+0-71mlir/test/Conversion/TosaToLinalg/tosa-to-linalg.mlir
+0-70mlir/lib/Conversion/TosaToLinalg/TosaToLinalg.cpp
+0-9mlir/test/Conversion/TosaToLinalg/tosa-to-linalg-invalid.mlir
+0-1503 files

LLVM/project 27862ae — llvm/lib/Transforms/Scalar GVN.cpp, llvm/test/Transforms/GVN/PRE pre-aliasning-path.ll pre-loop-load.ll

[GVN] Use willNotFreeBetween in loop-load PRE (#228033)

`canBeFreed` used by GVN is very conservative. For an argument, one of
the cases where it returns false is when the argument has
`nofree`/`readonly` and `noalias`. A pointer that may alias a clobber in
the loop is never `noalias`, so PRE bails out even when nothing on the
path can free the object.

This uses `willNotFreeBetween` alongside `canBeFreed`. The header load
has already dereferenced `LoadPtr` on an iteration, so only a
deallocation between that load and the reload can make the same address
unsafe to read again.

The non-linear walk this relies on landed in #223580. The existing limit
of 32 instructions per query is left unchanged.

This is the first of two patches, split as suggested in review of my
earlier combined change in #227983. The follow-up will allow a blocker
inside an inner loop; that change is a no-op in many cases without this

    [7 lines not shown]
DeltaFile
+98-12llvm/test/Transforms/GVN/PRE/pre-loop-load.ll
+14-6llvm/test/Transforms/GVN/PRE/pre-aliasning-path.ll
+7-1llvm/lib/Transforms/Scalar/GVN.cpp
+2-2llvm/test/Transforms/PhaseOrdering/X86/ptrtoaddr-ptrtoint.ll
+121-214 files

LLVM/project e883060 — llvm/test/CodeGen/AMDGPU frem.ll, llvm/test/CodeGen/AMDGPU/GlobalISel srem.i64.ll sdiv.i64.ll

Merge branch 'main' into users/kparzysz/gfortran-1
DeltaFile
+1,609-1,040llvm/test/CodeGen/AMDGPU/frem.ll
+986-1,144llvm/test/CodeGen/AMDGPU/GlobalISel/sdiv.i64.ll
+914-1,072llvm/test/CodeGen/AMDGPU/GlobalISel/srem.i64.ll
+608-584llvm/test/CodeGen/NVPTX/arbitrary-fp-to-float.ll
+1,023-147llvm/test/Transforms/LoopVectorize/smax-idx.ll
+581-521llvm/test/CodeGen/X86/vector-compress.ll
+5,721-4,508678 files not shown
+27,558-15,114684 files

LLVM/project 7f0cd81 — llvm/test/Transforms/LoopVectorize select-umin-last-index.ll select-umax-last-index.ll

[LV] Allow out-of-loop backedge users for min/max value of argmin/argmax. (#220351)

The only user of multi-use reductions is LoopVectorize and its
argmin/argmax transform only requires a single user of the backedge
value in the loop and already works properly for additional users
outside of the loop.

Refine the legality check if we have invalid uses, to skip uses outside
the loop, and only bail out if there are any in-loop users that is not
the phi itself.

PR: https://github.com/llvm/llvm-project/pull/220351
DeltaFile
+1,023-147llvm/test/Transforms/LoopVectorize/smax-idx.ll
+90-11llvm/test/Transforms/LoopVectorize/select-smax-last-index.ll
+48-13llvm/test/Transforms/LoopVectorize/select-umin-first-index.ll
+46-11llvm/test/Transforms/LoopVectorize/select-umin-last-index.ll
+46-11llvm/test/Transforms/LoopVectorize/select-umax-last-index.ll
+46-11llvm/test/Transforms/LoopVectorize/select-smin-last-index.ll
+1,299-2041 files not shown
+1,305-2077 files

LLVM/project 1c694b9 — llvm/lib/Target/AMDGPU GCNHazardRecognizer.h GCNHazardRecognizer.cpp, llvm/test/CodeGen/AMDGPU mai-hazards-gfx942.mir mai-hazards-gfx90a.mir

[AMDGPU] Measure MFMA hazard windows at their own producers (#225989)

Depends on #225988.

The recogniser pads with s_nop when an MFMA result is read or
overwritten too
early. The padding is window - distance, where the window depends on the
producer's shape. Two defects:

- The distance comes from the nearest matching producer, the window from
a
pointer a predicate stored as a side effect. With several producers in
   range they describe different instructions and the padding can be too
   small.
- The backward walk records visited blocks without their distance, so a
block
reached the long way first is not re-examined when a shorter path
appears.
   The answer then depends on block layout.

    [11 lines not shown]
DeltaFile
+674-0llvm/test/CodeGen/AMDGPU/mai-hazards-gfx90a.mir
+285-129llvm/lib/Target/AMDGPU/GCNHazardRecognizer.cpp
+136-0llvm/test/CodeGen/AMDGPU/mai-hazards-gfx942.mir
+10-0llvm/lib/Target/AMDGPU/GCNHazardRecognizer.h
+1,105-1294 files

LLVM/project da3277f — llvm/docs AMDGPUUsage.rst, llvm/test/CodeGen/AMDGPU frem.ll

Rebase, address comments

Created using spr 1.3.7
DeltaFile
+1,609-1,040llvm/test/CodeGen/AMDGPU/frem.ll
+986-1,144llvm/test/CodeGen/AMDGPU/GlobalISel/sdiv.i64.ll
+914-1,072llvm/test/CodeGen/AMDGPU/GlobalISel/srem.i64.ll
+608-584llvm/test/CodeGen/NVPTX/arbitrary-fp-to-float.ll
+581-521llvm/test/CodeGen/X86/vector-compress.ll
+448-433llvm/docs/AMDGPUUsage.rst
+5,146-4,794652 files not shown
+24,603-14,582658 files

LLVM/project 2037f0e —

[SLP][NFC]Add an extra test with the reodering of the aggregate bitpacks, NFC



Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/229380
DeltaFile
+0-00 files

LLVM/project c4c287a —

[SLP][NFC]Add an extra test with the reodering of the aggregate bitpacks, NFC



Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/229380
DeltaFile
+0-00 files

LLVM/project 57f2572 — mlir/lib/Conversion/TosaToLinalg TosaToLinalg.cpp, mlir/test/Conversion/TosaToLinalg tosa-to-linalg-invalid.mlir tosa-to-linalg.mlir

[mlir][tosa] Lower ROW_GATHER to Linalg (#225417)

Lower ROW_GATHER to a linalg.generic that maps each expanded output row
back to its source index slot and consecutive row offset. Support
dynamic output dimensions and both i32 and i64 indices.

Assisted-by: Codex
DeltaFile
+71-0mlir/test/Conversion/TosaToLinalg/tosa-to-linalg.mlir
+70-0mlir/lib/Conversion/TosaToLinalg/TosaToLinalg.cpp
+9-0mlir/test/Conversion/TosaToLinalg/tosa-to-linalg-invalid.mlir
+150-03 files

LLVM/project 9a39608 — llvm/test/Transforms/SLPVectorizer/X86 bitpacked-aggregate.ll

[SLP][NFC]Add an extra test with the reodering of the aggregate bitpacks, NFC



Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/229380
DeltaFile
+54-0llvm/test/Transforms/SLPVectorizer/X86/bitpacked-aggregate.ll
+54-01 files

LLVM/project f30e753 — bolt/test high-segments.s, bolt/test/X86 high-segments.s

[BOLT][AArch64] Port X86 --custom-allocation-vma test to AArch64 (#229104)

**Before**: #136385 introduced the `--custom-allocation-vma` flag for
BOLT to be able to specify a suitable location to place rewritten
binary. This was accompanied by a corresponding test for `x86`.

**After**: This PR ports the existing test `high-segments.s` to AArch64,
improving test coverage.

Assisted-by: Codex
DeltaFile
+0-46bolt/test/X86/high-segments.s
+33-0bolt/test/high-segments.s
+33-462 files

LLVM/project 4e75021 — llvm/lib/CodeGen MachineOutliner.cpp, llvm/test/CodeGen/AArch64 machine-outliner-bundle.mir machine-outliner-operand-flags.mir

Process instructions inside bundles
DeltaFile
+47-3llvm/test/CodeGen/AArch64/machine-outliner-operand-flags.mir
+4-2llvm/lib/CodeGen/MachineOutliner.cpp
+1-1llvm/test/CodeGen/AArch64/machine-outliner-bundle.mir
+52-63 files

LLVM/project 1b825bd — llvm/test/CodeGen/AArch64 machine-outliner-operand-flags.mir

Fix comments
DeltaFile
+3-3llvm/test/CodeGen/AArch64/machine-outliner-operand-flags.mir
+3-31 files

LLVM/project 1c49151 — llvm/test/CodeGen/ARM machine-outliner-stack-fixup-thumb.mir

Update CodeGen/ARM/machine-outliner-stack-fixup-thumb.mir
DeltaFile
+1-1llvm/test/CodeGen/ARM/machine-outliner-stack-fixup-thumb.mir
+1-11 files

LLVM/project b6be3ae — llvm/lib/CodeGen MachineOutliner.cpp, llvm/test/CodeGen/AArch64 machine-outliner-operand-flags.mir

[AArch64][PAC] Reset `killed` operand flags in outlined functions

Presently, MachineOutliner does not take `killed` operand flags into
account when merging instruction sequences. While it sounds perfectly
reasonable not to inhibit merging of the instruction sequences that
only differ in `killed` flags (for N flags there is technically 2^N
valid ways to drop some subset of them), copying these flags from
an arbitrarily chosen representative instruction may result in
incorrect codegen of PAuth-related pseudo instructions on AArch64.

To keep `killed` flags conservatively correct as if `OUTLINED_FUNCTION`s
are virtually re-inserted at every call site, this patch takes the
simplest approach of resetting every `killed` flag inside the
outlined functions.
DeltaFile
+82-0llvm/test/CodeGen/AArch64/machine-outliner-operand-flags.mir
+1-0llvm/lib/CodeGen/MachineOutliner.cpp
+83-02 files

LLVM/project 22ddf4a — llvm/test/Transforms/SLPVectorizer/X86 bitpacked-aggregate.ll

[𝘀𝗽𝗿] initial version

Created using spr 1.3.7
DeltaFile
+54-0llvm/test/Transforms/SLPVectorizer/X86/bitpacked-aggregate.ll
+54-01 files

LLVM/project 0c53d0f — llvm/lib/Target/X86 X86ISelLowering.cpp, llvm/test/CodeGen/X86 avx512fp16-combine-fmsubadd.ll

[X86] Require contract on FMUL while folding FMADDSUB to VFMULC (#229309)

FMADDSUB/FMSUBADD has no FMF, so we check its only operand (third; FMUL) which still has it.
DeltaFile
+45-0llvm/test/CodeGen/X86/avx512fp16-combine-fmsubadd.ll
+1-1llvm/lib/Target/X86/X86ISelLowering.cpp
+46-12 files

LLVM/project 08483c5 — llvm/lib/Target/RISCV RISCVISelLowering.cpp

RISCV: Stop setting kill flags on virtual registers before FinalizeISel (#229013)

Kill flags have no remaining use before register allocation and are stripped by 
LiveIntervals.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+2-2llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+2-21 files

LLVM/project 4fadb2e — utils/bazel/llvm-project-overlay/mlir BUILD.bazel

[Bazel] Fixes fc90bd3 (#229344)

This fixes fc90bd3cfb0bab14a2bf7cbd33f5f1e904c7e133 (#228459).

Buildkite error link:
https://buildkite.com/llvm-project/upstream-bazel/builds?commit=fc90bd3cfb0bab14a2bf7cbd33f5f1e904c7e133

Co-authored-by: Google Bazel Bot <google-bazel-bot at google.com>
DeltaFile
+1-0utils/bazel/llvm-project-overlay/mlir/BUILD.bazel
+1-01 files

LLVM/project 94eb65e — clang/cmake/modules CMakeLists.txt ClangConfig.cmake.in, llvm/docs ReleaseNotes.md

[CMake] Make find_package(Clang) load MLIR when ClangIR uses a full MLIR (#228366)

With CLANG_ENABLE_CIR=ON, Clang's exported libraries link MLIR. When
MLIR is listed in LLVM_ENABLE_PROJECTS it installs its own package and
owns every target exported from mlir/. ClangTargets.cmake references
those targets without defining them, and ClangConfig.cmake never loaded
the MLIR package, so an external find_package(Clang) failed with:

```
  The following imported targets are referenced, but are missing:
  MLIRIR MLIRPass MLIRAnalysis ... MLIRSupport MLIRTransforms ...
```

unless the consumer happened to call find_package(MLIR) first.

When MLIR is enabled only implicitly as a ClangIR dependency
(LLVM_DEPENDENCY_ONLY_PROJECTS), no MLIR package exists and Clang
already promotes the reachable MLIR closure into its own exports, so
that mode worked. The two modes therefore need different plumbing but

    [131 lines not shown]
DeltaFile
+53-0clang/cmake/modules/ClangConfig.cmake.in
+36-0clang/cmake/modules/CMakeLists.txt
+18-4mlir/cmake/modules/CMakeLists.txt
+10-0llvm/docs/ReleaseNotes.md
+6-0mlir/cmake/modules/MLIRConfig.cmake.in
+123-45 files

LLVM/project 07e9854 — llvm/lib/Target/RISCV RISCVISelLowering.cpp

RISCV: Stop setting kill flags on virtual registers before FinalizeISel

Kill flags have no remaining use before register allocation and are
stripped by LiveIntervals.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+2-2llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+2-21 files

LLVM/project fab6b85 — llvm/lib/Target/AMDGPU AMDGPUCodeGenPrepare.cpp, llvm/test/CodeGen/AMDGPU amdgpu-codegenprepare-fmul-legacy.ll

[AMDGPU] Fold fmul with zero checks to fmul.legacy

Recognize f32 multiplication with zero checks in AMDGPUCodeGenPrepare,
including fabs/fneg wrappers.

Example:
```
  (x == 0 || y == 0) ? +0.0 : x * y
==>
  llvm.amdgcn.fmul.legacy(x, y)
```
DeltaFile
+21-107llvm/test/CodeGen/AMDGPU/amdgpu-codegenprepare-fmul-legacy.ll
+77-0llvm/lib/Target/AMDGPU/AMDGPUCodeGenPrepare.cpp
+98-1072 files

LLVM/project b6396d7 — llvm/test/CodeGen/AMDGPU amdgpu-codegenprepare-fmul-legacy.ll

[AMDGPU] Add fmul.legacy select fold tests. NFC

Add AMDGPUCodeGenPrepare tests for folding a select guarded by zero
checks into llvm.amdgcn.fmul.legacy, including negative tests and min
cases for future work.
DeltaFile
+1,030-0llvm/test/CodeGen/AMDGPU/amdgpu-codegenprepare-fmul-legacy.ll
+1,030-01 files

LLVM/project 92d7d31 — llvm/lib/Target/AArch64 AArch64Combine.td, llvm/lib/Target/AArch64/GISel AArch64PostLegalizerCombiner.cpp

[AArch64][SPIRV][GlobalISel] Migrate target-specific wip_match_opcode combines to MIR-pattern (#222903)

It replaces the `wip_match_opcode` with declarative MIR-pattern match
roots across the AArch64 and SPIRV target-specific GICombines.

Some important points to consider :

- **shuffle_vector_lowering:** these `G_SHUFFLE_VECTOR` rules are order-
sensitive. `fullrev` previously gained priority through a nested
`G_IMPLICIT_DEF` match; that predicate is relocated to C++ so all rules
are equal-priority and `fullrev` is ordered last, preserving the
original dispatch sequence and codegen.
- **SPIRV intrinsic roots:** the matrix/length/distance rules match and
apply on target intrinsics (`int_spv_*`, `int_matrix_*`), requiring the
`IntrinsicsSPIRV` enum in the combiner translation unit.
- **Deferred:** `vector_unmerge_lowering` and `unmerge_ext_to_unmerge`
remain on `wip_match_opcode`; their `G_UNMERGE_VALUES` variadic-def
roots are not yet expressible as MIR patterns.
DeltaFile
+247-410llvm/test/CodeGen/AArch64/neon-shuffle-vector-tbl.ll
+62-49llvm/lib/Target/AArch64/AArch64Combine.td
+0-63llvm/lib/Target/AArch64/GISel/AArch64PostLegalizerCombiner.cpp
+50-0llvm/test/CodeGen/AArch64/GlobalISel/postlegalizercombiner-mul-identity.mir
+0-33llvm/lib/Target/SPIRV/SPIRVCombinerHelper.cpp
+8-11llvm/lib/Target/SPIRV/SPIRVCombine.td
+367-5665 files not shown
+376-58811 files

LLVM/project 08652b2 — llvm/lib/Target/AArch64 AArch64TargetTransformInfo.h AArch64TargetTransformInfo.cpp, llvm/test/CodeGen/AArch64 zext-to-tbl.ll

[AArch64][LSR] Prefer pointer IVs for SVE accesses that need splitting (#228510)

Currently, LSR always prefers scaled accesses but that ends up doing
more work as only the first access can used the scaled register. Later
split accesses need to compute the base then another `mul vl` offset.

Assisted-by: Codex
DeltaFile
+187-3llvm/test/Transforms/LoopStrengthReduce/AArch64/vscale-fixups.ll
+38-0llvm/lib/Target/AArch64/AArch64TargetTransformInfo.cpp
+16-16llvm/test/CodeGen/AArch64/zext-to-tbl.ll
+5-0llvm/lib/Target/AArch64/AArch64TargetTransformInfo.h
+246-194 files

LLVM/project 54f957e — llvm/lib/Target/AArch64 AArch64InstrInfo.td, llvm/lib/Target/AArch64/AsmParser AArch64AsmParser.cpp

[AArch64][llvm] Armv9.8-A: Add support for FEAT_RLCS (release consistency scoping)

Add support for FEAT_RLCS (release consistency scoping) instructions:
  - SRLS
  - SLBND
DeltaFile
+33-0llvm/test/MC/AArch64/armv9.8a-rlcs.s
+27-0llvm/test/MC/AArch64/armv9.8a-rlcs-diagnostics.s
+3-1llvm/lib/Target/AArch64/AsmParser/AArch64AsmParser.cpp
+3-0llvm/lib/Target/AArch64/AArch64InstrInfo.td
+66-14 files