[ValueTracking] Compute known bits of and/or recurrences from start and step (#226164)
For simple phi recurrences of the form `%iv = %iv op %step` we currently
only derive trailing zero bits for and/or, in a case shared with
add/sub/mul. This moves and/or into their own case and propagates full
known bits:
* or: bits that are zero in both the start value and the step stay zero,
and bits that are one in the start value stay one.
* and: bits that are zero in the start value stay zero, and bits that
are one in both the start value and the step stay one.
The step only applies from the second iteration on, so every fact must
also hold for the start value alone. This subsumes the trailing-zeros
rule for these operations. The nsw handling of the add/sub/mul case
never applied to them.
The two AMDGPU tests are adjusted to use an opaque start value for their
recurrences, so that the improved known bits don't fold away the
[4 lines not shown]
[AArch64][PAC] Reset `killed` operand flags in outlined functions (#221041)
Presently, MachineOutliner does not take `killed` operand flags into
account when merging instruction sequences. While it sounds perfectly
reasonable not to inhibit merging of the instruction sequences that only
differ in `killed` flags (for N flags there is technically 2^N valid
ways to drop some subset of them), copying these flags from an
arbitrarily chosen representative instruction may result in incorrect
codegen of PAuth-related pseudo instructions on AArch64.
To keep `killed` flags conservatively correct as if `OUTLINED_FUNCTION`s
are virtually re-inserted at every call site, this patch takes the
simplest approach of resetting every `killed` flag inside the outlined
functions.
[flang][OpenMP] Perform checks on features from different versions (#229116)
The check for DETACH and MERGEABLE used together was only done when the
OpenMP version was set to the versions that allow both clauses. The use
of these clauses is accepted with a warning in older versions as well,
so make sure to perform the checks for all versions.
[clang/wsm] Check for suppression section before computing presumed loc (#229158)
With --warning-suppression-mappings=, every
DiagnosticIDs::getDiagnosticSeverity() call for a diagnostic that isn't
ignored calls WarningsSpecialCaseList::isDiagSuppressed(), so it's
called fairly often.
It seems reasonable to assume that the warning suppression list has few
entries compared to all the diagnostics clang knows about. So checking
if a diag ID is in the list is a) fast and b) rejects most DiagIds.
So check if DiagId is in DiagToSection before calling getPresumedLoc, as
the latter is somewhat expensive.
For 60 random Chromium TUs (linux x64, -O2, with Chromium's suppression
mapping file) picked with probability proportional to their compile
time, sum over all TUs, mean of two runs:
CPU time: 192.7 s => 191.4 s, -0.7% (runs differ by up to 0.7%)
[2 lines not shown]
[GVN] Use willNotFreeBetween in loop-load PRE (#228033)
`canBeFreed` used by GVN is very conservative. For an argument, one of
the cases where it returns false is when the argument has
`nofree`/`readonly` and `noalias`. A pointer that may alias a clobber in
the loop is never `noalias`, so PRE bails out even when nothing on the
path can free the object.
This uses `willNotFreeBetween` alongside `canBeFreed`. The header load
has already dereferenced `LoadPtr` on an iteration, so only a
deallocation between that load and the reload can make the same address
unsafe to read again.
The non-linear walk this relies on landed in #223580. The existing limit
of 32 instructions per query is left unchanged.
This is the first of two patches, split as suggested in review of my
earlier combined change in #227983. The follow-up will allow a blocker
inside an inner loop; that change is a no-op in many cases without this
[7 lines not shown]
[LV] Allow out-of-loop backedge users for min/max value of argmin/argmax. (#220351)
The only user of multi-use reductions is LoopVectorize and its
argmin/argmax transform only requires a single user of the backedge
value in the loop and already works properly for additional users
outside of the loop.
Refine the legality check if we have invalid uses, to skip uses outside
the loop, and only bail out if there are any in-loop users that is not
the phi itself.
PR: https://github.com/llvm/llvm-project/pull/220351
[AMDGPU] Measure MFMA hazard windows at their own producers (#225989)
Depends on #225988.
The recogniser pads with s_nop when an MFMA result is read or
overwritten too
early. The padding is window - distance, where the window depends on the
producer's shape. Two defects:
- The distance comes from the nearest matching producer, the window from
a
pointer a predicate stored as a side effect. With several producers in
range they describe different instructions and the padding can be too
small.
- The backward walk records visited blocks without their distance, so a
block
reached the long way first is not re-examined when a shorter path
appears.
The answer then depends on block layout.
[11 lines not shown]
[mlir][tosa] Lower ROW_GATHER to Linalg (#225417)
Lower ROW_GATHER to a linalg.generic that maps each expanded output row
back to its source index slot and consecutive row offset. Support
dynamic output dimensions and both i32 and i64 indices.
Assisted-by: Codex
[BOLT][AArch64] Port X86 --custom-allocation-vma test to AArch64 (#229104)
**Before**: #136385 introduced the `--custom-allocation-vma` flag for
BOLT to be able to specify a suitable location to place rewritten
binary. This was accompanied by a corresponding test for `x86`.
**After**: This PR ports the existing test `high-segments.s` to AArch64,
improving test coverage.
Assisted-by: Codex
[AArch64][PAC] Reset `killed` operand flags in outlined functions
Presently, MachineOutliner does not take `killed` operand flags into
account when merging instruction sequences. While it sounds perfectly
reasonable not to inhibit merging of the instruction sequences that
only differ in `killed` flags (for N flags there is technically 2^N
valid ways to drop some subset of them), copying these flags from
an arbitrarily chosen representative instruction may result in
incorrect codegen of PAuth-related pseudo instructions on AArch64.
To keep `killed` flags conservatively correct as if `OUTLINED_FUNCTION`s
are virtually re-inserted at every call site, this patch takes the
simplest approach of resetting every `killed` flag inside the
outlined functions.
[X86] Require contract on FMUL while folding FMADDSUB to VFMULC (#229309)
FMADDSUB/FMSUBADD has no FMF, so we check its only operand (third; FMUL) which still has it.
RISCV: Stop setting kill flags on virtual registers before FinalizeISel (#229013)
Kill flags have no remaining use before register allocation and are stripped by
LiveIntervals.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[CMake] Make find_package(Clang) load MLIR when ClangIR uses a full MLIR (#228366)
With CLANG_ENABLE_CIR=ON, Clang's exported libraries link MLIR. When
MLIR is listed in LLVM_ENABLE_PROJECTS it installs its own package and
owns every target exported from mlir/. ClangTargets.cmake references
those targets without defining them, and ClangConfig.cmake never loaded
the MLIR package, so an external find_package(Clang) failed with:
```
The following imported targets are referenced, but are missing:
MLIRIR MLIRPass MLIRAnalysis ... MLIRSupport MLIRTransforms ...
```
unless the consumer happened to call find_package(MLIR) first.
When MLIR is enabled only implicitly as a ClangIR dependency
(LLVM_DEPENDENCY_ONLY_PROJECTS), no MLIR package exists and Clang
already promotes the reachable MLIR closure into its own exports, so
that mode worked. The two modes therefore need different plumbing but
[131 lines not shown]
RISCV: Stop setting kill flags on virtual registers before FinalizeISel
Kill flags have no remaining use before register allocation and are
stripped by LiveIntervals.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[AMDGPU] Fold fmul with zero checks to fmul.legacy
Recognize f32 multiplication with zero checks in AMDGPUCodeGenPrepare,
including fabs/fneg wrappers.
Example:
```
(x == 0 || y == 0) ? +0.0 : x * y
==>
llvm.amdgcn.fmul.legacy(x, y)
```
[AMDGPU] Add fmul.legacy select fold tests. NFC
Add AMDGPUCodeGenPrepare tests for folding a select guarded by zero
checks into llvm.amdgcn.fmul.legacy, including negative tests and min
cases for future work.
[AArch64][SPIRV][GlobalISel] Migrate target-specific wip_match_opcode combines to MIR-pattern (#222903)
It replaces the `wip_match_opcode` with declarative MIR-pattern match
roots across the AArch64 and SPIRV target-specific GICombines.
Some important points to consider :
- **shuffle_vector_lowering:** these `G_SHUFFLE_VECTOR` rules are order-
sensitive. `fullrev` previously gained priority through a nested
`G_IMPLICIT_DEF` match; that predicate is relocated to C++ so all rules
are equal-priority and `fullrev` is ordered last, preserving the
original dispatch sequence and codegen.
- **SPIRV intrinsic roots:** the matrix/length/distance rules match and
apply on target intrinsics (`int_spv_*`, `int_matrix_*`), requiring the
`IntrinsicsSPIRV` enum in the combiner translation unit.
- **Deferred:** `vector_unmerge_lowering` and `unmerge_ext_to_unmerge`
remain on `wip_match_opcode`; their `G_UNMERGE_VALUES` variadic-def
roots are not yet expressible as MIR patterns.
[AArch64][LSR] Prefer pointer IVs for SVE accesses that need splitting (#228510)
Currently, LSR always prefers scaled accesses but that ends up doing
more work as only the first access can used the scaled register. Later
split accesses need to compute the base then another `mul vl` offset.
Assisted-by: Codex