[mlir][tosa] Lower ROW_GATHER to Linalg (#225417)
Lower ROW_GATHER to a linalg.generic that maps each expanded output row
back to its source index slot and consecutive row offset. Support
dynamic output dimensions and both i32 and i64 indices.
Assisted-by: Codex
[BOLT][AArch64] Port X86 --custom-allocation-vma test to AArch64 (#229104)
**Before**: #136385 introduced the `--custom-allocation-vma` flag for
BOLT to be able to specify a suitable location to place rewritten
binary. This was accompanied by a corresponding test for `x86`.
**After**: This PR ports the existing test `high-segments.s` to AArch64,
improving test coverage.
Assisted-by: Codex
[AArch64][PAC] Reset `killed` operand flags in outlined functions
Presently, MachineOutliner does not take `killed` operand flags into
account when merging instruction sequences. While it sounds perfectly
reasonable not to inhibit merging of the instruction sequences that
only differ in `killed` flags (for N flags there is technically 2^N
valid ways to drop some subset of them), copying these flags from
an arbitrarily chosen representative instruction may result in
incorrect codegen of PAuth-related pseudo instructions on AArch64.
To keep `killed` flags conservatively correct as if `OUTLINED_FUNCTION`s
are virtually re-inserted at every call site, this patch takes the
simplest approach of resetting every `killed` flag inside the
outlined functions.
[X86] Require contract on FMUL while folding FMADDSUB to VFMULC (#229309)
FMADDSUB/FMSUBADD has no FMF, so we check its only operand (third; FMUL) which still has it.
RISCV: Stop setting kill flags on virtual registers before FinalizeISel (#229013)
Kill flags have no remaining use before register allocation and are stripped by
LiveIntervals.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[CMake] Make find_package(Clang) load MLIR when ClangIR uses a full MLIR (#228366)
With CLANG_ENABLE_CIR=ON, Clang's exported libraries link MLIR. When
MLIR is listed in LLVM_ENABLE_PROJECTS it installs its own package and
owns every target exported from mlir/. ClangTargets.cmake references
those targets without defining them, and ClangConfig.cmake never loaded
the MLIR package, so an external find_package(Clang) failed with:
```
The following imported targets are referenced, but are missing:
MLIRIR MLIRPass MLIRAnalysis ... MLIRSupport MLIRTransforms ...
```
unless the consumer happened to call find_package(MLIR) first.
When MLIR is enabled only implicitly as a ClangIR dependency
(LLVM_DEPENDENCY_ONLY_PROJECTS), no MLIR package exists and Clang
already promotes the reachable MLIR closure into its own exports, so
that mode worked. The two modes therefore need different plumbing but
[131 lines not shown]
RISCV: Stop setting kill flags on virtual registers before FinalizeISel
Kill flags have no remaining use before register allocation and are
stripped by LiveIntervals.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[AMDGPU] Fold fmul with zero checks to fmul.legacy
Recognize f32 multiplication with zero checks in AMDGPUCodeGenPrepare,
including fabs/fneg wrappers.
Example:
```
(x == 0 || y == 0) ? +0.0 : x * y
==>
llvm.amdgcn.fmul.legacy(x, y)
```
[AMDGPU] Add fmul.legacy select fold tests. NFC
Add AMDGPUCodeGenPrepare tests for folding a select guarded by zero
checks into llvm.amdgcn.fmul.legacy, including negative tests and min
cases for future work.
[AArch64][SPIRV][GlobalISel] Migrate target-specific wip_match_opcode combines to MIR-pattern (#222903)
It replaces the `wip_match_opcode` with declarative MIR-pattern match
roots across the AArch64 and SPIRV target-specific GICombines.
Some important points to consider :
- **shuffle_vector_lowering:** these `G_SHUFFLE_VECTOR` rules are order-
sensitive. `fullrev` previously gained priority through a nested
`G_IMPLICIT_DEF` match; that predicate is relocated to C++ so all rules
are equal-priority and `fullrev` is ordered last, preserving the
original dispatch sequence and codegen.
- **SPIRV intrinsic roots:** the matrix/length/distance rules match and
apply on target intrinsics (`int_spv_*`, `int_matrix_*`), requiring the
`IntrinsicsSPIRV` enum in the combiner translation unit.
- **Deferred:** `vector_unmerge_lowering` and `unmerge_ext_to_unmerge`
remain on `wip_match_opcode`; their `G_UNMERGE_VALUES` variadic-def
roots are not yet expressible as MIR patterns.
[AArch64][LSR] Prefer pointer IVs for SVE accesses that need splitting (#228510)
Currently, LSR always prefers scaled accesses but that ends up doing
more work as only the first access can used the scaled register. Later
split accesses need to compute the base then another `mul vl` offset.
Assisted-by: Codex
[LAA] Avoid duplicate runtime-check group members (#226829)
Found while working on #226816 and extracted into a separate PR for
easier review.
A read-modify-write pointer has separate read and write entries in the
dependence candidates. Both entries can map back to the same
runtime-check pointer, causing its index to be added to a checking group
twice.
Key the runtime-pointer map by MemAccessInfo, including the read/write
distinction, so each runtime pointer appears only once in a checking
group.
Duplicate members also consume the memory-check merge budget. With a low
merge threshold, the fix reduces two runtime checks to one. It also
restores the singleton-group invariant used by #226816.
Developed with assistance from OpenAI Codex.
[MLIR][ODS] Require the value after oilist keyword (#229334)
The generated parser treated optional attribute inside an oilist clause
as optional even after its keyword was parsed
As a result, a keyword without a value was accepted, but the attribute
was simply dropped
Prerequisite for https://github.com/llvm/llvm-project/pull/228977
[lldb][NativePDB] Use host arch for thread locals test (#229340)
The test used to use `--arch=64`, but technically, it doesn't care about
the architecture, so it was removed. This intends to fix the
lldb-remote-linux-win buildbot:
<lab.llvm.org/buildbot/#/builders/197/builds/19717>.
AMDGPU: Improve optimization remarks for atomic lowering
Improve the phrasing when the atomicrmw legalization emits the hardware
instruction. Replace the unhelpful "due to an unsafe request" wording in the
optimization remark with the actual reason the native instruction was legal
to use.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>