[LV] Use SCEV loop-uniformity for outer-loop branch legality (#199632)
This patch refactors the outer-loop vectorization branch legality checks
to reason about conditional branches directly instead of using the old
recursive inner-loop shape check.
The new check allows conditional branches when their condition is
either:
- loop-invariant with respect to the vectorized outer loop, or
- a compare whose operands are both SCEV loop-uniform with respect to
the vectorized outer loop.
Divergent conditional branches are still rejected, now with a more
specific diagnostic.
[KnownFPClass][NFC] Update ATTR values for atan2 tests (#224797)
Ran the following command since it was not run for
https://github.com/llvm/llvm-project/pull/223176
```
llvm/utils/update_test_checks.py \
--opt-binary build/bin/opt \
llvm/test/Transforms/Attributor/nofpclass-atan2.ll
```
[mlir][arith] Handle unsigned moduli in int-range optimizations (#224933)
`DeleteTrivialRem` reads constant moduli as signed values, causing
`remui` operations with sign-bit-set moduli to be rejected. Keep the
modulus as an `APInt` and apply signedness according to the remainder
operation.
Fixes #224630
[RISCV][P-ext] Remove riscv_pmulh(u)intrinsics. (#227846)
These are redundant with the llvm.smulh/umulh intrinsics that were added
recently.
Strangely we don't have clang IRgen tests for these intrinsics/builtins,
but we do have a cross-project test for assembly.
[flang][openacc] Erase unused stack allocations in compute regions (#227807)
ACCEraseUnusedKernelAllocations only deleted unused fir.allocmem. A
dynamic fir.alloca, memref.alloca, or memref.alloc inside
acc.compute_region has the same problem: fir.declare's debug effect and
the matching free keep it alive through ordinary dead-code elimination,
and lowering turns it into a checked device malloc.
Delete those allocations when they have no uses, or when every use is
fir.freemem, memref.dealloc, a view such as fir.convert, or fir.declare.
A load, store, or other memory use still keeps the allocation.
This can happen when using stack arrays flags which replace the
fir.allocmem
[AMDGPU] Use isGFX125xOnly as the assembler predicate for tensor load/store (#227887)
The VIMAGE_TENSOR gfx1250 real instructions are only available on
GFX125x, so predicate the assembler on isGFX125xOnly rather than on the
HasTDMInsts feature.
[Attributor][NFC] rename fadd_double --> fadd_self (#227931)
I have renamed `fadd_double` to `fadd_self` in `nofpclass-fadd-fsub.ll`
to make it clear that it refers to doubling `x += x` and **not** the
`double` type.
This makes it consistent with other tests that use the name
`fadd_double` to refer to the `double` type.
[Option] Declare library command line options in TableGen (#226087)
Implement the first step of
https://discourse.llvm.org/t/rfc-declare-library-command-line-options-in-tablegen-one-struct-per-library/91877:
the TableGen backend, the cl:: dispatch, and LLVMCGData's 13 cl::opts as
the first migrated library.
A .td with an `OptionsStruct` def declares a library's options with
`BoolField` (`-x`, `-x=<bool>`) and `ValueField` (`-x=v`,
`-x v`). `-gen-opt-parser-defs` generates a struct with one member per
option, a `Global` instance, the option table, and `apply(const Arg &)`.
A member is named after its option (`-codegen-data-generate` sets
`codegen_data_generate`) unless the defm names it.
`cl::ParseCommandLineOptions` keeps owning argv: a static
`opt::RegisterLibraryOptions<T>` registers the struct as a
`cl::LibraryOptions`, and an argument naming none of cl::'s options is
dispatched to the library that declares it. `-help-hidden` lists library
options (`let Hidden = 0 in` also lists them in `-help`),
[7 lines not shown]
[InstCombine] Fold select of pow into select of the differing operand
When both arms of a select are calls to `llvm.pow` with one use that
differ in exactly one operand, sink the select into that operand:
$$
\mathrm{select}(c,\ x^{y},\ x^{z}) \rightarrow x^{\mathrm{select}(c,\ y,\ z)}
$$
$$
\mathrm{select}(c,\ x^{z},\ y^{z}) \rightarrow \mathrm{select}(c,\ x,\ y)^{z}
$$
This removes one `pow` call. FMF are intersected and debug locations
are merged. The transform is skipped when a differing operand is a
constant, since that may enable a cheaper lowering (e.g.
$x^2 \rightarrow x \cdot x$).
[InstCombine] Add tests for select of pow with one differing operand
Add tests for select between two `llvm.pow` calls whose operands differ in
exactly one position, covering scalar and vector types, FMF, !prof
metadata, constant operands, and negative cases (multiple differing
operands, extra uses, mismatched intrinsics).
[flang-rt] Consider NaN and the signedness of zero for PRODUCT (#226918)
Currently, the result of PRODUCT can be calculated in three places:
1. Constant folding in Semantics
2. The runtime library
3. Inlined code
However, only the runtime ignores NaN and the signedness of zero. This patch
fixes this discrepancy.
Fixes #211437
---------
Co-authored-by: Eugene Epshteyn <eepshteyn at nvidia.com>
[NVPTX] Properly support acquire/release/acq_rel atomics pre-SM70 by emitting membars (#222449)
acquire/release/acq_rel are not supported semantics on pre-SM70 atomics.
We need to implement these by relaxing the atomic to "relaxed" and then
surounding it with membars. This makes the atomic sequentially
consistent, which is a superset of acquire-release behavior. We already
do this for pre-SM70 atomic that cmpxchg expand, just not for atomics
that otherwise are natively supported.
[mlir][sparse] Register bufferization dialect in stage pass (#227550)
StageSparseOperations may create bufferization.dealloc_tensor while
staging conversions. Register the dialect as a pass dependency and add a
standalone regression test.
[mlir][bufferization] Make equivalent result removal order-independent (#227752)
After bufferization, in-place tensor results commonly become memref
results that are equivalent to function arguments. These results are
redundant and should be removed before later calling-convention
conversions such as buffer-results-to-out-params.
DropEquivalentBufferResults currently visits each function once in
module order. If a caller precedes its callee, rewriting the callee can
expose an equivalent result in the caller after it has already been
visited. This makes the transformation depend on function order and may
require running the pass repeatedly.
Use a caller-driven worklist to revisit affected functions until no more
results can be dropped. Refresh stored call operations after rewriting
them so recursive call graphs can also converge safely. Termination is
guaranteed because each propagation step is triggered by removing a
function result.
[4 lines not shown]
[Github] Automatically derive python version in release-binaries (#227795)
Makes updating the python version easier, especially for renovate.
Assisted By: LLM
[MLIR][Bufferization] Fix IdentityLayoutMap allocation at function boundaries (#227253)
Resolves silent data corruption when passing subviews across function
boundaries under `LayoutMapOption::IdentityLayoutMap`.
### The Problem
When the `IdentityLayoutMap` option is specified for function boundary
bufferization, all function parameters are expected to have a fully
contiguous, zero-offset layout. However, if a caller passes a non-unit
stride or offset view (e.g. the result of `tensor.extract_slice`), the
bufferization pass incorrectly lowered this to a `memref.cast`.
Since `memref.cast` strips layout metadata but leaves the base pointer
unchanged, this caused silent wrong-value loads in the callee (reading
from offset 0 regardless of the actual dynamic offset).
### The Solution
This patch intercepts the operand materialization logic in
`FuncBufferizableOpInterfaceImpl.cpp` (specifically during `CallOp`
[19 lines not shown]
[mlir][ValueBounds] Skip analysis for identical slice components (#226894)
One-Shot Bufferize repeatedly compares subset slices whose offsets,
sizes, and strides often reuse the same SSA values and attributes. Avoid
constructing a ValueBounds constraint set when the two OpFoldResults are
already identical, while preserving the existing solver fallback for
distinct values.
[CIR][SYCL] Enable relocatable device code for SYCL (#226596)
Emit `sycl_external` functions with sycl-module-id, allow -fgpu-rdc
mangling, and embed offload objects in the host.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[lldb-dap][test] Let tests run under both stdio and server adapter modes (#227435)
Add create_debug_adapter(), which picks stdio or server mode based on
self.run_as_server. This allows tests to run under both modes when
toggling `LLDBDAP_RUN_AS_SERVER`, rather than being pinned to stdio.
[mlir][linalg] Document and diagnose pack/unpack memref limits (#225773)
Scoped down from the [original
RFC](https://discourse.llvm.org/t/rfc-transformation-support-for-linalg-pack-linalg-unpack-on-memrefs/91832)
per discussion in #225650.
- Replace/add the `// TODO: Support Memref Pack/UnPackOp...` comment
across all sites with a comment stating the actual invariant, pointing
to #225650 for the reasoning.
- Document the invariant in the `Linalg_PackOp`/`Linalg_UnPackOp`
descriptions.
- Emit a dedicated diagnostic from `structured.pack`, `lower_pack`,
`lower_unpack`, and tiling when the target has memref operands, instead
of a generic/silent failure.
- Add test coverage for the new diagnostics, previously untested.
---------
Signed-off-by: Federico Bruzzone <federico.bruzzone.i at gmail.com>
[SlotIndexes] Add queries for stale indexes
An erased instruction leaves its index list entry in place, making the
index indistinguishable from a block boundary entry. Add
isBlockBoundaryIndex() and isStaleIndex() to tell the two apart, and
canonicalizeIndex() to resolve a stale index to the closest preceding
instruction's register slot, or the block start if none survives.
NFC. No caller yet. LiveDebugVariables is next.