[InstCombine] Fix profile propagation in zext-bool-add-sub.ll (#227949)
Mark the profiles for the created selects as unknown as we cannot know
anything about the distribution of the condition in the general case.
[flang][NFC] Correct a stale comment on loop reclassification
Two places weaken the classification now, so calling this one "the one
place" is out of date.
[flang] Let a directive keep the loop it owns when its body branches
A loop whose branching is confined to its body keeps its structured
form, but the construct holding it stayed Unstructured. A directive does
not merely contain such a loop, it owns it, and its lowering reads the
construct's own classification to decide whether the loop op carries its
bounds. The directive was left with a bounds-free loop that nothing
could partition, and the loop it owns became a second one nested inside.
Reclassify a directive construct once the loops it holds no longer need
it to stay Unstructured. Children are visited first, so those loops have
already been reclassified by the time the construct is reached. A
construct whose branching leaves it is untouched, as is one holding a
branch of its own.
Taking a loop over also means genFIR(DoConstruct) -- where a plain loop
folds a body whose branching stays inside it into a region -- never runs
for that loop, so fold its body through the same helper. A construct
that takes over no loop, acc data or acc parallel without a loop
[5 lines not shown]
[flang][NFC] Split the OpenACC construct lowering into two lanes (#227706)
genFIR(OpenACCConstruct) decided twice, in three places, whether the
construct it lowers is structured, and reassigned the evaluation it
works from halfway through: before the descent that evaluation is the
construct, after it the loop the directive absorbs. Everything
downstream had to know which one it was holding.
Give each form its own function and leave genFIR to choose between them.
One lane allocates the exit selector, lowers the evaluations the
construct holds, and emits the jump table; the other reads the collapse
clauses, descends to the absorbed depth, and lowers what is inside it.
The prologue and epilogue are short enough to state in both rather than
share.
[RISCV][P-ext] Prevent accidental matches in riscv_packed_simd.c. NFC (#227830)
The function name is printeded multiple times in the output. We need to
make sure we are matching an instruction mnemonic. The way other
existing test cases do this is by checking for a space after the
instruction name. We don't need to do this if the mnemonic contains a
period since those are replaced with underscores in the function name.
[CIR][EH] Fix scope for partial array cleanup (#227838)
There was a bug in CIR where if an array whose elements required
destruction was initialized with an ILE, we weren't properly closing the
EH cleanup scope after the initialization completed, so it enclosed the
rest of the function. The result was that if anything later in the
function threw an exception, it would trigger both the normal
destruction of the array and the leftover EH "partial" cleanup, leading
to a double-free.
This change fixes that problem by introducing a CleanupDeactivationScope
around the init list processing (where we already had a MissingFeature
marker saying this was needed). The EH cleanup scope is now closed when
the CleanupDeactivationScope object goes out of scope.
This change also caused some observable changes to existing tests where
we were previously behaving incorrectly.
Assisted-by: Cursor / various models
[flang][OpenACC] Preserve DO CONCURRENT independence in kernels loops (#227775)
`DO CONCURRENT` asserts that its iterations may execute independently.
When it is directly associated with a combined OpenACC `KERNELS LOOP`,
Flang currently lowers the loop as `auto`, unless an explicit `seq`,
`auto`, or `independent` clause is present. This patch adds a
default-enabled extension that preserves the `DO CONCURRENT`
independence assertion by lowering the loop as `independent`. This
behavior is OpenACC-conforming. Explicit loop parallelism clauses
continue to take precedence.
The extension can be disabled with:
`-fno-openacc-acc-kernels-do-concurrent-independent`
This patch also documents the extension and adds lowering tests for its
enabled and disabled behavior.
[LV] Use SCEV loop-uniformity for outer-loop branch legality (#199632)
This patch refactors the outer-loop vectorization branch legality checks
to reason about conditional branches directly instead of using the old
recursive inner-loop shape check.
The new check allows conditional branches when their condition is
either:
- loop-invariant with respect to the vectorized outer loop, or
- a compare whose operands are both SCEV loop-uniform with respect to
the vectorized outer loop.
Divergent conditional branches are still rejected, now with a more
specific diagnostic.
[KnownFPClass][NFC] Update ATTR values for atan2 tests (#224797)
Ran the following command since it was not run for
https://github.com/llvm/llvm-project/pull/223176
```
llvm/utils/update_test_checks.py \
--opt-binary build/bin/opt \
llvm/test/Transforms/Attributor/nofpclass-atan2.ll
```
[mlir][arith] Handle unsigned moduli in int-range optimizations (#224933)
`DeleteTrivialRem` reads constant moduli as signed values, causing
`remui` operations with sign-bit-set moduli to be rejected. Keep the
modulus as an `APInt` and apply signedness according to the remainder
operation.
Fixes #224630
[RISCV][P-ext] Remove riscv_pmulh(u)intrinsics. (#227846)
These are redundant with the llvm.smulh/umulh intrinsics that were added
recently.
Strangely we don't have clang IRgen tests for these intrinsics/builtins,
but we do have a cross-project test for assembly.
[flang][openacc] Erase unused stack allocations in compute regions (#227807)
ACCEraseUnusedKernelAllocations only deleted unused fir.allocmem. A
dynamic fir.alloca, memref.alloca, or memref.alloc inside
acc.compute_region has the same problem: fir.declare's debug effect and
the matching free keep it alive through ordinary dead-code elimination,
and lowering turns it into a checked device malloc.
Delete those allocations when they have no uses, or when every use is
fir.freemem, memref.dealloc, a view such as fir.convert, or fir.declare.
A load, store, or other memory use still keeps the allocation.
This can happen when using stack arrays flags which replace the
fir.allocmem
[AMDGPU] Use isGFX125xOnly as the assembler predicate for tensor load/store (#227887)
The VIMAGE_TENSOR gfx1250 real instructions are only available on
GFX125x, so predicate the assembler on isGFX125xOnly rather than on the
HasTDMInsts feature.
[Attributor][NFC] rename fadd_double --> fadd_self (#227931)
I have renamed `fadd_double` to `fadd_self` in `nofpclass-fadd-fsub.ll`
to make it clear that it refers to doubling `x += x` and **not** the
`double` type.
This makes it consistent with other tests that use the name
`fadd_double` to refer to the `double` type.
[Option] Declare library command line options in TableGen (#226087)
Implement the first step of
https://discourse.llvm.org/t/rfc-declare-library-command-line-options-in-tablegen-one-struct-per-library/91877:
the TableGen backend, the cl:: dispatch, and LLVMCGData's 13 cl::opts as
the first migrated library.
A .td with an `OptionsStruct` def declares a library's options with
`BoolField` (`-x`, `-x=<bool>`) and `ValueField` (`-x=v`,
`-x v`). `-gen-opt-parser-defs` generates a struct with one member per
option, a `Global` instance, the option table, and `apply(const Arg &)`.
A member is named after its option (`-codegen-data-generate` sets
`codegen_data_generate`) unless the defm names it.
`cl::ParseCommandLineOptions` keeps owning argv: a static
`opt::RegisterLibraryOptions<T>` registers the struct as a
`cl::LibraryOptions`, and an argument naming none of cl::'s options is
dispatched to the library that declares it. `-help-hidden` lists library
options (`let Hidden = 0 in` also lists them in `-help`),
[7 lines not shown]
[InstCombine] Fold select of pow into select of the differing operand
When both arms of a select are calls to `llvm.pow` with one use that
differ in exactly one operand, sink the select into that operand:
$$
\mathrm{select}(c,\ x^{y},\ x^{z}) \rightarrow x^{\mathrm{select}(c,\ y,\ z)}
$$
$$
\mathrm{select}(c,\ x^{z},\ y^{z}) \rightarrow \mathrm{select}(c,\ x,\ y)^{z}
$$
This removes one `pow` call. FMF are intersected and debug locations
are merged. The transform is skipped when a differing operand is a
constant, since that may enable a cheaper lowering (e.g.
$x^2 \rightarrow x \cdot x$).
[InstCombine] Add tests for select of pow with one differing operand
Add tests for select between two `llvm.pow` calls whose operands differ in
exactly one position, covering scalar and vector types, FMF, !prof
metadata, constant operands, and negative cases (multiple differing
operands, extra uses, mismatched intrinsics).
[flang-rt] Consider NaN and the signedness of zero for PRODUCT (#226918)
Currently, the result of PRODUCT can be calculated in three places:
1. Constant folding in Semantics
2. The runtime library
3. Inlined code
However, only the runtime ignores NaN and the signedness of zero. This patch
fixes this discrepancy.
Fixes #211437
---------
Co-authored-by: Eugene Epshteyn <eepshteyn at nvidia.com>