[flang][OpenACC] Preserve DO CONCURRENT independence in kernels loops (#227775)
`DO CONCURRENT` asserts that its iterations may execute independently.
When it is directly associated with a combined OpenACC `KERNELS LOOP`,
Flang currently lowers the loop as `auto`, unless an explicit `seq`,
`auto`, or `independent` clause is present. This patch adds a
default-enabled extension that preserves the `DO CONCURRENT`
independence assertion by lowering the loop as `independent`. This
behavior is OpenACC-conforming. Explicit loop parallelism clauses
continue to take precedence.
The extension can be disabled with:
`-fno-openacc-acc-kernels-do-concurrent-independent`
This patch also documents the extension and adds lowering tests for its
enabled and disabled behavior.
x11/glcapsviewer: allow to build with future CMake versions
Drop cmake_policy(VERSION) as cmake_minimum_required(VERSION)
command implicitly calls it too.
[LV] Use SCEV loop-uniformity for outer-loop branch legality (#199632)
This patch refactors the outer-loop vectorization branch legality checks
to reason about conditional branches directly instead of using the old
recursive inner-loop shape check.
The new check allows conditional branches when their condition is
either:
- loop-invariant with respect to the vectorized outer loop, or
- a compare whose operands are both SCEV loop-uniform with respect to
the vectorized outer loop.
Divergent conditional branches are still rejected, now with a more
specific diagnostic.
[KnownFPClass][NFC] Update ATTR values for atan2 tests (#224797)
Ran the following command since it was not run for
https://github.com/llvm/llvm-project/pull/223176
```
llvm/utils/update_test_checks.py \
--opt-binary build/bin/opt \
llvm/test/Transforms/Attributor/nofpclass-atan2.ll
```
[mlir][arith] Handle unsigned moduli in int-range optimizations (#224933)
`DeleteTrivialRem` reads constant moduli as signed values, causing
`remui` operations with sign-bit-set moduli to be rejected. Keep the
modulus as an `APInt` and apply signedness according to the remainder
operation.
Fixes #224630
Report crash-looping app containers as crashed
## Problem
A container that keeps crashing under the catalog's default `unless-stopped` restart policy is reported by Docker as `restarting`, not `exited`. The container state mapping had no case for it, so it fell through to `exited`. A multi-container app with a healthy sibling then showed as Running, and a single-container app showed as Stopped, which hid its workloads and blocked logs, upgrade and rollback.
## Solution
Map `restarting` to the existing `crashed` container state, so the app is reported as Crashed through the existing aggregation. Added a unit test for the Docker status to container state mapping, which had no coverage.
[RISCV][P-ext] Remove riscv_pmulh(u)intrinsics. (#227846)
These are redundant with the llvm.smulh/umulh intrinsics that were added
recently.
Strangely we don't have clang IRgen tests for these intrinsics/builtins,
but we do have a cross-project test for assembly.
[flang][openacc] Erase unused stack allocations in compute regions (#227807)
ACCEraseUnusedKernelAllocations only deleted unused fir.allocmem. A
dynamic fir.alloca, memref.alloca, or memref.alloc inside
acc.compute_region has the same problem: fir.declare's debug effect and
the matching free keep it alive through ordinary dead-code elimination,
and lowering turns it into a checked device malloc.
Delete those allocations when they have no uses, or when every use is
fir.freemem, memref.dealloc, a view such as fir.convert, or fir.declare.
A load, store, or other memory use still keeps the allocation.
This can happen when using stack arrays flags which replace the
fir.allocmem
sftp: be stricter in accepting paths returned by the server for
SSH_FXP_REALPATH or SSH2_FXP_READDIR replies, as these can be
used in some situations to decide the destination path for recursive
transfers.
Report and patch from Junghoon Cho
[AMDGPU] Use isGFX125xOnly as the assembler predicate for tensor load/store (#227887)
The VIMAGE_TENSOR gfx1250 real instructions are only available on
GFX125x, so predicate the assembler on isGFX125xOnly rather than on the
HasTDMInsts feature.
[Attributor][NFC] rename fadd_double --> fadd_self (#227931)
I have renamed `fadd_double` to `fadd_self` in `nofpclass-fadd-fsub.ll`
to make it clear that it refers to doubling `x += x` and **not** the
`double` type.
This makes it consistent with other tests that use the name
`fadd_double` to refer to the `double` type.
[Option] Declare library command line options in TableGen (#226087)
Implement the first step of
https://discourse.llvm.org/t/rfc-declare-library-command-line-options-in-tablegen-one-struct-per-library/91877:
the TableGen backend, the cl:: dispatch, and LLVMCGData's 13 cl::opts as
the first migrated library.
A .td with an `OptionsStruct` def declares a library's options with
`BoolField` (`-x`, `-x=<bool>`) and `ValueField` (`-x=v`,
`-x v`). `-gen-opt-parser-defs` generates a struct with one member per
option, a `Global` instance, the option table, and `apply(const Arg &)`.
A member is named after its option (`-codegen-data-generate` sets
`codegen_data_generate`) unless the defm names it.
`cl::ParseCommandLineOptions` keeps owning argv: a static
`opt::RegisterLibraryOptions<T>` registers the struct as a
`cl::LibraryOptions`, and an argument naming none of cl::'s options is
dispatched to the library that declares it. `-help-hidden` lists library
options (`let Hidden = 0 in` also lists them in `-help`),
[7 lines not shown]
[InstCombine] Fold select of pow into select of the differing operand
When both arms of a select are calls to `llvm.pow` with one use that
differ in exactly one operand, sink the select into that operand:
$$
\mathrm{select}(c,\ x^{y},\ x^{z}) \rightarrow x^{\mathrm{select}(c,\ y,\ z)}
$$
$$
\mathrm{select}(c,\ x^{z},\ y^{z}) \rightarrow \mathrm{select}(c,\ x,\ y)^{z}
$$
This removes one `pow` call. FMF are intersected and debug locations
are merged. The transform is skipped when a differing operand is a
constant, since that may enable a cheaper lowering (e.g.
$x^2 \rightarrow x \cdot x$).
[InstCombine] Add tests for select of pow with one differing operand
Add tests for select between two `llvm.pow` calls whose operands differ in
exactly one position, covering scalar and vector types, FMF, !prof
metadata, constant operands, and negative cases (multiple differing
operands, extra uses, mismatched intrinsics).