LLVM/project 15aa5fe — llvm/lib/Target/AMDGPU AMDGPUInstCombineIntrinsic.cpp, llvm/test/Transforms/InstCombine/AMDGPU mbcnt.ll

Revert peephole optimization for folding a zero mask to base in mbcnt
DeltaFile
+0-14llvm/test/Transforms/InstCombine/AMDGPU/mbcnt.ll
+0-4llvm/lib/Target/AMDGPU/AMDGPUInstCombineIntrinsic.cpp
+0-182 files

LLVM/project 7870b69 — llvm/lib/Analysis ValueTracking.cpp, llvm/lib/Target/AMDGPU AMDGPUInstCombineIntrinsic.cpp

[AMDGPU] Fix amdgcn.mbcnt known bits conflicting with the range attribute
DeltaFile
+20-6llvm/test/Transforms/InstCombine/AMDGPU/mbcnt.ll
+3-2llvm/lib/Analysis/ValueTracking.cpp
+4-0llvm/lib/Target/AMDGPU/AMDGPUInstCombineIntrinsic.cpp
+27-83 files

LLVM/project 9152ed6 — llvm/test/Transforms/InstCombine/AMDGPU mbcnt.ll

Remove confusing comments about other add/shfl cases
DeltaFile
+1-4llvm/test/Transforms/InstCombine/AMDGPU/mbcnt.ll
+1-41 files

LLVM/project f3db644 — llvm/test/Transforms/InstCombine/AMDGPU mbcnt.ll

Precommit test cases for mbcnt intrinsic calls being incorrectly folded to 1
DeltaFile
+49-0llvm/test/Transforms/InstCombine/AMDGPU/mbcnt.ll
+49-01 files

LLVM/project ccac700 — llvm/lib/Transforms/InstCombine InstCombineAddSub.cpp, llvm/test/Transforms/InstCombine zext-bool-add-sub.ll

[InstCombine] Fix profile propagation in zext-bool-add-sub.ll (#227949)

Mark the profiles for the created selects as unknown as we cannot know
anything about the distribution of the condition in the general case.
DeltaFile
+15-4llvm/lib/Transforms/InstCombine/InstCombineAddSub.cpp
+10-5llvm/test/Transforms/InstCombine/zext-bool-add-sub.ll
+0-1llvm/utils/profcheck-xfail.txt
+25-103 files

LLVM/project 27ade7b — flang/lib/Lower PFTBuilder.cpp

[flang][NFC] Correct a stale comment on loop reclassification

Two places weaken the classification now, so calling this one "the one
place" is out of date.
DeltaFile
+1-2flang/lib/Lower/PFTBuilder.cpp
+1-21 files

LLVM/project 3d302b4 — flang/lib/Lower PFTBuilder.cpp, flang/test/Lower/OpenACC acc-unstructured.f90 acc-unstructured-combined-construct.f90

[flang] Let a directive keep the loop it owns when its body branches

A loop whose branching is confined to its body keeps its structured
form, but the construct holding it stayed Unstructured. A directive does
not merely contain such a loop, it owns it, and its lowering reads the
construct's own classification to decide whether the loop op carries its
bounds. The directive was left with a bounds-free loop that nothing
could partition, and the loop it owns became a second one nested inside.

Reclassify a directive construct once the loops it holds no longer need
it to stay Unstructured. Children are visited first, so those loops have
already been reclassified by the time the construct is reached. A
construct whose branching leaves it is untouched, as is one holding a
branch of its own.

Taking a loop over also means genFIR(DoConstruct) -- where a plain loop
folds a body whose branching stays inside it into a region -- never runs
for that loop, so fold its body through the same helper. A construct
that takes over no loop, acc data or acc parallel without a loop

    [5 lines not shown]
DeltaFile
+160-0flang/test/Lower/OpenACC/acc-unstructured-loop-construct.f90
+8-124flang/test/Lower/OpenACC/Todo/acc-unstructured-loop-construct.f90
+66-0flang/test/Lower/OpenACC/acc-directive-loop-bounds.f90
+45-0flang/test/Lower/OpenACC/acc-unstructured-combined-construct.f90
+23-16flang/test/Lower/OpenACC/acc-unstructured.f90
+33-4flang/lib/Lower/PFTBuilder.cpp
+335-1444 files not shown
+387-17410 files

LLVM/project 6ad1b7b — flang/lib/Lower Bridge.cpp

[flang][NFC] Split the OpenACC construct lowering into two lanes (#227706)

genFIR(OpenACCConstruct) decided twice, in three places, whether the
construct it lowers is structured, and reassigned the evaluation it
works from halfway through: before the descent that evaluation is the
construct, after it the loop the directive absorbs. Everything
downstream had to know which one it was holding.

Give each form its own function and leave genFIR to choose between them.
One lane allocates the exit selector, lowers the evaluations the
construct holds, and emits the jump table; the other reads the collapse
clauses, descends to the absorbed depth, and lowers what is inside it.
The prologue and epilogue are short enough to state in both rather than
share.
DeltaFile
+149-101flang/lib/Lower/Bridge.cpp
+149-1011 files

LLVM/project ba8dc06 — llvm/test/tools/llubi intr_int_arith.ll, llvm/tools/llubi/lib Interpreter.cpp

[llubi] Add support for pext/pdep (#227803)
DeltaFile
+18-0llvm/test/tools/llubi/intr_int_arith.ll
+8-0llvm/tools/llubi/lib/Interpreter.cpp
+26-02 files

LLVM/project 0eb3614 — cross-project-tests/intrinsic-header-tests riscv_packed_simd.c

[RISCV][P-ext] Prevent accidental matches in riscv_packed_simd.c. NFC (#227830)

The function name is printeded multiple times in the output. We need to
make sure we are matching an instruction mnemonic. The way other
existing test cases do this is by checking for a space after the
instruction name. We don't need to do this if the mnemonic contains a
period since those are replaced with underscores in the function name.
DeltaFile
+70-74cross-project-tests/intrinsic-header-tests/riscv_packed_simd.c
+70-741 files

LLVM/project 1f52221 — clang/include/clang/CIR MissingFeatures.h, clang/lib/CIR/CodeGen CIRGenExprAggregate.cpp CIRGenDecl.cpp

[CIR][EH] Fix scope for partial array cleanup (#227838)

There was a bug in CIR where if an array whose elements required
destruction was initialized with an ILE, we weren't properly closing the
EH cleanup scope after the initialization completed, so it enclosed the
rest of the function. The result was that if anything later in the
function threw an exception, it would trigger both the normal
destruction of the array and the leftover EH "partial" cleanup, leading
to a double-free.

This change fixes that problem by introducing a CleanupDeactivationScope
around the init list processing (where we already had a MissingFeature
marker saying this was needed). The EH cleanup scope is now closed when
the CleanupDeactivationScope object goes out of scope.

This change also caused some observable changes to existing tests where
we were previously behaving incorrectly.

Assisted-by: Cursor / various models
DeltaFile
+170-81clang/test/CIR/CodeGen/partial-array-cleanup.cpp
+6-16clang/test/CIR/CodeGen/new-array-in-ternary.cpp
+3-3clang/lib/CIR/CodeGen/CIRGenExprAggregate.cpp
+3-3clang/lib/CIR/CodeGen/CIRGenDecl.cpp
+0-1clang/include/clang/CIR/MissingFeatures.h
+182-1045 files

LLVM/project 76e096c — cross-project-tests CMakeLists.txt

[cross-project-tests] Add llvm-readobj as a dependency (#227950)

cross-project-tests/riscv/lto-inline-asm-abi.c uses it

Hopefully this will fix failures like
https://github.com/llvm/llvm-project/actions/runs/36810997644
DeltaFile
+1-0cross-project-tests/CMakeLists.txt
+1-01 files

LLVM/project 9a66279 — clang/lib/CIR/Dialect/IR CIRTypes.cpp, clang/test/CIR/CodeGenSYCL address-space-lang.cpp

[CIR][SYCL] Map SYCL address spaces to CIR language address spaces (#226598)
DeltaFile
+104-0clang/test/CIR/CodeGenSYCL/address-space-lang.cpp
+8-8clang/lib/CIR/Dialect/IR/CIRTypes.cpp
+112-82 files

LLVM/project 6f761a9 — flang/docs OpenACC-extensions.md, flang/include/flang/Lower LoweringOptions.def

[flang][OpenACC] Preserve DO CONCURRENT independence in kernels loops (#227775)

`DO CONCURRENT` asserts that its iterations may execute independently.
When it is directly associated with a combined OpenACC `KERNELS LOOP`,
Flang currently lowers the loop as `auto`, unless an explicit `seq`,
`auto`, or `independent` clause is present. This patch adds a
default-enabled extension that preserves the `DO CONCURRENT`
independence assertion by lowering the loop as `independent`. This
behavior is OpenACC-conforming. Explicit loop parallelism clauses
continue to take precedence.

The extension can be disabled with:

`-fno-openacc-acc-kernels-do-concurrent-independent`

This patch also documents the extension and adds lowering tests for its
enabled and disabled behavior.
DeltaFile
+16-4flang/lib/Lower/OpenACC.cpp
+20-0flang/test/Lower/OpenACC/acc-kernels-do-concurrent-independent-flag.f90
+11-0flang/docs/OpenACC-extensions.md
+6-0flang/lib/Frontend/CompilerInvocation.cpp
+6-0flang/include/flang/Lower/LoweringOptions.def
+2-2flang/test/Lower/OpenACC/acc-do-concurrent-locality.f90
+61-61 files not shown
+64-67 files

LLVM/project edb0ee6 — llvm/cmake/modules MLGOLower.cmake, llvm/lib/Analysis CMakeLists.txt

[mlgo] Allow passing pre-emitc-ed models
DeltaFile
+79-58llvm/cmake/modules/MLGOLower.cmake
+98-0llvm/lib/Analysis/models/inline-oz-test-model.inc
+49-0llvm/lib/Analysis/models/regalloc-eviction-test-model.inc
+13-12llvm/lib/CodeGen/CMakeLists.txt
+13-12llvm/lib/Analysis/CMakeLists.txt
+5-11llvm/unittests/Analysis/MLGOUtilsTest.cpp
+257-9320 files not shown
+321-19226 files

LLVM/project c4ff7a7 — llvm/lib/Transforms/Vectorize LoopVectorizationLegality.cpp, llvm/test/Transforms/LoopVectorize outer_loop_early_exit.ll explicit_outer_nonuniform_inner.ll

[LV] Use SCEV loop-uniformity for outer-loop branch legality (#199632)

This patch refactors the outer-loop vectorization branch legality checks
to reason about conditional branches directly instead of using the old
recursive inner-loop shape check.

The new check allows conditional branches when their condition is
either:

- loop-invariant with respect to the vectorized outer loop, or
- a compare whose operands are both SCEV loop-uniform with respect to
the vectorized outer loop.

Divergent conditional branches are still rejected, now with a more
specific diagnostic.
DeltaFile
+30-108llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp
+58-2llvm/test/Transforms/LoopVectorize/explicit_outer_uniform_diverg_branch.ll
+34-0llvm/test/Transforms/LoopVectorize/outer_loop_inner_loop_exits.ll
+3-3llvm/test/Transforms/LoopVectorize/explicit_outer_nonuniform_inner.ll
+1-1llvm/test/Transforms/LoopVectorize/outer_loop_early_exit.ll
+126-1145 files

LLVM/project 419fa6a — llvm/cmake/modules MLGOLower.cmake, llvm/lib/Analysis CMakeLists.txt

[mlgo] Allow passing pre-emitc-ed models
DeltaFile
+79-58llvm/cmake/modules/MLGOLower.cmake
+56-0llvm/lib/Analysis/models/inline-oz-test-model.inc
+49-0llvm/lib/Analysis/models/regalloc-eviction-test-model.inc
+13-12llvm/lib/CodeGen/CMakeLists.txt
+13-12llvm/lib/Analysis/CMakeLists.txt
+1-14llvm/lib/CodeGen/MLRegAllocEvictAdvisor.cpp
+211-9620 files not shown
+278-18826 files

LLVM/project 3919526 — llvm/test/Transforms/Attributor nofpclass-atan2.ll

[KnownFPClass][NFC] Update ATTR values for atan2 tests (#224797)

Ran the following command since it was not run for
https://github.com/llvm/llvm-project/pull/223176
```
llvm/utils/update_test_checks.py \
  --opt-binary build/bin/opt \
  llvm/test/Transforms/Attributor/nofpclass-atan2.ll
```
DeltaFile
+48-48llvm/test/Transforms/Attributor/nofpclass-atan2.ll
+48-481 files

LLVM/project 78b41db — mlir/lib/Dialect/Arith/Transforms IntRangeOptimizations.cpp, mlir/test/Dialect/Arith int-range-opts.mlir

[mlir][arith] Handle unsigned moduli in int-range optimizations (#224933)

`DeleteTrivialRem` reads constant moduli as signed values, causing
`remui` operations with sign-bit-set moduli to be rejected. Keep the
modulus as an `APInt` and apply signedness according to the remainder
operation.

Fixes #224630
DeltaFile
+39-0mlir/test/Dialect/Arith/int-range-opts.mlir
+18-11mlir/lib/Dialect/Arith/Transforms/IntRangeOptimizations.cpp
+57-112 files

LLVM/project 8bac179 — utils/bazel/llvm-project-overlay/flang/lib/Optimizer/OpenACC/Transforms BUILD.bazel

[Bazel] Fixes 6a7945b (#227939)

This fixes 6a7945b744e251cd6c843521a5e918e364a49dc5 (#227807).

Buildkite error link:
https://buildkite.com/llvm-project/upstream-bazel/builds?commit=6a7945b744e251cd6c843521a5e918e364a49dc5

Co-authored-by: Google Bazel Bot <google-bazel-bot at google.com>
DeltaFile
+1-0utils/bazel/llvm-project-overlay/flang/lib/Optimizer/OpenACC/Transforms/BUILD.bazel
+1-01 files

LLVM/project 640e1c6 — clang/lib/CodeGen/TargetBuiltins RISCV.cpp, llvm/include/llvm/IR IntrinsicsRISCV.td

[RISCV][P-ext] Remove riscv_pmulh(u)intrinsics. (#227846)

These are redundant with the llvm.smulh/umulh intrinsics that were added
recently.

Strangely we don't have clang IRgen tests for these intrinsics/builtins,
but we do have a cross-project test for assembly.
DeltaFile
+0-10llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+4-4llvm/test/CodeGen/RISCV/rvp-simd-64.ll
+2-2llvm/test/CodeGen/RISCV/rvp-simd-32.ll
+2-2clang/lib/CodeGen/TargetBuiltins/RISCV.cpp
+0-2llvm/include/llvm/IR/IntrinsicsRISCV.td
+8-205 files

LLVM/project 6a7945b — flang/include/flang/Optimizer/OpenACC Passes.td, flang/lib/Optimizer/OpenACC/Transforms CMakeLists.txt ACCEraseUnusedKernelAllocations.cpp

[flang][openacc] Erase unused stack allocations in compute regions (#227807)

ACCEraseUnusedKernelAllocations only deleted unused fir.allocmem. A
dynamic fir.alloca, memref.alloca, or memref.alloc inside
acc.compute_region has the same problem: fir.declare's debug effect and
the matching free keep it alive through ordinary dead-code elimination,
and lowering turns it into a checked device malloc.
Delete those allocations when they have no uses, or when every use is
fir.freemem, memref.dealloc, a view such as fir.convert, or fir.declare.
A load, store, or other memory use still keeps the allocation.

This can happen when using stack arrays flags which replace the
fir.allocmem
DeltaFile
+39-19flang/lib/Optimizer/OpenACC/Transforms/ACCEraseUnusedKernelAllocations.cpp
+43-0flang/test/Fir/OpenACC/acc-erase-unused-kernel-allocations.mlir
+10-8flang/include/flang/Optimizer/OpenACC/Passes.td
+1-0flang/lib/Optimizer/OpenACC/Transforms/CMakeLists.txt
+93-274 files

LLVM/project 15d8013 — llvm/lib/Target/AMDGPU MIMGInstructions.td

[AMDGPU] Use isGFX125xOnly as the assembler predicate for tensor load/store (#227887)

The VIMAGE_TENSOR gfx1250 real instructions are only available on
GFX125x, so predicate the assembler on isGFX125xOnly rather than on the
HasTDMInsts feature.
DeltaFile
+1-1llvm/lib/Target/AMDGPU/MIMGInstructions.td
+1-11 files

LLVM/project d87168d — llvm/test/Transforms/Attributor nofpclass-fadd-fsub.ll

[Attributor][NFC] rename fadd_double --> fadd_self (#227931)

I have renamed `fadd_double` to `fadd_self` in `nofpclass-fadd-fsub.ll`
to make it clear that it refers to doubling `x += x` and **not** the
`double` type.

This makes it consistent with other tests that use the name
`fadd_double` to refer to the `double` type.
DeltaFile
+88-88llvm/test/Transforms/Attributor/nofpclass-fadd-fsub.ll
+88-881 files

LLVM/project 14304b9 — clang/test/CodeGen/RISCV rvp-intrinsics.c

[RISCV] Re-generate rvp-intrinsics.c using common prefix. NFC (#227836)

The common prefix was added while widening unzip was in review.
DeltaFile
+168-360clang/test/CodeGen/RISCV/rvp-intrinsics.c
+168-3601 files

LLVM/project 17fdd96 — llvm/include/llvm/Option LibraryOptions.h, llvm/lib/CGData StableFunctionMap.cpp

[Option] Declare library command line options in TableGen (#226087)

Implement the first step of
https://discourse.llvm.org/t/rfc-declare-library-command-line-options-in-tablegen-one-struct-per-library/91877:
the TableGen backend, the cl:: dispatch, and LLVMCGData's 13 cl::opts as
the first migrated library.

A .td with an `OptionsStruct` def declares a library's options with
`BoolField` (`-x`, `-x=<bool>`) and `ValueField` (`-x=v`,
`-x v`). `-gen-opt-parser-defs` generates a struct with one member per
option, a `Global` instance, the option table, and `apply(const Arg &)`.
A member is named after its option (`-codegen-data-generate` sets
`codegen_data_generate`) unless the defm names it.

`cl::ParseCommandLineOptions` keeps owning argv: a static
`opt::RegisterLibraryOptions<T>` registers the struct as a
`cl::LibraryOptions`, and an argument naming none of cl::'s options is
dispatched to the library that declares it. `-help-hidden` lists library
options (`let Hidden = 0 in` also lists them in `-help`),

    [7 lines not shown]
DeltaFile
+149-2llvm/utils/TableGen/OptionParserEmitter.cpp
+141-0llvm/unittests/Support/CommandLineTest.cpp
+115-0llvm/unittests/Option/LibraryOptionsTest.cpp
+95-0llvm/include/llvm/Option/LibraryOptions.h
+92-2llvm/lib/Support/CommandLine.cpp
+10-47llvm/lib/CGData/StableFunctionMap.cpp
+602-5118 files not shown
+919-8624 files

LLVM/project 8b1d8d1 — llvm/lib/Transforms/InstCombine InstCombineSelect.cpp, llvm/test/Transforms/InstCombine intrinsic-select.ll

[InstCombine] Fold select of pow into select of the differing operand

When both arms of a select are calls to `llvm.pow` with one use that
differ in exactly one operand, sink the select into that operand:

$$
\mathrm{select}(c,\ x^{y},\ x^{z}) \rightarrow x^{\mathrm{select}(c,\ y,\ z)}
$$

$$
\mathrm{select}(c,\ x^{z},\ y^{z}) \rightarrow \mathrm{select}(c,\ x,\ y)^{z}
$$

This removes one `pow` call. FMF are intersected and debug locations
are merged. The transform is skipped when a differing operand is a
constant, since that may enable a cheaper lowering (e.g.
$x^2 \rightarrow x \cdot x$).
DeltaFile
+24-32llvm/test/Transforms/InstCombine/intrinsic-select.ll
+35-0llvm/lib/Transforms/InstCombine/InstCombineSelect.cpp
+59-322 files

LLVM/project 75e37b3 — llvm/lib/Transforms/InstCombine InstCombineSelect.cpp, llvm/test/Transforms/InstCombine intrinsic-select.ll

Update for comments
DeltaFile
+15-20llvm/test/Transforms/InstCombine/intrinsic-select.ll
+4-0llvm/lib/Transforms/InstCombine/InstCombineSelect.cpp
+19-202 files

LLVM/project 9747c7f — llvm/test/CodeGen/AMDGPU simplify-libcalls.ll amdgpu-simplify-libcall-pow.ll

Fix lit test issue.
DeltaFile
+18-36llvm/test/CodeGen/AMDGPU/amdgpu-simplify-libcall-rootn.ll
+9-11llvm/test/CodeGen/AMDGPU/amdgpu-simplify-libcall-pow.ll
+2-4llvm/test/CodeGen/AMDGPU/simplify-libcalls.ll
+29-513 files

LLVM/project 6fa6d14 — llvm/lib/Transforms/InstCombine InstCombineSelect.cpp, llvm/test/Transforms/InstCombine select-fcmp-fmul-zero-absorbing-value.ll

Use TriviallyVectorizable
DeltaFile
+10-4llvm/lib/Transforms/InstCombine/InstCombineSelect.cpp
+2-3llvm/test/Transforms/InstCombine/select-fcmp-fmul-zero-absorbing-value.ll
+12-72 files