LLVM/project 446c201 — llvm/lib/CodeGen/SelectionDAG DAGCombiner.cpp, llvm/test/CodeGen/X86 fpenv.ll

[DAGCombiner] Don't fold GET_FPENV_MEM into a store whose address depends on it (#228839)

visitGET_FPENV_MEM folds `get_fpenv_mem tmp; load tmp; store -> p` into
`get_fpenv_mem p`, and the new node replaces the original one. When the
store address depends on the original node, the DAG becomes cyclic. This
happens for `ret i256 @llvm.get.fpenv.i256()`: the hidden sret pointer is
read with a CopyFromReg that is chained after the load from the stack
temporary. Type legalization then stops with "Operand not processed?",
and
release builds crash.

Skip the fold when the store address has the original node as a
predecessor.

Fixes #228828

Co-Authored-By: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+100-0llvm/test/CodeGen/X86/fpenv.ll
+6-0llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp
+106-02 files

LLVM/project 669dd4b — llvm/test/Transforms/SLPVectorizer/AArch64 bitpacked-aggregate.ll

[SLP][NFC]Add a test for big endian target per review, NFC



Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/229132
DeltaFile
+42-0llvm/test/Transforms/SLPVectorizer/AArch64/bitpacked-aggregate.ll
+42-01 files

LLVM/project b273349 — llvm/test/Transforms/SLPVectorizer/AArch64 bitpacked-aggregate.ll

[𝘀𝗽𝗿] initial version

Created using spr 1.3.7
DeltaFile
+42-0llvm/test/Transforms/SLPVectorizer/AArch64/bitpacked-aggregate.ll
+42-01 files

LLVM/project 2835f7b — llvm/test/Transforms/InstCombine bitinsert-bitextract.ll

[InstCombine] Pre-commit tests for full-width `bitinsert`/`bitextract`
DeltaFile
+152-0llvm/test/Transforms/InstCombine/bitinsert-bitextract.ll
+152-01 files

LLVM/project 4fe36ce — llvm/lib/Transforms/InstCombine InstCombineInternal.h InstructionCombining.cpp, llvm/test/Transforms/InstCombine bitinsert-bitextract.ll

[InstCombine] Fold full-width `bitinsert` and `bitextract` to `bitcast`
DeltaFile
+33-0llvm/lib/Transforms/InstCombine/InstructionCombining.cpp
+8-10llvm/test/Transforms/InstCombine/bitinsert-bitextract.ll
+2-0llvm/lib/Transforms/InstCombine/InstCombineInternal.h
+43-103 files

OPNSense/core 2320a83 — src/opnsense/scripts/syslog lockout_handler

lockout: type safety fix.
DeltaFile
+1-1src/opnsense/scripts/syslog/lockout_handler
+1-11 files

LLVM/project f59949b — llvm/lib/Target/AMDGPU SMInstructions.td VOPInstructions.td, llvm/lib/Target/AMDGPU/Utils AMDGPUBaseInfo.cpp

[AMDGPU] Turn tablegen tables with single bool into filtered lists. NFC (#227746)
DeltaFile
+6-18llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.cpp
+9-6llvm/lib/Target/AMDGPU/VOPInstructions.td
+3-2llvm/lib/Target/AMDGPU/SMInstructions.td
+18-263 files

LLVM/project 3c32a0e — llvm/lib/IR Instructions.cpp, llvm/test/Transforms/GVN pr207532.ll

[IR] Fold `bitcast` pairs through bytes with `ptrtoint`/`inttoptr`
DeltaFile
+22-13llvm/test/Transforms/InstCombine/byte-cast-pairs.ll
+11-8llvm/lib/IR/Instructions.cpp
+1-1llvm/test/Transforms/InstSimplify/byte-cast-pairs.ll
+1-1llvm/test/Transforms/GVN/pr207532.ll
+35-234 files

LLVM/project c8d992a — llvm/docs AMDGPUUsage.rst, llvm/lib/Target/AMDGPU AMDGPULowerIntrinsics.cpp SIMemoryLegalizer.cpp

Re-apply demotion logic

This reverts commit 992e6130ed441b82490f9a6cf01f5bba38293e4e.
DeltaFile
+1,324-824llvm/test/CodeGen/AMDGPU/memory-legalizer-single-wave-workgroup-memops.ll
+33-27llvm/test/CodeGen/AMDGPU/memory-legalizer-single-wave-workgroup-lds-dma.ll
+20-7llvm/lib/Target/AMDGPU/SIMemoryLegalizer.cpp
+10-0llvm/docs/AMDGPUUsage.rst
+3-4llvm/test/CodeGen/AMDGPU/memory-legalizer-single-wave-workgroup-lds-dma-async.ll
+2-4llvm/lib/Target/AMDGPU/AMDGPULowerIntrinsics.cpp
+1,392-8662 files not shown
+1,400-8668 files

LLVM/project bcaf89d — llvm/test/Transforms/LoopVectorize fold-epilogue-tail.ll

[LV] Improve epilogue tail-folding test coverage (NFC) (#228549)

Add a test where the trip count is a multiple of the main loop VF, so no
iterations are left for the epilogue.

Rewrite @early_exit so that it can be vectorized as an early-exit loop.
The epilogue tail-folding checks will run only after the main loop plan
is built, so the loop must be vectorizable to reach the early-exit
remark.

Also rename blocks and values for consistency.
DeltaFile
+71-45llvm/test/Transforms/LoopVectorize/fold-epilogue-tail.ll
+71-451 files

LLVM/project 992e613 — llvm/docs AMDGPUUsage.rst, llvm/lib/Target/AMDGPU AMDGPULowerIntrinsics.cpp SIMemoryLegalizer.cpp

Split demotion logic out
DeltaFile
+828-1,328llvm/test/CodeGen/AMDGPU/memory-legalizer-single-wave-workgroup-memops.ll
+27-33llvm/test/CodeGen/AMDGPU/memory-legalizer-single-wave-workgroup-lds-dma.ll
+7-20llvm/lib/Target/AMDGPU/SIMemoryLegalizer.cpp
+0-10llvm/docs/AMDGPUUsage.rst
+4-3llvm/test/CodeGen/AMDGPU/memory-legalizer-single-wave-workgroup-lds-dma-async.ll
+4-2llvm/lib/Target/AMDGPU/AMDGPULowerIntrinsics.cpp
+870-1,3962 files not shown
+870-1,4048 files

LLVM/project decb108 — llvm/lib/Transforms/Vectorize VPlanTransforms.cpp

[LV] NFC: Move PHI/result update out of transformToPartialReduction. (#228436)

That lets `transformToPartialReduction` focus on just creating a partial
reduction expression for each link in the chain, whereas
`createPartialReductions` updates the PHI and ReductionResult to
complete the work for the chain. It removes the need to know that the
partial reduction is part of a 'chain' in `transformToPartialReduction`.

I'll rebase this with the better terminology after #222377 gets merged.
DeltaFile
+50-40llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
+50-401 files

LLVM/project e05016c — llvm/include/llvm/Analysis InstSimplifyFolder.h InstructionSimplify.h, llvm/include/llvm/IR PatternMatch.h

[InstSimplify] Simplify `bitinsert` and `bitextract`
DeltaFile
+89-0llvm/lib/Analysis/InstructionSimplify.cpp
+18-43llvm/test/Transforms/InstSimplify/bitinsert-bitextract.ll
+29-0llvm/unittests/IR/PatternMatch.cpp
+15-0llvm/include/llvm/IR/PatternMatch.h
+8-0llvm/include/llvm/Analysis/InstructionSimplify.h
+2-4llvm/include/llvm/Analysis/InstSimplifyFolder.h
+161-476 files

LLVM/project 59f16d0 — llvm/test/Transforms/InstSimplify bitinsert-bitextract.ll

[InstSimplify] Pre-commit tests for `bitinsert` and `bitextract`




DeltaFile
+331-0llvm/test/Transforms/InstSimplify/bitinsert-bitextract.ll
+331-01 files

FreeBSD/ports c3216c6 — net/kea-devel Makefile distinfo

net/kea-devel: Update to 3.3.2
DeltaFile
+41-40net/kea-devel/pkg-plist
+3-3net/kea-devel/distinfo
+1-2net/kea-devel/Makefile
+45-453 files

LLVM/project c0a1634 —

[mlir][sparse] Order the inferred lvlToDim results by dimension (#227674)
DeltaFile
+0-00 files

FreeNAS/freenas 6869f21 — src/middlewared/middlewared/plugins/webshare sharing.py, src/middlewared/middlewared/test/integration/assets webshare.py

ruff
DeltaFile
+63-55tests/api2/test_webshare_sharing.py
+48-33tests/api2/test_webshare_homedir.py
+7-3src/middlewared/middlewared/plugins/webshare/sharing.py
+1-5src/middlewared/middlewared/test/integration/assets/webshare.py
+119-964 files

FreeNAS/freenas 1fc82d2 — src/middlewared/middlewared/plugins/failover_ nftables.py

Fix failover firewall to match VIPs as destination

drop_all matched the VIPs with saddr, so client traffic to a VIP was
never dropped. Match daddr, and only tcp/udp so that ICMP and IPv6
neighbor discovery keep working. This is the CORE behavior: pf dropped
inbound tcp/udp to the VIPs, ssh and web UI excepted, to hold clients
off a controller that was not ready to serve.
DeltaFile
+7-4src/middlewared/middlewared/plugins/failover_/nftables.py
+7-41 files

LLVM/project 737211f — clang/include/clang/CIR/Dialect/IR CIRDialect.td, clang/lib/CIR/Dialect/IR CIRDialect.cpp

[mlir][CIR] Use OpFoldResults for the cir.scope fold

The CIR dialect now sets the `useOpFoldResults` bit. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each CIR op that does not have
exactly one fixed result. `cir.scope` is the only such op with a fold.
Its fold now returns the yielded value directly. The behavior does not
change. The existing test `clang/test/CIR/Transforms/canonicalize.cir`
covers this fold.

Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
DeltaFile
+2-4clang/lib/CIR/Dialect/IR/CIRDialect.cpp
+2-0clang/include/clang/CIR/Dialect/IR/CIRDialect.td
+4-42 files

LLVM/project c6d4f85 — mlir/include/mlir/Dialect/Vector/IR Vector.td, mlir/lib/Dialect/Vector/IR VectorOps.cpp

[mlir][vector] Use OpFoldResults for vector folds

The vector dialect now sets the `useOpFoldResults` bit. ODS then
declares `OpFoldResults fold(FoldAdaptor)` for each vector op that does
not have exactly one fixed result. This patch moves the five folds of
such ops to the new form: `to_elements`, `transfer_write`, `store`,
`masked_store`, and `mask`.

The behavior of the folds does not change, except in the case below. A
fold that returned success with an empty vector now returns `success()`,
which is an in-place change. A fold that filled the vector now returns
its values.

The all-true fold of `vector.mask` moves the masked op out of the
region, and the terminator then has null operands. So this fold must
replace every result, and the driver erases the op. A mask without
results has nothing to replace, so the fold reports the move as an
in-place change.


    [27 lines not shown]
DeltaFile
+65-0mlir/test/Dialect/Vector/canonicalize.mlir
+22-17mlir/lib/Dialect/Vector/IR/VectorOps.cpp
+1-0mlir/include/mlir/Dialect/Vector/IR/Vector.td
+88-173 files

LLVM/project afa9f2b — clang/lib/CIR/Dialect/IR CIRDialect.cpp, clang/test/CIR/Transforms canonicalize.cir

[mlir] Deprecate the legacy fold APIs with a results vector

The legacy `Operation::fold` overloads with a
`SmallVectorImpl<OpFoldResult> &` parameter drop partial folds. The
overloads that return `OpFoldResults` keep them.

This patch moves every in-tree caller of the legacy overloads to the
overloads that return `OpFoldResults`, except the unit tests of the
legacy APIs. `OpFoldResultsTest.cpp` suppresses the deprecation warnings
for these tests. The test dialect keeps a legacy fold trait for the fold
tests, so `TestOps.cpp` suppresses the warning too. Then the patch marks
these APIs as deprecated: the two legacy `Operation::fold` overloads,
the legacy general `foldTrait` form, the legacy fold hook overloads of
`DynamicOpDefinition`, and the `DynamicOpDefinition::LegacyFoldHookFn`
alias. ODS cannot put an attribute on an interface method, so only the
documentation marks the legacy `DialectFoldInterface::fold` method as
deprecated.

The new code in `cir::CastOp::fold` also fixes a crash. The fold

    [21 lines not shown]
DeltaFile
+14-11mlir/lib/IR/ExtensibleDialect.cpp
+17-0clang/test/CIR/Transforms/canonicalize.cir
+11-4mlir/lib/IR/Operation.cpp
+6-8mlir/lib/Dialect/Affine/IR/AffineOps.cpp
+11-2mlir/include/mlir/IR/ExtensibleDialect.h
+6-5clang/lib/CIR/Dialect/IR/CIRDialect.cpp
+65-3010 files not shown
+110-5016 files

LLVM/project 6ed10a9 — mlir/docs Canonicalization.md, mlir/docs/DefiningDialects _index.md

[mlir][tblgen] Warn about the deprecated multi-result fold form

The legacy form `LogicalResult fold(FoldAdaptor,
SmallVectorImpl<OpFoldResult> &)` will be removed. This patch makes
`mlir-tblgen -gen-op-decls` warn when a dialect keeps `useOpFoldResults`
at 0 and has an op with `hasFolder` that does not have exactly one fixed
result. The warning points at the dialect definition. A note names the
op that uses the legacy form and comes first in name order.

The warning comes once for each dialect in one `-gen-op-decls` run. A
dialect with ops in more than one `.td` file can warn once for each file
that contains such an op. `-gen-op-defs` does not warn.

No in-tree dialect warns, because each in-tree dialect with such an op
already sets the bit.

The diagnostic follows the `-on-deprecated` option: `none` silences it,
`warn` (the default) warns, and `error` reports an error and fails the
run. `MlirTblgenMain.h` exposes the option value through

    [3 lines not shown]
DeltaFile
+88-0mlir/test/mlir-tblgen/op-fold-results-deprecation.td
+45-3mlir/tools/mlir-tblgen/OpDefinitionsGen.cpp
+6-3mlir/lib/Tools/mlir-tblgen/MlirTblgenMain.cpp
+6-0mlir/include/mlir/Tools/mlir-tblgen/MlirTblgenMain.h
+3-1mlir/docs/Canonicalization.md
+4-0mlir/docs/DefiningDialects/_index.md
+152-76 files

LLVM/project 71911e9 — mlir/lib/Dialect/Math/IR MathOps.cpp, mlir/lib/Dialect/Shape/IR Shape.cpp

[mlir] Use OpFoldResults in more upstream dialects

The scf, shape, sparse_tensor, gpu, math, and builtin dialects now set
the `useOpFoldResults` bit. ODS then declares `OpFoldResults
fold(FoldAdaptor)` for each of their ops that does not have exactly one
fixed result. This patch moves the seven folds of such ops to the new
form: `scf.if`, `shape.split_at`, `sparse_tensor.crd_translate`,
`gpu.memcpy`, `gpu.memset`, `math.sincos`, and
`builtin.unrealized_conversion_cast`.

The behavior of the folds does not change, except in the case below.
Each fold builds the same normalized `OpFoldResults` object that the
legacy adapter built from the old return. The current tests of each
dialect cover these folds.

In a graph region or in an unreachable block, an operand that
`unrealized_conversion_cast` or `sparse_tensor.crd_translate` forwards
can be another result of the same op that the same fold replaces. The
new form drops the replacements of such a fold. For a forward chain, the

    [22 lines not shown]
DeltaFile
+37-0mlir/test/Dialect/Builtin/canonicalize.mlir
+16-0mlir/test/Dialect/SparseTensor/fold.mlir
+6-9mlir/lib/Dialect/SparseTensor/IR/SparseTensorDialect.cpp
+4-9mlir/lib/IR/BuiltinDialect.cpp
+3-7mlir/lib/Dialect/Math/IR/MathOps.cpp
+3-5mlir/lib/Dialect/Shape/IR/Shape.cpp
+69-308 files not shown
+79-3614 files

LLVM/project a8c4a64 — mlir/include/mlir/Dialect/Linalg/IR LinalgBase.td, mlir/lib/Dialect/Linalg/IR LinalgOps.cpp

[mlir][linalg] Use OpFoldResults for linalg folds

The linalg dialect now sets the `useOpFoldResults` bit. ODS then
declares `OpFoldResults fold(FoldAdaptor)` for each linalg op that does
not have exactly one fixed result. This patch moves the eleven folds of
such ops in `LinalgOps.cpp` to the new form. It also changes the fold
template of `mlir-linalg-ods-yaml-gen`, so the generated named ops use
the new form too.

The behavior of the folds does not change. A fold that returned the
result of `memref::foldMemRefCast` now returns the same in-place state
through the `LogicalResult` constructor. The folds of `transpose`,
`pack`, and `unpack` return their replacement value directly.

The yaml-gen test now checks the generated fold definition.

A build directory that uses a native `mlir-linalg-ods-yaml-gen` does
not regenerate `LinalgNamedStructuredOps.yamlgen.cpp.inc` when the tool
changes. Delete that file before the build.

    [2 lines not shown]
DeltaFile
+19-36mlir/lib/Dialect/Linalg/IR/LinalgOps.cpp
+1-2mlir/tools/mlir-linalg-ods-gen/mlir-linalg-ods-yaml-gen.cpp
+3-0mlir/test/mlir-linalg-ods-gen/test-linalg-ods-yaml-gen.yaml
+1-0mlir/include/mlir/Dialect/Linalg/IR/LinalgBase.td
+24-384 files

LLVM/project 7f705e9 — mlir/include/mlir/Dialect/MemRef/IR MemRefBase.td, mlir/lib/Dialect/MemRef/IR MemRefOps.cpp

[mlir][memref] Use OpFoldResults for memref folds

The memref dialect now sets `useOpFoldResults`. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each memref op that does not have
exactly one fixed result. This patch moves the seven folds of such ops
to the new form: `copy`, `dealloc`, `dma_start`, `dma_wait`,
`extract_strided_metadata`, `prefetch`, and `store`. The behavior of
these folds does not change, except for `extract_strided_metadata`.

The `memref.extract_strided_metadata` fold created `arith.constant` ops
with its own builder and replaced the uses of the constant results
itself. No driver saw these changes. The fold now returns a partial
fold: one replacement for each constant result, and an in-place mark
when it removes a `memref.cast` source. This patch removes the helper
`replaceConstantUsesOf`, which has no other user.

DialectConversion folds an op before it applies the patterns. So the
conversion did not see the new constants and left them unconverted.
Now the conversion legalizes the new constants of the partial fold, as

    [29 lines not shown]
DeltaFile
+24-59mlir/lib/Dialect/MemRef/IR/MemRefOps.cpp
+25-0mlir/test/Conversion/MemRefToLLVM/memref-to-llvm.mlir
+1-0mlir/include/mlir/Dialect/MemRef/IR/MemRefBase.td
+50-593 files

LLVM/project e02a45e — mlir/include/mlir/Dialect/Arith/IR ArithBase.td, mlir/lib/Dialect/Arith/IR ArithOps.cpp

[mlir][arith] Use OpFoldResults for arith folds

The arith dialect now sets `useOpFoldResults`. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each arith op that does not have
exactly one fixed result. This patch moves the four folds of such ops to
the new form: `addui_extended`, `subui_extended`, `mulsi_extended`, and
`mului_extended`. The behavior of these folds does not change, with two
exceptions.

The `arith.mulsi_extended` fold now replaces the low result of
`mulsi_extended(x, 1)` by `x`, and keeps the high result. This is also
correct for i1, where the constant `true` is -1. The i1 tests in
`Arith/canonicalize.mlir` change their expected output.

In a graph region or in an unreachable block, the identity fold of
`addui_extended`, `subui_extended`, or `mului_extended` can forward an
operand that is the other result of the same op. The fold also replaces
that result, so the new form drops the replacements of the fold. The
legacy form gave a correct result in this case, because it first moved

    [9 lines not shown]
DeltaFile
+22-52mlir/lib/Dialect/Arith/IR/ArithOps.cpp
+73-0mlir/test/Dialect/Arith/fold-partial.mlir
+2-2mlir/test/Dialect/Arith/canonicalize.mlir
+1-0mlir/include/mlir/Dialect/Arith/IR/ArithBase.td
+98-544 files

LLVM/project cb902ac — mlir/docs Canonicalization.md, mlir/docs/DefiningDialects _index.md

[mlir] Apply partial folds in the greedy pattern rewrite driver

The greedy driver now applies the `OpFoldResults` of a fold:
- A fold that replaces every result erases the op. The driver
  materializes every result, also a result without uses, because
  `replaceOp` gives the listeners a value for each result.
- A partial fold replaces the uses of each replaced result and keeps
  the op. A replaced result without uses gets no constant. The driver
  puts the op on the worklist again.
- The materialization of constants is all-or-nothing. When one
  constant fails, the driver inserts no constant and applies no
  replacement. An in-place change still counts.
- A fold that keeps every result and has no in-place mark fails.

The driver uses `OpBuilder::materializeFoldResults` for the constants.

When `MLIR_ENABLE_EXPENSIVE_PATTERN_API_CHECKS` is on, the driver
compares the fingerprints of the op before and after the fold. A fold
that changes the op and returns failure is a fatal error. A fold that

    [10 lines not shown]
DeltaFile
+54-54mlir/lib/Transforms/Utils/GreedyPatternRewriteDriver.cpp
+98-0mlir/test/Transforms/test-canonicalize.mlir
+61-7mlir/docs/Canonicalization.md
+14-3mlir/docs/DefiningDialects/_index.md
+12-0mlir/test/Transforms/test-canonicalize-unmarked-in-place-fold.mlir
+9-2mlir/docs/Traits/_index.md
+248-661 files not shown
+252-677 files

LLVM/project 838288c — mlir/include/mlir/IR Builders.h, mlir/test/Transforms test-operation-folder.mlir

[mlir] Apply partial folds in OpBuilder::createOrFold

The multi-result `createOrFold` now uses the new `OpBuilder::tryFold`
overload and `materializeFoldResults`. When the fold replaces only some
results, the new op stays, and the results of `createOrFold` mix the
replacement values and the kept results of the op. Before this patch,
such a fold counted as an in-place fold or as a failure. The new op has
no uses yet, so `createOrFold` materializes every replaced result.

The zero-result `createOrFold` also uses the new overload. A fold of a
zero-result op can only change the op in place, so its behavior does
not change.

A new test pattern builds `test.op_partial_fold` with the multi-result
`createOrFold`. The new test in `test-operation-folder.mlir` checks
that the op stays next to the replacement values.

Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
DeltaFile
+24-11mlir/include/mlir/IR/Builders.h
+23-0mlir/test/Transforms/test-operation-folder.mlir
+18-1mlir/test/lib/Dialect/Test/TestPatterns.cpp
+7-0mlir/test/lib/Dialect/Test/TestOps.td
+72-124 files

LLVM/project e89b6ef — mlir/lib/Dialect/Shape/IR Shape.cpp

[mlir][shape] Report the in-place change of the assuming_all fold

The `shape.assuming_all` fold visits its inputs in reverse order. It
erases each constant input from the operands, and it returns null at
the first input that is not constant. If the fold already erased an
input at that point, it changed the op but reported a failure, so no
driver knew about the change. The fold now returns the result of the op
in this case, which is the in-place signal of the single-result fold
form.

No test fails on main. The later patch "[mlir] Apply partial folds in
the greedy pattern rewrite driver" adds a fingerprint check of each fold
to the greedy driver. The check runs only in a build with
`-DMLIR_ENABLE_EXPENSIVE_PATTERN_API_CHECKS=ON`. In such a build of that
patch, with this fix reverted, the existing test
`Dialect/Shape/canonicalize.mlir` fails:

  $ mlir-opt -split-input-file -allow-unregistered-dialect \
      -canonicalize="test-convergence" \

    [4 lines not shown]
DeltaFile
+9-4mlir/lib/Dialect/Shape/IR/Shape.cpp
+9-41 files

LLVM/project 4793993 — mlir/include/mlir/Dialect/Affine/IR AffineOps.td, mlir/lib/Dialect/Affine/IR AffineOps.cpp

[mlir][affine] Use OpFoldResults for affine folds

The affine dialect now sets `useOpFoldResults`. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each affine op that does not have
exactly one fixed result. This patch moves the eight folds of such ops
to the new form: `dma_start`, `dma_wait`, `for`, `if`, `store`,
`prefetch`, `parallel`, and `delinearize_index`. The behavior of these
folds does not change, with two exceptions.

The `affine.delinearize_index` fold now replaces each result whose basis
element is 1 with the constant 0, and keeps the other results. In
`Tensor/bubble-up-extract-slice-op.mlir`, these results now fold to a
constant 0. The new test `Affine/fold-partial.mlir` runs
`-test-single-fold` and `-sccp`, because `-canonicalize` also runs
`DropUnitExtentBasis`, which hides a broken fold.

In a graph region or in an unreachable block, an init of a zero-trip
`affine.for` can be another result of the same loop. The fold also
replaces that result, so the new form drops the replacements of the

    [21 lines not shown]
DeltaFile
+43-41mlir/lib/Dialect/Affine/IR/AffineOps.cpp
+74-0mlir/test/Dialect/Affine/fold-partial.mlir
+37-0mlir/test/Dialect/Affine/canonicalize.mlir
+6-4mlir/test/Dialect/Tensor/bubble-up-extract-slice-op.mlir
+1-0mlir/include/mlir/Dialect/Affine/IR/AffineOps.td
+161-455 files