LLVM/project 2641277 — mlir/docs Canonicalization.md, mlir/docs/DefiningDialects _index.md

[mlir][tblgen] Warn about the deprecated multi-result fold form

The legacy form `LogicalResult fold(FoldAdaptor,
SmallVectorImpl<OpFoldResult> &)` will be removed. This patch makes
`mlir-tblgen -gen-op-decls` warn when a dialect keeps `useOpFoldResults`
at 0 and has an op with `hasFolder` that does not have exactly one fixed
result. The warning points at the dialect definition. A note names the
first op that uses the legacy form.

The warning comes once for each dialect in one `-gen-op-decls` run. A
dialect with ops in more than one `.td` file can warn once for each file
that contains such an op. `-gen-op-defs` does not warn.

No in-tree dialect warns, because each in-tree dialect with such an op
already sets the bit.

The diagnostic follows the `-on-deprecated` option: `none` silences it,
`warn` (the default) warns, and `error` reports an error and fails the
run. `MlirTblgenMain.h` exposes the option value through

    [3 lines not shown]
DeltaFile
+88-0mlir/test/mlir-tblgen/op-fold-results-deprecation.td
+45-3mlir/tools/mlir-tblgen/OpDefinitionsGen.cpp
+6-3mlir/lib/Tools/mlir-tblgen/MlirTblgenMain.cpp
+6-0mlir/include/mlir/Tools/mlir-tblgen/MlirTblgenMain.h
+3-1mlir/docs/Canonicalization.md
+4-0mlir/docs/DefiningDialects/_index.md
+152-76 files

LLVM/project 4aaa30c — mlir/include/mlir/Dialect/Vector/IR Vector.td, mlir/lib/Dialect/Vector/IR VectorOps.cpp

[mlir][vector] Use OpFoldResults for vector folds

The vector dialect now sets the `useOpFoldResults` bit. ODS then
declares `OpFoldResults fold(FoldAdaptor)` for each vector op that does
not have exactly one fixed result. This patch moves the five folds of
such ops to the new form: `to_elements`, `transfer_write`, `store`,
`masked_store`, and `mask`.

The behavior of the folds does not change, except in the graph-region
case below. A fold that returned success with an empty vector now
returns `success()`, which is an in-place change. A fold that filled the
vector now returns its values.

The all-true fold of `vector.mask` moves the masked op out of the
region, and the terminator then has null operands. So this fold must
replace every result, and the driver erases the op. A mask without
results has nothing to replace, so the fold reports the move as an
in-place change.


    [26 lines not shown]
DeltaFile
+65-0mlir/test/Dialect/Vector/canonicalize.mlir
+22-17mlir/lib/Dialect/Vector/IR/VectorOps.cpp
+1-0mlir/include/mlir/Dialect/Vector/IR/Vector.td
+88-173 files

LLVM/project 3b8c37d — mlir/lib/Dialect/Math/IR MathOps.cpp, mlir/lib/Dialect/Shape/IR Shape.cpp

[mlir] Use OpFoldResults in more upstream dialects

The scf, shape, sparse_tensor, gpu, math, and builtin dialects now set
the `useOpFoldResults` bit. ODS then declares `OpFoldResults
fold(FoldAdaptor)` for each of their ops that does not have exactly one
fixed result. This patch moves the seven folds of such ops to the new
form: `scf.if`, `shape.split_at`, `sparse_tensor.crd_translate`,
`gpu.memcpy`, `gpu.memset`, `math.sincos`, and
`builtin.unrealized_conversion_cast`.

The behavior of the folds does not change, except in the graph-region
case below. Each fold builds the same normalized `OpFoldResults` object
that the legacy adapter built from the old return. The existing tests of
each dialect cover these folds.

In a graph region, an operand that `unrealized_conversion_cast` or
`sparse_tensor.crd_translate` forwards can be another result of the
same op that the same fold replaces. The new form drops such a fold.
For a forward chain, the legacy form gave a correct result by the order

    [22 lines not shown]
DeltaFile
+37-0mlir/test/Dialect/Builtin/canonicalize.mlir
+16-0mlir/test/Dialect/SparseTensor/fold.mlir
+6-9mlir/lib/Dialect/SparseTensor/IR/SparseTensorDialect.cpp
+4-9mlir/lib/IR/BuiltinDialect.cpp
+3-7mlir/lib/Dialect/Math/IR/MathOps.cpp
+3-5mlir/lib/Dialect/Shape/IR/Shape.cpp
+69-308 files not shown
+79-3614 files

LLVM/project 245fd90 — mlir/include/mlir/Dialect/MemRef/IR MemRefBase.td, mlir/lib/Dialect/MemRef/IR MemRefOps.cpp

[mlir][memref] Use OpFoldResults for memref folds

The memref dialect now sets `useOpFoldResults`. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each memref op that does not have
exactly one fixed result. This patch moves the seven folds of such ops
to the new form: `copy`, `dealloc`, `dma_start`, `dma_wait`,
`extract_strided_metadata`, `prefetch`, and `store`. The behavior of
these folds does not change, except for `extract_strided_metadata`.

The `memref.extract_strided_metadata` fold created `arith.constant` ops
with its own builder and replaced the uses of the constant results
itself. No driver saw these changes. The fold now returns a partial
fold: one replacement for each constant result, and an in-place mark
when it removes a `memref.cast` source. This patch removes the helper
`replaceConstantUsesOf`, which has no other user.

DialectConversion folds an op before it applies the patterns. So the
conversion did not see the new constants and left them unconverted.
With pattern rollback on, the conversion now does not apply the partial

    [27 lines not shown]
DeltaFile
+24-59mlir/lib/Dialect/MemRef/IR/MemRefOps.cpp
+25-0mlir/test/Conversion/MemRefToLLVM/memref-to-llvm.mlir
+1-0mlir/include/mlir/Dialect/MemRef/IR/MemRefBase.td
+50-593 files

LLVM/project b630756 — clang/lib/CIR/Dialect/IR CIRDialect.cpp, clang/test/CIR/Transforms canonicalize.cir

[mlir] Deprecate the legacy fold APIs with a results vector

The legacy `Operation::fold` overloads with a
`SmallVectorImpl<OpFoldResult> &` parameter drop partial folds. The
overloads that return `OpFoldResults` keep them.

This patch moves every in-tree caller of the legacy overloads to the
overloads that return `OpFoldResults`, except the unit tests of the
legacy APIs. `OpFoldResultsTest.cpp` suppresses the deprecation warnings
for these tests. Then the patch marks these APIs as deprecated: the two
legacy `Operation::fold` overloads, the legacy general `foldTrait` form,
the legacy fold hook overloads of `DynamicOpDefinition`, and the
`DynamicOpDefinition::LegacyFoldHookFn` alias. ODS cannot put an
attribute on an interface method, so only the documentation marks the
legacy `DialectFoldInterface::fold` method as deprecated.

The new code in `cir::CastOp::fold` also fixes two bugs. The fold
crashed when the source of an integral cast was a block argument, and it
read the fold result of result 0 when the source was a different result

    [17 lines not shown]
DeltaFile
+14-11mlir/lib/IR/ExtensibleDialect.cpp
+17-0clang/test/CIR/Transforms/canonicalize.cir
+11-4mlir/lib/IR/Operation.cpp
+6-8mlir/lib/Dialect/Affine/IR/AffineOps.cpp
+11-2mlir/include/mlir/IR/ExtensibleDialect.h
+6-5clang/lib/CIR/Dialect/IR/CIRDialect.cpp
+65-309 files not shown
+103-4715 files

LLVM/project 170a405 — clang/include/clang/CIR/Dialect/IR CIRDialect.td, clang/lib/CIR/Dialect/IR CIRDialect.cpp

[mlir][CIR] Use OpFoldResults for the cir.scope fold

The CIR dialect now sets the `useOpFoldResults` bit. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each CIR op that does not have
exactly one fixed result. `cir.scope` is the only such op with a fold.
Its fold now returns the yielded value directly. The behavior does not
change. The existing test `clang/test/CIR/Transforms/canonicalize.cir`
covers this fold.

Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
DeltaFile
+2-4clang/lib/CIR/Dialect/IR/CIRDialect.cpp
+2-0clang/include/clang/CIR/Dialect/IR/CIRDialect.td
+4-42 files

LLVM/project 9484365 — mlir/include/mlir/Dialect/Linalg/IR LinalgBase.td, mlir/lib/Dialect/Linalg/IR LinalgOps.cpp

[mlir][linalg] Use OpFoldResults for linalg folds

The linalg dialect now sets the `useOpFoldResults` bit. ODS then
declares `OpFoldResults fold(FoldAdaptor)` for each linalg op that does
not have exactly one fixed result. This patch moves the eleven folds of
such ops in `LinalgOps.cpp` to the new form. It also changes the fold
template of `mlir-linalg-ods-yaml-gen`, so the generated named ops use
the new form too.

The behavior of the folds does not change. A fold that returned the
result of `memref::foldMemRefCast` now returns the same in-place state
through the `LogicalResult` constructor. The folds of `transpose`,
`pack`, and `unpack` return their replacement value directly.

The yaml-gen test now checks the generated fold definition.

A build directory that uses a native `mlir-linalg-ods-yaml-gen` does
not regenerate `LinalgNamedStructuredOps.yamlgen.cpp.inc` when the tool
changes. Delete that file before the build.

    [2 lines not shown]
DeltaFile
+19-36mlir/lib/Dialect/Linalg/IR/LinalgOps.cpp
+1-2mlir/tools/mlir-linalg-ods-gen/mlir-linalg-ods-yaml-gen.cpp
+3-0mlir/test/mlir-linalg-ods-gen/test-linalg-ods-yaml-gen.yaml
+1-0mlir/include/mlir/Dialect/Linalg/IR/LinalgBase.td
+24-384 files

LLVM/project 66fc0f2 — mlir/include/mlir/Dialect/Affine/IR AffineOps.td, mlir/lib/Dialect/Affine/IR AffineOps.cpp

[mlir][affine] Use OpFoldResults for affine folds

The affine dialect now sets `useOpFoldResults`. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each affine op that does not have
exactly one fixed result. This patch moves the eight folds of such ops
to the new form: `dma_start`, `dma_wait`, `for`, `if`, `store`,
`prefetch`, `parallel`, and `delinearize_index`. The behavior of these
folds does not change, with two exceptions.

The `affine.delinearize_index` fold now replaces each result whose basis
element is 1 with the constant 0, and keeps the other results. In
`Tensor/bubble-up-extract-slice-op.mlir`, these results now fold to a
constant 0. The new test `Affine/fold-partial.mlir` runs
`-test-single-fold` and `-sccp`, because `-canonicalize` also runs
`DropUnitExtentBasis`, which hides a broken fold.

In a graph region, an init of a zero-trip `affine.for` can be another
result of the same loop. The fold also replaces that result, so the new
form drops the fold. For a forward chain such as

    [21 lines not shown]
DeltaFile
+43-41mlir/lib/Dialect/Affine/IR/AffineOps.cpp
+74-0mlir/test/Dialect/Affine/fold-partial.mlir
+37-0mlir/test/Dialect/Affine/canonicalize.mlir
+6-4mlir/test/Dialect/Tensor/bubble-up-extract-slice-op.mlir
+1-0mlir/include/mlir/Dialect/Affine/IR/AffineOps.td
+161-455 files

LLVM/project e1a1dc3 — mlir/include/mlir/Dialect/Arith/IR ArithBase.td, mlir/lib/Dialect/Arith/IR ArithOps.cpp

[mlir][arith] Use OpFoldResults for arith folds

The arith dialect now sets `useOpFoldResults`. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each arith op that does not have
exactly one fixed result. This patch moves the four folds of such ops to
the new form: `addui_extended`, `subui_extended`, `mulsi_extended`, and
`mului_extended`. The behavior of these folds does not change, with two
exceptions.

The `arith.mulsi_extended` fold now replaces the low result of
`mulsi_extended(x, 1)` by `x`, and keeps the high result. This is also
correct for i1, where the constant `true` is -1. The i1 tests in
`Arith/canonicalize.mlir` change their expected output.

In a graph region, the identity fold of `addui_extended`,
`subui_extended`, or `mului_extended` can forward an operand that is the
other result of the same op. The fold also replaces that result, so the
new form drops the fold. The legacy form gave a correct result in this
case, because it first moved the uses to the forwarded result, and then

    [7 lines not shown]
DeltaFile
+22-52mlir/lib/Dialect/Arith/IR/ArithOps.cpp
+73-0mlir/test/Dialect/Arith/fold-partial.mlir
+2-2mlir/test/Dialect/Arith/canonicalize.mlir
+1-0mlir/include/mlir/Dialect/Arith/IR/ArithBase.td
+98-544 files

LLVM/project 18439f5 — mlir/docs Canonicalization.md, mlir/docs/DefiningDialects _index.md

[mlir] Apply partial folds in the greedy pattern rewrite driver

The greedy driver now applies the `OpFoldResults` of a fold:
- A fold that replaces every result erases the op. A result without
  uses gets no constant.
- A partial fold replaces the uses of each replaced result and keeps
  the op. The driver puts the op on the worklist again.
- The materialization of constants is all-or-nothing. When one
  constant fails, the driver erases only the constants of this call
  and applies no replacement. An in-place change still counts.
- A fold that keeps every result and has no in-place mark fails.

The driver uses `detail::materializeFoldResults` for the constants.

When `MLIR_ENABLE_EXPENSIVE_PATTERN_API_CHECKS` is on, the driver
compares the fingerprints of the op before and after the fold. A fold
that changes the op and returns failure is a fatal error. A fold that
changes an op that stays, without an in-place mark, is also a fatal
error.

    [8 lines not shown]
DeltaFile
+53-54mlir/lib/Transforms/Utils/GreedyPatternRewriteDriver.cpp
+98-0mlir/test/Transforms/test-canonicalize.mlir
+57-7mlir/docs/Canonicalization.md
+14-3mlir/docs/DefiningDialects/_index.md
+12-0mlir/test/Transforms/test-canonicalize-unmarked-in-place-fold.mlir
+7-2mlir/docs/Traits/_index.md
+241-661 files not shown
+245-677 files

LLVM/project be7a5f7 — mlir/include/mlir/Transforms FoldUtils.h, mlir/lib/Transforms/Utils FoldUtils.cpp

[mlir] Apply partial folds in OperationFolder

`OperationFolder::tryToFold` now applies the `OpFoldResults` of a fold.
A fold that replaces every result erases the op. A partial fold replaces
the uses of each replaced result, keeps the op, and sets
`inPlaceUpdate`. The folder materializes constants only for the replaced
results that have uses. If a constant fails to materialize, no result is
replaced, and an in-place change still counts. The private vector
overload of `tryToFold` has no callers after this change, so this patch
removes it.

This patch also fixes `OperationFolder::processFoldResults`. It moved a
reused constant to the front of the block as soon as it used the
constant for a result. If a later result then failed to materialize,
the cleanup erased all ops before the insertion point. These ops
included the moved constant, which still had uses. The folder now moves
the reused constants only after all results materialize.

The bug does not need partial folds. Without this patch, the new test

    [16 lines not shown]
DeltaFile
+98-0mlir/test/Transforms/constant-fold.mlir
+60-35mlir/lib/Transforms/Utils/FoldUtils.cpp
+9-13mlir/include/mlir/Transforms/FoldUtils.h
+167-483 files

LLVM/project c5e8a3b — mlir/include/mlir/IR Builders.h, mlir/lib/IR Builders.cpp

[mlir] Apply partial folds in OpBuilder::createOrFold

`OpBuilder::tryFold` gets a `FoldApplyMode` argument and an optional
in-place output. The default mode, `AllOrNothing`, keeps the old
behavior for existing callers: it applies a fold only if the fold
replaces every result. A partial fold, which only the new form can
return, counts as an in-place fold if it changed the op, and as a
failure otherwise. The `Partial` mode materializes each replaced result
and puts the result of the op in the entry of each kept result.

The `PartialLiveOnly` mode also skips the replaced results without uses.
A later patch uses it in DialectConversion, which must legalize each new
op: the mode creates no constant that nothing uses, and a result without
uses can then hold an attribute that does not materialize.

The multi-result `createOrFold` uses the `Partial` mode. When the fold
keeps a result, the new op stays, and the results of `createOrFold` mix
the replacement values and the kept results of the op.


    [11 lines not shown]
DeltaFile
+94-41mlir/lib/IR/Builders.cpp
+68-8mlir/include/mlir/IR/Builders.h
+23-0mlir/test/Transforms/test-operation-folder.mlir
+18-1mlir/test/lib/Dialect/Test/TestPatterns.cpp
+7-0mlir/test/lib/Dialect/Test/TestOps.td
+210-505 files

LLVM/project ff65e7c — mlir/lib/Transforms/Utils DialectConversion.cpp, mlir/test/Transforms test-legalizer-fold-dead-result.mlir test-legalizer-partial-fold.mlir

[mlir] Apply partial folds in DialectConversion

`OperationLegalizer::legalizeWithFold` now applies a partial fold when
pattern rollback is off. The legalizer cannot roll back the replacement
of some results of an op that stays. So with rollback on, the legalizer
applies a fold only if it replaces every result or changes the op in
place, as before. Without rollback, the legalizer uses the
`PartialLiveOnly` mode of `OpBuilder::tryFold`, replaces the uses of
each replaced result, and tries to legalize the op again.

In a graph region, the replacement skips the uses that come before the
replacement value. So the legalizer counts progress by the use counts
of the replaced results. A fold that replaces all results still erases
the op, also when all of its results have no uses. A fold that gives no
replacement and no in-place change fails.

The new tests `test-legalizer-partial-fold.mlir` and
`test-legalizer-fold-dead-result.mlir` run both modes.

Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
DeltaFile
+118-0mlir/test/Transforms/test-legalizer-partial-fold.mlir
+65-10mlir/lib/Transforms/Utils/DialectConversion.cpp
+36-0mlir/test/Transforms/test-legalizer-fold-dead-result.mlir
+219-103 files

LLVM/project 831f488 — mlir/lib/Analysis/DataFlow ConstantPropagationAnalysis.cpp, mlir/test/Transforms sccp.mlir

[mlir] Use the replaced results of a partial fold in SCCP

`SparseConstantPropagation` now uses the `OpFoldResults` form of
`Operation::fold`. It merges the value of each replaced result into the
lattice of that result, and it sets each kept result to the entry
state. A kept result must not join with its own lattice, because that
leaves the lattice uninitialized. Before this patch, SCCP saw a partial
fold as a failure and set every result to the entry state.

The new test in `sccp.mlir` checks that SCCP uses the replaced results
and that the kept result stays overdefined.

Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
DeltaFile
+14-9mlir/lib/Analysis/DataFlow/ConstantPropagationAnalysis.cpp
+21-0mlir/test/Transforms/sccp.mlir
+35-92 files

LLVM/project 4e7dc26 — clang/lib/CIR/Dialect/IR CIRDialect.cpp, clang/test/CIR/Transforms canonicalize.cir

[mlir] Deprecate the legacy fold APIs with a results vector

The legacy `Operation::fold` overloads with a
`SmallVectorImpl<OpFoldResult> &` parameter drop partial folds. The
overloads that return `OpFoldResults` keep them.

This patch moves every in-tree caller of the legacy overloads to the
overloads that return `OpFoldResults`, except the unit tests of the
legacy APIs. `OpFoldResultsTest.cpp` suppresses the deprecation warnings
for these tests. Then the patch marks these APIs as deprecated: the two
legacy `Operation::fold` overloads, the legacy general `foldTrait` form,
the legacy fold hook overloads of `DynamicOpDefinition`, and the
`DynamicOpDefinition::LegacyFoldHookFn` alias. ODS cannot put an
attribute on an interface method, so only the documentation marks the
legacy `DialectFoldInterface::fold` method as deprecated.

The new code in `cir::CastOp::fold` also fixes two bugs. The fold
crashed when the source of an integral cast was a block argument, and it
read the fold result of result 0 when the source was a different result

    [17 lines not shown]
DeltaFile
+14-11mlir/lib/IR/ExtensibleDialect.cpp
+17-0clang/test/CIR/Transforms/canonicalize.cir
+11-4mlir/lib/IR/Operation.cpp
+6-8mlir/lib/Dialect/Affine/IR/AffineOps.cpp
+11-2mlir/include/mlir/IR/ExtensibleDialect.h
+6-5clang/lib/CIR/Dialect/IR/CIRDialect.cpp
+65-309 files not shown
+103-4715 files

LLVM/project c86e970 — mlir/docs Canonicalization.md, mlir/docs/DefiningDialects _index.md

[mlir][tblgen] Warn about the deprecated multi-result fold form

The legacy form `LogicalResult fold(FoldAdaptor,
SmallVectorImpl<OpFoldResult> &)` will be removed. This patch makes
`mlir-tblgen -gen-op-decls` warn when a dialect keeps `useOpFoldResults`
at 0 and has an op with `hasFolder` that does not have exactly one fixed
result. The warning points at the dialect definition. A note names the
first op that uses the legacy form.

The warning comes once for each dialect in one `-gen-op-decls` run. A
dialect with ops in more than one `.td` file can warn once for each file
that contains such an op. `-gen-op-defs` does not warn.

No in-tree dialect warns, because each in-tree dialect with such an op
already sets the bit.

The diagnostic follows the `-on-deprecated` option: `none` silences it,
`warn` (the default) warns, and `error` reports an error and fails the
run. `MlirTblgenMain.h` exposes the option value through

    [3 lines not shown]
DeltaFile
+88-0mlir/test/mlir-tblgen/op-fold-results-deprecation.td
+45-3mlir/tools/mlir-tblgen/OpDefinitionsGen.cpp
+6-3mlir/lib/Tools/mlir-tblgen/MlirTblgenMain.cpp
+6-0mlir/include/mlir/Tools/mlir-tblgen/MlirTblgenMain.h
+3-1mlir/docs/Canonicalization.md
+4-0mlir/docs/DefiningDialects/_index.md
+152-76 files

LLVM/project 6a8be9f — mlir/include/mlir/Dialect/Arith/IR ArithBase.td, mlir/lib/Dialect/Arith/IR ArithOps.cpp

[mlir][arith] Use OpFoldResults for arith folds

The arith dialect now sets `useOpFoldResults`. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each arith op that does not have
exactly one fixed result. This patch moves the four folds of such ops to
the new form: `addui_extended`, `subui_extended`, `mulsi_extended`, and
`mului_extended`. The behavior of these folds does not change, with two
exceptions.

The `arith.mulsi_extended` fold now replaces the low result of
`mulsi_extended(x, 1)` by `x`, and keeps the high result. This is also
correct for i1, where the constant `true` is -1. The i1 tests in
`Arith/canonicalize.mlir` change their expected output.

In a graph region, the identity fold of `addui_extended`,
`subui_extended`, or `mului_extended` can forward an operand that is the
other result of the same op. The fold also replaces that result, so the
new form drops the fold. The legacy form gave a correct result in this
case, because it first moved the uses to the forwarded result, and then

    [7 lines not shown]
DeltaFile
+22-52mlir/lib/Dialect/Arith/IR/ArithOps.cpp
+73-0mlir/test/Dialect/Arith/fold-partial.mlir
+2-2mlir/test/Dialect/Arith/canonicalize.mlir
+1-0mlir/include/mlir/Dialect/Arith/IR/ArithBase.td
+98-544 files

LLVM/project 853f406 — mlir/lib/Dialect/Math/IR MathOps.cpp, mlir/lib/Dialect/Shape/IR Shape.cpp

[mlir] Use OpFoldResults in more upstream dialects

The scf, shape, sparse_tensor, gpu, math, and builtin dialects now set
the `useOpFoldResults` bit. ODS then declares `OpFoldResults
fold(FoldAdaptor)` for each of their ops that does not have exactly one
fixed result. This patch moves the seven folds of such ops to the new
form: `scf.if`, `shape.split_at`, `sparse_tensor.crd_translate`,
`gpu.memcpy`, `gpu.memset`, `math.sincos`, and
`builtin.unrealized_conversion_cast`.

The behavior of the folds does not change, except in the graph-region
case below. Each fold builds the same normalized `OpFoldResults` object
that the legacy adapter built from the old return. The existing tests of
each dialect cover these folds.

In a graph region, an operand that `unrealized_conversion_cast` or
`sparse_tensor.crd_translate` forwards can be another result of the
same op that the same fold replaces. The new form drops such a fold.
For a forward chain, the legacy form gave a correct result by the order

    [22 lines not shown]
DeltaFile
+37-0mlir/test/Dialect/Builtin/canonicalize.mlir
+16-0mlir/test/Dialect/SparseTensor/fold.mlir
+6-9mlir/lib/Dialect/SparseTensor/IR/SparseTensorDialect.cpp
+4-9mlir/lib/IR/BuiltinDialect.cpp
+3-7mlir/lib/Dialect/Math/IR/MathOps.cpp
+3-5mlir/lib/Dialect/Shape/IR/Shape.cpp
+69-308 files not shown
+79-3614 files

LLVM/project 0ce8e06 — mlir/include/mlir/Dialect/Linalg/IR LinalgBase.td, mlir/lib/Dialect/Linalg/IR LinalgOps.cpp

[mlir][linalg] Use OpFoldResults for linalg folds

The linalg dialect now sets the `useOpFoldResults` bit. ODS then
declares `OpFoldResults fold(FoldAdaptor)` for each linalg op that does
not have exactly one fixed result. This patch moves the eleven folds of
such ops in `LinalgOps.cpp` to the new form. It also changes the fold
template of `mlir-linalg-ods-yaml-gen`, so the generated named ops use
the new form too.

The behavior of the folds does not change. A fold that returned the
result of `memref::foldMemRefCast` now returns the same in-place state
through the `LogicalResult` constructor. The folds of `transpose`,
`pack`, and `unpack` return their replacement value directly.

The yaml-gen test now checks the generated fold definition.

A build directory that uses a native `mlir-linalg-ods-yaml-gen` does
not regenerate `LinalgNamedStructuredOps.yamlgen.cpp.inc` when the tool
changes. Delete that file before the build.

    [2 lines not shown]
DeltaFile
+19-36mlir/lib/Dialect/Linalg/IR/LinalgOps.cpp
+1-2mlir/tools/mlir-linalg-ods-gen/mlir-linalg-ods-yaml-gen.cpp
+3-0mlir/test/mlir-linalg-ods-gen/test-linalg-ods-yaml-gen.yaml
+1-0mlir/include/mlir/Dialect/Linalg/IR/LinalgBase.td
+24-384 files

LLVM/project df00449 — mlir/include/mlir/Dialect/Vector/IR Vector.td, mlir/lib/Dialect/Vector/IR VectorOps.cpp

[mlir][vector] Use OpFoldResults for vector folds

The vector dialect now sets the `useOpFoldResults` bit. ODS then
declares `OpFoldResults fold(FoldAdaptor)` for each vector op that does
not have exactly one fixed result. This patch moves the five folds of
such ops to the new form: `to_elements`, `transfer_write`, `store`,
`masked_store`, and `mask`.

The behavior of the folds does not change, except in the graph-region
case below. A fold that returned success with an empty vector now
returns `success()`, which is an in-place change. A fold that filled the
vector now returns its values.

The all-true fold of `vector.mask` moves the masked op out of the
region, and the terminator then has null operands. So this fold must
replace every result, and the driver erases the op. A mask without
results has nothing to replace, so the fold reports the move as an
in-place change.


    [26 lines not shown]
DeltaFile
+65-0mlir/test/Dialect/Vector/canonicalize.mlir
+22-17mlir/lib/Dialect/Vector/IR/VectorOps.cpp
+1-0mlir/include/mlir/Dialect/Vector/IR/Vector.td
+88-173 files

LLVM/project dfe2bc4 — clang/include/clang/CIR/Dialect/IR CIRDialect.td, clang/lib/CIR/Dialect/IR CIRDialect.cpp

[mlir][CIR] Use OpFoldResults for the cir.scope fold

The CIR dialect now sets the `useOpFoldResults` bit. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each CIR op that does not have
exactly one fixed result. `cir.scope` is the only such op with a fold.
Its fold now returns the yielded value directly. The behavior does not
change. The existing test `clang/test/CIR/Transforms/canonicalize.cir`
covers this fold.

Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
DeltaFile
+2-4clang/lib/CIR/Dialect/IR/CIRDialect.cpp
+2-0clang/include/clang/CIR/Dialect/IR/CIRDialect.td
+4-42 files

LLVM/project 007e5ac — mlir/include/mlir/Dialect/MemRef/IR MemRefBase.td, mlir/lib/Dialect/MemRef/IR MemRefOps.cpp

[mlir][memref] Use OpFoldResults for memref folds

The memref dialect now sets `useOpFoldResults`. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each memref op that does not have
exactly one fixed result. This patch moves the seven folds of such ops
to the new form: `copy`, `dealloc`, `dma_start`, `dma_wait`,
`extract_strided_metadata`, `prefetch`, and `store`. The behavior of
these folds does not change, except for `extract_strided_metadata`.

The `memref.extract_strided_metadata` fold created `arith.constant` ops
with its own builder and replaced the uses of the constant results
itself. No driver saw these changes. The fold now returns a partial
fold: one replacement for each constant result, and an in-place mark
when it removes a `memref.cast` source. This patch removes the helper
`replaceConstantUsesOf`, which has no other user.

DialectConversion folds an op before it applies the patterns. So the
conversion did not see the new constants and left them unconverted.
With pattern rollback on, the conversion now does not apply the partial

    [27 lines not shown]
DeltaFile
+24-59mlir/lib/Dialect/MemRef/IR/MemRefOps.cpp
+25-0mlir/test/Conversion/MemRefToLLVM/memref-to-llvm.mlir
+1-0mlir/include/mlir/Dialect/MemRef/IR/MemRefBase.td
+50-593 files

LLVM/project 1c6391d — mlir/include/mlir/Dialect/Affine/IR AffineOps.td, mlir/lib/Dialect/Affine/IR AffineOps.cpp

[mlir][affine] Use OpFoldResults for affine folds

The affine dialect now sets `useOpFoldResults`. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each affine op that does not have
exactly one fixed result. This patch moves the eight folds of such ops
to the new form: `dma_start`, `dma_wait`, `for`, `if`, `store`,
`prefetch`, `parallel`, and `delinearize_index`. The behavior of these
folds does not change, with two exceptions.

The `affine.delinearize_index` fold now replaces each result whose basis
element is 1 with the constant 0, and keeps the other results. In
`Tensor/bubble-up-extract-slice-op.mlir`, these results now fold to a
constant 0. The new test `Affine/fold-partial.mlir` runs
`-test-single-fold` and `-sccp`, because `-canonicalize` also runs
`DropUnitExtentBasis`, which hides a broken fold.

In a graph region, an init of a zero-trip `affine.for` can be another
result of the same loop. The fold also replaces that result, so the new
form drops the fold. For a forward chain such as

    [21 lines not shown]
DeltaFile
+43-41mlir/lib/Dialect/Affine/IR/AffineOps.cpp
+74-0mlir/test/Dialect/Affine/fold-partial.mlir
+37-0mlir/test/Dialect/Affine/canonicalize.mlir
+6-4mlir/test/Dialect/Tensor/bubble-up-extract-slice-op.mlir
+1-0mlir/include/mlir/Dialect/Affine/IR/AffineOps.td
+161-455 files

LLVM/project b8d2d7f — mlir/lib/Analysis/DataFlow ConstantPropagationAnalysis.cpp, mlir/test/Transforms sccp.mlir

[mlir] Use the replaced results of a partial fold in SCCP

`SparseConstantPropagation` now uses the `OpFoldResults` form of
`Operation::fold`. It merges the value of each replaced result into the
lattice of that result, and it sets each kept result to the entry
state. A kept result must not join with its own lattice, because that
leaves the lattice uninitialized. Before this patch, SCCP saw a partial
fold as a failure and set every result to the entry state.

The new test in `sccp.mlir` checks that SCCP uses the replaced results
and that the kept result stays overdefined.

Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
DeltaFile
+14-9mlir/lib/Analysis/DataFlow/ConstantPropagationAnalysis.cpp
+21-0mlir/test/Transforms/sccp.mlir
+35-92 files

LLVM/project 31e9a0b — mlir/include/mlir/Transforms FoldUtils.h, mlir/lib/Transforms/Utils FoldUtils.cpp

[mlir] Apply partial folds in OperationFolder

`OperationFolder::tryToFold` now applies the `OpFoldResults` of a fold.
A fold that replaces every result erases the op. A partial fold replaces
the uses of each replaced result, keeps the op, and sets
`inPlaceUpdate`. The folder materializes constants only for the replaced
results that have uses. If a constant fails to materialize, no result is
replaced, and an in-place change still counts. The private vector
overload of `tryToFold` has no callers after this change, so this patch
removes it.

This patch also fixes `OperationFolder::processFoldResults`. It moved a
reused constant to the front of the block as soon as it used the
constant for a result. If a later result then failed to materialize,
the cleanup erased all ops before the insertion point. These ops
included the moved constant, which still had uses. The folder now moves
the reused constants only after all results materialize.

The bug does not need partial folds. Without this patch, the new test

    [16 lines not shown]
DeltaFile
+98-0mlir/test/Transforms/constant-fold.mlir
+60-35mlir/lib/Transforms/Utils/FoldUtils.cpp
+9-13mlir/include/mlir/Transforms/FoldUtils.h
+167-483 files

LLVM/project ce33fc4 — mlir/lib/Transforms/Utils DialectConversion.cpp, mlir/test/Transforms test-legalizer-fold-dead-result.mlir test-legalizer-partial-fold.mlir

[mlir] Apply partial folds in DialectConversion

`OperationLegalizer::legalizeWithFold` now applies a partial fold when
pattern rollback is off. The legalizer cannot roll back the replacement
of some results of an op that stays. So with rollback on, the legalizer
applies a fold only if it replaces every result or changes the op in
place, as before. Without rollback, the legalizer uses the
`PartialLiveOnly` mode of `OpBuilder::tryFold`, replaces the uses of
each replaced result, and tries to legalize the op again.

In a graph region, the replacement skips the uses that come before the
replacement value. So the legalizer counts progress by the use counts
of the replaced results. A fold that replaces all results still erases
the op, also when all of its results have no uses. A fold that gives no
replacement and no in-place change fails.

The new tests `test-legalizer-partial-fold.mlir` and
`test-legalizer-fold-dead-result.mlir` run both modes.

Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
DeltaFile
+118-0mlir/test/Transforms/test-legalizer-partial-fold.mlir
+65-10mlir/lib/Transforms/Utils/DialectConversion.cpp
+36-0mlir/test/Transforms/test-legalizer-fold-dead-result.mlir
+219-103 files

LLVM/project d55e5b9 — mlir/include/mlir/IR Builders.h, mlir/lib/IR Builders.cpp

[mlir] Apply partial folds in OpBuilder::createOrFold

`OpBuilder::tryFold` gets a `FoldApplyMode` argument and an optional
in-place output. The default mode, `AllOrNothing`, keeps the old
behavior for existing callers: it applies a fold only if the fold
replaces every result. A partial fold, which only the new form can
return, counts as an in-place fold if it changed the op, and as a
failure otherwise. The `Partial` mode materializes each replaced result
and puts the result of the op in the entry of each kept result.

The `PartialLiveOnly` mode also skips the replaced results without uses.
A later patch uses it in DialectConversion, which must legalize each new
op: the mode creates no constant that nothing uses, and a result without
uses can then hold an attribute that does not materialize.

The multi-result `createOrFold` uses the `Partial` mode. When the fold
keeps a result, the new op stays, and the results of `createOrFold` mix
the replacement values and the kept results of the op.


    [11 lines not shown]
DeltaFile
+90-41mlir/lib/IR/Builders.cpp
+68-8mlir/include/mlir/IR/Builders.h
+23-0mlir/test/Transforms/test-operation-folder.mlir
+18-1mlir/test/lib/Dialect/Test/TestPatterns.cpp
+7-0mlir/test/lib/Dialect/Test/TestOps.td
+206-505 files

LLVM/project 0697b95 — mlir/docs Canonicalization.md, mlir/docs/DefiningDialects _index.md

[mlir] Apply partial folds in the greedy pattern rewrite driver

The greedy driver now applies the `OpFoldResults` of a fold:
- A fold that replaces every result erases the op. A result without
  uses gets no constant.
- A partial fold replaces the uses of each replaced result and keeps
  the op. The driver puts the op on the worklist again.
- The materialization of constants is all-or-nothing. When one
  constant fails, the driver erases only the constants of this call
  and applies no replacement. An in-place change still counts.
- A fold that keeps every result and has no in-place mark fails.

The driver uses `detail::materializeFoldResults` for the constants.

When `MLIR_ENABLE_EXPENSIVE_PATTERN_API_CHECKS` is on, the driver
compares the fingerprints of the op before and after the fold. A fold
that changes the op and returns failure is a fatal error. A fold that
changes an op that stays, without an in-place mark, is also a fatal
error.

    [8 lines not shown]
DeltaFile
+53-54mlir/lib/Transforms/Utils/GreedyPatternRewriteDriver.cpp
+98-0mlir/test/Transforms/test-canonicalize.mlir
+57-7mlir/docs/Canonicalization.md
+14-3mlir/docs/DefiningDialects/_index.md
+12-0mlir/test/Transforms/test-canonicalize-unmarked-in-place-fold.mlir
+7-2mlir/docs/Traits/_index.md
+241-661 files not shown
+245-677 files

LLVM/project 65a8290 — flang/include/flang/Evaluate tools.h, flang/lib/Evaluate tools.cpp

[flang][cuda] Transfer device data used in host operations (#228292)

An assignment to host data from an expression with device data is
lowered
as an implicit data transfer only when the expression also contains a
constant or a host symbol. Constants in subscripts are ignored, so
expressions such as a(3)*a(4), -a(3), -d1, or a conversion a(3) from
real(8) to real(4) were lowered as plain transfers. The scalar result
was
then computed on the host directly from device memory and stored to the
left-hand side, and array expressions failed the cuf.data_transfer
verifier.

Treat any right-hand side with device data that is not a variable as an
implicit transfer. The device data is copied to a host temporary and the
expression is evaluated on the host. An assignment of such an expression
to device data is now diagnosed as an unsupported transfer, as a = a +
10
already is.
DeltaFile
+94-12flang/test/Lower/CUDA/cuda-data-transfer.cuf
+6-2flang/lib/Evaluate/tools.cpp
+2-2flang/include/flang/Evaluate/tools.h
+2-0flang/test/Semantics/CUDA/cuf18.cuf
+104-164 files

LLVM/project f94b1c9 — mlir/include/mlir/IR OpDefinition.h, mlir/include/mlir/Interfaces FoldInterfaces.h DialectFoldInterface.td

[mlir] Add the OpFoldResults form to fold traits and DialectFoldInterface

The previous patch added `OpFoldResults fold(FoldAdaptor)` for ops. The
fold traits and `DialectFoldInterface` still use only the legacy vector
form, which cannot express a partial fold.

A trait can now define `static OpFoldResults foldTrait(Operation *,
ArrayRef<Attribute>)`. The trait fold dispatch detects this form by its
return type and calls it directly. This form has the same parameters as
the single-result trait form, so the single-result detector excludes it.
The legacy trait form still works through the strict adapter.
`IsCommutative` and `CastOpInterface` now use the new form.

`DialectFoldInterface` gets a new `fold` method that returns
`OpFoldResults`. Its default implementation calls the legacy `fold`
method through the strict adapter, so existing dialects keep their
behavior. `Operation::fold` now calls the new method.

The rule that a replacement does not name a result that the same fold

    [9 lines not shown]
DeltaFile
+281-3mlir/unittests/IR/OpFoldResultsTest.cpp
+28-10mlir/include/mlir/IR/OpDefinition.h
+12-6mlir/lib/IR/Operation.cpp
+17-0mlir/include/mlir/Interfaces/DialectFoldInterface.td
+11-1mlir/include/mlir/Interfaces/FoldInterfaces.h
+4-7mlir/lib/Interfaces/CastInterfaces.cpp
+353-272 files not shown
+358-338 files