[LV] Improve epilogue tail-folding test coverage (NFC) (#228549)
Add a test where the trip count is a multiple of the main loop VF, so no
iterations are left for the epilogue.
Rewrite @early_exit so that it can be vectorized as an early-exit loop.
The epilogue tail-folding checks will run only after the main loop plan
is built, so the loop must be vectorizable to reach the early-exit
remark.
Also rename blocks and values for consistency.
[LV] NFC: Move PHI/result update out of transformToPartialReduction. (#228436)
That lets `transformToPartialReduction` focus on just creating a partial
reduction expression for each link in the chain, whereas
`createPartialReductions` updates the PHI and ReductionResult to
complete the work for the chain. It removes the need to know that the
partial reduction is part of a 'chain' in `transformToPartialReduction`.
I'll rebase this with the better terminology after #222377 gets merged.
[mlir][CIR] Use OpFoldResults for the cir.scope fold
The CIR dialect now sets the `useOpFoldResults` bit. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each CIR op that does not have
exactly one fixed result. `cir.scope` is the only such op with a fold.
Its fold now returns the yielded value directly. The behavior does not
change. The existing test `clang/test/CIR/Transforms/canonicalize.cir`
covers this fold.
Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
[mlir][vector] Use OpFoldResults for vector folds
The vector dialect now sets the `useOpFoldResults` bit. ODS then
declares `OpFoldResults fold(FoldAdaptor)` for each vector op that does
not have exactly one fixed result. This patch moves the five folds of
such ops to the new form: `to_elements`, `transfer_write`, `store`,
`masked_store`, and `mask`.
The behavior of the folds does not change, except in the case below. A
fold that returned success with an empty vector now returns `success()`,
which is an in-place change. A fold that filled the vector now returns
its values.
The all-true fold of `vector.mask` moves the masked op out of the
region, and the terminator then has null operands. So this fold must
replace every result, and the driver erases the op. A mask without
results has nothing to replace, so the fold reports the move as an
in-place change.
[27 lines not shown]
[mlir] Deprecate the legacy fold APIs with a results vector
The legacy `Operation::fold` overloads with a
`SmallVectorImpl<OpFoldResult> &` parameter drop partial folds. The
overloads that return `OpFoldResults` keep them.
This patch moves every in-tree caller of the legacy overloads to the
overloads that return `OpFoldResults`, except the unit tests of the
legacy APIs. `OpFoldResultsTest.cpp` suppresses the deprecation warnings
for these tests. The test dialect keeps a legacy fold trait for the fold
tests, so `TestOps.cpp` suppresses the warning too. Then the patch marks
these APIs as deprecated: the two legacy `Operation::fold` overloads,
the legacy general `foldTrait` form, the legacy fold hook overloads of
`DynamicOpDefinition`, and the `DynamicOpDefinition::LegacyFoldHookFn`
alias. ODS cannot put an attribute on an interface method, so only the
documentation marks the legacy `DialectFoldInterface::fold` method as
deprecated.
The new code in `cir::CastOp::fold` also fixes a crash. The fold
[21 lines not shown]
[mlir][tblgen] Warn about the deprecated multi-result fold form
The legacy form `LogicalResult fold(FoldAdaptor,
SmallVectorImpl<OpFoldResult> &)` will be removed. This patch makes
`mlir-tblgen -gen-op-decls` warn when a dialect keeps `useOpFoldResults`
at 0 and has an op with `hasFolder` that does not have exactly one fixed
result. The warning points at the dialect definition. A note names the
op that uses the legacy form and comes first in name order.
The warning comes once for each dialect in one `-gen-op-decls` run. A
dialect with ops in more than one `.td` file can warn once for each file
that contains such an op. `-gen-op-defs` does not warn.
No in-tree dialect warns, because each in-tree dialect with such an op
already sets the bit.
The diagnostic follows the `-on-deprecated` option: `none` silences it,
`warn` (the default) warns, and `error` reports an error and fails the
run. `MlirTblgenMain.h` exposes the option value through
[3 lines not shown]
[mlir] Use OpFoldResults in more upstream dialects
The scf, shape, sparse_tensor, gpu, math, and builtin dialects now set
the `useOpFoldResults` bit. ODS then declares `OpFoldResults
fold(FoldAdaptor)` for each of their ops that does not have exactly one
fixed result. This patch moves the seven folds of such ops to the new
form: `scf.if`, `shape.split_at`, `sparse_tensor.crd_translate`,
`gpu.memcpy`, `gpu.memset`, `math.sincos`, and
`builtin.unrealized_conversion_cast`.
The behavior of the folds does not change, except in the case below.
Each fold builds the same normalized `OpFoldResults` object that the
legacy adapter built from the old return. The current tests of each
dialect cover these folds.
In a graph region or in an unreachable block, an operand that
`unrealized_conversion_cast` or `sparse_tensor.crd_translate` forwards
can be another result of the same op that the same fold replaces. The
new form drops the replacements of such a fold. For a forward chain, the
[22 lines not shown]
[mlir][linalg] Use OpFoldResults for linalg folds
The linalg dialect now sets the `useOpFoldResults` bit. ODS then
declares `OpFoldResults fold(FoldAdaptor)` for each linalg op that does
not have exactly one fixed result. This patch moves the eleven folds of
such ops in `LinalgOps.cpp` to the new form. It also changes the fold
template of `mlir-linalg-ods-yaml-gen`, so the generated named ops use
the new form too.
The behavior of the folds does not change. A fold that returned the
result of `memref::foldMemRefCast` now returns the same in-place state
through the `LogicalResult` constructor. The folds of `transpose`,
`pack`, and `unpack` return their replacement value directly.
The yaml-gen test now checks the generated fold definition.
A build directory that uses a native `mlir-linalg-ods-yaml-gen` does
not regenerate `LinalgNamedStructuredOps.yamlgen.cpp.inc` when the tool
changes. Delete that file before the build.
[2 lines not shown]
[mlir][memref] Use OpFoldResults for memref folds
The memref dialect now sets `useOpFoldResults`. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each memref op that does not have
exactly one fixed result. This patch moves the seven folds of such ops
to the new form: `copy`, `dealloc`, `dma_start`, `dma_wait`,
`extract_strided_metadata`, `prefetch`, and `store`. The behavior of
these folds does not change, except for `extract_strided_metadata`.
The `memref.extract_strided_metadata` fold created `arith.constant` ops
with its own builder and replaced the uses of the constant results
itself. No driver saw these changes. The fold now returns a partial
fold: one replacement for each constant result, and an in-place mark
when it removes a `memref.cast` source. This patch removes the helper
`replaceConstantUsesOf`, which has no other user.
DialectConversion folds an op before it applies the patterns. So the
conversion did not see the new constants and left them unconverted.
Now the conversion legalizes the new constants of the partial fold, as
[29 lines not shown]
[mlir][arith] Use OpFoldResults for arith folds
The arith dialect now sets `useOpFoldResults`. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each arith op that does not have
exactly one fixed result. This patch moves the four folds of such ops to
the new form: `addui_extended`, `subui_extended`, `mulsi_extended`, and
`mului_extended`. The behavior of these folds does not change, with two
exceptions.
The `arith.mulsi_extended` fold now replaces the low result of
`mulsi_extended(x, 1)` by `x`, and keeps the high result. This is also
correct for i1, where the constant `true` is -1. The i1 tests in
`Arith/canonicalize.mlir` change their expected output.
In a graph region or in an unreachable block, the identity fold of
`addui_extended`, `subui_extended`, or `mului_extended` can forward an
operand that is the other result of the same op. The fold also replaces
that result, so the new form drops the replacements of the fold. The
legacy form gave a correct result in this case, because it first moved
[9 lines not shown]
[mlir] Apply partial folds in the greedy pattern rewrite driver
The greedy driver now applies the `OpFoldResults` of a fold:
- A fold that replaces every result erases the op. The driver
materializes every result, also a result without uses, because
`replaceOp` gives the listeners a value for each result.
- A partial fold replaces the uses of each replaced result and keeps
the op. A replaced result without uses gets no constant. The driver
puts the op on the worklist again.
- The materialization of constants is all-or-nothing. When one
constant fails, the driver inserts no constant and applies no
replacement. An in-place change still counts.
- A fold that keeps every result and has no in-place mark fails.
The driver uses `OpBuilder::materializeFoldResults` for the constants.
When `MLIR_ENABLE_EXPENSIVE_PATTERN_API_CHECKS` is on, the driver
compares the fingerprints of the op before and after the fold. A fold
that changes the op and returns failure is a fatal error. A fold that
[10 lines not shown]
[mlir] Apply partial folds in OpBuilder::createOrFold
The multi-result `createOrFold` now uses the new `OpBuilder::tryFold`
overload and `materializeFoldResults`. When the fold replaces only some
results, the new op stays, and the results of `createOrFold` mix the
replacement values and the kept results of the op. Before this patch,
such a fold counted as an in-place fold or as a failure. The new op has
no uses yet, so `createOrFold` materializes every replaced result.
The zero-result `createOrFold` also uses the new overload. A fold of a
zero-result op can only change the op in place, so its behavior does
not change.
A new test pattern builds `test.op_partial_fold` with the multi-result
`createOrFold`. The new test in `test-operation-folder.mlir` checks
that the op stays next to the replacement values.
Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
[mlir][shape] Report the in-place change of the assuming_all fold
The `shape.assuming_all` fold visits its inputs in reverse order. It
erases each constant input from the operands, and it returns null at
the first input that is not constant. If the fold already erased an
input at that point, it changed the op but reported a failure, so no
driver knew about the change. The fold now returns the result of the op
in this case, which is the in-place signal of the single-result fold
form.
No test fails on main. The later patch "[mlir] Apply partial folds in
the greedy pattern rewrite driver" adds a fingerprint check of each fold
to the greedy driver. The check runs only in a build with
`-DMLIR_ENABLE_EXPENSIVE_PATTERN_API_CHECKS=ON`. In such a build of that
patch, with this fix reverted, the existing test
`Dialect/Shape/canonicalize.mlir` fails:
$ mlir-opt -split-input-file -allow-unregistered-dialect \
-canonicalize="test-convergence" \
[4 lines not shown]
[mlir][affine] Use OpFoldResults for affine folds
The affine dialect now sets `useOpFoldResults`. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each affine op that does not have
exactly one fixed result. This patch moves the eight folds of such ops
to the new form: `dma_start`, `dma_wait`, `for`, `if`, `store`,
`prefetch`, `parallel`, and `delinearize_index`. The behavior of these
folds does not change, with two exceptions.
The `affine.delinearize_index` fold now replaces each result whose basis
element is 1 with the constant 0, and keeps the other results. In
`Tensor/bubble-up-extract-slice-op.mlir`, these results now fold to a
constant 0. The new test `Affine/fold-partial.mlir` runs
`-test-single-fold` and `-sccp`, because `-canonicalize` also runs
`DropUnitExtentBasis`, which hides a broken fold.
In a graph region or in an unreachable block, an init of a zero-trip
`affine.for` can be another result of the same loop. The fold also
replaces that result, so the new form drops the replacements of the
[21 lines not shown]
[mlir] Apply partial folds in DialectConversion
`OperationLegalizer::legalizeWithFold` now uses the new
`OpBuilder::tryFold` overload and `materializeFoldResults`, and it
applies partial folds. The legalizer replaces the uses of each replaced
result that has uses, legalizes the new constants, and tries to
legalize the op again.
With pattern rollback, `replaceAllUsesWith` only records a value
mapping, which the conversion applies when it commits and drops on a
rollback. So the legalizer skips a result that an earlier fold already
replaced, and a replacement that leads back to its result, which would
make the mapping cyclic. Each new mapping counts as progress. If a new
constant does not legalize, the legalizer rolls the fold back, as for a
full fold. A full fold with rollback still materializes every result.
Without rollback, the legalizer materializes only the replaced results
that have uses, also for a full fold, because it must legalize each new
op. So it creates no constant that nothing uses, and a result without
[12 lines not shown]
[mlir] Use the replacements of a partial fold in SCCP
`SparseConstantPropagation` now uses the `OpFoldResults` form of
`Operation::fold`. It merges each replacement into the lattice of its
result, and it sets each kept result to the entry state. A kept result
must not join with its own lattice, because that leaves the lattice
uninitialized. Before this patch, SCCP saw a partial fold as a failure
and set every result to the entry state.
The new test in `sccp.mlir` checks that SCCP uses the replacements and
that the kept result stays at the entry state.
Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
[mlir] Deprecate the OpBuilder::tryFold overload with a results vector
The `OpBuilder::tryFold` overload with a `SmallVectorImpl<Value> &`
parameter drops partial folds. The overload that returns
`OpFoldResults`, together with `materializeFoldResults`, keeps them.
No in-tree code calls the old overload anymore, so this patch marks it
as deprecated. The unit test of the old overload suppresses the
deprecation warning.
Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
[mlir] Add an OpBuilder::tryFold overload that returns OpFoldResults
`OpBuilder::tryFold(op, results, constants)` applies a fold only if the
fold replaces every result, so it cannot give a partial fold to the
caller. This patch adds two `OpBuilder` functions that let a caller
apply a partial fold:
- `OpFoldResults tryFold(Operation *op)` folds the op and returns the
fold result. Like the old overload, it skips constants and repeats a
fold that only changes the op in place. It creates no constant.
- `materializeFoldResults(op, foldResults, liveOnly)` returns one value
per result: the replacement of a replaced result, or null for a kept
result. It turns each attribute into a new constant at the insertion
point. If `liveOnly` is set, a replaced result without uses gets null
and no constant. If one constant fails to materialize, the function
inserts no constant and fails.
The old overload now uses the two new functions. A fold that replaces
only some results, which only the new fold form can return, counts as
an in-place fold if it changed the op in place, and as a failure
[11 lines not shown]
[mlir] Apply partial folds in OperationFolder
`OperationFolder::tryToFold` now applies the `OpFoldResults` of a fold.
A fold that replaces every result erases the op. It materializes every
result, also a result without uses, because `replaceOp` gives the
listener a value for each result. A partial fold replaces the uses of
each replaced result, keeps the op, and sets `inPlaceUpdate`. It
materializes constants only for the replaced results that have uses. If
a constant fails to materialize, no result is replaced, and an in-place
change still counts. The private vector overload of `tryToFold` has no
callers after this change, so this patch removes it.
This patch also fixes `OperationFolder::processFoldResults`. It moved a
reused constant to the front of the block as soon as it used the
constant for a result. If a later result then failed to materialize,
the cleanup erased all ops before the insertion point. These ops
included the moved constant, which still had uses. The folder now moves
the reused constants only after all results materialize.
[17 lines not shown]
[flang] Use the OpFoldResults overload of OpBuilder::tryFold
`getAttrIfConstant` in `LLVMInsertChainFolder.cpp` folds the op that
defines a value and reads the constant of the value from the fold. It
used the old `OpBuilder::tryFold` overload, which materializes a
constant for each attribute replacement. The function only read the
value of the new `llvm.mlir.constant`, so these constants stayed in the
IR without uses.
The function now uses the new overload, which returns the fold result
and creates no constant. It accepts the same constants as before. The
fold must replace every result. A `Value` replacement counts if
`llvm.mlir.constant` defines it. An attribute replacement counts if the
op is in the LLVM dialect and `LLVM::ConstantOp::isBuildableWith`
accepts the attribute. Only in this case did the old overload create an
`llvm.mlir.constant`. One difference remains: the other results of the
op no longer have to materialize.
Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
[mlir] Add the useOpFoldResults ODS dialect bit
The new `OpFoldResults fold(FoldAdaptor)` form needs a matching ODS
declaration. This patch adds the dialect bit `useOpFoldResults`. The bit
is off by default. When a dialect sets it, `genFolderDecls` declares
`OpFoldResults fold(FoldAdaptor adaptor)` for each op that does not have
exactly one fixed result. An op with exactly one fixed result keeps the
`OpFoldResult` form. A dialect without the bit keeps the legacy
declaration.
The `test` dialect sets the bit. Its three folds of ops that do not have
exactly one fixed result now use the new form. The behavior of these
folds does not change, except in two cases. In a graph region or in an
unreachable block, the new form drops a forwarding fold that names a
result that the same fold replaces. And the fold of
`TestOpWithVariadicResultsAndFolder` with no operands now fails. Before,
it reported an in-place change, but it changed nothing. The greedy
driver then put the op on the worklist again and folded it again, so
canonicalize did not stop. A new test in `test-canonicalize.mlir`
[18 lines not shown]
[mlir] Add the OpFoldResults op fold form
An op can now define `OpFoldResults fold(FoldAdaptor)`. With this form,
a fold can replace some results of a multi-result op and keep the
others. The drivers still call the legacy `Operation::fold` overloads,
so they do not apply a partial fold yet. This form has the same
parameters as the single-result form, so the single-result detectors
also match it, and `getFoldHookFn` excludes it from the single-result
forms. `getFoldHookFn` selects the new form if the op defines it, else
the single-result form, else the legacy vector form. The new form then
folds the traits if its own fold replaces no result.
A replacement in the `OpFoldResults` of the new form can be another
result of the op only if the fold keeps that result. The hook normalizes
the `OpFoldResults` first, so replacement i that is result i still means
"keep". If a replacement names a result that the same fold also
replaces, the outcome would depend on the order of the replacements, so
the hook drops the replacements. A forwarding fold can return such a
replacement where an op can use its own results: in a graph region or
[21 lines not shown]
[mlir] Add the OpFoldResults form to foldTrait and DialectFoldInterface
The patch "[mlir] Add the OpFoldResults op fold form" added
`OpFoldResults fold(FoldAdaptor)` for ops. The fold traits and
`DialectFoldInterface` still use only the legacy vector form, which
cannot express a partial fold.
A trait can now define `static OpFoldResults foldTrait(Operation *,
ArrayRef<Attribute>)`. The trait fold dispatch detects this form by its
return type and calls it directly. This form has the same parameters as
the single-result trait form, so the single-result detector excludes it.
The legacy trait form still works through the strict adapter.
`IsCommutative` and `CastOpInterface` now use the new form.
`DialectFoldInterface` gets a new `fold` method that returns
`OpFoldResults`. Its default implementation calls the legacy `fold`
method through the strict adapter, so existing dialects keep their
behavior. `Operation::fold` now calls the new method.
[20 lines not shown]
[mlir] Return OpFoldResults from the type-erased fold hooks
The type-erased fold hooks return `LogicalResult` and fill a vector of
`OpFoldResult`. This form cannot express a fold that replaces only some
results of an op. This patch changes the hooks to return
`OpFoldResults`. The drivers do not use partial folds yet.
The patch keeps every existing fold signature and wraps each one in an
adapter that returns `OpFoldResults`:
- `OpFoldResult fold(FoldAdaptor)` and the single-result `foldTrait`: a
null result is a failure, the op's own result is an in-place change,
and any other value replaces the result.
- `LogicalResult fold(FoldAdaptor, SmallVectorImpl<OpFoldResult> &)` and
the general `foldTrait`: `detail::convertLegacyFoldResults` applies
the strict legacy contract. Failure stays failure, success with an
empty vector is an in-place change, and success with one entry per
result replaces every result.
- The legacy fold hooks of `DynamicOpDefinition` and
`DialectFoldInterface` use the same strict conversion.
[63 lines not shown]
[offload][omp] Read the kernel environment from the device image
This should be an optimization for #222606. The kernel environment
doesn't actually need to be loaded back from the device since it
shouldn't be different compared to the time where it was transferred to
the device as part of the image.
Claude assisted with this patch.
[cmake] Build libxml from source in release (#221365)
With 23.x we started linking a static libxml in release builds to avoid
requiring that to be installed, and the varying names of that library on
different linux distros. That broke more than it fixed because the
static libxml in our docker images are built targeting ICU, and ICU's
shared library version changes quickly, so even between 2 ubuntu
versions the binaries will not work.
With this change we now build from source in the release build, which is
already what windows does, and we disable ICU / iconv since those don't
seem vital for the windows manifest use case. This seems easier than
building a libxml out of band in our CI images that do the release
builds.
Fixes https://github.com/llvm/llvm-project/issues/215764
Assisted-By: codex
(cherry picked from commit 5ae71ec799917c67e9e72510e68831602fc60699)