[DAGCombiner] Don't fold GET_FPENV_MEM into a store whose address depends on it (#228839)
visitGET_FPENV_MEM folds `get_fpenv_mem tmp; load tmp; store -> p` into
`get_fpenv_mem p`, and the new node replaces the original one. When the
store address depends on the original node, the DAG becomes cyclic. This
happens for `ret i256 @llvm.get.fpenv.i256()`: the hidden sret pointer is
read with a CopyFromReg that is chained after the load from the stack
temporary. Type legalization then stops with "Operand not processed?",
and
release builds crash.
Skip the fold when the store address has the original node as a
predecessor.
Fixes #228828
Co-Authored-By: Claude Opus 5.5 <noreply at anthropic.com>
[LV] Improve epilogue tail-folding test coverage (NFC) (#228549)
Add a test where the trip count is a multiple of the main loop VF, so no
iterations are left for the epilogue.
Rewrite @early_exit so that it can be vectorized as an early-exit loop.
The epilogue tail-folding checks will run only after the main loop plan
is built, so the loop must be vectorizable to reach the early-exit
remark.
Also rename blocks and values for consistency.
[LV] NFC: Move PHI/result update out of transformToPartialReduction. (#228436)
That lets `transformToPartialReduction` focus on just creating a partial
reduction expression for each link in the chain, whereas
`createPartialReductions` updates the PHI and ReductionResult to
complete the work for the chain. It removes the need to know that the
partial reduction is part of a 'chain' in `transformToPartialReduction`.
I'll rebase this with the better terminology after #222377 gets merged.
Fix failover firewall to match VIPs as destination
drop_all matched the VIPs with saddr, so client traffic to a VIP was
never dropped. Match daddr, and only tcp/udp so that ICMP and IPv6
neighbor discovery keep working. This is the CORE behavior: pf dropped
inbound tcp/udp to the VIPs, ssh and web UI excepted, to hold clients
off a controller that was not ready to serve.
[mlir][CIR] Use OpFoldResults for the cir.scope fold
The CIR dialect now sets the `useOpFoldResults` bit. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each CIR op that does not have
exactly one fixed result. `cir.scope` is the only such op with a fold.
Its fold now returns the yielded value directly. The behavior does not
change. The existing test `clang/test/CIR/Transforms/canonicalize.cir`
covers this fold.
Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
[mlir][vector] Use OpFoldResults for vector folds
The vector dialect now sets the `useOpFoldResults` bit. ODS then
declares `OpFoldResults fold(FoldAdaptor)` for each vector op that does
not have exactly one fixed result. This patch moves the five folds of
such ops to the new form: `to_elements`, `transfer_write`, `store`,
`masked_store`, and `mask`.
The behavior of the folds does not change, except in the case below. A
fold that returned success with an empty vector now returns `success()`,
which is an in-place change. A fold that filled the vector now returns
its values.
The all-true fold of `vector.mask` moves the masked op out of the
region, and the terminator then has null operands. So this fold must
replace every result, and the driver erases the op. A mask without
results has nothing to replace, so the fold reports the move as an
in-place change.
[27 lines not shown]
[mlir] Deprecate the legacy fold APIs with a results vector
The legacy `Operation::fold` overloads with a
`SmallVectorImpl<OpFoldResult> &` parameter drop partial folds. The
overloads that return `OpFoldResults` keep them.
This patch moves every in-tree caller of the legacy overloads to the
overloads that return `OpFoldResults`, except the unit tests of the
legacy APIs. `OpFoldResultsTest.cpp` suppresses the deprecation warnings
for these tests. The test dialect keeps a legacy fold trait for the fold
tests, so `TestOps.cpp` suppresses the warning too. Then the patch marks
these APIs as deprecated: the two legacy `Operation::fold` overloads,
the legacy general `foldTrait` form, the legacy fold hook overloads of
`DynamicOpDefinition`, and the `DynamicOpDefinition::LegacyFoldHookFn`
alias. ODS cannot put an attribute on an interface method, so only the
documentation marks the legacy `DialectFoldInterface::fold` method as
deprecated.
The new code in `cir::CastOp::fold` also fixes a crash. The fold
[21 lines not shown]
[mlir][tblgen] Warn about the deprecated multi-result fold form
The legacy form `LogicalResult fold(FoldAdaptor,
SmallVectorImpl<OpFoldResult> &)` will be removed. This patch makes
`mlir-tblgen -gen-op-decls` warn when a dialect keeps `useOpFoldResults`
at 0 and has an op with `hasFolder` that does not have exactly one fixed
result. The warning points at the dialect definition. A note names the
op that uses the legacy form and comes first in name order.
The warning comes once for each dialect in one `-gen-op-decls` run. A
dialect with ops in more than one `.td` file can warn once for each file
that contains such an op. `-gen-op-defs` does not warn.
No in-tree dialect warns, because each in-tree dialect with such an op
already sets the bit.
The diagnostic follows the `-on-deprecated` option: `none` silences it,
`warn` (the default) warns, and `error` reports an error and fails the
run. `MlirTblgenMain.h` exposes the option value through
[3 lines not shown]
[mlir] Use OpFoldResults in more upstream dialects
The scf, shape, sparse_tensor, gpu, math, and builtin dialects now set
the `useOpFoldResults` bit. ODS then declares `OpFoldResults
fold(FoldAdaptor)` for each of their ops that does not have exactly one
fixed result. This patch moves the seven folds of such ops to the new
form: `scf.if`, `shape.split_at`, `sparse_tensor.crd_translate`,
`gpu.memcpy`, `gpu.memset`, `math.sincos`, and
`builtin.unrealized_conversion_cast`.
The behavior of the folds does not change, except in the case below.
Each fold builds the same normalized `OpFoldResults` object that the
legacy adapter built from the old return. The current tests of each
dialect cover these folds.
In a graph region or in an unreachable block, an operand that
`unrealized_conversion_cast` or `sparse_tensor.crd_translate` forwards
can be another result of the same op that the same fold replaces. The
new form drops the replacements of such a fold. For a forward chain, the
[22 lines not shown]
[mlir][linalg] Use OpFoldResults for linalg folds
The linalg dialect now sets the `useOpFoldResults` bit. ODS then
declares `OpFoldResults fold(FoldAdaptor)` for each linalg op that does
not have exactly one fixed result. This patch moves the eleven folds of
such ops in `LinalgOps.cpp` to the new form. It also changes the fold
template of `mlir-linalg-ods-yaml-gen`, so the generated named ops use
the new form too.
The behavior of the folds does not change. A fold that returned the
result of `memref::foldMemRefCast` now returns the same in-place state
through the `LogicalResult` constructor. The folds of `transpose`,
`pack`, and `unpack` return their replacement value directly.
The yaml-gen test now checks the generated fold definition.
A build directory that uses a native `mlir-linalg-ods-yaml-gen` does
not regenerate `LinalgNamedStructuredOps.yamlgen.cpp.inc` when the tool
changes. Delete that file before the build.
[2 lines not shown]
[mlir][memref] Use OpFoldResults for memref folds
The memref dialect now sets `useOpFoldResults`. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each memref op that does not have
exactly one fixed result. This patch moves the seven folds of such ops
to the new form: `copy`, `dealloc`, `dma_start`, `dma_wait`,
`extract_strided_metadata`, `prefetch`, and `store`. The behavior of
these folds does not change, except for `extract_strided_metadata`.
The `memref.extract_strided_metadata` fold created `arith.constant` ops
with its own builder and replaced the uses of the constant results
itself. No driver saw these changes. The fold now returns a partial
fold: one replacement for each constant result, and an in-place mark
when it removes a `memref.cast` source. This patch removes the helper
`replaceConstantUsesOf`, which has no other user.
DialectConversion folds an op before it applies the patterns. So the
conversion did not see the new constants and left them unconverted.
Now the conversion legalizes the new constants of the partial fold, as
[29 lines not shown]
[mlir][arith] Use OpFoldResults for arith folds
The arith dialect now sets `useOpFoldResults`. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each arith op that does not have
exactly one fixed result. This patch moves the four folds of such ops to
the new form: `addui_extended`, `subui_extended`, `mulsi_extended`, and
`mului_extended`. The behavior of these folds does not change, with two
exceptions.
The `arith.mulsi_extended` fold now replaces the low result of
`mulsi_extended(x, 1)` by `x`, and keeps the high result. This is also
correct for i1, where the constant `true` is -1. The i1 tests in
`Arith/canonicalize.mlir` change their expected output.
In a graph region or in an unreachable block, the identity fold of
`addui_extended`, `subui_extended`, or `mului_extended` can forward an
operand that is the other result of the same op. The fold also replaces
that result, so the new form drops the replacements of the fold. The
legacy form gave a correct result in this case, because it first moved
[9 lines not shown]
[mlir] Apply partial folds in the greedy pattern rewrite driver
The greedy driver now applies the `OpFoldResults` of a fold:
- A fold that replaces every result erases the op. The driver
materializes every result, also a result without uses, because
`replaceOp` gives the listeners a value for each result.
- A partial fold replaces the uses of each replaced result and keeps
the op. A replaced result without uses gets no constant. The driver
puts the op on the worklist again.
- The materialization of constants is all-or-nothing. When one
constant fails, the driver inserts no constant and applies no
replacement. An in-place change still counts.
- A fold that keeps every result and has no in-place mark fails.
The driver uses `OpBuilder::materializeFoldResults` for the constants.
When `MLIR_ENABLE_EXPENSIVE_PATTERN_API_CHECKS` is on, the driver
compares the fingerprints of the op before and after the fold. A fold
that changes the op and returns failure is a fatal error. A fold that
[10 lines not shown]
[mlir] Apply partial folds in OpBuilder::createOrFold
The multi-result `createOrFold` now uses the new `OpBuilder::tryFold`
overload and `materializeFoldResults`. When the fold replaces only some
results, the new op stays, and the results of `createOrFold` mix the
replacement values and the kept results of the op. Before this patch,
such a fold counted as an in-place fold or as a failure. The new op has
no uses yet, so `createOrFold` materializes every replaced result.
The zero-result `createOrFold` also uses the new overload. A fold of a
zero-result op can only change the op in place, so its behavior does
not change.
A new test pattern builds `test.op_partial_fold` with the multi-result
`createOrFold`. The new test in `test-operation-folder.mlir` checks
that the op stays next to the replacement values.
Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
[mlir][shape] Report the in-place change of the assuming_all fold
The `shape.assuming_all` fold visits its inputs in reverse order. It
erases each constant input from the operands, and it returns null at
the first input that is not constant. If the fold already erased an
input at that point, it changed the op but reported a failure, so no
driver knew about the change. The fold now returns the result of the op
in this case, which is the in-place signal of the single-result fold
form.
No test fails on main. The later patch "[mlir] Apply partial folds in
the greedy pattern rewrite driver" adds a fingerprint check of each fold
to the greedy driver. The check runs only in a build with
`-DMLIR_ENABLE_EXPENSIVE_PATTERN_API_CHECKS=ON`. In such a build of that
patch, with this fix reverted, the existing test
`Dialect/Shape/canonicalize.mlir` fails:
$ mlir-opt -split-input-file -allow-unregistered-dialect \
-canonicalize="test-convergence" \
[4 lines not shown]
[mlir][affine] Use OpFoldResults for affine folds
The affine dialect now sets `useOpFoldResults`. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each affine op that does not have
exactly one fixed result. This patch moves the eight folds of such ops
to the new form: `dma_start`, `dma_wait`, `for`, `if`, `store`,
`prefetch`, `parallel`, and `delinearize_index`. The behavior of these
folds does not change, with two exceptions.
The `affine.delinearize_index` fold now replaces each result whose basis
element is 1 with the constant 0, and keeps the other results. In
`Tensor/bubble-up-extract-slice-op.mlir`, these results now fold to a
constant 0. The new test `Affine/fold-partial.mlir` runs
`-test-single-fold` and `-sccp`, because `-canonicalize` also runs
`DropUnitExtentBasis`, which hides a broken fold.
In a graph region or in an unreachable block, an init of a zero-trip
`affine.for` can be another result of the same loop. The fold also
replaces that result, so the new form drops the replacements of the
[21 lines not shown]