SelectionDAG: Stop emitting kill flags in InstrEmitter
These is no point to maintaining these before register allocation.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[MLIR][Remark] Report why the LLVM remark streamer could not be created
LLVMRemarkStreamer::createToFile dropped the reason it failed: a file that
could not be opened and a serializer that could not be created both turned
into a bare failure, and enableOptimizationRemarksWithLLVMStreamer failed
without a diagnostic.
- createToFile opens the file with mlir::openOutputFile and, like it, takes
an optional errorMessage out-parameter. The serializer error is reported
the same way.
- enableOptimizationRemarksWithLLVMStreamer emits that message as an error
on the context.
Assisted-by: Claude Code (Claude Opus 5.5)
[InstCombine] Fix profile propagation in ldexp.ll (#229908)
In these cases we just reuse the same condition, so we can directly
propagate the metadata.
[InstCombine] Mark profiles unknown in matchSelectFromAndOr (#229963)
In the general case we have no idea of the distribution of the
synthesized condition, so just mark the profile unknown.
Add a generator for Unicode properties. (#229072)
Over the past few years, I have been maintaining the unicode tables
mostly by hand with the help of an unwieldy python script that only
existed on my machine, and sed.
So I figured throwing an AI at the issue would be an improvement over
the status quo.
It took a lot of iteration and refinement but this is very much
AI-produced code.
This introduce a C++ program that parses the UCD data (in the xml format
which i find easier to read) and generate the tables.
I added some documentation.
As it stands, we have 3 different tools to update Unicode data, and some
that still needs manual work.
I plan to consolidate that a bit in future PRs, whenever I have the
[3 lines not shown]
[libc++] Reorder link flags in exception_guard.odr.sh.cpp (#229589)
Moving `%{link_flags}` after the test object files resolves undefined
reference errors to ex: `strcmp` when using linkers that are more strict
about the command line order of objects.
[MLIR][Remark] Report why the LLVM remark streamer could not be created
LLVMRemarkStreamer::createToFile returned FailureOr and dropped the reason:
a file that could not be opened and a serializer that could not be created
both turned into a bare failure, and enableOptimizationRemarksWithLLVMStreamer
failed without a diagnostic.
- createToFile returns llvm::Expected with the open or serializer error.
- enableOptimizationRemarksWithLLVMStreamer emits
"cannot initialize remark output '<file>': <reason>" on the context.
- The output file is opened with the same flags as
llvm::setupLLVMOptimizationRemarks: OF_TextWithCRLF for YAML, OF_None for
bitstream. Before, both used OF_Text, so on Windows YAML remark files now
get CRLF line endings, as LLVM's own remark files do.
Assisted-by: Claude Code (Claude Opus 5.5)
[SLP] NFC: make candidate state ownership explicit
BoUpSLP contains a mixture of state which relates to the current
candidate and state which persists beyond its lifetime. Gather this
state onto a new CanidateState object which belongs to BoUpSLP.
I think this is compatible with the existing modularisation proposal.
Assisted-by: codex
[MLIR][Remark] Look through nested locations for the remark file location
Remark::generateRemark() and Remark::print() only handled a remark whose
location is a FileLineColLoc. Remarks emitted at a NameLoc, FusedLoc or
CallSiteLoc, which is how debug locations usually look, reached the LLVM
remark serializer as "<unknown file>" and printed without a location.
Use findInstanceOf<FileLineColLoc>() in both places, which takes the first
file location in a preorder walk (the callee for a CallSiteLoc).
Assisted-by: Claude Code (Claude Opus 5.5)
[VPlan] Re-run makeScalarizationDecisions after call widening. (#229925)
Widened calls may only use the first lane of some operands, e.g. uniform
and linear parameters of vector variants. Those operands are only known
to be single-scalar once the calls have been widened, so run
makeScalarizationDecisions again after makeCallWideningDecisions.
Re-run scalarization if any calls were widened, so catch additional
opportunities. On its own it is not super useful currently, but enables
a few follow ups (https://github.com/llvm/llvm-project/pull/229099,
BuildStructVector removal)
[libc++] Fix ill-formed return in function_ref's static-call path (#228868)
In function_ref, the current `__statically_callable` branch is ill-formed when
the lambda's return type `_Rp` is void but the target's static operator()
returns non-void. This patch discards the return value when `_Rp` is
void to make it behave the same as `invoke_r<_Rp>`.
Fixes #228717
[MLIR][Remark] Look through nested locations for the remark file location
Remark::generateRemark() and Remark::print() only handled a remark whose
location is a FileLineColLoc. Remarks emitted at a NameLoc, FusedLoc or
CallSiteLoc, which is how debug locations usually look, reached the LLVM
remark serializer as "<unknown file>" and printed without a location.
Use findInstanceOf<FileLineColLoc>() in both places, which takes the first
file location in a preorder walk (the callee for a CallSiteLoc).
Assisted-by: Claude Code (Claude Opus 5.5)
[flang] Keep loops structured when a backward GOTO in their body can be left
A DO loop whose body branched backward, closing a cycle within the body,
was always lowered as unstructured, even when control could leave that
cycle and finish the iteration. Such loops lost their structured form: an
OpenACC loop directive could not take them over, so an independent loop of
this shape ran sequentially.
Reject a backward GOTO only when its cycle can never be left for the end
of the body. A cycle that can be left keeps the loop structured, with the
branching confined to its body, as is already done for other branches that
stay inside the body.
[flang][mlir][OpenMP] Report Fortran names for privatized target maps… (#229823)
(repush of https://github.com/llvm/llvm-project/pull/228195, added test
`target-firstprivate-mapnames.f90` failed on systems that did not have
the flang-rt built on the gpu, so instead split the test into two
separate compile only checks `target-firstprivate-mapnames.f90` and
`omptarget-map-names.mlir`. Actual source / compiler changes are the
same from https://github.com/llvm/llvm-project/pull/228195)
Privatized, firstprivate and implicitly captured variables offloaded to
a target device were reported as "unknown" in LIBOMPTARGET_INFO debug
output. This PR addresses this by adding the variable names to the
offload mapping information:
- `MapsForPrivatizedSymbols` now sets the name on the omp.map.info it
creates,
recovering it from the hlfir.declare uniq name or, for anonymously boxed
values (e.g. a firstprivate array), from a NameLoc on the descriptor's
[64 lines not shown]
[IR][RFC] Add @llvm.mask.beforefirst intrinsic (#203874)
In the loop vectorizer we support early exit vectorization with side
effects by masking instructions after the first taken exit.
To do so it emits a mask that masks off every lane after the first exit
condition is met:
```llvm
%cttz = call i64 @llvm.experimental.cttz.elts(<vscale x 4 x i1> %exitcond, /*isZeroPoison*/ i1 false)
%mask = call <vscale x 4 x i1> @llvm.get.active.lane.mask(i64 0, i64 %cttz)
```
Which in turn is equivalent to
```llvm
%cttz = call i64 @llvm.experimental.cttz.elts(<vscale x 4 x i1> %exitcond, i1 false)
%cttz.head = insertelement <vscale x 4 x i64> poison, i64 %cttz, i32 0
%cttz.splat = shufflevector <vscale x 4 x i64> %cttz.head, <vscale x 4 x 64> poison, <vscale x 4 x i32> zeroinitializer
[35 lines not shown]
[mlir][CIR] Use OpFoldResults for the cir.scope fold
The CIR dialect now sets the `useOpFoldResults` bit. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each CIR op that does not have
exactly one fixed result. `cir.scope` is the only such op with a fold.
Its fold now returns the yielded value directly. The behavior does not
change. The existing test `clang/test/CIR/Transforms/canonicalize.cir`
covers this fold.
Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
[mlir][tblgen] Warn about the deprecated multi-result fold form
The legacy form `LogicalResult fold(FoldAdaptor,
SmallVectorImpl<OpFoldResult> &)` will be removed. This patch makes
`mlir-tblgen -gen-op-decls` warn when a dialect keeps `useOpFoldResults`
at 0 and has an op with `hasFolder` that does not have exactly one fixed
result. The warning points at the dialect definition. A note names the
op that uses the legacy form and comes first in name order.
The warning comes once for each dialect in one `-gen-op-decls` run. A
dialect with ops in more than one `.td` file can warn once for each file
that contains such an op. `-gen-op-defs` does not warn.
No in-tree dialect warns, because each in-tree dialect with such an op
already sets the bit.
The diagnostic follows the `-on-deprecated` option: `none` silences it,
`warn` (the default) warns, and `error` reports an error and fails the
run. `MlirTblgenMain.h` exposes the option value through
[5 lines not shown]
[mlir] Deprecate the legacy fold APIs with a results vector
The legacy `Operation::fold` overloads with a
`SmallVectorImpl<OpFoldResult> &` parameter drop partial folds. The
overloads that return `NormalizedOpFoldResults` keep them.
This patch moves every in-tree caller of the legacy overloads to the
overloads that return `NormalizedOpFoldResults`, except the unit tests
of the legacy APIs. `OpFoldResultsTest.cpp` suppresses the deprecation
warnings for these tests. The test dialect keeps a legacy fold trait for
the fold tests, so `TestOps.cpp` suppresses the warning too. Then the
patch marks these APIs as deprecated: the two legacy `Operation::fold`
overloads, the legacy general `foldTrait` form, the legacy fold hook
overloads of `DynamicOpDefinition`, and the
`DynamicOpDefinition::LegacyFoldHookFn` alias. ODS cannot put an
attribute on an interface method, so only the documentation marks the
legacy `DialectFoldInterface::fold` method as deprecated.
The new code in `cir::CastOp::fold` also fixes a crash. The fold
[21 lines not shown]
[mlir] Use OpFoldResults in more upstream dialects
The scf, shape, sparse_tensor, gpu, math, and builtin dialects now set
the `useOpFoldResults` bit. ODS then declares `OpFoldResults
fold(FoldAdaptor)` for each of their ops that does not have exactly one
fixed result. This patch moves the seven folds of such ops to the new
form: `scf.if`, `shape.split_at`, `sparse_tensor.crd_translate`,
`gpu.memcpy`, `gpu.memset`, `math.sincos`, and
`builtin.unrealized_conversion_cast`.
The behavior of the folds does not change, except in the case below.
Each fold builds the same normalized `OpFoldResults` object that the
legacy adapter built from the old return. The current tests of each
dialect cover these folds.
In a graph region or in an unreachable block, an operand that
`unrealized_conversion_cast` or `sparse_tensor.crd_translate` forwards
can be another result of the same op that the same fold replaces. The
new form drops the replacements of such a fold. For a forward chain, the
[24 lines not shown]