[NFCI][llvm-jitlink] Return 1 instead of exit(1) on entry point error (#228311)
In main(), if EntryPoint lookup fails, llvm-jitlink was calling
exit(1). Because exit() is noreturn, the stack is not unwound and
destructors for Session S and its associated objects are not run.
LSan detects these in-flight allocations (such as dlsym buffers
in InProcessDylibManager) as memory leaks at atexit, causing
ExecutionEngine/JITLink/x86-64/COFF_directive_alternatename_fail.s
to fail on ASan/LSan buildbots.
Returning 1 from main() allows S and other local variables to be cleanly
destructed before process termination.
For https://github.com/llvm/llvm-project/issues/186534
Fixes https://lab.llvm.org/buildbot/#/builders/169/builds/27104
Log:
https://lab.llvm.org/buildbot/#/builders/169/builds/27104/steps/11/logs/stdio
Assisted-by: Gemini
[clang] Fix an accepts-invalid related to concepts and parameter packs (#227494)
HashParameterMapping computes the keys of
UnsubstitutedConstraintSatisfactionCache. In a fold expression, it
hashes the current element for every use of the expanded pack. But in
P...[sizeof(P)], P refers to the whole pack, not the current element. So
two checks with the same element at the same position got the same key
even if the indexed pack differs, and the second one used the first
one's result. This accepted e.g.
template <class T> concept two_bytes = sizeof(T) == 2;
template <typename... P>
void f() requires(two_bytes<P...[sizeof(P)]> && ...) {}
void g() {
f<char, short, short>();
f<char, int, short, short, short>(); // P...[sizeof(char)] is int
}
[4 lines not shown]
[flang][cuda] Diagnose host reads of device data (#228287)
Host code may only use device data in a data transfer assignment or as
an
actual argument. When device data appeared anywhere else, such as in an
IF
condition, flang accepted it silently and lowered it to a plain host
load of
device memory:
subroutine s(h, a)
double precision :: h, a(10)
attributes(device) :: a
if (a(3) > 2.5d0) h = 1.0d0
end subroutine
The CUDA checker now emits an error when host code reads data with the
DEVICE or CONSTANT attribute in a scalar expression. These expressions
include IF and ELSE IF conditions, DO bounds, DO WHILE, SELECT CASE
[12 lines not shown]
[mlir][scf] Fold Self-addition of the induction variable into the loop range (#228725)
scf-for-loop-range-folding only folded an arith op into the loop range
if the induction variable had a single use (hasOneUse()). That counts
operand slots rather than users, so `arith.addi %i, %i`, a single
operation that uses the induction variable twice, was rejected.
now we check whether the induction variable is used in a single
operation.
Fixes #228474
Assisted-by: Claude
[RegAllocFast] Lower tied operands, absorbing TwoAddressInstructionPass (#225316)
TwoAddressInstructionPass inserts a copy for every tied operand it
cannot rewrite, which register allocation then tries to fold away. Teach
RegAllocFast to lower tied operands itself so that the pipeline can skip
the pass:
* A tied use that dies at the instruction takes over the tied def's
register when its class contains it, and is otherwise copied into it,
with the instruction's other reads of the value following the copy.
* Copy hints follow ties as chain links, so argument copies feeding
two-address chains still fold.
* REG_SEQUENCE and INSERT_SUBREG expand to subregister COPYs; an undef
REG_SEQUENCE source needs one only where a use reads its lane.
PHIElimination still runs, so the allocator's input is not SSA. The
lowering keys on the TiedOpsRewritten property, which the allocator now
sets itself, so partial pipelines (-run-pass, -start-before) follow the
MIR they are given.
[10 lines not shown]
[AMDGPU] Fix missed WMMA C-operand co-exec hazard
The gfx1250 WMMA co-execution hazard check treats only A, B and the
SWMMAC index as registers the in-flight MMA still reads. C (src2 of a
non-SWMMAC WMMA) is missing, so a VALU scheduled into the MMA's shadow
can clobber C and the MMA consumes the new value.
This is latent while C is tied to vdst, since the existing D check then
covers it. It miscompiles where the tie does not hold: for
v_wmma_bf16f32_16x16x32_bf16, whose D is narrower than C, and for the
_threeaddr form of any WMMA.
[tsan]: fix Go race syso build on s390x with GCC (#225217)
Native GCC on s390x defines __GCC_HAVE_SYNC_COMPARE_AND_SWAP_16, which
causes the SpinMutex func_cas overload for a128 to be skipped. The
generic template emits a __sync_val_compare_and_swap_16 libcall instead
of inlining CDSG (GCC cannot prove alignment) and that symbol
is not available without libatomic, which the Go race syso does not
link.
Extend the SpinMutex condition to also cover SANITIZER_GO builds,
regardless of __GCC_HAVE_SYNC_COMPARE_AND_SWAP_16. This is safe since
all atomic accesses in Go go through the TSan trampolines.
Verified with s390x-linux-gnu-g++ (GCC 12): without the fix the syso
contains an unresolved __sync_val_compare_and_swap_16 reference; with
the fix all three __tsan_go_atomic128_* symbols are present and no
libatomic references remain.
Tell the target when the users of a cost context stay scalar
Add a hint to getArithmeticInstrCost that says whether the users of the
context use the priced operation or see its lanes extracted. SLP passes
it for vector binops, the users stay scalar when the user node is a
gather or is not vectorized.
[Mips][RuntimeDyld] Reserve t9 for static JIT stubs (#228594)
MIPS PIC functions expect their entry address in t9 so their prologue
can compute gp. However, R_MIPS_26 is used for both function calls and
jumps within a function. A local jump is not a call boundary, so t9 may
still hold a live value there even though it is caller-saved.
Commit 458a983df422df85363ec40b1f1556f55b238aac switched these stubs
from t9 to at to preserve live t9 values across local jumps. That breaks
calls from JIT code with the static relocation model into PIC functions,
because the callee receives the wrong t9 and computes the wrong gp.
Revert that change so every MIPS stub materializes its destination in
t9. Reserve t9 in codegen to avoid stubs corrupting live values in
registers.
Fixes #228414.
[mlir][scf] Specialize loops with constant affine.min operands (#228909)
`scf-for-loop-specialization` only inspected affine map results, so it
missed constants supplied through map operands. Canonicalize the map and
operands before finding constant results.
Fixes #228470
[MLIR] ValueBoundsOpInterface: slice size from source size (#221756)
Improve ValueBoundsOpInterface for tensor.extract_slice by adding an
upper bound of `ceil((sourceSize - offset) / stride)`.
[LV] Avoid repeatedly expanding ignored operand graphs. (NFC) (#228944)
Add a set to track the already processed operations in DeadOps, to avoid
repeatedly processing the same instruction chains for ops with shared
operands.
Without the fix, the newly added test takes a long time to compile.
Addresses part of https://github.com/llvm/llvm-project/issues/228403.
[KnownFPClass] Add output/input denormal helpers (#225575)
Added `applyInputDenormalMode(KnownSrc, Mode)` and
`applyOutputDenormalMode(KnownSrc, Mode)`. These helpers allow us to
correctly handle the output/input denormal mode correctly every single
time, while also being much easier to use.
Example usage:
```c++
KnownFPClass func(const KnownFPClass& KnownSrc_, DenormalMode Mode) {
KnownFPClass KnownSrc = applyInputDenormalMode(KnownSrc_, Mode);
KnownFPClass Known;
// We can treat everything here as if it were IEEE.
return applyOutputDenormalMode(Known, Mode);
}
```
[2 lines not shown]
[mlir][ArmSME][NFC] Improving the readability of outer product fusion (#227140)
This NFC improves the readability of the outer-product fusion in ArmSME.
---------
Signed-off-by: Federico Bruzzone <federico.bruzzone.i at gmail.com>
[LV] Add tests for reductions of ext(mul) with narrow wrapping mul (NFC) (#228943)
Add tests for partial reductions and in-loop multiply-accumulate
reductions of ext(mul(ext(A), ext(B))) where the narrow mul may or may
not wrap with respect to the extend kinds.
Some of those combinations are currently miscompiled.
https://alive2.llvm.org/ce/z/Ys8bNn
[LV] Use a ResumeForEpilogue marker for the main vector TC (NFC) (#228942)
Previously the vector trip count of the main loop was recovered via IR
based pattern matching.
Instead, add a ResumeForEpilogue marker for the main plan's vector trip
count to its middle block, which records the generated value, and pass
it to addMinimumVectorEpilogueIterationCheck as a VPValue. The check is
only reached after the main vector loop, so it uses that value directly
rather than the resume phi merging it with the bypass value.
This makes matching more robust and is another step towards modeling the
full epilogue skeleton in VPlan.
[ADT] Remove zombie-instance handling from SmallDenseMapStorage (NFC) (#228925)
This patch removes the redundant storage.Large.NumBuckets == 0 check
and the two dead stores to the field from SmallDenseMapStorage, and
simplifies DenseMapStorage::maybeMoveFast to use direct assignment
instead of swap.
SmallDenseMapStorage used to have a zombie state:
!Small && storage.Large.NumBuckets == 0
in grow() so that ~DenseMapBase(), which calls deallocateBuckets(),
could safely destruct the moved-from temporary DenseMapBase instance.
Likewise, DenseMapStorage::maybeMoveFast used swap(Other) so that the
temporary instance would not deallocate the moved-from buffer upon
destruction.
With #228802, grow() directly uses StorageT for the temporary storage,
which has a trivial destructor, so neither workaround is needed.
Assisted-by: Antigravity
[llvm-objcopy][COFF] Keep COMDAT section definition symbols when stripping (#228783)
--strip-unneeded and --discard-all removed the section definition symbol of
a COMDAT section if no relocation referenced it directly, which is the
common case since relocations target the COMDAT leader. The aux record of
that symbol carries the COMDAT selection; without it lld discards the
section as a leaderless COMDAT and silently drops relocations against it.
This broke e.g. .refptr.* sections in MinGW objects, producing executables
that load garbage instead of the referenced pointer.
Assisted-by: Claude Opus 5.5
[mlir][tblgen] Warn about the deprecated multi-result fold form
The legacy form `LogicalResult fold(FoldAdaptor,
SmallVectorImpl<OpFoldResult> &)` will be removed. This patch makes
`mlir-tblgen -gen-op-decls` warn when a dialect keeps `useOpFoldResults`
at 0 and has an op with `hasFolder` that does not have exactly one fixed
result. The warning points at the dialect definition. A note names the
op that uses the legacy form and comes first in name order.
The warning comes once for each dialect in one `-gen-op-decls` run. A
dialect with ops in more than one `.td` file can warn once for each file
that contains such an op. `-gen-op-defs` does not warn.
No in-tree dialect warns, because each in-tree dialect with such an op
already sets the bit.
The diagnostic follows the `-on-deprecated` option: `none` silences it,
`warn` (the default) warns, and `error` reports an error and fails the
run. `MlirTblgenMain.h` exposes the option value through
[6 lines not shown]
[mlir] Deprecate the legacy fold APIs with a results vector
The legacy `Operation::fold` overloads with a
`SmallVectorImpl<OpFoldResult> &` parameter drop partial folds. The
overloads that return `OpFoldResults` keep them.
This patch moves every in-tree caller of the legacy overloads to the
overloads that return `OpFoldResults`, except the unit tests of the
legacy APIs. `OpFoldResultsTest.cpp` suppresses the deprecation warnings
for these tests. The test dialect keeps a legacy fold trait for the fold
tests, so `TestOps.cpp` suppresses the warning too. Then the patch marks
these APIs as deprecated: the two legacy `Operation::fold` overloads,
the legacy general `foldTrait` form, the legacy fold hook overloads of
`DynamicOpDefinition`, and the `DynamicOpDefinition::LegacyFoldHookFn`
alias. ODS cannot put an attribute on an interface method, so only the
documentation marks the legacy `DialectFoldInterface::fold` method as
deprecated.
The new code in `cir::CastOp::fold` also fixes a crash. The fold
[22 lines not shown]
[mlir][CIR] Use OpFoldResults for the cir.scope fold
The CIR dialect now sets the `useOpFoldResults` bit. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each CIR op that does not have
exactly one fixed result. `cir.scope` is the only such op with a fold.
Its fold now returns the yielded value directly. The behavior does not
change. The existing test `clang/test/CIR/Transforms/canonicalize.cir`
covers this fold.
Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
[mlir] Use OpFoldResults in more upstream dialects
The scf, shape, sparse_tensor, gpu, math, and builtin dialects now set
the `useOpFoldResults` bit. ODS then declares `OpFoldResults
fold(FoldAdaptor)` for each of their ops that does not have exactly one
fixed result. This patch moves the seven folds of such ops to the new
form: `scf.if`, `shape.split_at`, `sparse_tensor.crd_translate`,
`gpu.memcpy`, `gpu.memset`, `math.sincos`, and
`builtin.unrealized_conversion_cast`.
The behavior of the folds does not change, except in the case below.
Each fold builds the same normalized `OpFoldResults` object that the
legacy adapter built from the old return. The current tests of each
dialect cover these folds.
In a graph region or in an unreachable block, an operand that
`unrealized_conversion_cast` or `sparse_tensor.crd_translate` forwards
can be another result of the same op that the same fold replaces. The
new form drops the replacements of such a fold. For a forward chain, the
[25 lines not shown]
[mlir][linalg] Use OpFoldResults for linalg folds
The linalg dialect now sets the `useOpFoldResults` bit. ODS then
declares `OpFoldResults fold(FoldAdaptor)` for each linalg op that does
not have exactly one fixed result. This patch moves the eleven folds of
such ops in `LinalgOps.cpp` to the new form. It also changes the fold
template of `mlir-linalg-ods-yaml-gen`, so the generated named ops use
the new form too.
The behavior of the folds does not change. A fold that returned the
result of `memref::foldMemRefCast` now returns the same in-place state
through the `LogicalResult` constructor. The folds of `transpose`,
`pack`, and `unpack` return their replacement value directly.
The yaml-gen test now checks the generated fold definition.
A build directory that uses a native `mlir-linalg-ods-yaml-gen` does
not regenerate `LinalgNamedStructuredOps.yamlgen.cpp.inc` when the tool
changes. Delete that file before the build.
[5 lines not shown]
[mlir][vector] Use OpFoldResults for vector folds
The vector dialect now sets the `useOpFoldResults` bit. ODS then
declares `OpFoldResults fold(FoldAdaptor)` for each vector op that does
not have exactly one fixed result. This patch moves the five folds of
such ops to the new form: `to_elements`, `transfer_write`, `store`,
`masked_store`, and `mask`.
The behavior of the folds does not change, except in the case below. A
fold that returned success with an empty vector now returns `success()`,
which is an in-place change. A fold that filled the vector now returns
its values.
The all-true fold of `vector.mask` moves the masked op out of the
region, and the terminator then has null operands. So this fold must
replace every result, and the driver erases the op. A mask without
results has nothing to replace, so the fold reports the move as an
in-place change.
[30 lines not shown]
[mlir][memref] Use OpFoldResults for memref folds
The memref dialect now sets `useOpFoldResults`. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each memref op that does not have
exactly one fixed result. This patch moves the seven folds of such ops
to the new form: `copy`, `dealloc`, `dma_start`, `dma_wait`,
`extract_strided_metadata`, `prefetch`, and `store`. The behavior of
these folds does not change, except for `extract_strided_metadata`.
The `memref.extract_strided_metadata` fold created `arith.constant` ops
with its own builder and replaced the uses of the constant results
itself. No driver saw these changes. The fold now returns a partial
fold: one replacement for each constant result, and an in-place mark
when it removes a `memref.cast` source. This patch removes the helper
`replaceConstantUsesOf`, which has no other user.
DialectConversion folds an op before it applies the patterns. So the
conversion did not see the new constants and left them unconverted.
Now the conversion legalizes the new constants of the partial fold, as
[32 lines not shown]