[mlir][nvgpu] Align integer WGMMA type checks with i8 operands (#212215)
Integer WGMMA uses i8 operands with an i32 accumulator, but the NVGPU
verifier and K-shape selection still checked for i16.
Changes both checks to i8 and adds tests for rejected i16 inputs and the
existing i8 "not supported yet" limitation.
[NFC][TSan] Move ObtainCurrentStack out of ThreadRegistryLock scope
ObtainCurrentStack only reads the current thread's shadow stack and
allocates a VarSizeStackTrace buffer, which does not require
ThreadRegistryLock (and already runs outside ThreadRegistryLock in
ReportRace, SignalUnsafeCall, and ReportErrnoSpoiling).
Move ObtainCurrentStack (and dummy_pc in ReportDeadlock) before
ScopedReport in ReportMutexHeldWrongContext, ReportMutexMisuse,
ReportDeadlock, and ReportDestroyLocked so the stack trace buffers also
outlive ScopedReport and OutputReport.
Assisted-by: Gemini
Pull Request: https://github.com/llvm/llvm-project/pull/228794
[NFC][TSan] Move RestoreStack before ScopedReport in ReportDestroyLocked
Run RestoreStack in its own lock scope before constructing ScopedReport
in ReportDestroyLocked (matching ReportRace). This avoids acquiring
ScopedErrorReportLock or symbolizing the current stack if RestoreStack
fails, and avoids holding slot_lock and slot_mtx while populating the
ScopedReport.
Assisted-by: Gemini
Pull Request: https://github.com/llvm/llvm-project/pull/228643
[TSan] Defer symbolization in AddStack and AddSleep to SymbolizeStackElems
PR #151495 delayed symbolization for memory accesses, locations,
threads, and mutexes until OutputReport (after ThreadRegistryLock is
released), but missed ScopedReport::AddStack and ScopedReport::AddSleep.
As a result, ReportRace (via AddSleep) and ReportMutexMisuse,
ReportDeadlock, and ReportDestroyLocked (via AddStack) still invoked the
symbolizer while holding ThreadRegistryLock.
Store the unsymbolized stack traces and sleep stack ID in ReportDesc and
symbolize them in ScopedReport::SymbolizeStackElems().
Assisted-by: Gemini
Pull Request: https://github.com/llvm/llvm-project/pull/228795
[NFC][TSan] Allocate ScopedReport as a stack variable
Now that ScopedReport is constructed before acquiring ThreadRegistryLock
or slot locks across all reporting functions, it no longer needs to be
constructed inside the lock scope via placement new on __builtin_alloca
storage.
Declare ScopedReport as a normal stack variable before the lock scope
and remove the manual destructor calls.
Assisted-by: Gemini
Pull Request: https://github.com/llvm/llvm-project/pull/228637
[NFC][TSan] Move ObtainCurrentStack out of ThreadRegistryLock scope
ObtainCurrentStack only reads the current thread's shadow stack and
allocates a VarSizeStackTrace buffer, which does not require
ThreadRegistryLock (and already runs outside ThreadRegistryLock in
ReportRace, SignalUnsafeCall, and ReportErrnoSpoiling).
Move ObtainCurrentStack (and dummy_pc in ReportDeadlock) before
ScopedReport in ReportMutexHeldWrongContext, ReportMutexMisuse,
ReportDeadlock, and ReportDestroyLocked so the stack trace buffers also
outlive ScopedReport and OutputReport.
Assisted-by: Gemini
Pull Request: https://github.com/llvm/llvm-project/pull/228794
[TSan] Lock ScopedErrorReportLock before slot and thread_registry locks
OutputReport runs while ScopedErrorReportLock is held after slot_mtx and
thread_registry have been unlocked. Because code executed during
OutputReport (symbolizer, callbacks, or signal handlers) can acquire
slot_mtx or thread_registry, ScopedErrorReportLock must precede slot and
thread_registry locks in the lock hierarchy to avoid AB-BA deadlocks
between concurrent reports or fork().
- Move ScopedErrorReportLock::Lock() before slot.mtx, thread_registry,
and slot_mtx in ForkBefore (and unlock in reverse order in ForkAfter).
- Replace ctx->thread_registry.CheckLocked() in ScopedReportBase's
constructor with CheckedMutex::CheckNoLocks(), and add CheckLocked() to
AddThread(const ThreadContext *) and CheckNoLocks() to OutputReport.
- Construct ScopedReport before acquiring ThreadRegistryLock across all
reporting functions, and close the RestoreStack lock scope before
constructing ScopedReport in ReportRace.
Assisted-by: Gemini
[2 lines not shown]
[NFC][TSan] Move RestoreStack before ScopedReport in ReportDestroyLocked
Run RestoreStack in its own lock scope before constructing ScopedReport
in ReportDestroyLocked (matching ReportRace). This avoids acquiring
ScopedErrorReportLock or symbolizing the current stack if RestoreStack
fails, and avoids holding slot_lock and slot_mtx while populating the
ScopedReport.
Assisted-by: Gemini
Pull Request: https://github.com/llvm/llvm-project/pull/228643
[NFC][TSan] Use in-class member initializers in tsan_report.h
Use in-class member initializers for all structs and classes in
tsan_report.h and default constructors and destructor in
tsan_report.cpp.
Assisted-by: Gemini
Pull Request: https://github.com/llvm/llvm-project/pull/228780
[orc-rt] Make noErrors report errors as test failures (#228807)
noErrors used cantFail, so an unexpected error aborted the whole test
binary without naming the test that hit it, and with assertions disabled
was dropped silently. Use EXPECT_THAT_ERROR instead, so the error is
reported as a failure of the current test and the remaining tests still
run.
rpc_generic.c: Initialize "cp" to shut the compiler up
This patch does not fix any semantics issue.
MFC after: 3 months
Fixes: 884ee8d6c9b4 ("nfscl: Add some glue for client side NFS over RDMA")
[NFC][TSan] Merge ScopedReportBase into ScopedReport (#228640)
ScopedReport is the only subclass of ScopedReportBase and nothing uses
ScopedReportBase directly. Merge ScopedReportBase into ScopedReport.
Assisted-by: Gemini
[SLP]Fix crash on deleted main op of copyable state in reductions
Vectorizing one group of reduced values can erase the main operation of
a later group's copyable state, leaving it with dropped operands.
Fixes #228768
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/228806
[orc-rt] Use the Error matchers in SPSWrapperFunctionTest (#228804)
Use the Error matchers introduced in 4c8a437d0487 to clean up error
tests in SPSWrapperFunctionTest.
Re-enable Rescan when a WiFi scan fails
A failure in networkdictionary() inside the rescan thread, or the
WiFi card vanishing from the new scan, left the Rescan button greyed
out for good. Restore it in a finally block and skip the access point
list update when the card is gone.
[NFC][TSan] Move ObtainCurrentStack out of ThreadRegistryLock scope
ObtainCurrentStack only reads the current thread's shadow stack and
allocates a VarSizeStackTrace buffer, which does not require
ThreadRegistryLock (and already runs outside ThreadRegistryLock in
ReportRace, SignalUnsafeCall, ReportErrnoSpoiling, and
ReportDestroyLocked).
Move ObtainCurrentStack (and dummy_pc in ReportDeadlock) before
ScopedReport in ReportMutexHeldWrongContext, ReportMutexMisuse, and
ReportDeadlock so the stack trace buffers also outlive ScopedReport and
OutputReport.
Assisted-by: Gemini
Pull Request: https://github.com/llvm/llvm-project/pull/228794
[NFC][TSan] Move RestoreStack before ScopedReport in ReportDestroyLocked
Run RestoreStack in its own lock scope before constructing ScopedReport
in ReportDestroyLocked (matching ReportRace). This avoids acquiring
ScopedErrorReportLock or symbolizing the current stack if RestoreStack
fails, and avoids holding slot_lock and slot_mtx while populating the
ScopedReport.
Assisted-by: Gemini
Pull Request: https://github.com/llvm/llvm-project/pull/228643
[NFC][TSan] Merge ScopedReportBase into ScopedReport
ScopedReport is the only subclass of ScopedReportBase and nothing uses
ScopedReportBase directly. Merge ScopedReportBase into ScopedReport.
Assisted-by: Gemini
Pull Request: https://github.com/llvm/llvm-project/pull/228640
www/deno: map deno compile payload instead of copying it
Executables produced by `deno compile` read their embedded payload into
memory at startup on FreeBSD, so a 200 MB payload made a trivial app
start with 431 MB RSS instead of 48 MB. Map the payload section from
the executable instead, so it is paged in on demand like on other
platforms.
[CodeGen] Declare command line options in TableGen
Move the cl::opts of TargetPassConfig.cpp and CodeGenPrepare.cpp into
CodeGenOptions.td, private to lib/CodeGen; other files will follow.
-enable-machine-outliner (cl::ValueOptional) and -regalloc
(RegisterPassParser) stay cl::opt.
CodeGenPrepare and its addressing-mode helpers hold
`const CodeGenOptions &Opts`; TargetPassConfig functions read
CodeGenOptions::Global. getCGPassBuilderOption() converts
std::optional<bool> members to the cl::boolOrDefault fields of the
public CGPassBuilderOption. -basic-block-section-match-infer, which was
not cl::Hidden, is now listed by -help-hidden only.
Aided by Opus 5.5
[mlir][tblgen] Warn about the deprecated multi-result fold form
The legacy form `LogicalResult fold(FoldAdaptor,
SmallVectorImpl<OpFoldResult> &)` will be removed. This patch makes
`mlir-tblgen -gen-op-decls` warn when a dialect keeps `useOpFoldResults`
at 0 and has an op with `hasFolder` that does not have exactly one fixed
result. The warning points at the dialect definition. A note names the
first op that uses the legacy form.
The warning comes once for each dialect in one `-gen-op-decls` run. A
dialect with ops in more than one `.td` file can warn once for each file
that contains such an op. `-gen-op-defs` does not warn.
No in-tree dialect warns, because each in-tree dialect with such an op
already sets the bit.
The diagnostic follows the `-on-deprecated` option: `none` silences it,
`warn` (the default) warns, and `error` reports an error and fails the
run. `MlirTblgenMain.h` exposes the option value through
[3 lines not shown]
[mlir] Deprecate the legacy fold APIs with a results vector
The legacy `Operation::fold` overloads with a
`SmallVectorImpl<OpFoldResult> &` parameter drop partial folds. The
overloads that return `OpFoldResults` keep them.
This patch moves every in-tree caller of the legacy overloads to the
overloads that return `OpFoldResults`, except the unit tests of the
legacy APIs. `OpFoldResultsTest.cpp` suppresses the deprecation warnings
for these tests. Then the patch marks these APIs as deprecated: the two
legacy `Operation::fold` overloads, the legacy general `foldTrait` form,
the legacy fold hook overloads of `DynamicOpDefinition`, and the
`DynamicOpDefinition::LegacyFoldHookFn` alias. ODS cannot put an
attribute on an interface method, so only the documentation marks the
legacy `DialectFoldInterface::fold` method as deprecated.
The new code in `cir::CastOp::fold` also fixes two bugs. The fold
crashed when the source of an integral cast was a block argument, and it
read the fold result of result 0 when the source was a different result
[17 lines not shown]
[mlir][CIR] Use OpFoldResults for the cir.scope fold
The CIR dialect now sets the `useOpFoldResults` bit. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each CIR op that does not have
exactly one fixed result. `cir.scope` is the only such op with a fold.
Its fold now returns the yielded value directly. The behavior does not
change. The existing test `clang/test/CIR/Transforms/canonicalize.cir`
covers this fold.
Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
[mlir] Use OpFoldResults in more upstream dialects
The scf, shape, sparse_tensor, gpu, math, and builtin dialects now set
the `useOpFoldResults` bit. ODS then declares `OpFoldResults
fold(FoldAdaptor)` for each of their ops that does not have exactly one
fixed result. This patch moves the seven folds of such ops to the new
form: `scf.if`, `shape.split_at`, `sparse_tensor.crd_translate`,
`gpu.memcpy`, `gpu.memset`, `math.sincos`, and
`builtin.unrealized_conversion_cast`.
The behavior of the folds does not change, except in the graph-region
case below. Each fold builds the same normalized `OpFoldResults` object
that the legacy adapter built from the old return. The existing tests of
each dialect cover these folds.
In a graph region, an operand that `unrealized_conversion_cast` or
`sparse_tensor.crd_translate` forwards can be another result of the
same op that the same fold replaces. The new form drops such a fold.
For a forward chain, the legacy form gave a correct result by the order
[22 lines not shown]
[mlir][linalg] Use OpFoldResults for linalg folds
The linalg dialect now sets the `useOpFoldResults` bit. ODS then
declares `OpFoldResults fold(FoldAdaptor)` for each linalg op that does
not have exactly one fixed result. This patch moves the eleven folds of
such ops in `LinalgOps.cpp` to the new form. It also changes the fold
template of `mlir-linalg-ods-yaml-gen`, so the generated named ops use
the new form too.
The behavior of the folds does not change. A fold that returned the
result of `memref::foldMemRefCast` now returns the same in-place state
through the `LogicalResult` constructor. The folds of `transpose`,
`pack`, and `unpack` return their replacement value directly.
The yaml-gen test now checks the generated fold definition.
A build directory that uses a native `mlir-linalg-ods-yaml-gen` does
not regenerate `LinalgNamedStructuredOps.yamlgen.cpp.inc` when the tool
changes. Delete that file before the build.
[2 lines not shown]
[mlir][vector] Use OpFoldResults for vector folds
The vector dialect now sets the `useOpFoldResults` bit. ODS then
declares `OpFoldResults fold(FoldAdaptor)` for each vector op that does
not have exactly one fixed result. This patch moves the five folds of
such ops to the new form: `to_elements`, `transfer_write`, `store`,
`masked_store`, and `mask`.
The behavior of the folds does not change, except in the graph-region
case below. A fold that returned success with an empty vector now
returns `success()`, which is an in-place change. A fold that filled the
vector now returns its values.
The all-true fold of `vector.mask` moves the masked op out of the
region, and the terminator then has null operands. So this fold must
replace every result, and the driver erases the op. A mask without
results has nothing to replace, so the fold reports the move as an
in-place change.
[26 lines not shown]
[mlir][memref] Use OpFoldResults for memref folds
The memref dialect now sets `useOpFoldResults`. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each memref op that does not have
exactly one fixed result. This patch moves the seven folds of such ops
to the new form: `copy`, `dealloc`, `dma_start`, `dma_wait`,
`extract_strided_metadata`, `prefetch`, and `store`. The behavior of
these folds does not change, except for `extract_strided_metadata`.
The `memref.extract_strided_metadata` fold created `arith.constant` ops
with its own builder and replaced the uses of the constant results
itself. No driver saw these changes. The fold now returns a partial
fold: one replacement for each constant result, and an in-place mark
when it removes a `memref.cast` source. This patch removes the helper
`replaceConstantUsesOf`, which has no other user.
DialectConversion folds an op before it applies the patterns. So the
conversion did not see the new constants and left them unconverted.
With pattern rollback on, the conversion now does not apply the partial
[27 lines not shown]