[AMDGPU] Fix missed WMMA C-operand co-exec hazard
The gfx1250 WMMA co-execution hazard check treats only A, B and the
SWMMAC index as registers the in-flight MMA still reads. C (src2 of a
non-SWMMAC WMMA) is missing, so a VALU scheduled into the MMA's shadow
can clobber C and the MMA consumes the new value.
This is latent while C is tied to vdst, since the existing D check then
covers it. It miscompiles where the tie does not hold: for
v_wmma_bf16f32_16x16x32_bf16, whose D is narrower than C, and for the
_threeaddr form of any WMMA.
18471 simnet rejects its maximum MTU and reports an inaccurate MTU range
Reviewed by: Robert Mustacchi <rm at fingolfin.org>
Approved by: Dan McDonald <danmcd at oxide.computer>
[tsan]: fix Go race syso build on s390x with GCC (#225217)
Native GCC on s390x defines __GCC_HAVE_SYNC_COMPARE_AND_SWAP_16, which
causes the SpinMutex func_cas overload for a128 to be skipped. The
generic template emits a __sync_val_compare_and_swap_16 libcall instead
of inlining CDSG (GCC cannot prove alignment) and that symbol
is not available without libatomic, which the Go race syso does not
link.
Extend the SpinMutex condition to also cover SANITIZER_GO builds,
regardless of __GCC_HAVE_SYNC_COMPARE_AND_SWAP_16. This is safe since
all atomic accesses in Go go through the TSan trampolines.
Verified with s390x-linux-gnu-g++ (GCC 12): without the fix the syso
contains an unresolved __sync_val_compare_and_swap_16 reference; with
the fix all three __tsan_go_atomic128_* symbols are present and no
libatomic references remain.
Tell the target when the users of a cost context stay scalar
Add a hint to getArithmeticInstrCost that says whether the users of the
context use the priced operation or see its lanes extracted. SLP passes
it for vector binops, the users stay scalar when the user node is a
gather or is not vectorized.
[Mips][RuntimeDyld] Reserve t9 for static JIT stubs (#228594)
MIPS PIC functions expect their entry address in t9 so their prologue
can compute gp. However, R_MIPS_26 is used for both function calls and
jumps within a function. A local jump is not a call boundary, so t9 may
still hold a live value there even though it is caller-saved.
Commit 458a983df422df85363ec40b1f1556f55b238aac switched these stubs
from t9 to at to preserve live t9 values across local jumps. That breaks
calls from JIT code with the static relocation model into PIC functions,
because the callee receives the wrong t9 and computes the wrong gp.
Revert that change so every MIPS stub materializes its destination in
t9. Reserve t9 in codegen to avoid stubs corrupting live values in
registers.
Fixes #228414.
[mlir][scf] Specialize loops with constant affine.min operands (#228909)
`scf-for-loop-specialization` only inspected affine map results, so it
missed constants supplied through map operands. Canonicalize the map and
operands before finding constant results.
Fixes #228470
[MLIR] ValueBoundsOpInterface: slice size from source size (#221756)
Improve ValueBoundsOpInterface for tensor.extract_slice by adding an
upper bound of `ceil((sourceSize - offset) / stride)`.
[LV] Avoid repeatedly expanding ignored operand graphs. (NFC) (#228944)
Add a set to track the already processed operations in DeadOps, to avoid
repeatedly processing the same instruction chains for ops with shared
operands.
Without the fix, the newly added test takes a long time to compile.
Addresses part of https://github.com/llvm/llvm-project/issues/228403.
[KnownFPClass] Add output/input denormal helpers (#225575)
Added `applyInputDenormalMode(KnownSrc, Mode)` and
`applyOutputDenormalMode(KnownSrc, Mode)`. These helpers allow us to
correctly handle the output/input denormal mode correctly every single
time, while also being much easier to use.
Example usage:
```c++
KnownFPClass func(const KnownFPClass& KnownSrc_, DenormalMode Mode) {
KnownFPClass KnownSrc = applyInputDenormalMode(KnownSrc_, Mode);
KnownFPClass Known;
// We can treat everything here as if it were IEEE.
return applyOutputDenormalMode(Known, Mode);
}
```
[2 lines not shown]
[mlir][ArmSME][NFC] Improving the readability of outer product fusion (#227140)
This NFC improves the readability of the outer-product fusion in ArmSME.
---------
Signed-off-by: Federico Bruzzone <federico.bruzzone.i at gmail.com>
[LV] Add tests for reductions of ext(mul) with narrow wrapping mul (NFC) (#228943)
Add tests for partial reductions and in-loop multiply-accumulate
reductions of ext(mul(ext(A), ext(B))) where the narrow mul may or may
not wrap with respect to the extend kinds.
Some of those combinations are currently miscompiled.
https://alive2.llvm.org/ce/z/Ys8bNn
[LV] Use a ResumeForEpilogue marker for the main vector TC (NFC) (#228942)
Previously the vector trip count of the main loop was recovered via IR
based pattern matching.
Instead, add a ResumeForEpilogue marker for the main plan's vector trip
count to its middle block, which records the generated value, and pass
it to addMinimumVectorEpilogueIterationCheck as a VPValue. The check is
only reached after the main vector loop, so it uses that value directly
rather than the resume phi merging it with the bypass value.
This makes matching more robust and is another step towards modeling the
full epilogue skeleton in VPlan.