[AMDGPU] Fix missed WMMA C-operand co-exec hazard
The gfx1250 WMMA co-execution hazard check treats only A, B and the
SWMMAC index as registers the in-flight MMA still reads. C (src2 of a
non-SWMMAC WMMA) is missing, so a VALU scheduled into the MMA's shadow
can clobber C and the MMA consumes the new value.
This is latent while C is tied to vdst, since the existing D check then
covers it. It miscompiles where the tie does not hold: for
v_wmma_bf16f32_16x16x32_bf16, whose D is narrower than C, and for the
_threeaddr form of any WMMA.
[InstCombine] Treat `asin`, `asinh`, `atan` and `cbrt` as odd-functions (#227336)
Add `asin`, `asinh`, `atan` and `cbrt` into the odd-functions list. They
should be treated as odd-functions now.
For #227011
[MachineOutliner] Attribute outlined calls to the candidate's last call (#229260)
An outlined call stands in for a whole candidate but can carry only one
debug location, so no choice is correct for every instruction it
replaces. Prioritize the location that keeps unwinding as if the code
were not outlined: a return address is symbolized at the preceding
instruction, so unwinding through a call in the candidate resolves the
caller frame at the outlined call. Using the candidate's first location
could attribute that frame to an unrelated inlined callee, as seen in
ASan reports.
Prefer the last call's location. A candidate ending in a call may be
outlined as a thunk whose tail call returns directly past the outlined
call, making the backtrace match the unoutlined code exactly. With
multiple calls, the location can still be exact for only one of them.
Keep the first location when the candidate has no call, its last call
has no line, or the replacement is a tail branch that nothing returns
to.
Follow-up to llvm/llvm-project#224189.
[AsmPrinter] Fix direct callee operand lookup in CallGraphSection (#218534)
In handleCallsiteForCallgraph, direct callee operand was assumed to
always be at index 0 (MI.getOperand(0)). While true for x86_64, on
ARM Thumb/Thumb-2 (e.g. tBL), operand 0 represent predicate condition
code not the callee operand. Use getCalleeOperand to correctly retrieve
the operand.
Assisted-by: Gemini
[WebKit Checkers] Allow a view into a temporary CanBorrow prvalue without a Borrow<T> (#229495)
Common example:
```
for (auto& x : copyToVector(y)) {
}
```
This is safe without any explicit Borrow<T> because it is syntactically
impossible to name the Vector, so we must have exclusive access to it,
with no other pointers/references/views that can invalidate it.
The one edge case to consider is that the iterator itself might hold a
pointer to the Vector, and vend an API that can invalidate the Vector.
(A pre-existing regression test covers this case.)
To decide that, `isSafeExpr` now also receives the sink type (the type
of the variable, parameter, or lambda capture that receives the value)
[2 lines not shown]
[CIR] Correct 'null' init for array types with non-zero init (#229543)
The member pointers are supposed to be initialized to -1, so an array of
them or a record of them needs to be initialized properly to -1. This
patch makes sure we look through array types/etc to get the correct
initialization.
Also, quite a few places were using 'getZeroAttr' when they meant 'null
init', so this changes that as well.
[RISCV][NFC] Simplify isSupportedStackID (#229117)
The switch case approach meant that we had to add a case every time a
new `TargetStackID` gets added. The original switch was added in
https://reviews.llvm.org/D94465 at which time I guess we had only a few
valid cases.
[CIR] Skip 'dead' branches when emitting an 'if' statement (#229553)
At one point, we actively decided not to skip these, as it would
possibly be useful for static-analysis. However, we're finding that this
is actually taken advantage of in quite a few places (particularly
things that call undefined things in the false branch), so we are
going revert our previous decision and do the FE level omission.
This functionality could potentially be restored in the future, but we
probably would want a CIRSimplify patch to do the dead-branch
elimination that runs all the time, but that would require better
constant folding in CIR.
[msan] Correctly handle llvm::fake_use (as a no-op) (#229591)
The fake_use intrinsic (used by -fextend-variable-liveness) was being
strictly handled (i.e., check that the parameter is fully initialized),
which led to false positives
(https://github.com/llvm/llvm-project/issues/225425). This patch solves
the issue by silently not instrumenting fake_use, since they are, by
definition, not real uses and therefore cannot lead to
use-of-uninitialized-memory.
Fixes: #225425
[libc][math] Improve accuracy for exp*f float-only implementations. (#227881)
Pure Estrin's scheme pushes the rounding errors a bit more than 1 ULP on
non-FMA targets for these functions.
[SelectionDAG] Add ATOMIC_LOAD_FMAXIMUMNUM/FMINIMUMNUM to SDNode::getOperationName. (#229566)
Reorder the other FP min/max nodes to locally match the order in
ISDOpcodes.h.
[SelectionDAG] Add static_assert for ISD::BUILTIN_OP_END to SDNode::getOperationName to encourage updating when opcodes are added. (#229573)
Also add missing case for DEACTIVATION_SYMBOL.
GlobalISel: Use integer types when splitting loads in lowerLoad (#229544)
lowerLoad built the split pieces using the destination type with the
element size changed, which preserved floating-point types. An
unaligned f64 load was decomposed into G_ZEXTLOAD, G_SHL and G_OR on
f32 and f64, which then crashed in AMDGPU RegBankLegalize. Build the
pieces as integers and bitcast to a non-integer result type, as
lowerStore already does.
The pieces are now consistently integer typed, which allows more
constants to be CSEd in the existing tests.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[CIR] Implement ctor-try-body rethrow (#229436)
This came up in a test suite, but we weren't properly re-throwing
exceptions when they were in the body of a constructor's try-body. This
patch mirrors classic-codegen's behavior reasonably well, implementing
this behavior properly.
Side note: this mirrors classic codegen's behavior of including
dtor-try-bodies too, but that isn't implemented yet, so those will just
be an NYI for now.
[Bazel][mlir] Fixes build failures (#229605)
Fixes build failures following commit 4c520f2ae12f and commit
155462f440ff where ROCDLTargetInfo was introduced and referenced across
ROCDL/AMDGPU passes.
[lldb] Add LanguageRuntime::FixupVariableLocation (NFC) (#229602)
Some languages store a variable in a location whose indirection level is
only known at runtime. Swift, for example, emits resilient globals into
a fixed-size buffer; a value that doesn't fit is boxed on the heap and
the buffer holds a pointer to the box. Whether the value fits depends on
the runtime layout of a type, so the compiler can't encode it in the
DWARF location expression.
This patch adds a LanguageRuntime hook that ValueObjectVariable calls
after evaluating a variable's location, so the runtime can adjust it.
The default implementation does nothing.
Assisted-by: Claude
[Bazel] Add Support dependency to amdgpu_tests (#229597)
Fixes amdgpu_tests bazel layering failure introduced in commit
45ebdb6a55db, where AMDGPUUtilsTest.cpp includes
llvm/Support/Compiler.h.
[Bazel][libc] Update platform_file to depend on dup3 (#229595)
Fixes bazel build failure introduced in commit 64d83dcba1cb, where
file.cpp switched from using dup2 to dup3.
AMDGPU/GlobalISel: Use integer types when narrowing loads and stores
The narrowScalar mutation for loads and stores produced untyped scalars.
For FP-typed values this resulted in an untyped G_OR in the lowerLoad
expansion for unaligned private accesses, which failed to select.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[AMDGPU] Fix true16 losing 16-bit subreg operands
Folding a true16 v2s copy such as %2:sreg_32 = COPY %1.lo16 rewrites
its users to read %1.lo16 and relies on legalizeOperandsVALUt16 to
legalize the narrower operand. PHI and REG_SEQUENCE operands have no
register class, so it skipped them, leaving a 16-bit input in a VGPR_32
PHI or a 32-bit REG_SEQUENCE slot. DetectDeadLanes then marked the PHI
input undef and the defining load was deleted.
Widen such operands with a REG_SEQUENCE in legalizeOperandsVALUt16,
which runs both when the copy is folded and when the user is moved to
the VALU. This miscompiled uniform i16 loads feeding PHIs on gfx1250.
Change-Id: I9ddee5b11ff8503f4b2b1d1c4c776a12b270d968
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[RISCV] Fold (sub 0, (srl (and X, (1 << ShAmt)), ShAmt)) -> (sra (shl X, ShAmt2), bits-1) (#229509)
Improves codegen of is.fpclass+select. The AND to test bits of
is.fpclass may get turned into and+srl while the select emits a
neg. Isel will turn the and+srl into shl+srl but it's too late to
fold the neg.
Assisted-by: Claude
AMDGPU/GlobalISel: Use integer types when narrowing loads and stores
The narrowScalar mutation for loads and stores produced untyped scalars.
For FP-typed values this resulted in an untyped G_OR in the lowerLoad
expansion for unaligned private accesses, which failed to select.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
GlobalISel: Use integer types when splitting loads in lowerLoad
lowerLoad built the split pieces using the destination type with the
element size changed, which preserved floating-point types. An
unaligned f64 load was decomposed into G_ZEXTLOAD, G_SHL and G_OR on
f32 and f64, which then crashed in AMDGPU RegBankLegalize. Build the
pieces as integers and bitcast to a non-integer result type, as
lowerStore already does.
The pieces are now consistently integer typed, which allows more
constants to be CSEd in the existing tests.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>