AMDGPU: Give v_cvt_sr_pk_bf16_f32 its own subtarget feature
v_cvt_sr_pk_bf16_f32 was gated on bf16-cvt-insts, but that feature is
also present on gfx950 where the (non-sr) v_cvt_pk_bf16_f32 was first
added. The stochastic-rounding v_cvt_sr_pk_bf16_f32 was only added
for gfx1250 and has no gfx950 encoding, so it would mis-select and
later hit the "Invalid opcode" assert. Introduce cvt-sr-pk-bf16-f32-inst,
currently added to gfx13 and 125*
Co-authored-by: Claude (Claude-Opus-4.8)
[clang-format] Prevent re-assigning type on finalized tokens (#210763)
Prevents a finalized token inside modifyContext from being reassigned
through the setType member function by checking if the token is
finalized.
This ensures ill-defined code like does not trigger an assertion failure
during reformatting.
Fixes #210509
[Offload][Test] Add llvm bin in the PATH for offload-unit suite (#213149)
This PR makes the offload-unit suite put llvm bin directory on PATH so
that tests need lld can find. It fixes the issue exposed in:
https://github.com/llvm/llvm-project/pull/212860
[SBVec] Add top-down vectorization to the unified Sandbox Vectorizer
Extend the Sandbox Vectorizer's `bottom-up-vec` pass so a single
implementation can vectorize in either direction, and add the top-down
strategy that walks def-use chains forward from a seed.
Direction selection
--------------------
The pass direction is chosen from the Region's auxiliary pass argument:
"bottom-up" (or empty, the default) and "top-down" map onto a
SchedDirection, and any other value is rejected with a fatal usage error.
The vectorizer always runs in the same direction as the scheduler.
Top-down traversal
------------------
Bottom-up starts from a seed slice (e.g. stores to consecutive addresses)
and recurses into operands. Top-down instead starts from a seed of
consecutive loads and recurses into *users*:
[32 lines not shown]
[SandboxIR] Fix notifyEraseInstr to skip scheduled neighbors
Guard both loops with !PredN->scheduled() / !SuccN->scheduled() so
scheduled neighbors are left untouched, and add a unit test that erases
a node with one scheduled and one unscheduled predecessor to cover the
fix.
[AMDGPU] Use a single SubtargetPredicate for fp8/bf8 -> f32 conversions (#212888)
Add FeatureCvtFP8SDWASrcSel for the form that takes the byte from the SDWA src0_sel field and FeatureCvtFP8ByteSel for the form that takes it from a byte_sel operand. Both imply FeatureFP8ConversionInsts, and a subtarget provides one of them, so the multiclass-generated HasCvtFP8SDWASrcSel and HasCvtFP8ByteSel replace the GFX9 and GFX11Plus predicates outright.
Assisted-By: Claude Opus 5
[flang] Add a pass to get OpenACC device ptr for CUDA kernel (#212299)
When a CUDA kernel is launched inside an OpenACC data region, it does
not properly get the device pointers and instead uses host data. This PR
adds a pass to get the device pointers set up by the OpenACC data
construct and pass them explicitly to the CUDA kernel.
Note: A possibly non-contiguous array argument is currently not supported. This pass will skip these cases.
---------
Co-authored-by: Yebin Chon <ychon at nvidia.com>
[mlir][OpenACC] Atomicize contended shared array reduction updates (#212971)
Example:
```fortran
!$acc parallel loop gang reduction(+:a)
do i = 1, N
!$acc loop worker reduction(+:a)
do j = 1, M
a(i) = a(i) + b(j,i)
end do
end do
```
In this code, the worker accumulator is block-shared, so every worker
updates
the same element with a plain read-modify-write and all but one partial
is lost.
Fix: make in-place updates of a block-shared array accumulator atomic,
[7 lines not shown]