[AMDGPU] Reserve ENABLE_WAVEFRONT_SIZE32 on gfx125
Wave32-only targets (gfx1250, gfx1250-strict, gfx1251, gfx12-5-generic)
reserve the kernel descriptor's ENABLE_WAVEFRONT_SIZE32; it must be 0.
The compiler set it to 1.
The wave size is selectable only on targets with both supports-wave32
and supports-wave64. Elsewhere, clear the bit, stop printing
.amdhsa_wavefront_size32, and reject the directive in the assembler.
The disassembler now rejects a set bit on every target without a
selectable wave size. This replaces the gfx9-only check and also covers
GFX6-GFX8, matching the documentation.
Document which targets select the wave size through this bit.
Change-Id: I0c3942272f03ff5aba04a15ca063b7a02e6fed06
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
DAG: Fix error message casing to match style policy
The developer guidelines suggests that diagnostic messages should start with a
lowercase letter and not end in a period
Co-authored-by: Claude <noreply at anthropic.com>
[offload][omp] Load and resolve device binaries through liboffload
Migrate DeviceTy::loadBinary and global/kernel symbol resolution off
GenericPluginTy::load_binary/get_global/get_function onto liboffload's
Program/Symbol API, encapsulated in a new ProgramTy abstraction that wraps
an ol_program_handle_t. Kernel symbol resolution still needs the plugin's
opaque GenericKernelTy* handle for the legacy launch path, obtained via a
temporary __ol_tgt_GetKernelFromSymbol helper rather than new public
liboffload API surface. Removes the now-dead __tgt_device_binary type and
the corresponding GenericPluginTy methods and exports entries.
[offload][omp] Query device info directly through liboffload
Route DeviceTy::getInfo through olGetDeviceInfo instead of the
plugin's obtain_device_info, and drop the now-unused
GenericPluginTy::obtain_device_info wrapper and its liboffload
export.
[offload][omp] Manage memory allocation through liboffload
Migrate DeviceTy::allocData/deleteData off GenericPluginTy::data_alloc/
data_delete onto liboffload's olMemAlloc*/olMemFree, migrate
targetLockExplicit/targetUnlockExplicit off data_lock/data_unlock onto
olMemRegister/olMemUnregister (fixing a latent bug where these passed
the OpenMP-visible device number instead of the plugin device id), and
migrate DeviceTy::isAccessiblePtr onto a new olMemIsAccessible API
(added with a unit test) since liboffload had no equivalent for
querying accessibility of arbitrary, not-necessarily-liboffload-
allocated pointers. Removes the now-dead GenericPluginTy::data_alloc/
data_delete/data_lock/data_unlock/is_accessible_ptr wrappers and their
exports entries.
[offload][omp] Remove data_fence
All supported backends execute enqueued work on a given queue in
submission order (CUDA, AMDGPU, and Host queues are always in-order;
Level Zero's default and non-default in-order/synchronous command
modes are as well), so the explicit data-fence used to order a
pointer-attachment after prior data transfers is unnecessary. Remove
the DeviceTy/GenericDeviceTy/GenericPluginTy dataFence chain and the
now-unused olQueueBarrier liboffload API added to support it.
[clangd] Handle template template parameter pack indexes correctly in TargetFinder::VisitTemplateSpecializationType (#228996)
The code was already handling template template parameters, but it used
TemplateName::getAsTemplateDecl() which does not work for a template
template parameter referenced inside a pack index.
Fixes https://github.com/clangd/clangd/issues/2712
[mlir][OpenMP] Allow multi-block omp.iterator regions
Frontends lower a locator inside an omp.iterator region like any other
expression, so the region can contain control flow. For example, Flang
lowers LEN_TRIM in `depend(iterator(i=1:n), in: a(i+len_trim(s)))` to a
loop, and later passes inline SUM or COUNT as loops. After control-flow
conversion the region has several blocks, which the single-block
omp.iterator rejects:
%it = omp.iterator(%i: i64) = (%c1 to %n step %c1) {
llvm.br ^scan(%slen : i64)
^scan(%k: i64): // LEN_TRIM skips trailing blanks
...
llvm.cond_br %blank, ^scan(%km1 : i64), ^done
^done:
... // address of a(i+len_trim(s))
omp.yield(%addr : !llvm.ptr)
} -> !omp.iterated<!llvm.ptr>
[9 lines not shown]
[SPIRV] Lower OpenCL sub_group_barrier to OpControlBarrier (#228963)
Recognize the one- and two-argument forms of sub_group_barrier.
Support constant flags and scopes, and diagnose runtime operands instead
of asserting.
Runtime operand lowering remains unsupported.
[Clang][Sema] Refactor checks on for/while loops (NFC) (#226101)
This is an NFC refactor of `clang/lib/Sema/SemaStmt.cpp` based on
@Sirraide's suggestion in
https://github.com/llvm/llvm-project/pull/225748#discussion_r4083103871.
- For both `Sema::ActOnForStmt` and `Sema::ActOnWhileStmt`,
`CommaVisitor` visiting and empty loop handling have been factored out
into a new `CheckConditionalLoop()` function to reduce duplication.
- Also added clarifying comments to explain the need for
`setHasEmptyLoopBodies()` usage in the specific cases of `for`/`while`
loops.
This PR is intended to be merged prior to #225748, which will benefit
from this refactor by having the redundant-defer checks for both
`for`/`while` loops in the same `CheckConditionalLoop()` function
instead of duplicating them.
[SLP] Add new test to expose SLP bug in canBuildSplitNode() (#229190)
Test for #220014.
Creates a tree from a store chain. Afterwards, analyzes a reduction for
possibly being a split-node. Decides a split node would not be
profitable because the shuffle to combine the nodes would be primitively
expensive. However, a shuffle is not needed since the split node is
feeding a reduction. Because the tree wasn't clear from the store chain,
`canBuildSplitNode` incorrectly thinks that this is feeding a store.
[ProfCheck] Remove fixed tests (#229192)
These tests are passing in the profcheck configuration now. Remove them
so that we can detect any regressions. Also sort one or two tests that
were not alphabetized properly.