[MLIR][ODS] Require the value after oilist keyword (#229334)
The generated parser treated optional attribute inside an oilist clause
as optional even after its keyword was parsed
As a result, a keyword without a value was accepted, but the attribute
was simply dropped
Prerequisite for https://github.com/llvm/llvm-project/pull/228977
[lldb][NativePDB] Use host arch for thread locals test (#229340)
The test used to use `--arch=64`, but technically, it doesn't care about
the architecture, so it was removed. This intends to fix the
lldb-remote-linux-win buildbot:
<lab.llvm.org/buildbot/#/builders/197/builds/19717>.
[SPIRV] Fix invalid SPIR-V for loops with switch headers (#229289)
For non-shader targets without SPV_INTEL_unstructured_loop_controls, the
backend uses OpLoopMerge to carry loop unroll hints. When the loop
header ends with a switch, it emits OpLoopMerge after OpSwitch,
producing invalid SPIR-V.
OpLoopMerge must immediately precede OpBranch or OpBranchConditional
[1]. Emit it only when the loop header ends with one of these branches.
Non-shader loops do not require OpLoopMerge, so skipping it drops only
the unroll hint and preserves the loop's behavior.
[1]
https://registry.khronos.org/SPIR-V/specs/unified1/SPIRV.html#OpLoopMerge
---------
Co-authored-by: Arseniy Obolenskiy <gooddoog at student.su>
[SPIRV] Default OpenCL 2.0 atomic RMW builtins to Device scope (#229285)
Per the OpenCL C 2.0 specification (section 6.13.11), non-explicit
OpenCL 2.0 atomic RMW builtins (`atomic_fetch_<key>` and
`atomic_exchange`) default to `memory_order_seq_cst` and
`memory_scope_device`, and the 3-argument `_explicit` overloads without
an explicit `memory_scope` argument default to `memory_scope_device`.
Previously, `buildAtomicRMWInst` defaulted both legacy OpenCL 1.x atomic
builtins (`atom_*`, `atomic_add`, etc.) and OpenCL 2.0 atomic RMW
builtins to `Workgroup` scope and `None` (Relaxed) memory semantics when
explicit scope/order arguments were omitted, unlike
`buildAtomicLoadInst`, `buildAtomicStoreInst`, and
`buildAtomicCompareExchangeInst`.
Distinguish OpenCL 2.0 atomic RMW builtins from OpenCL 1.x legacy
atomics so that OpenCL 2.0 builtins default to `Device` scope and
`SequentiallyConsistent` + storage-class memory semantics.
[Offload][AMDGPU] Wire HSA profiling into GenericProfiler abstraction
Add device profiling infrastructure to the AMDGPU plugin so that the
GenericProfiler can receive nanosecond-accurate kernel execution and
data transfer timestamps from the HSA runtime.
Key changes:
- Add ProfilingInfoTy struct to transport HSA profiling data
- Add timeKernelInNsAsync/timeDataTransferInNsAsync callbacks that
extract dispatch/copy times from HSA signals and call
handleKernelCompletion/handleDataTransfer on the profiler
- Add getOrNullProfilerSpecificData helper to extract ProfilerData
from AsyncInfoWrapperTy
- Add getDeviceTimeStamp() override using hsa_system_get_info
- Add getSystemTimestampInNs() for HSA system timestamp queries
- Add schedProfilerKernelTiming/schedProfilerDataTransferTiming to
StreamSlotTy for scheduling profiler callbacks on stream slots
- Thread ProfilerSpecificData through pushKernelLaunch,
pushMemoryCopyH2DAsync, pushMemoryCopyD2HAsync, pushMemoryCopyD2DAsync
[6 lines not shown]
[Offload] Add GenericProfilerTy abstraction and APITypes extensions
Introduce GenericProfilerTy alongside the existing OMPT callback dispatch.
The weak profiler factory returns a no-op implementation, so the new hooks
are silent while the established callback path continues to handle OMPT
device events.
Co-Authored-By: Dhruva Chakrabarti <dhruva.chakrabarti at amd.com>
Co-Authored-By: Michael Halkenhauser <michaelgerald.halkenhauser at amd.com>
Assisted-by: Claude Code
[Offload] Drop OMPT-specific hooks from GenericProfilerTy
getTraceRecordManager() made the generic profiler interface name
OmptTracingBufferMgr, an OMPT implementation detail, and
getProfilerSpecificData() only existed to hand out OMPT event data.
The OMPT profiler will own its trace buffer manager privately and
libomptarget will reach it through an OMPT-typed accessor instead, so
neither hook belongs in the language-agnostic interface.
Assisted-by: Claude Code
[Offload] Drop device lifecycle hooks from GenericProfilerTy
handleInit, handleDeinit and handleLoadBinary were meant to let the
plugins report device initialization, finalization and image loading to
the profiler. Since #221726 the plugins no longer dispatch these events;
libomptarget emits the corresponding OMPT callbacks itself, so the hooks
have no callers left.
Assisted-by: Claude Code
[mlir][MemRef] Move the implementation of Mem2Reg interface from Vector to MemRef (nfc) (#228459)
Moves the `MemRefSlotInterface` implementations for MemRef dialect ops
from:
* `lib/Dialect/Vector/Transforms`
to:
* `lib/Dialect/MemRef/Transforms`.
This is preferable because external interface models for MemRef dialect
ops should, where possible, be owned by the MemRef dialect rather than by
another dialect.
This direction was not pursued in the PR that introduced the affected
implementations (#211880), in order to:
* "(...) avoid cyclic dependence between memref and vector."
See:
* https://github.com/llvm/llvm-project/pull/211880#discussion_r3709075872
However, the `MemRef` transformation library
[32 lines not shown]
clang: Set the exception model module flag for IR inputs (#228372)
Previously the -exception-model cc1 flag did nothing for IR inputs. Record the
module flag in the IR input path too, next to the existing triple override. A
model contradicting one the input module already records is rejected matching
err_data_layout_mismatch and the conflict llc reports for the same pair of
inputs.
This is to keep the wasm-eh.ll test passing when the TargetOptions field for
ExecptionModel is removed. I'm not sure if this behavior really needs to be
preserved given the bitcode producers should now materialize the flag in the
IR.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[SPIRV] Handle loads from indirect args in the presence of `SPV_KHR_untyped_pointers` (#229184)
The initial bring-up of `SPV_KHR_untyped_pointers` support appears to
have missed a spot. When dealing with indirect args (`byval`/`byref`) we
want to retain the pointee type / issue a typed pointer (this is handled
on the arg typing side). However, the pointer type assignment transform
can work back from a loaded type, and erroneously assign that type as
the dominant type. This leads to broken SPIR-V, because needed bitcasts
from the indirect type to the load are not emitted (they'd be spurious
due to the lost typing). The patch corrects it by adjusting how we work
out pointer typing when dealing with a load instruction.
[AMDGPU] Fix amdgcn.mbcnt known bits conflicting with the range attribute (#227725)
`computeKnownBits` applied the `mbcnt` `popcount` bound to `Known` and
added the base to compute the known bits of `mbcnt`'s result. But
`Known` may not be always empty and can already have the correct call
range (base + count) computed by `instCombineIntrinsic`.
The incorrect re-application of bound and re-addition with the base
results in some of the known-one bits now marked as zeroes, causing
conflicts. And `KnownBits::add` turns a conflict into zero, so users
folded as if the intrinsic returned `0`.
So, the `popcount` bound should be computed in a fresh `KnownBits` and
`union`ed with the existing `Known` result instead.