[GVN] Move `ValueTable` out of `GVN.h` (NFC) (#226233)
* Rename `GVNPass::ValueTable` to `GVNValueTable`, and move it out to
the `llvm` namespace and to its own file `GVNValueTable.h`.
* Move `GVNPass::Expression` into `llvm::GVNValueTable`.
* Move `GVNHoistPass` and `GVNSinkPass` to their own headers.
With these changes `GVHoist.cpp` and `GVNSink.cpp` no longer need to
include `GVN.h` or peek into `GVNPass` internals.
X86: Fix missing ... separator between functions in mir test
The update_mir_test_checks output is incomplete so the later functions
here were really manually checked.
[Openmp] Add OMPT device tracing support tests
These tests concern functionality used by parts of the OMPT device
tracing implementation.
Assisted-by: Claude Code
[Offload][OMPT] Route device callbacks through OmptProfilerTy
Compile and activate the complete OMPT profiler backend now that its tracing dependencies are available, and remove the superseded direct callback dispatch from PluginInterface in the same transition.
Assisted-by: Claude Code
[Offload][OMPT] Add tracing orchestration and libomptarget integration
Wire the OMPT device tracing subsystem into libomptarget, completing the tracing pipeline from record production through buffer management to tool delivery. Include the synchronous fallback and initialization fixes in the integration that requires them.
Assisted-by: Claude Code
[Offload][OMPT] Add tracing extensions to Interface.h
Extend the OMPT Interface class with trace record creation methods and
the async-aware TracerInterfaceRAII for device tracing support.
Key changes to Interface.h:
- Add includes for OmptEventInfoTy.h, APITypes.h, GenericProfiler.h
- Add extern TracingActive flag and isTracingEnabled() declaration
- Add trace start/stop methods for all target operations:
startTargetDataAllocTrace, stopTargetDataAllocTrace,
startTargetDataSubmitTrace, startTargetDataDeleteTrace,
startTargetDataRetrieveTrace, stopTargetDataMovementTraceAsync,
startTargetSubmitTrace, stopTargetSubmitTraceAsync, and
corresponding methods for enter/exit/update/target regions
- Add getTraceGenerators<>() template overloads parallel to
existing getCallbacks<>()
- Add private helpers: setTraceRecordCommon, setTraceRecordTargetDataOp,
setTraceRecordTargetKernel, setTraceRecordTarget, announceTargetRegion
- Add TracerInterfaceRAII class: async-aware RAII that checks
[12 lines not shown]
[Offload][OMPT] Add OmptTracingBufferMgr trace buffer infrastructure
Add the trace buffer management subsystem for OMPT device tracing. This
manages allocation, population, and flushing of trace record buffers
between OpenMP worker threads and helper threads.
New files:
- OmptTracingBuffer.h: OmptTracingBufferMgr class with Buffer metadata,
FlushInfo, TraceRecord types, thread-local per-device buffer pointers,
helper thread management, and cursor-based record allocation
- OmptTracingBuffer.cpp: Full implementation of buffer lifecycle:
assignCursor (lock-free fast path for existing buffers, locked path
for new allocations), triggerFlushOnBufferFull, driveCompletion
(helper thread main loop), invokeCallbacks, flushBuffer (dispatches
buffer-completion callbacks for ranges of ready records),
flushAllBuffers, helper thread start/shutdown/flush-and-shutdown
- OmptTracing.h: Libomptarget-side declarations for tracing state,
buffer management callback registration, device tracing control,
timestamp/clock correlation, and thread-local trace record fields.
[17 lines not shown]
[Offload][OMPT] Call device tracing entry points directly
The plugin-side tracing code reached the libomptarget-side entry points by
dlopen'ing "libomptarget.so" and dlsym'ing each libomptarget_ompt_* symbol
on first use. That indirection dates back to when libomptarget and the
plugins were separate shared objects. They are not: PluginOmpt is linked
into omptarget itself, so the library was dlopen'ing itself to look up its
own symbols. The callback path was already converted to direct calls; the
tracing path should never have reintroduced the pattern.
Declare the entry points in OmptCommonDefs.h and call them directly. This
removes ParentLibrary, ensureFuncPtrLoaded(), the generated function
pointers, and the libomptarget_ompt_*_t typedefs (two of which were
duplicates, and three of which were never used).
Calling directly also means the symbols no longer have to be exported, so
no version script entry is needed for them.
The per-entry-point mutexes are left untouched here: they currently wrap
[4 lines not shown]
[Offload][OMPT] Add OMPT profiler implementation
Add the OMPT-specific profiler types and plugin tracing implementation without compiling or activating them yet. The existing callback path remains authoritative until the tracing interface and orchestration dependencies are available.
Assisted-by: Claude Code
[Offload][AMDGPU] Wire HSA profiling into GenericProfiler abstraction
Add device profiling infrastructure to the AMDGPU plugin so that the
GenericProfiler can receive nanosecond-accurate kernel execution and
data transfer timestamps from the HSA runtime.
Key changes:
- Add ProfilingInfoTy struct to transport HSA profiling data
- Add timeKernelInNsAsync/timeDataTransferInNsAsync callbacks that
extract dispatch/copy times from HSA signals and call
handleKernelCompletion/handleDataTransfer on the profiler
- Add getOrNullProfilerSpecificData helper to extract ProfilerData
from AsyncInfoWrapperTy
- Add getDeviceTimeStamp() override using hsa_system_get_info
- Add getSystemTimestampInNs() for HSA system timestamp queries
- Add schedProfilerKernelTiming/schedProfilerDataTransferTiming to
StreamSlotTy for scheduling profiler callbacks on stream slots
- Thread ProfilerSpecificData through pushKernelLaunch,
pushMemoryCopyH2DAsync, pushMemoryCopyD2HAsync, pushMemoryCopyD2DAsync
[6 lines not shown]
[Offload] Add GenericProfilerTy abstraction and APITypes extensions
Introduce GenericProfilerTy alongside the existing OMPT callback dispatch.
The weak profiler factory returns a no-op implementation, so the new hooks
are silent while the established callback path continues to handle OMPT
device events.
Co-Authored-By: Dhruva Chakrabarti <dhruva.chakrabarti at amd.com>
Co-Authored-By: Michael Halkenhauser <michaelgerald.halkenhauser at amd.com>
Assisted-by: Claude Code
ARM: Mark __chkstk call's implicit LR def dead
The WIN__CHKSTK custom inserter already marks the clobbered R12 and CPSR
defs dead, but left the call's implicit-def $lr for later inference.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
AArch64/GlobalISel: Drop redundant IMPLICIT_DEFs before the ptrauth pseudos
MOVaddrPAC, LOADgotPAC, AUTPAC and AUTx16x17 all fully define X16/X17;
none of them read X17, and the only X16 read is the one the surrounding
code sets up with a COPY.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
Mips: Mark the $gp setup copy dead when erasing the call's $gp use
MipsOptimizePICCall drops the implicit $gp operand from a call when the
lazy binding stub for the callee has already run. The copy that set $gp
up for that call then has no reader left, but nothing flagged it, so the
MIR carried a live def until a later liveness recomputation cleaned it
up.
The pass already walks each block in order, so track the reaching
definition of $gp as it goes and mark it dead when the use is erased.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
[mlir] Own the registrations of SourceMgrDiagnosticVerifierHandler (#226258)
`SourceMgrDiagnosticVerifierHandler::registerInContext` registered a
callback capturing `this` and dropped the handler id, so nothing could
erase it. The constructor did the same in its own context. After the
verifier was destroyed, the next diagnostic called into a dead object.
mlir-opt's per-buffer contexts die first, which hid the bug.
Make both registrations owned:
- The constructor installs its callback through `setHandler`, replacing
the base class's printing handler. That handler was unreachable anyway:
the verifier consumes every diagnostic and reports unexpected ones
itself. The base destructor now erases it.
- `registerInContext` returns a scoped registration. Callers destroy it
before the verifier and the context.
Erasing from every registered context in the destructor would not work:
mlir-opt's per-buffer contexts are already gone by then.
[8 lines not shown]
[AMDGPU] Route no-modifier reg-or-inline AsmParser operands through HwMode predicate
Convert the reg-or-inline operands with no modifiers (MFMA VGPR/AGPR
sources, VCSrc, v_pk_mov_b32, VOP scalar f64) from the fixed-class
isRegOrInlineNoMods to the HwMode-aware isRegOrInlineNoModsByHwMode, so an
odd-aligned tuple is rejected at the offending operand column instead of by
the validateVGPRAlign catch-all.
Co-Authored-By: Claude <noreply at anthropic.com>
[AMDGPU] Update no-modifier operand tests for the dropped align diagnostic
The no-modifier reg-or-inline operands routed through the HwMode
predicate now report a misaligned tuple as a plain invalid operand,
matching the diagnostic dropped earlier in the stack.
[compiler-rt][ARM] Make interworking test compatible with BTI (#226154)
In the aeabi_cmpflags_interwork_test.c test wrappers, we should return
through lr instead of r3 so that runs with BTI enabled don't fail. As
returns through lr don't require a landing pad, and returns through r3
do.
[SCEV] - Add positive-stride predicate for backedge-taken count. (#222261)
When `howManyLessThans` encounters a loop with an unknown stride that
cannot be proven finite (no `mustprogress` or side-effect-free
guarantee), SCEV currently returns `CouldNotCompute` for the
backedge-taken count. This blocks downstream consumers like the loop
vectorizer from optimizing such loops.
This patch relaxes the requirement by allowing a predicated
backedge-taken count when `AllowPredicates` is true. Instead of
requiring `loopIsFiniteByAssumption(L)` unconditionally, we add a
`Compare predicate: stride sgt 0` when finiteness cannot be proven.
A positive stride guarantees forward progress, making the BTC formula
correct. The predicate is emitted as a runtime check by consumers
(e.g., the loop vectorizer generates a guard branch before the vector
loop).
[17 lines not shown]
[docs] Pass -c when generating a PCH file (#226850)
Without options like -fsyntax-only/-E/-S/-c, the driver runs in link
mode, and a header-only command line produces a PCH only incidentally: a
linker option such as -lm or -Wl,..., including one from a config file,
adds a link job. Use -c in the PCH examples.
Also fix two broken examples: -ignore-pch is a driver option (cc1
rejects -Xclang -ignore-pch), and the relocatable PCH example is
missing -o. In LibASTImporter.md, generate the C++ AST files with
-emit-ast instead of treating .cpp files as headers.
LLM-aided