[Clang] prevent instantiation of invalid friend function templates (#227038)
Fixes #226673
---
This patch addresses the issue in which an invalid friend function
template is instantiated as a valid declaration, leading to a crash upon
a later redeclaration.
---
This fixes a regression introduced by
https://github.com/llvm/llvm-project/pull/216555
[Clang][docs] Markup corrections for attribute documentation (#226729)
The conversion of Clang attribute documentation from reST to MyST
markdown did not transform markup correctly in all cases. Additionally,
some PRs that were originally authored for reST were merged after the
conversion resulting in reST formatted markdown getting re-introduced.
This change corrects all such cases the author was able to identify,
applies markup that was previously missing, corrects a few typos along
the way, and even adds a missing attribute documentation reference that
has prevented the `#pragma clang loop` `pipeline` and
`pipeline_initiation_interval` documentation from being published since
2019.
[mlir][OpenACC] Fix tile clause element bounds for descending loops (#227185)
acc.loops could have negative steps, but ACCLoopTiling computed each
element loop's bound assuming an ascending one. Both halves of that
computation were wrong for a descending dimension: the inclusive edge
subtracted one instead of adding one, and the bound was clamped with min
instead of max.
For `DO j = n, floor, -1` under `tile(32)`, the last tile's bound came
out as `min(floor, j - 33)`, which selects `j - 33` whenever it falls
below the floor. The element loop then ran past the end of the iteration
space, reading and writing outside the array.
With this change, we pick the clamp and the inclusive edge from the
step's direction. When the sign is only known at runtime, we compute
both and select on it.
The tiling tests had no coverage of negative steps or of inclusive upper
bounds. This patch adds four new tests: descending, descending with an
inclusive bound, ascending with an inclusive bound, and a runtime-signed
step.
[MLIR][Remark] Make remark reporting thread-safe and order final remarks by source position
Passes nested under the multithreaded pass manager report into the same
RemarkEngine from worker threads, but nothing in the engine or the
policies was synchronized. RemarkEmittingPolicyFinal inserted into its
map from several threads at once, and under RemarkEmittingPolicyAll the
streamer, including the LLVM remark serializer, was called from several
threads at once. ThreadSanitizer reports both races in the new unit tests.
RemarkEngine now holds a lock around every call into the policy. It is a
recursive llvm::sys::SmartMutex, the same type as the DiagnosticEngine's
lock. report() takes it, and so does a new finalizePolicy(), which the
engine destructor and mlir-opt now use instead of calling finalize() on
the policy directly. For the All policy the streamer and the diagnostic
printer also run under the lock, so custom policies and streamers need no
lock of their own.
With the race fixed, the final policy's first-report order is still not
deterministic: which thread reports first depends on scheduling.
[16 lines not shown]
[SPIRV] LLVM IR DIScope nodes are now emitted after their dependencies.
This allows supporting the following cases:
A pointer, array, typedef, or function type whose operand is a composite.
A nested composite member, and a typedef whose base is another typedef.
A type scoped in a function, lexical block, or namespace, and a function declaration whose parent is a composite.
Cycles are not fully supported. A back edge is not emitted: by design, a composite member that refers back to its own type is dropped, and a cycle that cannot be cut emits nothing.
Emission functions now return an EmitResult, which indicates the status of the emission.
Emitted and Unsupported results are cached. InProgress is not.
For struct S { S *p; }, the edge from p back to S returns InProgress, so member p is dropped and S is still emitted.
S* is left uncached, so a later call emits it once S exists. Caching that InProgress result would completely drop S*.
[offload][omp] Load and resolve device binaries through liboffload
Migrate DeviceTy::loadBinary and global/kernel symbol resolution off
GenericPluginTy::load_binary/get_global/get_function onto liboffload's
Program/Symbol API, encapsulated in a new ProgramTy abstraction that wraps
an ol_program_handle_t. Kernel symbol resolution still needs the plugin's
opaque GenericKernelTy* handle for the legacy launch path, obtained via a
temporary __ol_tgt_GetKernelFromSymbol helper rather than new public
liboffload API surface. Removes the now-dead __tgt_device_binary type and
the corresponding GenericPluginTy methods and exports entries.
[offload][omp] Route RPC callback registration through liboffload
Replace __tgt_register_rpc_callback's direct iteration over plugins
with olIteratePlatforms + olPlatformRegisterRPCCallback, and drop the
now-unused RPCServerTy::registerCallback export. Move the
initialized/has-devices guard that used to live in libomptarget into
olPlatformRegisterRPCCallback_impl.
[offload][omp] Query device info directly through liboffload
Route DeviceTy::getInfo through olGetDeviceInfo instead of the
plugin's obtain_device_info, and drop the now-unused
GenericPluginTy::obtain_device_info wrapper and its liboffload
export.
[offload][omp] Manage memory allocation through liboffload
Migrate DeviceTy::allocData/deleteData off GenericPluginTy::data_alloc/
data_delete onto liboffload's olMemAlloc*/olMemFree, migrate
targetLockExplicit/targetUnlockExplicit off data_lock/data_unlock onto
olMemRegister/olMemUnregister (fixing a latent bug where these passed
the OpenMP-visible device number instead of the plugin device id), and
migrate DeviceTy::isAccessiblePtr onto a new olMemIsAccessible API
(added with a unit test) since liboffload had no equivalent for
querying accessibility of arbitrary, not-necessarily-liboffload-
allocated pointers. Removes the now-dead GenericPluginTy::data_alloc/
data_delete/data_lock/data_unlock/is_accessible_ptr wrappers and their
exports entries.
[offload][omp] Remove data_fence
All supported backends execute enqueued work on a given queue in
submission order (CUDA, AMDGPU, and Host queues are always in-order;
Level Zero's default and non-default in-order/synchronous command
modes are as well), so the explicit data-fence used to order a
pointer-attachment after prior data transfers is unnecessary. Remove
the DeviceTy/GenericDeviceTy/GenericPluginTy dataFence chain and the
now-unused olQueueBarrier liboffload API added to support it.
[MISched](NFC) Factor control dependency construction state and logic out of `buildSchedGraph` (#226533)
Continuing the groundwork for a proper implementation of #205689, factor
the control dependency construction state and logic out of
`ScheduleDAGInstrs` into a friend class. This will also enable us to
further break up `buildSchedGraph` to make, e.g., the barrier chain
handling easier to comprehend.
[ADT] Move DenseMapIterator above DenseMapBase (NFC) (#227208)
This patch moves the definition of DenseMapIterator right above
DenseMapBase so that all classes in DenseMap.h are defined in the
bottom-up order without forward declarations:
- DenseMapStorage
- SmallDenseMapStorage
- DenseMapIterator
- DenseMapBase
- DenseMap
- SmallDenseMap
Assisted-by: Antigravity
[GitHub] Fix test-suite.yml FFmpeg build with AArch64 (#227346)
On AArch64 there are assembly files that need CMAKE_ASM_COMPILER_TARGET
and CMAKE_ASM_FLAGS_INIT set