[CIR][CodeGen] Share isStandardLibraryRTTIDescriptor
Deduplicates `isStandardLibraryRTTIDescriptor` between CIR and classic CodeGen
into `ItaniumCXXABIUtils.h`, taking the classic implementation. The two copies
have been equivalent since #227781 filled in the builtin types CIR was missing.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share canUseSingleInheritance
Deduplicates `canUseSingleInheritance` between CIR and classic CodeGen into
`ItaniumCXXABIUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share requiresAMDGPUProtectedVisibility
Deduplicates `requiresAMDGPUProtectedVisibility` between CIR and classic CodeGen
into `TargetUtils.h`. The shared version takes a bool for "currently hidden" in
place of the `llvm::GlobalValue` and `cir::VisibilityKind` the two callers
passed.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share hasExtraNeonArgument
Deduplicates `hasExtraNeonArgument` between CIR and classic CodeGen into
`TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Arm SME inlinability check
Deduplicates `ArmSMEInlinability` and `getArmSMEInlinability` between CIR and
classic CodeGen into a new `TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Retire the CodeGenUtils.h catch-all header
Moves the last helpers out of `CodeGenUtils.h` into `ClassUtils.h`,
`ModuleUtils.h`, `TargetUtils.h` and `FunctionUtils.h` (`checkTargetFeatures`,
since it came from CodeGenFunction.cpp) and deletes the header. Only moves code
already on main, so it can be dropped on its own.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen] Share the EH personality selection logic
Deduplicates the EH personality selection (`getEHPersonality`,
`getCXXEHPersonality`) between CIR and classic CodeGen, taking the classic
implementation. CIR's copy lacked the z/OS, Wasm and GNUstep-on-CygMing cases;
none are reachable in CIR today, so no test changes.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Itanium __vmi_class_type_info flags computation (#227256)
Deduplicates the `__vmi_class_type_info` and `__base_class_type_info` flags and
`computeVMIClassTypeInfoFlags` between CIR and classic CodeGen into
`ItaniumCXXABIUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[OpenMP] Tile/stripe followed by reverse (#222401)
`Clang` rejects `#pragma omp tile` or `#pragma omp stripe` immediately
followed by `#pragma omp reverse` on the same loop nest, emitting the
error:
`error: statement after '#pragma omp tile' must be a for loop`
See https://godbolt.org/z/d1fsfaMe9
The correct behavior is to accept it. Bare`reverse`, bare
`tile`/`stripe`, and `tile` followed by `stripe` all work. Only
`tile`/`stripe` and `reverse` fails.
This PR fixes the issue.
[AArch64] Use multi-vector intrinsics for masked load/store users of predicate-as-counter
If the user of the original wide mask is a masked load or store
intrinsic (matching the element size of the predicate-as-counter),
rewrite it directly to a masked multi-vector load/store.
This avoids materializing the vector mask and is easier to handle here
than later (e.g. in SelectionDAG), since we do not need to match the
concatenation of all `pext` segments of the predicate-as-counter.
Assisted-by: Codex
[AArch64] Avoid materializing full masks for extractelement users of predicate-as-counter
If the user of the original wide mask is an `extractelement` and the
index is known to be within the first segment of the
predicate-as-counter, replace it with `extractelement(pext(counter, 0))`.
This avoids materializing the vector mask and produces a form that can
be folded into a conditional branch when the predicate-as-counter is
produced by a `whilelo`.
Assisted-by: Codex
[Offload] Avoid global mutex destructors (#221471)
Replace namespace-scope std::mutex instances that trigger
-Wglobal-constructors with destruction-free storage.
This allows the offload libraries to build with
-Werror=global-constructors.
AMDGPU/GlobalISel: Fix selecting i16 sext atomic loads on gfx6/gfx7 (#229723)
An atomic global sextload from i16 to i32 was incorrectly matched to
buffer_load_sbyte. The correct pattern to buffer_load_sshort is already
there, so remove the incorrect duplicate which was arbitrarily preferred.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[offload][omp] Load and resolve device binaries through liboffload
Migrate DeviceTy::loadBinary and global/kernel symbol resolution off
GenericPluginTy::load_binary/get_global/get_function onto liboffload's
Program/Symbol API, encapsulated in a new ProgramTy abstraction that wraps
an ol_program_handle_t. Kernel symbol resolution still needs the plugin's
opaque GenericKernelTy* handle for the legacy launch path, obtained via a
temporary __ol_tgt_GetKernelFromSymbol helper rather than new public
liboffload API surface. Removes the now-dead __tgt_device_binary type and
the corresponding GenericPluginTy methods and exports entries.
[AArch64][FMV] Skip unreachable versions before emitting feature checks (#229120)
Currently there are checks emitted for unreachable features whose
results are then ignored. This results in a lot of unnecessary IR being
emitted. This patch removes that code, by moving the check for
unreachable version to the top of the loop.
Share one no-loop eligibility check between exec mode and body
The kernel is tagged SPMD_NO_LOOP by a single check, and the body is
emitted loop-free exactly when the kernel carries that tag. Add tests
for the clauses that block promotion.
[clang][OpenMP] Add no-loop SPMD kernel promotion
A target teams distribute parallel for that is guaranteed a thread for
every iteration does not need the loop around its body. Flang already
drops it and runs the region as a no-loop kernel.
Enable the same optimization for Clang through mirroring Flang's MLIR
promotion using OpenMPIRBuilder. The kernel is tagged SPMD_NO_LOOP, so
the runtime sizes the grid to the iteration space, and the body is
emitted without a loop around it. The canonical loop it consumes is
reconstructed in the no-loop branch rather than taken from an
OMPCanonicalLoop node, so the promotion does not require
-fopenmp-enable-irbuilder.
Restrict offload entry creation to module level finalize, preventing
asserts on missing offload entries from nested CodeGenFunction
finalizing before module completion.
[lldb] Move DemangledNameInfo cache out of Mangled class (#225332)
This patch moves the per-Mangled DemangledInfoCache into a global
cache with a size limit. The motivation for this is:
1. It makes `Mangled` a smaller which saves memory as its a frequently
heap-allocated object.
2. The current Mangled class is sometimes accessed via the SB API from
multiple threads, and the stateful cache in Symbol causes crashes due
to the inherent race conditions of this approach.
3. It avoids potentially large memory usage in case many symbols have
their DemangledNameInfo computed and cached.
The size of the global cache can be changed via the new
'symbols.demangled-name-info-cache-size' setting. Setting this to 0
disables the cache.
[10 lines not shown]
[lldb] Keep the inferior's terminal open until LLDB read its output (#229387)
lldb gives the inferior the secondary side of a pseudo terminal for its
stdio and drains the primary side (what we get back from from
posix_openpt) via ThreadedCommunication's read thread.
On Darwin, it seems that closing the last secondary descriptor clears
the terminal's contents. Any output the inferior wrote but that the read
thread has not consumed yet is then destroyed alongside the inferior.
There isn't really any guarantee that we read the inferior's output
before it dies (and closes its secondary side of the terminal). So
reading stdout is currently racy in LLDB. This is also why
TestTargetAPI.py randomly fails to read the output:
```
File ".../lldb/test/API/python_api/target/TestTargetAPI.py", line 189
self.assertIn("arg: foo", output)
AssertionError: 'arg: foo' not found in ''
[19 lines not shown]
[examples] Set target triple on modules in Kaleidoscope. (#229666)
If no triple is set then the object format component of the Module's
triple will be set by getDefaultFormat in Triple.cpp, typically to ELF.
In AArch64AsmPrinter::emitStartOfAsmFile, the "ELF" default will cause
control to fall through the following check:
if (!TT.isOSBinFormatELF())
return;
and into a region that assumes an MCTargetStreamer. On Darwin/arm64,
which does not set an MCTargetStreamer, this will cause a nullptr
access.
Fix the issue by setting the Module's triple to the host triple as
reported by the KaleidoscopeJIT object.
rdar://189147321