[CIR][SYCL] Emit SYCL device globals in the global address space (#229293)
ClangIR now puts SYCL device globals, static locals
and string literals with no explicit address space in
sycl_global, matching classic CodeGen, instead of
reporting them as not yet implemented.
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
RuntimeLibcalls: Add provider libraries for targets to reference
Add shared LibcallLibrary defs for targets to reference instead of listing
impls directly: compiler-rt, libm and libc. Define various OS specific library
variants.
ARM, Lanai and SPIRV are migrated to the new organization here. The remaining
targets' SystemRuntimeLibrary bodies are stubbed to (add) and filled in by
pending per-target changes.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
RuntimeLibcalls: Add baseline DeclareRuntimeLibcalls tests
Add per-target/OS/ABI tests for the runtime libcall sets that lacked coverage so
upcoming LibcallLibrary refactors don't accidentally change behavior.
update_test_checks does not support checking function declarations, so these
were generated by a bespoke script. In the future it would be better to add a
flag to update_test_checks.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
[libc++abi] Split catch_info from __dynamic_cast_info (#223744)
A lot of fields are only used for either exception handling or the
implementation of `dynamic_cast`, but not both. Splitting them makes it
clear which algorithm is using which data, and it reduces the memory
footprint a bit.
[X86] Keep scalar bf16/f16 selects in vector registers (#224218)
Scalar `select` on a soft-promoted `bf16`/`f16` currently lowers to a CMOV,
which extracts both operands out of XMM into GPRs and moves the result
back. Instead, blend them in vector register.
I do blending in `v8i16` rather than `v8f16` because a `v8f16` VSELECT can
fail to select on subtargets that fall back to BLENDV when there is no
VBLENDVPH.
Right now I skip the transform when the operands are constants(no extra
cost being in GPR), a compare feeds the condition (CMOV reuses the flags
so hard to tell if worth it), both operands bit cast from GPR or f16 is
legal and a mask register select can be selected.
AI was used for review and bug fixing as well as writing comments.
[IRBuilder] Make IRBuilder inserter non-virtual (#229700)
Currently IRBuilderBase stores a reference to a virtual Inserter.
Replace this with a single function_ref callback, which is populated by
a method on `IRBuilder<>`, which always has a known Inserter.
This is a small compile-time improvement.
Based on discussion in https://github.com/llvm/llvm-project/pull/229437.
[LLVMABI] Address reviewer feedback
Give every abi::Type an ABI size, the size Clang's getTypeSize gives it,
and use it wherever the x86-64 classifier and isSingleElementStruct
mirror a getTypeSize call. The X86 width helpers are removed.
Assisted-by: Claude Code / Claude Opus 5.5
[X86][TTI]Model the sqrt folded into the reciprocal estimate as free (#229482)
The scalar sqrt, divided into with the reciprocal allowed, is folded
with
the division into the rsqrt* estimate sequence. Report no cost for it
when
the type has the estimate and the function does not disable it, so the
vectorizers do not pack it and break the folding.
Assisted-by: Cursor
RuntimeLibcalls: Add function signatures to RTLIB. (#226264)
This PR adds function signatures to `RuntimeLibcall` and maps C-type
specified signatures into IR types in
`RuntimeLibcalls::getFunctionTy()`.
Since llvm/ABI is not available now (it is only supporting ararch64 and
x86), loose mapping from C-types to IR types is used.
[clang][NVPTX] Emit !atomic.ignore.denormal.mode for CUDA atomics
CUDA's atomicAdd() family is defined in terms of PTX atom.add, whose
denormal behavior is fixed by the hardware. Without any annotation the
backend has to assume the function's denormal mode must be honored and
expands these into CAS loops whenever the two disagree. Mark them with
!atomic.ignore.denormal.mode so the native instruction is used.
That covers the __nvvm_atom_*_add_gen_f builtins that atomicAdd(),
atomicAdd_block() and atomicAdd_system() are written in terms of, plus
C11/C++11 atomics under -fatomic-ignore-denormal-mode and the
[[clang::atomic(ignore_denormal_mode)]] attribute, which requires
teaching the NVPTX target about AtomicOptions.
The condition for when the metadata is meaningful is now shared with the
AMDGPU and SPIR-V targets in addAtomicIgnoreDenormalModeMetadata(). It
takes an AllowHalf flag because whether f16 denormals are observable is
target specific: PTX exposes no FTZ control for f16 operations, so
atom.add.f16 never flushes and the opt-in is meaningful there, whereas
[3 lines not shown]