[libc] Allow hermetic tests to use external startup objects (#212557)
Adds LLVM_LIBC_HERMETIC_TEST_USE_INTERNAL_STARTUP to control whether
libc hermetic tests link against LLVM libc’s own crt1.o.
When the option is OFF, hermetic tests no longer require or link
libc.startup.${LIBC_TARGET_OS}.crt1. This allows baremetal
configurations to provide startup externally through
LIBC_TEST_LINK_OPTIONS_DEFAULT while still running hermetic tests.
For baremetal external-startup mode, the patch also adds the required
process startup/termination dependencies and ensures exit uses the
public packaging object, since external startup code may reference the
public C exit symbol.
Tested on Arm Toolchain for Embedded:
https://github.com/arm/arm-toolchain
Assisted-by: Codex
RegisterCoalescer: Keep remat def dead if it's a copy destination superregister
This is a refinement of #226037, which was too strict.
When rematerializing into a physical register that is not exactly the copy's
destination, the def should only stay live if it is a sub-register of the copy
destination, i.e. part of the live value. Checking register unit coverage also
kept the def live when it is a super-register with the same units as the copy
destination, such as $rax for a copy into $eax on x86_64:
dead $rax = MOV64ri32 -11, implicit-def $eax
Only the $eax part is used, so the $rax def is dead. This matches what
LiveVariables produces for a full def with partial uses.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[mlir] Avoid parallel code in single-threaded builds (#229794)
We hit this while building MLIR 23.1.2 for wasm32 and wasm64 in
emscripten-forge (this is an effort towards
[RFC](https://discourse.llvm.org/t/rfc-bringing-interactive-mlir-compilation-and-execution-to-webassembly/91566/3))
with `LLVM_ENABLE_THREADS=OFF`. `libMLIRIR.a` still retained
`pthread_create` and `pthread_setspecific` references from libc++ async
machinery, causing our package checks to fail.
`failableParallelForEach` only checks whether threading is enabled at
runtime, so the unused thread-pool code still reaches code generation.
We also need to check `LLVM_ENABLE_THREADS`, allowing the compiler to
discard that path in single-threaded builds !
[IR] Deprecate Instruction::getStableDebugLoc() (#230016)
In favor of getDebugLoc(). Since the migration from debuginfo intrinsics
to records, these are equivalent.
virtio_gpu: use an X8R8G8B8 resource on big-endian guests
The virtio-gpu resource formats are defined by byte order in memory,
while vt(4) and the X server (through vt_fb and the fb mmap) write
native 32-bit 0x00RRGGBB pixels into the shadow framebuffer. On a
big-endian guest such as powerpc64 those pixels land in memory as
00 RR GG BB, which the host interprets under B8G8R8X8 with red and
blue swapped and the padding byte taken as blue: white renders yellow
and blue renders black.
Request X8R8G8B8 on big-endian guests instead, which is also what the
Linux driver does (DRM_FORMAT_HOST_XRGB8888). Reproduced with a colour
test pattern on a powerpc64 QEMU pseries guest using the identically
coded out-of-tree fork of this driver (graphics/virtio-gpu-qemu-kmod);
little-endian guests are unchanged.
Differential Revision: https://reviews.freebsd.org/D60085
Approved by: jhibbits
MFC after: 1 week
[mlir][acc] Declare the memory effects of acc.atomic.write
The operation wrote the location designated by `x` but declared no memory
effects, so it did not implement MemoryEffectOpInterface at all and every
memory analysis had to treat it as unknown. fir::AliasAnalysis::getModRef
returns ModAndRef for such an operation, so a loop containing one blocked
the hoisting of any load, however unrelated.
Declare the write on `x`. `expr` is a value rather than a location and
carries no effect.
This completes the set: acc.atomic.update already declared MemRead and
MemWrite on `x`, acc.atomic.read declares its effects since #228899, and
acc.atomic.capture derives its effects from its body.
The relaxed-ordering contract is stated in the operation description and
tested the same way as for acc.atomic.read: an unrelated load is hoisted
out of a loop across the write, and a load of `x` is not.
[Offload][omp] Add own Error class to libomptarget (#229711)
Add a new class for Errors generated within libomptarget to break the
dependence with the Error class defined inside the plugins.
Assisted by Claude.
[AMDGPU] Fold add of a variable into a zero dot accumulator
When the dot intrinsic has a zero accumulator, no clamp and a single
add user, fold the add operand into the accumulator:
```llvm
%dot = call i32 @llvm.amdgcn.sdot4(i32 %a, i32 %b, i32 0, i1 false)
%r = add i32 %dot, %x
=>
%r = call i32 @llvm.amdgcn.sdot4(i32 %a, i32 %b, i32 %x, i1 false)
```
If %x is defined after the dot in the same block, the dot is moved
down to the add. The fold is skipped across blocks to avoid sinking
the dot into a loop.
MC: Error if target did not register null streamer
Without this the AsmPrinter would see a null pointer
for the target streamer. More than likely this is just
going to crash. Only 5 of the in tree targets properly
registered a null streamer before I started looking into this.
The rest had randomly, inconsistent null checks for the streamer.
There are still a few targets not registering this, which should
update (some of which seem to not make use of the TargetStreamer).
There's no reason to leave this as a silently hazardous edge case.
AMDGPU: Stop relying on -amdgpu-scalarize-global-loads=false in more tests (#230013)
Convert operation tests to functions taking their inputs as arguments, and index loads by
workitem id in kernels where the stores matter.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[Offload][omp] Use olGetDeviceInfo to get device info (#229693)
Use olGetDeviceInfo instead of obtain_device_info. Reimplement
libomptarget printDeviceInfo using olGetDeviceInfo. Remove from the
plugin interafce the print_device_info and obtain_device_info support.
Assisted by Claude.
[clangd] [C++20] [Modules] Don't reuse prebuilt module files on windows (#193426)
Fix https://github.com/clangd/clangd/issues/2497
Note that this only works for --experimental-modules-support. Otherwise,
it is a natural fallback by the design clang based tools.
[CIR][CodeGen] Share isStandardLibraryRTTIDescriptor
Deduplicates `isStandardLibraryRTTIDescriptor` between CIR and classic CodeGen
into `ItaniumCXXABIUtils.h`, taking the classic implementation. The two copies
have been equivalent since #227781 filled in the builtin types CIR was missing.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share requiresAMDGPUProtectedVisibility
Deduplicates `requiresAMDGPUProtectedVisibility` between CIR and classic CodeGen
into `TargetUtils.h`. The shared version takes a bool for "currently hidden" in
place of the `llvm::GlobalValue` and `cir::VisibilityKind` the two callers
passed.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share hasExtraNeonArgument
Deduplicates `hasExtraNeonArgument` between CIR and classic CodeGen into
`TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Arm SME inlinability check
Deduplicates `ArmSMEInlinability` and `getArmSMEInlinability` between CIR and
classic CodeGen into a new `TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen] Share the EH personality selection logic
Deduplicates the EH personality selection (`getEHPersonality`,
`getCXXEHPersonality`) between CIR and classic CodeGen, taking the classic
implementation. CIR's copy lacked the z/OS, Wasm and GNUstep-on-CygMing cases;
none are reachable in CIR today, so no test changes.
Assisted-by: Claude Code (Claude Fable 5.1).