[orc-rt] Use the Error matchers in StandaloneMachOUnwindInfoRegistrar… (#228679)
…Test
Use the Error matchers introduced in 4c8a437d0487 to clean up error
checks in StandaloneMachOUnwindInfoRegistrarTest.
ARM: Preserve the dead flag when activating the optional CPSR def (#227655)
This did not preserve the original dead flag, so it would be recomputed later by
LiveVariables or RegAllocFast.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[orc-rt] Use the Error matchers in SimpleRemoteCATest (#228672)
Use the Error matchers introduced in 4c8a437d0487 to clean up error
checks in SimpleRemoteCATest.
[webkit.UncountedLambdaCapturesChecker] Treat call_once arguments as noescape (#224492)
call_once does not store the lambda in heap so consider all its
arguments as noescape.
[GlobalISel] [AArch64] Extend SelectShiftMask to handle ADD/SUB mod patterns (#225842)
AArch64 shift instructions only use the low log2(ShiftWidth) bits of the
shift amount, so shifting by X+N where N == 0 mod ShiftWidth is
equivalent to shifting by X alone.
This extends both SelectionDAG (SelectShiftMask) and GlobalISel
(selectShiftMask) implementations.
Part of the incremental work tracked in #224245 to replace
tryShiftAmountMod with ComplexPattern-based isel.
Depends on #223136.
ORC: Fix flaky OrcLazy tests (#228619)
I've seen this fail a few too many times so just let AI deal with it. I
don't
know anything about orc, but extending lifetime of lock_guard seems
plausible.
Notify lookupInitSymbols CV while holding the mutex
The init-symbol lookup completion callback decremented Count under
LookupMutex but called CV.notify_one() after releasing it. The waiting
thread could observe Count == 0, return from lookupInitSymbols, and
destroy the stack-allocated mutex and condition variable before the
callback signalled it. With concurrent compile threads the callback runs
on a dispatcher thread, so the late notify wrote into reused stack
memory, e.g. during endSession right after deinitialize.
This caused intermittent crashes in
ExecutionEngine/OrcLazy/multiple-compile-threads-basic.ll on macOS
[6 lines not shown]
[webkit.UncountedLambdaCapturesChecker] Treat call_once arguments as noescape (#224492)
call_once does not store the lambda in heap so consider all its
arguments as noescape.
[GlobalISel] [AArch64] Extend SelectShiftMask to handle ADD/SUB mod patterns (#225842)
AArch64 shift instructions only use the low log2(ShiftWidth) bits of the
shift amount, so shifting by X+N where N == 0 mod ShiftWidth is
equivalent to shifting by X alone.
This extends both SelectionDAG (SelectShiftMask) and GlobalISel
(selectShiftMask) implementations.
Part of the incremental work tracked in #224245 to replace
tryShiftAmountMod with ComplexPattern-based isel.
Depends on #223136.
[AArch64][SelectionDAG] Avoid fold bitcast of scalar_to_vector to anyext in big-endian (#225442)
```
int_vt (bitcast (vec_vt (scalar_to_vector elt_vt:x)))
=> int_vt (any_extend elt_vt:x)
```
This pattern not legal in big-endian, so disable in big-endian.
Fix https://github.com/llvm/llvm-project/issues/225436
[llvm-objcopy,test] Reorganize reserved st_shndx tests (#228664)
Rename section-index-unsupported.test to reserved-shndx.test, fold
unsupported-machine-specific-shndx.test into it, and enhance it.
[ADT] Make protected members of DenseMapBase private (NFC) (#228651)
Neither DenseMap nor SmallDenseMap accesses these members after #227063.
This patch also moves getMemorySize into the main public section.
Assisted-by: Antigravity
[HIPStdPar] Report reachable C++ exceptions from GPU kernel (#228082)
HipStdPar removes all host functions which are unreachable but there can
be cases where some reachable function contains C++ exceptions it is
forced to device code and since GPU devices don't support C++ exceptions
it must be reported.
This change allows for any reachable C++ exception to be reported as
error.
Part of https://github.com/llvm/llvm-project/issues/221941
[CIR] Lowering for __builtin_reduce_assoc_fadd (#226095)
Added lowering for __builtin_reduce_assoc_fadd with the use of the new
implementation of CIR FastMathFlags, following same lowering path as
Classic Codegen.
Added tests for the same.
[SLP] Partially revert #224931 for alternate nodes
Do not pass scalar context to alternate node vector cost queries.
PR #224931 passes the scalar main/alternate op as the context instruction when
costing the vector ops of an alternate node. Those vector ops only feed the
lane-select shuffle, so use-based discounts derived from the scalar's users
do not apply. On AMDGPU this priced an [fmul | fadd] node's vector fmul as
fused (free) where the sole user of the scalar fmul is fadd/fsub, keeping
unprofitable subtrees vectorized. In a rocFFT kernel on gfx950 this raised
VGPR usage from 160 to 178 and caused a performance regression. Restore the
null context for these queries.
[SLP] Partially revert #224931 for alternate nodes
Do not pass scalar context to alternate node vector cost queries.
PR #224931 passes the scalar main/alternate op as the context instruction when
costing the vector ops of an alternate node. Those vector ops only feed the
lane-select shuffle, so use-based discounts derived from the scalar's users
do not apply. On AMDGPU this priced an [fmul | fadd] node's vector fmul as
fused (free) where the sole user of the scalar fmul is fadd/fsub, keeping
unprofitable subtrees vectorized. In a rocFFT kernel on gfx950 this raised
VGPR usage from 160 to 178 and caused a performance regression. Restore the
null context for these queries.
[libc][test] Fix death test timeouts and dead-code elimination in math tests. (#227966)
- In `ExecuteFunctionUnix.cpp`, call `prctl(PR_SET_DUMPABLE, 0)` in
child processes to prevent external core dump handlers (such as apport)
from intercepting expected crashes during death tests, eliminating 10s
poll timeouts under parallel lit runs.
- Close `pipe_fds[0]` properly in `invoke_in_subprocess` to avoid
leaking file descriptors.
- In math smoke test templates (`AddTest.h`, `SubTest.h`, `MulTest.h`,
`DivTest.h`), assign test function results to `[[maybe_unused]] volatile
OutType res` to prevent GCC from dead-code eliminating floating point
operations in `test_inexact_results` when FMA optimization is disabled.
Assisted-by: Gemini
[RISCV][P-ext] Combine repeated code. NFC (#228163)
The body of these 2 if statements are identical and the if conditions
are equivalent to checking if the VT is XLenVT.
[LoopCacheAnalysis] Guard against bit-width overflow in computeRefCost (#224427)
In `computeRefCost()`, the non-consecutive access path estimates the
cache cost by multiplying the trip counts across outer loop dimensions.
Each multiplication ends up calling `getExtendedType()` to obtain an
integer type twice as wide as the current one.
For a deeply-nested loop (such as the 20-dimensional array found in this
change), it is possible for the doubling to exceed `MAX_INT_BITS`, which
triggers the following assert in `IntegerType::get()`.
```
Assertion failed: NumBits <= MAX_INT_BITS && "bitwidth too large", file llvm-project/llvm/lib/IR/Type.cpp, line 340, static IntegerType *llvm::IntegerType::get(LLVMContext &, unsigned int)()
```
This change checks whether the current integer width already exceeds
`MAX_INT_BITS / 2`. If this is true, the doubled width would overflow,
so we return `CacheCostTy::getInvalid()` instead of asserting.
[libc][hdr] Add poll and socket overlay headers for GCC overlay builds. (#227961)
When building LLVM libc in overlay mode with gcc/g++, alias and macro
redirections in GNU libc headers (such as ppoll and socket functions)
cause conflict with libc header declarations.
Fixes in this PR:
- Add `poll_overlay.h` and `sys_socket_overlay.h` to wrap system headers
before libc macros and type declarations in overlay mode.
- Update proxy headers in `libc/hdr/` and `libc/hdr/types/`.
- Update CMake and Bazel build files accordingly.
Assisted-by: Gemini
[libc][cmake] Suppress -Wpsabi note on GCC and clean up resource-dir check. (#227968)
- Add `-Wno-psabi` to GCC compile options to suppress ABI change notes
for 80-bit long double unions.
- Pass `ERROR_QUIET` to clang-specific `--print-resource-dir` query in
`libc/CMakeLists.txt` so it does not output errors on GCC and MSVC.
Assisted-by: Gemini
Reland "[llvm-objcopy] Handle SHF_ALLOC relocation sections in relocatable files" (#228654)
This relands #227204, reverted by #228572 due to a UBSan failure in
compress-sections.s:
```
ELFObject.h:887:35: runtime error: load of value 190, which is not a valid value for type 'const bool'
```
--compress-sections on a relocation section creates a CompressedSection
that incorrectly keeps OriginalType SHT_RELA, so
RelocationSection::classof cast it and read
RelocationSectionBase::Dynamic. Set its OriginalType to SHT_PROGBITS,
which also fixes a TODO in Object::removeSections
Original description:
llvm-objcopy treats every SHF_ALLOC relocation section as dynamic and
requires its sh_link to refer to .dynsym. Linux livepatch modules
[14 lines not shown]