CodeGen: Prefer getting the Triple from the Module (#228620)
Take the triple from the contextual module rather than TargetMachine
when it's already readily available.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[clang] Refactor CRC32 helper for x86 CRC builtins (#225774)
Move the reflected CRC32 calculation used by the x86 CRC builtins into
llvm::calculateReflectedCRC32() and reuse it from both the AST
interpreter and constant expression evaluator.
This removes duplicated CRC32 implementation logic from Clang. This is
planned to be used for implementing const folding CRC instructions for
x86 and AArch64 in the future in LLVM.
Based on suggestion in
https://github.com/llvm/llvm-project/pull/219452#discussion_r4070846498
[orc-rt] Use the Error matchers in AllocActionTest (#228680)
Use the Error matchers introduced in 4c8a437d0487 to clean up error
checks in AllocActionTest.
CodeGen: Prefer getting the Triple from the Module
Continue replacing TargetMachine::getTargetTriple() with the module's
triple at sites where a Module is one hop away through an available
Function, GlobalValue or MachineModuleInfo.
Where the surrounding class already holds a Subtarget, use its triple
rather than routing through the Module.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
CodeGen: Prefer getting the Triple from the Module
Take the triple from the contextual module rather than TargetMachine
when it's already readily available.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[Flang][OpenMP] Reset REQUIRES directive with new program unit (#227829)
A REQUIRES directive with unified_address, unified_shared_memory, or
reverse_offload must appear lexically before any device construct or
device routine is scoped to a program unit.
CodeGen: Prefer getting the Triple from the Module when convenient (#228429)
Take the triple from the contextual module rather than TargetMachine when it's
already readily available.
[orc-rt] Use the Error matchers in StandaloneMachOUnwindInfoRegistrar… (#228679)
…Test
Use the Error matchers introduced in 4c8a437d0487 to clean up error
checks in StandaloneMachOUnwindInfoRegistrarTest.
ARM: Preserve the dead flag when activating the optional CPSR def (#227655)
This did not preserve the original dead flag, so it would be recomputed later by
LiveVariables or RegAllocFast.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[orc-rt] Use the Error matchers in SimpleRemoteCATest (#228672)
Use the Error matchers introduced in 4c8a437d0487 to clean up error
checks in SimpleRemoteCATest.
[webkit.UncountedLambdaCapturesChecker] Treat call_once arguments as noescape (#224492)
call_once does not store the lambda in heap so consider all its
arguments as noescape.
[GlobalISel] [AArch64] Extend SelectShiftMask to handle ADD/SUB mod patterns (#225842)
AArch64 shift instructions only use the low log2(ShiftWidth) bits of the
shift amount, so shifting by X+N where N == 0 mod ShiftWidth is
equivalent to shifting by X alone.
This extends both SelectionDAG (SelectShiftMask) and GlobalISel
(selectShiftMask) implementations.
Part of the incremental work tracked in #224245 to replace
tryShiftAmountMod with ComplexPattern-based isel.
Depends on #223136.
ORC: Fix flaky OrcLazy tests (#228619)
I've seen this fail a few too many times so just let AI deal with it. I
don't
know anything about orc, but extending lifetime of lock_guard seems
plausible.
Notify lookupInitSymbols CV while holding the mutex
The init-symbol lookup completion callback decremented Count under
LookupMutex but called CV.notify_one() after releasing it. The waiting
thread could observe Count == 0, return from lookupInitSymbols, and
destroy the stack-allocated mutex and condition variable before the
callback signalled it. With concurrent compile threads the callback runs
on a dispatcher thread, so the late notify wrote into reused stack
memory, e.g. during endSession right after deinitialize.
This caused intermittent crashes in
ExecutionEngine/OrcLazy/multiple-compile-threads-basic.ll on macOS
[6 lines not shown]
[webkit.UncountedLambdaCapturesChecker] Treat call_once arguments as noescape (#224492)
call_once does not store the lambda in heap so consider all its
arguments as noescape.
[GlobalISel] [AArch64] Extend SelectShiftMask to handle ADD/SUB mod patterns (#225842)
AArch64 shift instructions only use the low log2(ShiftWidth) bits of the
shift amount, so shifting by X+N where N == 0 mod ShiftWidth is
equivalent to shifting by X alone.
This extends both SelectionDAG (SelectShiftMask) and GlobalISel
(selectShiftMask) implementations.
Part of the incremental work tracked in #224245 to replace
tryShiftAmountMod with ComplexPattern-based isel.
Depends on #223136.
[AArch64][SelectionDAG] Avoid fold bitcast of scalar_to_vector to anyext in big-endian (#225442)
```
int_vt (bitcast (vec_vt (scalar_to_vector elt_vt:x)))
=> int_vt (any_extend elt_vt:x)
```
This pattern not legal in big-endian, so disable in big-endian.
Fix https://github.com/llvm/llvm-project/issues/225436
[llvm-objcopy,test] Reorganize reserved st_shndx tests (#228664)
Rename section-index-unsupported.test to reserved-shndx.test, fold
unsupported-machine-specific-shndx.test into it, and enhance it.
[ADT] Make protected members of DenseMapBase private (NFC) (#228651)
Neither DenseMap nor SmallDenseMap accesses these members after #227063.
This patch also moves getMemorySize into the main public section.
Assisted-by: Antigravity
[HIPStdPar] Report reachable C++ exceptions from GPU kernel (#228082)
HipStdPar removes all host functions which are unreachable but there can
be cases where some reachable function contains C++ exceptions it is
forced to device code and since GPU devices don't support C++ exceptions
it must be reported.
This change allows for any reachable C++ exception to be reported as
error.
Part of https://github.com/llvm/llvm-project/issues/221941
[CIR] Lowering for __builtin_reduce_assoc_fadd (#226095)
Added lowering for __builtin_reduce_assoc_fadd with the use of the new
implementation of CIR FastMathFlags, following same lowering path as
Classic Codegen.
Added tests for the same.
[SLP] Partially revert #224931 for alternate nodes
Do not pass scalar context to alternate node vector cost queries.
PR #224931 passes the scalar main/alternate op as the context instruction when
costing the vector ops of an alternate node. Those vector ops only feed the
lane-select shuffle, so use-based discounts derived from the scalar's users
do not apply. On AMDGPU this priced an [fmul | fadd] node's vector fmul as
fused (free) where the sole user of the scalar fmul is fadd/fsub, keeping
unprofitable subtrees vectorized. In a rocFFT kernel on gfx950 this raised
VGPR usage from 160 to 178 and caused a performance regression. Restore the
null context for these queries.
[SLP] Partially revert #224931 for alternate nodes
Do not pass scalar context to alternate node vector cost queries.
PR #224931 passes the scalar main/alternate op as the context instruction when
costing the vector ops of an alternate node. Those vector ops only feed the
lane-select shuffle, so use-based discounts derived from the scalar's users
do not apply. On AMDGPU this priced an [fmul | fadd] node's vector fmul as
fused (free) where the sole user of the scalar fmul is fadd/fsub, keeping
unprofitable subtrees vectorized. In a rocFFT kernel on gfx950 this raised
VGPR usage from 160 to 178 and caused a performance regression. Restore the
null context for these queries.
[libc][test] Fix death test timeouts and dead-code elimination in math tests. (#227966)
- In `ExecuteFunctionUnix.cpp`, call `prctl(PR_SET_DUMPABLE, 0)` in
child processes to prevent external core dump handlers (such as apport)
from intercepting expected crashes during death tests, eliminating 10s
poll timeouts under parallel lit runs.
- Close `pipe_fds[0]` properly in `invoke_in_subprocess` to avoid
leaking file descriptors.
- In math smoke test templates (`AddTest.h`, `SubTest.h`, `MulTest.h`,
`DivTest.h`), assign test function results to `[[maybe_unused]] volatile
OutType res` to prevent GCC from dead-code eliminating floating point
operations in `test_inexact_results` when FMA optimization is disabled.
Assisted-by: Gemini