[GlobalISel][NFC] Restore the extended LLT flag to its saved value in tests (#224897)
AArch64GISelMITest's setUp() constructs an AArch64TargetMachine, which
unconditionally enables the process-global extended LLT flag. Ending the
ExtLLT tests with setUseExtended(false) therefore disables the flag for
every test that runs afterwards in the same process.
Save the flag's value on entry and restore it on exit instead.
[lld][LoongArch] Prevent relaxation oscillation for LA.PCRel and CALL
Relaxation of pcalau12i+addi (relaxPCHi20Lo12, isInt<22>) and
call36/call30 (relaxMediumCall, isInt<28>) can oscillate: shrinking
one section moves a symbol, which flips isInt<N> for other sites and
changes bytesDropped again.
Follow the same approach as RISCV::relaxCall: after a few passes, do
not allow remove to increase beyond the previous pass's value
(cur - delta). Pass that cap as prevRemove into the two helpers;
range checks may still clear remove (0) when the target goes out of
range.
[CIR][AMDGPU] Implement __builtin_amdgcn_*_dpp* builtins (#226469)
This commit implements the `__builtin_amdgcn_update_dpp`,
`__builtin_amdgcn_mov_dpp`, and `__builtin_amdgcn_mov_dpp8` builtins in
CIR, closely matching the implementation in
CodeGenFunction::EmitAMDGPUBuiltinExpr from OGCG.
Assisted-by: Claude Sonnet 5
Signed-off-by: Steffen Holst Larsen <sholstla at amd.com>
Reland [flang] Support scoped LICM and OpenACC capture provenance (#225401) (#226909)
Follow compute-region capture operands when checking whether scalar and
scalar-descriptor loads are safe to speculate. Preserve the existing
optional, array-element, and loop-modification safety checks.
Add an optional only-inside operation-name selector while retaining
function-scoped alias analysis. Resolve the name once per function and
compare interned operation names during ancestor traversal. The default
continues to select all loops. An explicit name selects loops with a
matching ancestor within the function, including the function itself.
Cover capture safety, host exclusion, non-OpenACC and nested scopes,
loop boundaries, unmatched names, and the function boundary.
The motivation for this is to allow running LICM only on device relevant
loop at O0.
Reland #225401 with CMakeFiles.txt change to fix shared library builds.
[CodeExtractor][Verifier] Fix OoB read when a DIExpression is used multiple times (#226857)
#224360 made fixupDebugInfoPostExtraction reuse the existing
DIExpression, but this is not sound if the expression is referenced
multiple times, as occurs with cold/hot code splitting.
This PR is the trivial fix of restricting this change to only apply if
there is a single user of the expression.
I've also added an additional verifier guard to capture these failures.
Fixes #226848
AI usage: Claude used to find a way to construct
dbg-value-arg-index-out-of-range.ll so that I could add a verifier guard
that bypassed the other existing verifier guards.
AMDGPU: Make rewrite-vgpr-mfma-to-agpr-spill-multi-store.ll less allocator sensitive
This test is sensitive to the exact split and spills which occur, and disappeared
under a future upstream improvement. Use basic RA with a fixed occupancy since it more
stably produces the spill pattern.
Also add a codegen reference test for the same kernel, so future codegen improvements are
visible.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
[GlobalISel] Use getCmpLibcallReturnType() for FCMP libcalls (NFC) (#226813)
The GCC soft-float comparison routines return `CMPtype`, not always
`i32`.
https://gcc.gnu.org/onlinedocs/gccint/Soft-float-library-routines.html#Comparison-functions-1
> 3.2.3 Comparison functions
> There are two sets of basic comparison functions.
> ...
> Runtime Function: CMPtype __unordsf2 (float a, float b)
> Runtime Function: CMPtype __unorddf2 (double a, double b)
> Runtime Function: CMPtype __unordtf2 (long double a, long double b)
> ...
The FCMP libcall result was hardcoded to i32. Use
`getCmpLibcallReturnType()`, as SelectionDAG does.
NFC for in-tree targets.
Co-authored-by: Thorbjørn Ravn Andersen <tra at ravnand.dk>
[AMDGPU] Add wait states between different MFMAs sharing an accumulator (#218363)
A full-register src2/C read after an MFMA write emits no wait states,
relying on accumulator forwarding that only works while the chain stays
on one MFMA. Two different MFMAs sharing an accumulator instead need the
wait states of a partial overlap: on gfx950, v_mfma_f32_16x16x32_f16
then
v_mfma_f32_16x16x16_f16 on the same tuple needs 5 and got none. Compare
canonicalized opcodes so mac and register-bank forms still match.
Assisted-by: Claude Opus 5
[CGData] Stop exporting cl::opts. NFC (#226861)
LTO reads `-codegen-data-thinlto-two-rounds` through
`extern cl::opt<bool> CodeGenDataThinLTOTwoRounds`, and llvm-cgdata
assigns
`IndexedCodeGenDataLazyLoading`, exported from CodeGenDataReader.h. Add
`cgdata::thinLTOTwoRounds()` for LTO, and pass lazy loading to
`CodeGenDataReader::create` as a parameter, so that the options can be
file-local.
Aided by Opus 5.5
[ORC] Remove callSPSWrapper, callSPSWrapperAsync, and callWrapper (#226893)
All in-tree callers now use Proxy. Remove the SPS convenience call
methods from ExecutionSession and ExecutorProcessControl, along with the
unit tests that exercised them (Proxy dispatch is covered by
SPSProxySpecTest). Also remove the blocking callWrapper convenience
methods, which have no remaining users.
Add a "How to call functions in the executor" section to
llvm/docs/ORCv2.md describing Proxy, controller-interface descriptors
and sps::ProxySpec, how lookupAndApply and recordProxy resolve proxies,
and how to migrate from callSPSWrapper.
Clients calling these methods directly should migrate to Proxy; see "How
to call functions in the executor" in llvm/docs/ORCv2.md.
[orc-rt] Add initial C addressing regression tests. (#226901)
Add tests that check that JIT'd C code can address data, both in the
same object and in other objects. Each test covers a single construct
(e.g. a load of static data, or a pointer to an array element in another
object stored in initialized data), since linker/loader bugs usually
crash the JIT'd program, and a crash identifies only the failing test.
Each test runs at -O0 and -O2.
To support multi-object tests, add a split-file substitution and lit
feature, and accept .test files in languages/c.
Update the README's test conventions to match: one construct per test,
"Check that" and "Stresses:" header comments, -O0 and -O2 RUN lines, and
guidance on keeping constructs alive under optimization without hiding
the optimized lowering.
Assisted-by: Claude