[mlir][memref] Enforce consistent reinterpret_cast metadata (#217338)
This PR implements [[RFC] Clarify `memref.reinterpret_cast` Verification
of Dynamic
Metadata](https://discourse.llvm.org/t/rfc-clarify-memref-reinterpret-cast-verification-of-dynamic-metadata).
## Context
`memref.reinterpret_cast` describes its result through both its mixed
offset/size/stride metadata and its result `MemRefType`. Currently,
verification allowed a dynamic result-type correspond to static
operation metadata, introducing inconsistent descriptors.
For example, the verifier permits:
```mlir
%r = memref.reinterpret_cast %base
to offset: [0], sizes: [%n, 2], strides: [%s, 1] : memref<f32>
to memref<?x?xf32, strided<[?, ?], offset: ?>>
```
where the operation constructs a descriptor with a statically known
[32 lines not shown]
[lldb] Fix offload bundle unit tests on 32-bit systems (#226922)
DataExtractor::GetByteSize returns uint64_t for reasons I don't yet
understand. I might change it to size_t but let's unblock the build for
now.
Fixes #222362 / 08ee0ef5da10d27d5776df1e1ff81c46a38d9b5c.
[X86] Select standalone ~(-1 << n) masks as BZHI (#226158)
InstCombine canonicalizes (1 << n) - 1 to xor (shl -1, n), -1. When that
mask feeds an AND, X86DAGToDAGISel::matchBitExtract folds the whole
expression into BZHI, and the add form is selected as BZHI even on its
own because Select enters matchBitExtract from ISD::ADD. The xor form on
its own is not: a mask that is returned, stored, or consumed by anything
but AND is selected as mov -1; shlx; not.
Enter matchBitExtract from ISD::XOR too under BMI2, so the standalone
mask becomes mov -1; bzhi. That is one instruction shorter and avoids
the shlx+not dependency chain. BMI1-only targets keep the shift form,
since BEXTR would need the count moved into bits 15:8 first.
Assisted-by: Claude Code
[VPlan] Remove (X && Y) | (X && !Y) -> X combine. NFC (#219368)
We have smaller combines that can take care of this now that we process
recipes in a worklist after #213900
[GlobalISel][NFC] Restore the extended LLT flag to its saved value in tests (#224897)
AArch64GISelMITest's setUp() constructs an AArch64TargetMachine, which
unconditionally enables the process-global extended LLT flag. Ending the
ExtLLT tests with setUseExtended(false) therefore disables the flag for
every test that runs afterwards in the same process.
Save the flag's value on entry and restore it on exit instead.
[lld][LoongArch] Prevent relaxation oscillation for LA.PCRel and CALL
Relaxation of pcalau12i+addi (relaxPCHi20Lo12, isInt<22>) and
call36/call30 (relaxMediumCall, isInt<28>) can oscillate: shrinking
one section moves a symbol, which flips isInt<N> for other sites and
changes bytesDropped again.
Follow the same approach as RISCV::relaxCall: after a few passes, do
not allow remove to increase beyond the previous pass's value
(cur - delta). Pass that cap as prevRemove into the two helpers;
range checks may still clear remove (0) when the target goes out of
range.
[CIR][AMDGPU] Implement __builtin_amdgcn_*_dpp* builtins (#226469)
This commit implements the `__builtin_amdgcn_update_dpp`,
`__builtin_amdgcn_mov_dpp`, and `__builtin_amdgcn_mov_dpp8` builtins in
CIR, closely matching the implementation in
CodeGenFunction::EmitAMDGPUBuiltinExpr from OGCG.
Assisted-by: Claude Sonnet 5
Signed-off-by: Steffen Holst Larsen <sholstla at amd.com>
Reland [flang] Support scoped LICM and OpenACC capture provenance (#225401) (#226909)
Follow compute-region capture operands when checking whether scalar and
scalar-descriptor loads are safe to speculate. Preserve the existing
optional, array-element, and loop-modification safety checks.
Add an optional only-inside operation-name selector while retaining
function-scoped alias analysis. Resolve the name once per function and
compare interned operation names during ancestor traversal. The default
continues to select all loops. An explicit name selects loops with a
matching ancestor within the function, including the function itself.
Cover capture safety, host exclusion, non-OpenACC and nested scopes,
loop boundaries, unmatched names, and the function boundary.
The motivation for this is to allow running LICM only on device relevant
loop at O0.
Reland #225401 with CMakeFiles.txt change to fix shared library builds.
[CodeExtractor][Verifier] Fix OoB read when a DIExpression is used multiple times (#226857)
#224360 made fixupDebugInfoPostExtraction reuse the existing
DIExpression, but this is not sound if the expression is referenced
multiple times, as occurs with cold/hot code splitting.
This PR is the trivial fix of restricting this change to only apply if
there is a single user of the expression.
I've also added an additional verifier guard to capture these failures.
Fixes #226848
AI usage: Claude used to find a way to construct
dbg-value-arg-index-out-of-range.ll so that I could add a verifier guard
that bypassed the other existing verifier guards.
AMDGPU: Make rewrite-vgpr-mfma-to-agpr-spill-multi-store.ll less allocator sensitive
This test is sensitive to the exact split and spills which occur, and disappeared
under a future upstream improvement. Use basic RA with a fixed occupancy since it more
stably produces the spill pattern.
Also add a codegen reference test for the same kernel, so future codegen improvements are
visible.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
[GlobalISel] Use getCmpLibcallReturnType() for FCMP libcalls (NFC) (#226813)
The GCC soft-float comparison routines return `CMPtype`, not always
`i32`.
https://gcc.gnu.org/onlinedocs/gccint/Soft-float-library-routines.html#Comparison-functions-1
> 3.2.3 Comparison functions
> There are two sets of basic comparison functions.
> ...
> Runtime Function: CMPtype __unordsf2 (float a, float b)
> Runtime Function: CMPtype __unorddf2 (double a, double b)
> Runtime Function: CMPtype __unordtf2 (long double a, long double b)
> ...
The FCMP libcall result was hardcoded to i32. Use
`getCmpLibcallReturnType()`, as SelectionDAG does.
NFC for in-tree targets.
Co-authored-by: Thorbjørn Ravn Andersen <tra at ravnand.dk>
[AMDGPU] Add wait states between different MFMAs sharing an accumulator (#218363)
A full-register src2/C read after an MFMA write emits no wait states,
relying on accumulator forwarding that only works while the chain stays
on one MFMA. Two different MFMAs sharing an accumulator instead need the
wait states of a partial overlap: on gfx950, v_mfma_f32_16x16x32_f16
then
v_mfma_f32_16x16x16_f16 on the same tuple needs 5 and got none. Compare
canonicalized opcodes so mac and register-bank forms still match.
Assisted-by: Claude Opus 5
[CGData] Stop exporting cl::opts. NFC (#226861)
LTO reads `-codegen-data-thinlto-two-rounds` through
`extern cl::opt<bool> CodeGenDataThinLTOTwoRounds`, and llvm-cgdata
assigns
`IndexedCodeGenDataLazyLoading`, exported from CodeGenDataReader.h. Add
`cgdata::thinLTOTwoRounds()` for LTO, and pass lazy loading to
`CodeGenDataReader::create` as a parameter, so that the options can be
file-local.
Aided by Opus 5.5
[ORC] Remove callSPSWrapper, callSPSWrapperAsync, and callWrapper (#226893)
All in-tree callers now use Proxy. Remove the SPS convenience call
methods from ExecutionSession and ExecutorProcessControl, along with the
unit tests that exercised them (Proxy dispatch is covered by
SPSProxySpecTest). Also remove the blocking callWrapper convenience
methods, which have no remaining users.
Add a "How to call functions in the executor" section to
llvm/docs/ORCv2.md describing Proxy, controller-interface descriptors
and sps::ProxySpec, how lookupAndApply and recordProxy resolve proxies,
and how to migrate from callSPSWrapper.
Clients calling these methods directly should migrate to Proxy; see "How
to call functions in the executor" in llvm/docs/ORCv2.md.
[orc-rt] Add initial C addressing regression tests. (#226901)
Add tests that check that JIT'd C code can address data, both in the
same object and in other objects. Each test covers a single construct
(e.g. a load of static data, or a pointer to an array element in another
object stored in initialized data), since linker/loader bugs usually
crash the JIT'd program, and a crash identifies only the failing test.
Each test runs at -O0 and -O2.
To support multi-object tests, add a split-file substitution and lit
feature, and accept .test files in languages/c.
Update the README's test conventions to match: one construct per test,
"Check that" and "Stresses:" header comments, -O0 and -O2 RUN lines, and
guidance on keeping constructs alive under optimization without hiding
the optimized lowering.
Assisted-by: Claude
TargetMachine: Remove pointer-size query methods (#226404)
Remove the shim methods from the TargetMachine's copy of the DataLayout,
which will soon be eliminated. The Module owns the authoritative DataLayout,
so callers should read the value from the contextual Module.
Completely unreasonably, Mips's ABI name can change the pointer size which
we probably should just not support. Many other triple checks will never be
correct. This avoids potential mismatches in these contexts, but I still expect
this to be widely broken.
Some of the TargetLowering constructor changes and AMDGPULegalizerInfo
changes are kind of annoying. We could pass in the DataLayout through the
subtarget constructors but it didn't seem worth the effort and the information
should be derivable from the triple anyway.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
NVPTX: Drop LiveVariables from the register allocation pipeline
The optimized RegAlloc pipeline ran LiveVariables only to satisfy PHIElimination
and TwoAddressInstruction, both of which no longer need it. Remove the
LiveVariables run (and, in the new pass manager, the UnreachableMachineBlockElim
that was there only as a LiveVariables prerequisite).
Co-authored-by: Claude (Claude-Opus-4.8)
CodeGen: Drop the LiveVariables parameter from convertToThreeAddress
This was used for analysis updates, but now the analysis is being
removed.
Co-authored-by: Claude (Claude-Opus-4.8)
AMDGPU: Remove update-only LiveVariables maintenance from SILowerControlFlow
This was only maintained, never relied on. Part of staged LiveVariables
removal.
Co-authored-by: Claude (Claude-Opus-4.8)
CodeGen: Remove LiveVariables use from TwoAddressInstructionPass
Now that LiveIntervals is computed unconditionally before TwoAddressInstructions
in the pipeline, the pass no longer needs LiveVariables.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
CodeGen: Compute LiveIntervals before TwoAddressInstructions
TwoAddressInstructions is the traditional primary use of LiveVariables,
but it has gained a LiveIntervals path. By moving LiveIntervals earlier,
the default flips to rely on it instead of LiveVariables. The overall
test churn is mostly neutral, with more net wins than losses.
This should move before phi elimination. This is a staging move to
incrementally remove the LiveVariables support from TwoAddressInstructions,
and because the move to running LiveIntervals on SSA is a bigger leap.
Co-authored-by: Claude (Claude-Opus-4.8)