[MC] Add ELF::getSymbolInfo and fix bitwise OR between enumeration types. NFC (#226860)
`ELF::STT_FILE | ELF::STB_LOCAL` puts the binding in the type bits,
which only works because STB_LOCAL is 0.
It combines enumerators from two distinct unnamed enums, which is
deprecated in C++20 and ill-formed in C++26 (P2864R2).
[orc-rt] Prefix test tool output, update test. (#226862)
Prefix the triple, page-size, and cpu-features lines of the
orc-rt-process-info-check. The prefixes can be used to ensure that the
test doesn't try to match patterns against the wrong lines.
Also relaxes the feature string check to allow an empty feature string.
[AMDGPU][InstCombine] Fold zero dot operands to accumulator (#225003)
Fold AMDGPU dot intrinsics when either operand is zero.
`dot(a, 0) = 0` and `dot(0, b) = 0`, so replace the intrinsic with its
accumulator.
This avoids unrelated clamp and add/sub reassociation cases.
[SlotIndexes] Add queries for stale indexes
An erased instruction leaves its index list entry in place, making the
index indistinguishable from a block boundary entry. Add
isBlockBoundaryIndex() and isStaleIndex() to tell the two apart, and
canonicalizeIndex() to resolve a stale index to the closest preceding
instruction's register slot, or the block start if none survives.
NFC. No caller yet. LiveDebugVariables is next.
[LiveDebugVariables] Repair stale SlotIndexes
The analysis keeps its indexes from before the first register allocator
until DBG_VALUEs are emitted, by which point passes in between have
erased some of the instructions they point at. Resolve them at the
start of each allocator run and before emitting.
SlotIndexes can then reclaim the entries of erased instructions without
sparing the ones held here, which would have made generated code depend
on -g. Emitted locations are unchanged, except that intervals resolving
to one position now emit a single DBG_VALUE rather than identical
consecutive ones.
[Github] Add LLVM-libc's full-build to Bazel Checks workflow
This PR adds a job to run `bazelisk test --@llvm-project//libc:build_mode=full @llvm-project//libc/...`. It is part of an effort to support LLVM-libc's full build functionality in Bazel.
[bazel][libc] Refactor to allow copts under full-build mode
This PR moves away from having sets of predefined groups of copts and allows just directly using them. This lets the Bazel rules more directly parallel CMake.
[ORC] Call Perf/VTune support wrappers through Proxies (#226677)
Hold PerfSupportPlugin's start/end registration wrappers and
VTuneSupportPlugin's unregister wrapper as Proxy members, replacing
their callSPSWrapper calls. Each proxy is built in the constructor from
the address it already takes, so the constructor signatures are
unchanged. The Perf impl and VTune register wrappers are only used as
alloc-action tags and stay ExecutorAddrs.
[clang] Unique NamespaceAndPrefixStorages with a UniquingSet (NFC) (#224221)
This patch migrates NamespaceAndPrefixStorages in ASTContext from
llvm::FoldingSet to llvm::UniquingSet.
NamespaceAndPrefixStorage keys on a pair of const NamespaceBaseDecl *
and NestedNameSpecifier. Switching to UniquingSet allows us to look
up storages with a typed key, eliminating FoldingSetNodeID
serialization at lookup sites and removing
NamespaceAndPrefixStorage::Profile.
Assisted-by: Antigravity
[AMDGPU] Price scalar integer to fp casts by source width and sign
Scalar sources between a byte and 31 bits fell to the default cost of one
while the matching vector lanes were already priced, which skewed the
difference SLP weighs a bundle against. Such a source is extended before
the conversion, and what the extension takes depends on the width, on the
sign and on whether the subtarget has SDWA and 16 bit instructions.
Sources narrower than a byte are left alone, because their vector form is
not priced either.
[AMDGPU] Price narrow integer to bfloat vector casts
A vector lane of 9 to 15 or 17 to 31 bits converted to bfloat got the
generic cost, which leaves out the rounding. Such a lane is converted to
f32 first like any other narrow lane, so price it as the f32 conversion of
the same source plus the rounding. Lanes of 8 and 16 bits keep their cost.
[AMDGPU] Model the cost of the expanded integer to/from floating point casts
No instruction converts to or from a 64 bit integer, and narrow vector
lanes are converted one at a time. Price these expansions by the FP64
rate, sdwa and 16 bit instruction support, and price i33 to i63 like
i64 and bf16 like f32 plus rounding.
Assisted-by: Claude Code Opus 5
[NFC][AMDGPU] Add cost tests for narrow integer to fp casts
Covers integer sources from a byte to 31 bits converted to half, float,
bfloat and double, as vector lanes and as scalars, over the subtarget
combinations that change the expansion. The existing cast tests get the
same subtarget coverage and the cases they were missing. The costs
recorded here are the ones the model reports today.
[LV] Remove EpilogueLoopVectorizationInfo::EpilogueUF (NFC). (#226792)
The epilogue vector loop is always unrolled by 1: the only construction
site passes 1 and the constructor asserted it. Drop the field and use 1
directly at its users, and drop the now always 1 EpilogueUF parameter of
addMinimumVectorEpilogueIterationCheck.
Clean-up in preparation for removing/simplifying
EpilogueLoopVectorizationInfo.
CodeGen: Merge TargetLoweringObjectFile::getModuleMetadata into Initialize
getModuleMetadata had a single caller, which invoked it immediately after
Initialize. Pass the module to Initialize and fold it in.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
CodeGen: Initialize TargetLoweringObjectFile from MachineModuleInfo
MachineModuleInfo passes TLOF to the MCContext the TargetLoweringObjectFile
but nothing initialized it until the AsmPrinter pass ran, so every codegen pass
in between saw it uninitialized. Initialize it from MachineModuleInfo, and drop
the calls llc and SPIRVTranslate used to work around this.
SPIRVTranslate's MachineModuleInfoWrapperPass was never passed on to
addPassesToEmitFile, so it was initializing a throwaway context.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
[unittests] Fix a recent test on Windows with LLVM_WINDOWS_PREFER_FORWARD_SLASH (#224963)
This fixes errors like these:
```
Expected equality of these values:
Context.Path
Which is: "D:/a/llvm-mingw/llvm-mingw/llvm-project/build/unittests/Support/LLVMToolSession/./LLVMToolSessionTests.exe"
ExecutablePath.c_str()
Which is: "D:\\a\\llvm-mingw\\llvm-mingw\\llvm-project\\build\\unittests\\Support\\LLVMToolSession\\.\\LLVMToolSessionTests.exe"
```
[DenseMap] Share rehash and grow for relocatable bucket types. NFC (#225018)
DenseMap is heavily instantiated and moveFrom and grow are among the
largest code families. For trivially copy constructible and destructible
bucket types (also satisfied by std::pair), call an out-of-line
`growRelocatable` with size/align/hasher (dictionary passing style).
A pointer key hashed by its value hashes inline, which a null
BucketHasher asks for, sparing a call per entry.
DenseMapInfo<T *>::PointerValueHash names the class declaring it, so an
info that specializes or inherits it to hash the pointee -- Attributor's
InstExclusionSet, VPCSEDenseMapInfo, MachineInstrExpressionTrait --
keeps its own hash. Any other key hashes through a thunk, one per
(KeyT, KeyInfoT) whatever the map's value type.
SmallDenseMap's in-place rehash for remove_if shares the loop through
`rehashRelocatable`. It rehashes once rather than into a temporary map
and back, which can reorder a probe chain that wraps around the table.
[3 lines not shown]
[IR] Memory effects for floating-point operations
Floating-point operations in a strictfp function have side effects,
which are modeled using memory effects in the form of a read-write
access to "inaccessible memory". This helps maintain strictfp semantics
but may hinder optimizations. Floating-point operations may depend on
rounding mode or not - this fact may be used to reorder them in a more
optimal way. Similarly, functions that control floating-point
environment (like `set_rounding`, `set_fpmode` etc.) also have more
specific access than generic read-write. Also, "inaccessible memory" is
used in cases other than FP operation, this results in unnecessary
restrictions.
This change implements two new memory location to use instead of the
access to "inaccessible memory". The "fpcontrol" location is used to
represent access to floating-point control modes, of which only rounding
mode is currently supported. The other location, "fpstatus", represents
access to floating-point exceptions. Together they replace the use of
"inaccessible memory".
[21 lines not shown]