[HIP] Compute kernel handle dso_local as for a variable (#229060)
The host-side kernel handle is a global variable, but getKernelHandle
copied dso_local from the device stub while the stub was still a
function declaration. emitDeviceStub then updated the handle's linkage,
but not its dso_local. As a result, dso_local followed the rules for
function declarations rather than those for the handle:
- Under PIE, and with -fno-plt or -fno-direct-access-external-data in
static builds, defined handles were not dso_local, though a definition
in an executable can't be preempted. Every reference to them went
through the GOT.
- With static -fno-plt, a declared handle was not dso_local, though data
can use a copy relocation. -fno-plt only affects calls.
- On MinGW, a declared handle was dso_local, though a variable may be
auto-imported from a DLL.
Compute dso_local with CGM.setDSOLocal when the handle is created and
again once emitDeviceStub gives it its final linkage. The result is what
[5 lines not shown]
[mlir][ArmSME] Make tile allocation order deterministic (#229096)
`gatherTileLiveRanges` used a `DenseMap<Value, LiveRange>`, so tile-ID
tie-breaking depended on pointer hash order: _non-deterministic_ in
nature. This PR switches it to `MapVector` for insertion-order
iteration.
Also restores the _move required_ diagnostic on `overlapping_branches`
(dropped in #227153).
Additionally, it adds the forward declaration of `MapVector` in
[`mlir/Support/LLVM.h`](https://github.com/llvm/llvm-project/blob/7bd5b596c3d0afff84caebe3be6f0e61190e937b/mlir/include/mlir/Support/LLVM.h)
to avoid the prefix `llvm::`.
---------
Signed-off-by: Federico Bruzzone <federico.bruzzone.i at gmail.com>
AMDGPU: Fix true16 build_vector (0, x) pattern using a 16-bit shift operand
The real true16 pattern for (build_vector 0, VGPR_16:$x) fed the 16-bit
register directly to V_LSHLREV_B32, which takes a 32-bit operand. Widen
it with a REG_SEQUENCE first. This avoids redundant 16-bit moves in
SelectionDAG, and fixes a GlobalISel selection failure when the 16-bit
input is a G_TRUNC of a 32-bit value, as the shift's operand class
constrained the trunc result to vgpr_32.
I also don't know why this pattern is overcomplicating this. I would expect
true16 to literally translate build_vector to reg_sequence plus a materialize
of the 0.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
X86: Only drop EFLAGS def in convertToThreeAddress if MI defines EFLAGS (#230293)
convertToThreeAddress unconditionally removed the EFLAGS value at the converted
instruction's slot. For instructions that do not define EFLAGS, such as masked moves
converted to blends or the _NF variants, a live-through EFLAGS value could be
removed leaving a missing segment.
Reported in
https://github.com/llvm/llvm-project/pull/225174#issuecomment-6056927367
Co-authored-by: Claude (Claude-Opus-5.5)
AMDGPU: Use functions in more operation tests instead of kernel loads (#230195)
Stop relying on -amdgpu-scalarize-global-loads=false. Inputs are passed
as VGPR arguments, or inreg for SGPR operands. Kernels that check
multiple stores index their loads by workitem id.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
LiveVariables: Only visit tracked physical registers
Keep a bitvector of physical registers with a recorded def or use in
the current block. Register mask handling, the end of block scan, and
the per-block reset now only visit those registers instead of every
register. This is significant for targets with many registers, such as
AMDGPU.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
LiveVariables: Remove dead live-in handling and Defs plumbing (#230150)
No physical register is tracked at the start of a block, so handling the
block live-ins was a no-op. The Defs list was only appended for
instruction defs, which runOnInstr already collects.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[KnownBits] Improve minimum denominator bounds for sdiv (#227264)
This change detects a zero value from getSignedMinValue() and sets the
lowest unknown bit using countMinTrailingZeros() to get the minimum
non-zero denominator.
For example, a denominator such as 0b0?00 can represent {0, 4}. Since
zero is not a valid divisor, the minimum non-zero denominator should be
used for the sdiv.
Added regression test.
added lit test Transforms/InstCombine/sdiv-knownbits.ll (fails without
the patch, passes with it)
added unit test
ninja -C build check-llvm-unit
ninja -C build check-llvm
git diff --check
Formatted using git clang-format HEAD
[8 lines not shown]
[TargetLowering][X86] Prefer 'r' over 'm' for foldable "rm" inline asm operands
An "rm" (register-or-memory) inline asm operand has always resolved to
'm', because getConstraintPreferences() picks the most general
constraint present, and 'm' is more general than 'r'. That's safe, since
memory can't run out, but it forces a value that could stay in a
register through a stack slot even when there's no register pressure
(https://github.com/llvm/llvm-project/issues/20571).
Prefer 'r' instead where the register allocator can fold the register
back to a stack slot when it runs out of registers, and mark the
register operand foldable (InlineAsm::Flag::setRegMayBeFolded()) so it
does. Both allocators can: the greedy allocator folds an operand when it
spills its value, and the fast allocator folds operands up front when
the asm's register operands wouldn't fit.
ParseConstraints() sets AsmOperandInfo::MayFoldRegister for an operand
whose constraint codes include 'r' and 'm' and are otherwise only
immediate codes, above -O0, on a target that opts in through the new
[26 lines not shown]
[SLP][modularisation][NFC] Move loop trip-count helpers to SLPUtils
Move the BoUpSLP-independent helpers findInnermostNonInvariantLoop and
getLoopTripCount out of SLPVectorizer.cpp into the self-contained
SLPVectorizer/SLPUtils.{h,cpp} module. getLoopTripCount reads the file-local
LoopAwareTripCount cl::opt, which stays static in SLPVectorizer.cpp and is
passed to the moved helper as an explicit parameter. NFC.
Part of the SLPVectorizer.cpp modularization effort:
https://discourse.llvm.org/t/modularizing-slpvectorizer-cpp/90922
[XRay] Fix wrong entry sled placement after TRE (#230146)
This fixes issue #229740.
XRay entry sleds are currently inserted in the first non-empty block (a
precaution to skip completely empty functions).
The application of tail recursion elimination may leave the entry block
empty, thus causing `PATCHABLE_FUNCTION_ENTER` to be placed in the
succeeding block.
In the reproducer of the linked issue, this block happens to also be the
loop header of the transformed recursion.
This results in the emission of `ENTRY` events during every loop
iteration, followed by a single `EXIT`.
The fix is straightforward:
- Leave the existing look-up of the first non-empty block in place, to
prevent insertion into empty functions.
- If the function is non-empty, use the entry block instead for sled
insertion.
[7 lines not shown]
[SLP][modularisation][NFC] Move getReductionInstr to SLPReductionUtils, getAggregateSize to SLPUtils
Move the reduction-root helper getReductionInstr out of SLPVectorizer.cpp
into the self-contained SLPVectorizer/SLPReductionUtils.{h,cpp} module.
getAggregateSize is not reduction-specific: its only user is
findBuildAggregate and it only computes the element count of a homogeneous
aggregate from its type. Move it to the generic SLPVectorizer/SLPUtils
module instead.
Part of the SLPVectorizer.cpp modularization effort:
https://discourse.llvm.org/t/modularizing-slpvectorizer-cpp/90922
Assisted by AI
[Hexagon] Require asserts for autohvx xqf-postra-conv-double2.ll (#230348)
autohvx/xqf-postra-conv-double2.ll passes -debug-only=handle-qfp to llc,
which is only supported when LLVM is built with assertions enabled.
Without 'REQUIRES: asserts', the test fails in non-asserts builds with:
llc: Unknown command line argument '-debug-only=handle-qfp'.
Add 'REQUIRES: asserts' matching other tests in this directory.
Fixes #227779
[Passes] Port -print-pipeline-passes to PassesOptions (#230220)
-print-pipeline-passes takes an optional value (=text or =tree). Add
FlagOrEnumField, an EnumField whose bare option assigns a given
enumerator and never consumes the next argument, as cl::ValueOptional
does. BoolField and OptionalBoolField use the same BareValue.
clang, flang, LTO, llc, and opt read the format through
PassBuilder::getPrintPipelinePasses() instead of the exported cl::opt.
Aided by Opus 5.5
[SLP][modularisation][NFC] Move isFirstInsertElement/getDebugLocFromPHI to SLPUtils
Move the BoUpSLP-independent helpers isFirstInsertElement and
getDebugLocFromPHI out of SLPVectorizer.cpp into the self-contained
SLPVectorizer/SLPUtils.{h,cpp} module.
Part of the SLPVectorizer.cpp modularization effort:
https://discourse.llvm.org/t/modularizing-slpvectorizer-cpp/90922
[AMDGPU] Insert MFMA anti-hints in GCNPreRAOptimizations (#218075)
This PR adds rule-based anti-hints in the GCNPreRAOptimizations phase so
that
the register allocator prefers registers that avoid hazard nops.
Instructions
are classified into different classes and hazard windows are computed
from the
waitstate info. The rules are built to describe the hazardous situations
and
their hazard windows. Added an AntiHintEngine that uses these rules to
insert
anti-hint relations between vregs, which the register allocator later
uses to
prefer allocations that avoid the corresponding hazardous physical
registers.
## Stack
PR **4/4**. Depends on #218074. Next #226397
Reject AFFINITY strides and nested vector subscripts
OpenMP allows a stride expression in an array section only where a clause
explicitly permits one, and AFFINITY does not. Semantics accepted positive
strides and lowering only reported a non-fatal error, so compilation
succeeded or crashed on the bounds. Report a stride in the last part
reference of an AFFINITY locator as a semantic error, including an explicit
unit stride. DEPEND and REDUCTION fall under the same rule but currently
accept strides; tightening them is left to a separate change.
The vector-subscript TODO for AFFINITY only checked the last part
reference, so a(v)%field(1) reached an assertion. Check every part
reference instead.
[cmake] Limit `/Zc:dllexportInlines-` flag to C/C++ (#227791)
This adjusts the way the `/Zc:dllexportInlines-` flag is passed so it
only takes effect for C and C++ targets built with clang-cl. These
changes are needed to prevent downstream consumers like Swift from
breaking due to being passed the flag.
The effort to build LLVM as a dylib is tracked in #109483.
[clangd] Clear LLVM_LINK_COMPONENTS for test plugin modules (#230359)
clang-tools-extra/clangd/CMakeLists.txt sets directory-scoped
LLVM_LINK_COMPONENTS (including Support), which is inherited by
clang-tools-extra/clangd/test/plugins/CMakeLists.txt.
When building ClangdFeatureModuleExample and
ClangdStandaloneTweakExample (added in #226995), llvm_add_library
statically links libLLVMSupport.a into the plugin shared libraries
unless LLVM_LINK_COMPONENTS is cleared before creating the targets.
Loading these plugins into clangd then triggers an ASan ODR violation on
llvm::vfs::FileSystem::ID because both clangd and the plugin module
contain their own copy of VirtualFileSystem.cpp.o.
Clear LLVM_LINK_COMPONENTS before calling llvm_add_library so the
plugins resolve LLVM symbols from the clangd executable at load time,
and remove the dead post-target LLVM_LINK_COMPONENTS block.
Fixing https://lab.llvm.org/buildbot/#/builders/52/builds/20707
Assisted-by: Gemini
[SLP][modularisation][NFC] Move getReductionInstr/getAggregateSize to SLPReductionUtils
Move the BoUpSLP-independent helpers getReductionInstr and getAggregateSize
out of SLPVectorizer.cpp into the self-contained
SLPVectorizer/SLPReductionUtils.{h,cpp} module.
Part of the SLPVectorizer.cpp modularization effort:
https://discourse.llvm.org/t/modularizing-slpvectorizer-cpp/90922
[SLP][modularisation][NFC] Move isFirstInsertElement/getDebugLocFromPHI to SLPUtils
Move the BoUpSLP-independent helpers isFirstInsertElement and
getDebugLocFromPHI out of SLPVectorizer.cpp into the self-contained
SLPVectorizer/SLPUtils.{h,cpp} module.
Part of the SLPVectorizer.cpp modularization effort:
https://discourse.llvm.org/t/modularizing-slpvectorizer-cpp/90922
[SLP][modularisation][NFC] Move loop trip-count helpers to SLPUtils
Move the BoUpSLP-independent helpers findInnermostNonInvariantLoop and
getLoopTripCount out of SLPVectorizer.cpp into the self-contained
SLPVectorizer/SLPUtils.{h,cpp} module. getLoopTripCount reads the file-local
LoopAwareTripCount cl::opt, which stays static in SLPVectorizer.cpp and is
passed to the moved helper as an explicit parameter. NFC.
Part of the SLPVectorizer.cpp modularization effort:
https://discourse.llvm.org/t/modularizing-slpvectorizer-cpp/90922
[AMDGPU] Promote wmma_f32_16x16x128_f8f6f4 to scaled version on gfx1250-strict (#230253)
Added scale factor is 0 which does not change the result. "0" is a
special value that maps to the exponent "1.0" when passed as a constant.
[CIR] Lower non-power-of-two vectors in CallConvLowering (#230087)
Summary:
- The x86_64 classifier now sizes a vector at its ABI size (#227889), so
the bridge no longer needs to reject vectors whose width is not a power
of two.
- The SSEUP walk rounds vector widths the same way, so a union holding
such a vector is still refused at sizes with no coerce type.
- This issue came up while enabling `CallConvLowering` for AMDGPU in
#220197: the gfx950/gfx1250 `transpose-load` builtins return
three-element vectors, and
`CIR/CodeGenHIP/builtins-amdgcn-gfx950-read-tr.hip` and
`builtins-amdgcn-gfx1250-load-tr.hip` fail on the gate.
Related to issue: #220471
Assisted by : claude opus 5.5
[Clang] Mark scoped_atomics with !noalias.addrspace(private)
The HIP specification marks atomics on thread private memory as UB.
Scoped atomics used within a HIP context are also considered UB,
unless explicitley specified via a command line argument.
These are now annotated with !noalias.addrspace(5) for amdgpus,
to avoid an expensive runtime check.