X86: Only drop EFLAGS def in convertToThreeAddress if MI defines EFLAGS (#230293)
convertToThreeAddress unconditionally removed the EFLAGS value at the converted
instruction's slot. For instructions that do not define EFLAGS, such as masked moves
converted to blends or the _NF variants, a live-through EFLAGS value could be
removed leaving a missing segment.
Reported in
https://github.com/llvm/llvm-project/pull/225174#issuecomment-6056927367
Co-authored-by: Claude (Claude-Opus-5.5)
AMDGPU: Use functions in more operation tests instead of kernel loads (#230195)
Stop relying on -amdgpu-scalarize-global-loads=false. Inputs are passed
as VGPR arguments, or inreg for SGPR operands. Kernels that check
multiple stores index their loads by workitem id.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
LiveVariables: Only visit tracked physical registers
Keep a bitvector of physical registers with a recorded def or use in
the current block. Register mask handling, the end of block scan, and
the per-block reset now only visit those registers instead of every
register. This is significant for targets with many registers, such as
AMDGPU.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
LiveVariables: Remove dead live-in handling and Defs plumbing (#230150)
No physical register is tracked at the start of a block, so handling the
block live-ins was a no-op. The Defs list was only appended for
instruction defs, which runOnInstr already collects.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[KnownBits] Improve minimum denominator bounds for sdiv (#227264)
This change detects a zero value from getSignedMinValue() and sets the
lowest unknown bit using countMinTrailingZeros() to get the minimum
non-zero denominator.
For example, a denominator such as 0b0?00 can represent {0, 4}. Since
zero is not a valid divisor, the minimum non-zero denominator should be
used for the sdiv.
Added regression test.
added lit test Transforms/InstCombine/sdiv-knownbits.ll (fails without
the patch, passes with it)
added unit test
ninja -C build check-llvm-unit
ninja -C build check-llvm
git diff --check
Formatted using git clang-format HEAD
[8 lines not shown]
[TargetLowering][X86] Prefer 'r' over 'm' for foldable "rm" inline asm operands
An "rm" (register-or-memory) inline asm operand has always resolved to
'm', because getConstraintPreferences() picks the most general
constraint present, and 'm' is more general than 'r'. That's safe, since
memory can't run out, but it forces a value that could stay in a
register through a stack slot even when there's no register pressure
(https://github.com/llvm/llvm-project/issues/20571).
Prefer 'r' instead where the register allocator can fold the register
back to a stack slot when it runs out of registers, and mark the
register operand foldable (InlineAsm::Flag::setRegMayBeFolded()) so it
does. Both allocators can: the greedy allocator folds an operand when it
spills its value, and the fast allocator folds operands up front when
the asm's register operands wouldn't fit.
ParseConstraints() sets AsmOperandInfo::MayFoldRegister for an operand
whose constraint codes include 'r' and 'm' and are otherwise only
immediate codes, above -O0, on a target that opts in through the new
[26 lines not shown]
[SLP][modularisation][NFC] Move loop trip-count helpers to SLPUtils
Move the BoUpSLP-independent helpers findInnermostNonInvariantLoop and
getLoopTripCount out of SLPVectorizer.cpp into the self-contained
SLPVectorizer/SLPUtils.{h,cpp} module. getLoopTripCount reads the file-local
LoopAwareTripCount cl::opt, which stays static in SLPVectorizer.cpp and is
passed to the moved helper as an explicit parameter. NFC.
Part of the SLPVectorizer.cpp modularization effort:
https://discourse.llvm.org/t/modularizing-slpvectorizer-cpp/90922
[XRay] Fix wrong entry sled placement after TRE (#230146)
This fixes issue #229740.
XRay entry sleds are currently inserted in the first non-empty block (a
precaution to skip completely empty functions).
The application of tail recursion elimination may leave the entry block
empty, thus causing `PATCHABLE_FUNCTION_ENTER` to be placed in the
succeeding block.
In the reproducer of the linked issue, this block happens to also be the
loop header of the transformed recursion.
This results in the emission of `ENTRY` events during every loop
iteration, followed by a single `EXIT`.
The fix is straightforward:
- Leave the existing look-up of the first non-empty block in place, to
prevent insertion into empty functions.
- If the function is non-empty, use the entry block instead for sled
insertion.
[7 lines not shown]
Randonly, when starting a daemon that sets rc_bg=YES, we may end up with:
/etc/rc.d/foo: kill: 123456: no such process
There's a race in the rc.subr(8) system where we can end up killing a non
existing process (which PID may already be reused by another one).
The reason is that rc_alarm timer sends a SIGALRM to all children of rc_cmd to
detach rc_start. But the timer is also a child... and can then die while
rc_cmd() is about to kill it; that's when you get the error (as it's already
gone).
Re-implement so the timer ignore SIGALRM and let it get killed by rc_cmd.
My tests confirm this fixes the issue, so commit it early in the release
process to make sure this does not introduce any regression.
reported on bugs@ by adfsilva at gmail dot com
ok kn@
[SLP][modularisation][NFC] Move getReductionInstr to SLPReductionUtils, getAggregateSize to SLPUtils
Move the reduction-root helper getReductionInstr out of SLPVectorizer.cpp
into the self-contained SLPVectorizer/SLPReductionUtils.{h,cpp} module.
getAggregateSize is not reduction-specific: its only user is
findBuildAggregate and it only computes the element count of a homogeneous
aggregate from its type. Move it to the generic SLPVectorizer/SLPUtils
module instead.
Part of the SLPVectorizer.cpp modularization effort:
https://discourse.llvm.org/t/modularizing-slpvectorizer-cpp/90922
Assisted by AI
[Hexagon] Require asserts for autohvx xqf-postra-conv-double2.ll (#230348)
autohvx/xqf-postra-conv-double2.ll passes -debug-only=handle-qfp to llc,
which is only supported when LLVM is built with assertions enabled.
Without 'REQUIRES: asserts', the test fails in non-asserts builds with:
llc: Unknown command line argument '-debug-only=handle-qfp'.
Add 'REQUIRES: asserts' matching other tests in this directory.
Fixes #227779
[Passes] Port -print-pipeline-passes to PassesOptions (#230220)
-print-pipeline-passes takes an optional value (=text or =tree). Add
FlagOrEnumField, an EnumField whose bare option assigns a given
enumerator and never consumes the next argument, as cl::ValueOptional
does. BoolField and OptionalBoolField use the same BareValue.
clang, flang, LTO, llc, and opt read the format through
PassBuilder::getPrintPipelinePasses() instead of the exported cl::opt.
Aided by Opus 5.5
pcib: Only apply ARI translation to a bridge's own secondary bus
ARI changes RID interpretation only for the device on a downstream
port's secondary bus. pcib_xlate_ari() applied that translation to
every config access through an ARI-enabled bridge, including cycles
forwarded to a subordinate bus.
A non-zero slot on a subordinate bus then panics an INVARIANTS kernel
and is misrouted otherwise. Translate only when the access targets
this bridge's secondary bus.
Reviewed by: kib, jhb
Fixes: 55d3ea1731d1 ("Add support for PCIe ARI")
Sponsored by: AMD
Differential Revision: https://reviews.freebsd.org/D60033
(cherry picked from commit 10409ee40baf593b810cd0140f3c2f0d737c1a65)
ktls CBC decrypt: Avoid creating zero length iovec entries
If an mbuf's length in the chain for an encrypted TLS record exactly
matches the remaining length of header bytes to skip, skip the mbuf
entirely rather than adding a zero-length iovec entry.
Sponsored by: Chelsio Communications
(cherry picked from commit 1756d1cab3ce28085eec94a45c557e1af265ccd0)
[SLP][modularisation][NFC] Move isFirstInsertElement/getDebugLocFromPHI to SLPUtils
Move the BoUpSLP-independent helpers isFirstInsertElement and
getDebugLocFromPHI out of SLPVectorizer.cpp into the self-contained
SLPVectorizer/SLPUtils.{h,cpp} module.
Part of the SLPVectorizer.cpp modularization effort:
https://discourse.llvm.org/t/modularizing-slpvectorizer-cpp/90922
route/fib_algo: Respect immediate_sync in fd_ref_nhop
Now fd_ref_nhop() returns zero for cross family routes,
Do not schedule nhop references and try to rebuild it immediately
for connected and static routes.
PR: 298733
Fixes: 633438224304 ("route/fib_algo: Fix nexthop index ...")
(cherry picked from commit 3d030136ad1906e6d463673a8d2ef313eb054fb3)
[AMDGPU] Insert MFMA anti-hints in GCNPreRAOptimizations (#218075)
This PR adds rule-based anti-hints in the GCNPreRAOptimizations phase so
that
the register allocator prefers registers that avoid hazard nops.
Instructions
are classified into different classes and hazard windows are computed
from the
waitstate info. The rules are built to describe the hazardous situations
and
their hazard windows. Added an AntiHintEngine that uses these rules to
insert
anti-hint relations between vregs, which the register allocator later
uses to
prefer allocations that avoid the corresponding hazardous physical
registers.
## Stack
PR **4/4**. Depends on #218074. Next #226397