[SLP] Pass along stride information when costing strided loads and stores (#227823)
In cases of constant strides, can give a more accurate cost when
accounting for a known stride.
Currently, only RISCV backend makes use of this information.
Assisted By: Codex
[TargetLowering][X86] Prefer 'r' over 'm' for foldable "rm" inline asm operands
An "rm" (register-or-memory) inline asm operand has always resolved to
'm', because getConstraintPreferences() picks the most general
constraint present, and 'm' is more general than 'r'. That's safe, since
memory can't run out, but it forces a value that could stay in a
register through a stack slot even when there's no register pressure
(https://github.com/llvm/llvm-project/issues/20571).
Prefer 'r' instead where the register allocator can fold the register
back to a stack slot when it runs out of registers, and mark the
register operand foldable (InlineAsm::Flag::setRegMayBeFolded()) so it
does. Both allocators can: the greedy allocator folds an operand when it
spills its value, and the fast allocator folds operands up front when
the asm's register operands wouldn't fit.
ParseConstraints() sets AsmOperandInfo::MayFoldRegister for an operand
whose constraint codes are exactly {r, m}, above -O0, on a target that
opts in through the new supportsRegMemInlineAsmFolding() hook, which
[20 lines not shown]
[RegAllocFast] Fold foldable inline asm operands under register pressure
An inline asm register operand marked foldable (from an "rm" constraint)
may be replaced with a stack slot when the register allocator runs out
of registers. The greedy allocator does that when it spills the value.
The fast allocator can't: it assigns the operands of an instruction one
at a time, and folding replaces the instruction. So it reported
"inline assembly requires more registers than available" instead.
Before allocating a block, estimate from each inline asm's own operands
whether they fit in registers, and fold as many foldable registers as
needed to make them fit, so the common case without pressure still gets
a register. Values that only live across the asm don't count, since the
allocator spills them when it needs their registers. The estimate
follows allocateInstruction():
- First all defs get distinct registers, avoiding physreg defs; then all
uses, along with the defs still occupied while the uses are read
(early-clobber and tied defs, see isLiveThroughDef()), get distinct
[26 lines not shown]
[InstCombine] Fix profiles in sinkNotIntoOtherHandOfLogicalOp (#229555)
We sometimes create a new select with an inverted condition of the
original. In that case just copy and swap the weights.
[TargetLowering][X86] Prefer 'r' over 'm' for foldable "rm" inline asm operands
An "rm" (register-or-memory) inline asm operand has always resolved to
'm', because getConstraintPreferences() picks the most general
constraint present, and 'm' is more general than 'r'. That's safe, since
memory can't run out, but it forces a value that could stay in a
register through a stack slot even when there's no register pressure
(https://github.com/llvm/llvm-project/issues/20571).
Prefer 'r' instead where the register allocator can fold the register
back to a stack slot when it runs out of registers, and mark the
register operand foldable (InlineAsm::Flag::setRegMayBeFolded()) so it
does. Both allocators can: the greedy allocator folds an operand when it
spills its value, and the fast allocator folds operands up front when
the asm's register operands wouldn't fit.
ParseConstraints() sets AsmOperandInfo::MayFoldRegister for an operand
whose constraint codes are exactly {r, m}, above -O0, on a target that
opts in through the new supportsRegMemInlineAsmFolding() hook, which
[20 lines not shown]
[RegAllocFast] Fold foldable inline asm operands under register pressure
An inline asm register operand marked foldable (from an "rm" constraint)
may be replaced with a stack slot when the register allocator runs out
of registers. The greedy allocator does that when it spills the value.
The fast allocator can't: it assigns the operands of an instruction one
at a time, and folding replaces the instruction. So it reported
"inline assembly requires more registers than available" instead.
Before allocating a block, estimate from each inline asm's own operands
whether they fit in registers, and fold as many foldable registers as
needed to make them fit, so the common case without pressure still gets
a register. Values that only live across the asm don't count, since the
allocator spills them when it needs their registers. The estimate
follows allocateInstruction():
- First all defs get distinct registers, avoiding physreg defs; then all
uses, along with the defs still occupied while the uses are read
(early-clobber and tied defs, see isLiveThroughDef()), get distinct
[25 lines not shown]
[CodeGen] Report an error for a direct inline asm output in memory
An inline asm output returned by value has no memory to write to, yet a
constraint such as "=rm" picks memory, the most general constraint, as
does "=m". SelectionDAG asserted on that ("Can only indirectify direct
input operands!"), and GlobalISel dereferenced a null pointer. Clang
never emits such an output, since it passes the address of a memory
output, but other IR can. Report "cannot handle direct memory outputs
yet for constraint 'm'" instead, like the other inline asm errors there.
Assisted-by: Claude Opus 5.5
[SimplifyLibCalls] Fix profile propagation in replacePowWithSqrt (#229645)
In some cases we create a select conditioned on whether the input value
is equal to -infinity. We assume that this case is rare and thus mark
the true arm of the select unlikely.
[MachineScheduler] Add support for scheduling while in SSA (#161054)
Allow targets to add an MachineScheduler before PHI elimination, i.e. in
SSA mode.
Add initial support in AMDGPU backend for using SSA Machine Scheduler
instead of normal Machine Scheduler.
(This behaviour is disabled by default.)
Also add basic "kick the tyres" demonstrator tests.
This change is intended to support the introduction of a pre-RA spilling
pass which runs prior to PHI elimination. Machine scheduler has a
significant impact on register pressure and as such is best run before
this new spilling pass.
Co-authored-by: Konstantina Mitropoulou <KonstantinaMitropoulou at amd.com>
[RISCV] Add a RISCVISD::TUPLE_CAST to represent casting one tuple to another. (#228249)
We allow converting between fractional LMUL and LMUL=1, but the number
of fields must be the same.
Previously we used a TUPLE_INSERT, but TUPLE_INSERT's SDTypeProfile says
the input should be a vector not a tuple.
Since the register classes must match, I'm lowering this without a copy.
Found while trying to verify SDTypeIsVec.
Assisted-by: Claude
PR kern/60601 - Avoid 32 bit wraparound on ILP32 hosts
In tmpfs_bytes_max calculate avail_mem using 64 bit calculations,
rather than one small piece being 32 bit only (and subject to
simple overflow, entirely within reasonable values).
Reported and diagnosed by Hashimoto Kenichi in PR kern/60601
XXX - pullup -11 -10 (in a week or two).
routing: fix rtentry use-after-free in multipath route append
add_route_flags() drops the RIB lock and passes the existing entry,
rt_orig, to add_route_flags_mpath(). If a concurrent delete removes
the prefix in that window, the retry re-inserts rt_orig, which is
then freed while still linked, crashing later in rn_match().
Pass the new rt instead, return ENOENT when the prefix is gone and
RTM_F_CREATE is not set, and fix the rnd_orig NULL check.
Approved by: pouria
Fixes: c24a8f19c5d5 ("routing: fix rib_add_route_px()")
Differential Revision: https://reviews.freebsd.org/D60353
[KnownFPClass] Correct subnormal handling in ldexp (#225904)
Uses the `applyInputDenormalMode` and `applyOutputDenormalMode` helpers
to fix subnormal handling for `KnownFPClass::ldexp`.
I also skipped the "never-subnormal" deduction for `ppcf128`.
Similar to https://github.com/llvm/llvm-project/pull/215094, but
only focusing on fixing incorrect deductions.
tests/netinet6: fix ndp_del_gu_success flakiness
The test pinged an unanswered address and then deleted the resulting
INCOMPLETE neighbor entry.
The kernel frees that entry after about 3s, so on a loaded VM,
ndp -d could run too late and fail with ENOENT.
Configure 2001:db8::2 on epair0b so the ping gets a reply and the
entry becomes REACHABLE.
Approved by: pouria
Sponsored by: Netflix
Differential Revision: https://reviews.freebsd.org/D60348
devel/redasm: prepare the port for modern versions of CMake
Synchronize cmake_minimum_required(VERSION 3.10) across all
components and switch to new CMP0048 policy.
[InstCombine] Preserve profile metadata when folding fneg selects (#228754)
Reuse select metadata in the fneg transformations when profile metadata
fixes are enabled.
Remove obsolete profcheck XFAIL entries for the now-passing InstCombine
tests.
---------
Co-authored-by: Aiden Grossman <aidengrossman at google.com>
[SimplifyLibCalls] Add initial support for non-8-bit bytes
The patch makes CharWidth argument of `getStringLength` mandatory
and ensures the correct values are passed in most cases.
This is *not* a complete support for unusual byte widths in
SimplifyLibCalls since `getConstantStringInfo` returns false for those.
The code guarded by `getConstantStringInfo` returning true is unchanged
because the changes are currently not testable.
[ValueTracking] Make isBytewiseValue byte width agnostic
This is a simple change to show how easy it can be to support unusual
byte widths in the middle end.
[DataLayout] Add byte specification
This patch adds byte specification to data layout string.
The specification is `b:<size>`, where `<size>` is the size of a byte
in bits (later referred to as "byte width").
Limitations:
* The only values allowed for byte width are 8, 16, and 32.
16-bit bytes are popular, and my downstream target has 32-bit bytes.
These are the widths I'm going to add tests for in follow-up patches,
so this restriction only exists because other widths are untested.
* It is assumed that bytes are the same in all address spaces.
Supporting different byte widths in different address spaces would
require adding an address space argument to all DataLayout methods
that query ABI / preferred alignments because they return *byte*
alignments, and those will be different for different address spaces.
This is too much effort, but it can be done in the future if the need
arises, the specification reserves address space number before ':'.
[ValueTracking] Add CharWidth argument to getConstantStringInfo (NFC)
The method assumes that host chars and target chars have the same width.
Add a CharWidth argument so that it can bail out if the requested char
width differs from the host char width.
Alternatively, the check could be done at call sites, but this is more
error-prone.
In the future, this method will be replaced with a different one that
allows host/target chars to have different widths. The prototype will
be the same except that StringRef is replaced with something that is
byte width agnostic. Adding CharWidth argument now reduces the future
diff.
[IRBuilder] Add getByteTy and use it in CreatePtrAdd
The change requires DataLayout instance to be available, which, in turn,
requires insertion point to be set. In-tree tests detected only one case
when the function was called without setting an insertion point, it was
changed to create a constant expression directly.
[IR] Account for byte width in m_PtrAdd
The method has few uses yet, so just pass DL argument to it. The change
follows m_PtrToIntSameSize, and I don't see a better way of delivering
the byte width to the method.
[IR] Make @llvm.memset prototype byte width dependent
This patch changes the type of the value argument of @llvm.memset and
similar intrinsics from i8 to iN, where N is the byte width specified
in data layout string.
Note that the argument still has fixed type (not overloaded), but type
checker will complain if the type does not match the byte width.
Ideally, the type of the argument would be dependent on the address
space of the pointer argument. It is easy to do this (and I did it
downstream as a PoC), but since data layout string doesn't currently
allow different byte widths for different address spaces, I refrained
from doing it now.
[flang][NFC] Add missing license and fix short first line (#229647)
- Fix shorter first line than 80 characters
- Add missing license at the top of some files
Assisted-by: AI
[IndirectBrExpand] Preserve profile weights (#227784)
We can derive the weights for the created select instruction from the
weights of the indirectbr instructions that are used to compose it.
This is a no-op most of the time outside of some contexts, but
production users that utilize PGO do enable this (e.g., the kernel).