[SLP] Partially revert #224931 for alternate nodes
Do not pass scalar context to alternate node vector cost queries.
PR #224931 passes the scalar main/alternate op as the context instruction when
costing the vector ops of an alternate node. Those vector ops only feed the
lane-select shuffle, so use-based discounts derived from the scalar's users
do not apply. On AMDGPU this priced an [fmul | fadd] node's vector fmul as
fused (free) where the sole user of the scalar fmul is fadd/fsub, keeping
unprofitable subtrees vectorized. In a rocFFT kernel on gfx950 this raised
VGPR usage from 160 to 178 and caused a performance regression. Restore the
null context for these queries.
X86: Stop setting kill flags on virtual registers before FinalizeISel
There is no point in maintaining kill flags before register allocation
anymore.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[Clang][C23] Fix typedef-name after 'auto' as storage-class use (#224019)
In` C23`, `auto <typedef-name> <var>;` was misparsed as auto
type-inference on the `typedef`, then errored on the missing
initializer. Pre-C23 clang accepts it as a declaration of `<var>` with
the typedef's type (`auto` used as storage-class with an explicit
type-name).
```
typedef int x;
{ auto x y; } // valid: declares `y` of type `x`
```
Fix for https://github.com/llvm/llvm-project/issues/164930.
RISCV: Use hasOneNonDBGUse instead of kill flag for cascaded select (#228940)
The cascaded select lowering requires the first select's result to only be used by
the second select. Avoid depending on the kill flag which will not be set in the
future.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
NAS-144211 / 27.0.0 / Fix failover firewall rule for VIPs (by anodos325) (#19948)
The original work on failover firewall had a drop_all rule matched the
VIPs with saddr instead of daddr, so client traffic to a VIP was never
dropped. Match tcp/udp daddr for rule (allowing others to reduce blast
radius of change).
In CORE pf dropped inbound tcp/udp to the VIPs (with the exception of
ssh and web UI) to prevent client traffic before the controller was
ready to serve, so this is basically restoring CORE HA failover behavior
for services.
Original PR: https://github.com/truenas/middleware/pull/19940
Co-authored-by: Andrew Walker <andrew.walker at truenas.com>
SystemZ: Don't convert AND to RISBG if CC is live
The RISBG-type replacements either don't define CC or set it with
different semantics, so the conversion dropped or clobbered a CC value
that was still used.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[AMDGPU][SROA] Expand cast chain handling to floating point types
CreateBitPreservingCastChain does not produce inttoptr or ptrtoint for floating-point types. When promoting structs like { float, float } to <2 x float>, this can lead to pointers being bitcast directory to <2 x float> which is invalid. This change expands the use of the intermediate to these cases
[RISCV] Use vunzip instead of shift and truncate in the presence of Zvzip (#228298)
This patch simply hoists the vunzip generation logics inside
lowerVECTOR_SHUFFLE before the one that lowers deinterleave2 into VNSRL
(i.e. shift and truncate). The rationale behind this is that
extension-specific logics should generally happen before the generic
(RVV) cases. Plus, both of the options emit a single instruction and
it's unlikely vunzip will be slower than VNSRL.
I found this when I'm writing a DAG combine for Zvzip, so this patch
could potentially enable more Zvzip patterns in the future as well.
[amdgpu] Add HasBufferInv AMDGPUSubtargetFeature (#226254)
As suggested by @arsenm I've created a new `AMDGPUSubtargetFeature` to
test whether `buffer_inv` is available, as opposed to just using
`hasGFX940Insts` to gate whether or not it's available.
The aforementioned suggestion was in response to
https://github.com/llvm/llvm-project/pull/221115. I plan to use this
flag in that PR.
In addition to introducing this new feature, I use it in place of
`hasGFX940Insts` in `SIMemoryLegalizer` where it clearly is checking
whether it's safe to use `buffer_inv`. I didn't find any other obvious
place in the codebase which uses `hasGFX940Insts` as a predicate for
`buffer_inv`.
This was co-authored with codex. Every line was audited by me, though.
iflib: Make the deferral test in iflib_txd_db_check() an early return
Invert the test so the doorbell write is no longer nested inside the
conditional, and wrap its comments to 80 columns. Fix a typo in one of
them.
No functional change intended.
Reviewed by: kbowling
Sponsored by: Rubicon Communications, LLC ("Netgate")
Differential Revision: https://reviews.freebsd.org/D60371
[flang-rt][test] Modify requirements for Driver/safe-trampoline-gnustack.f90 (#227851)
Replace multiple UNSUPPORTED lines with a single REQUIRES because the
-fsafe-trampoline option is only supported on x86-64 and AArch64
targets.
This is known to fix an issue on riscv64-linux-gnu.
SystemZ: Remove stale CC live range when converting AND to RISBG (#229062)
convertToThreeAddress replaces an AND immediate with a dead CC def with
RISBMux neither of which defines CC. If the liveness was precomputed for the
physreg, the dead def was still incorrectly tracked in the LiveInterval.
Fixes #229045
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
Fix failover firewall to match VIPs as destination
drop_all matched the VIPs with saddr, so client traffic to a VIP was
never dropped. Match daddr, and only tcp/udp so that ICMP and IPv6
neighbor discovery keep working. This is the CORE behavior: pf dropped
inbound tcp/udp to the VIPs, ssh and web UI excepted, to hold clients
off a controller that was not ready to serve.
(cherry picked from commit 1fc82d23062afeba27617bc4b44ca130a0de9a06)
NAS-144211 / 28.0.0-BETA.1 / Fix failover firewall rule for VIPs (#19940)
The original work on failover firewall had a drop_all rule matched the
VIPs with saddr instead of daddr, so client traffic to a VIP was never
dropped. Match tcp/udp daddr for rule (allowing others to reduce blast
radius of change).
In CORE pf dropped inbound tcp/udp to the VIPs (with the exception of
ssh and web UI) to prevent client traffic before the controller was
ready to serve, so this is basically restoring CORE HA failover behavior
for services.
[SPIRV] Respect source alignment when lowering memory copies (#228859)
Lowering memcpy, memcpy.inline, and memmove to OpCopyMemory or
OpCopyMemorySized only considers the destination alignment. A single
Aligned memory operand applies to both pointers, so a destination
alignment greater than the source alignment overstates the latter.
For SPIR-V 1.4 and later, emit separate destination and source
alignments
when they differ. For earlier versions, use their minimum in the single
permitted memory operand mask. Keep one mask when the alignments match.
Consolidate both copy selectors in selectCopyMemory so they share the
alignment handling.
expat: update to 2.9.0.
Release 2.9.0 Mon October 5 2026
Security fixes:
#1392 CVE-2026-102633 -- Integer overflow in function
expat_realloc on 32bit platforms
#1393 CVE-2026-77214 -- Validate parameter `len` against available
buffer capacity in XML_ParseBuffer
Bug fixes:
#1387 lib: Handle OOM when copying encodingName in XML_ParserReset
New features:
#1327 lib: Introduce new "Properties API" to get and set
scalar properties for a single parser instance.
There are six new functions:
- XML_GetPropertyBool
- XML_GetPropertyDouble
- XML_GetPropertyUInt64
[50 lines not shown]
[Dexter] Add timeout and finishing guard for startup (#228636)
A process can be terminated without start. In that case, the DAP.py will
spin over there and causes test loops forever. Add a timeout and finish
guard in here to at least stop the test.
This is found when a test fails to start on FreeBSD.
[lldb][NativePDB] Specify `/machine` for TLS test (#229175)
Intends to fix the failure from #228220.
On ARM this tried to link the arm64 `msvcrtd.lib`, but it should've used
the x64 one.
[AMDGPU][NFC] Clean up packed convert tests
Intrinsic codegen tests need neither intrinsic declarations nor the
amdgpu_kernel / amdgpu_ps calling conventions. Drop them. Kernel tests
become plain functions returning the result, with inreg args for SGPR
operands instead of kernarg loads.
amdgpu_ps stays where <32 x float> arguments exceed the default calling
convention's argument registers.
Change-Id: I5fc578bd34499da6399bb4551936e3cf171a9454
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[msan] Handle NEON FP8 FDOT{2,4} lane intrinsic (#228600)
This is follow-up work to
https://github.com/llvm/llvm-project/pull/227850, which handled the
non-lane intrinsic using the generic dot-product handler.
This patch adds a separate handler for the lane variant of the
intrinsics. We do not reuse the generic dot-product handler because the
features and optimizations are largely non-overlapping (e.g., odd/even
lanes vs. numbered lanes, ZeroPurifies, EltSizeInBits).
[SLP][AMDGPU][NFC] Precommit test for alternate node fmul cost
The vector fmul of an alternate [fmul | fadd] node only feeds the
lane-select shuffle and cannot fuse with the fsub that uses the scalar
fmul, but it is currently priced as fused, so this is vectorized.