Merge tag 'net-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Paolo Abeni:
"Including fixes from Bluetooth, WiFi and netfilter.
We are actively retargeting several non-urgent fixes towards next,
but the traffic on the ML looks ever-increasing, and propagating the
push-back towards subsystems is not immediate.
No known outstanding regressions.
Current release - regressions:
- netfilter: nft_set_rbtree: skip transaction elements during GC
Previous releases - regressions:
- sched: cls_api: reclaim an empty proto on the error path
[59 lines not shown]
Merge tag 'sysctl-7.03-fixes-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/sysctl/sysctl
Pull sysctl fixes from Joel Granados:
"Fix sysctl jiffies conversions errors introduced in 2dc164a48e6f
("sysctl: Create converter functions with two new macros")
- Ensure that mult_hz does *not* wrap
- Ensure we pass just the magnitude for the negative branch in
proc_int_k2u_conv_kop"
* tag 'sysctl-7.03-fixes-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/sysctl/sysctl:
time/jiffies: Saturate in mult_hz() instead of wrapping
sysctl: Negate before converting in the int read path
Merge tag 'nf-26-09-30' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf
Pablo Neira Ayuso says:
====================
Netfilter/IPVS fixes for net
The following batch contains Netfilter fixes for net. This batch
fixes crashes as recent feature regression, one of the due to a
dependency that has been pulled into -stable:
1) Expand existing ipset fix for bitmap sets to disallow comments
updates from kernel-side adds, from Florian Westphal.
2) Drop flowtable reference if nf_ct_netns_get() fails, otherwise
flowtable cannot ever be removed, from Aohan Mei.
3) nft_rbtree GC should collect end elements that contained in
this transaction batch, new or deleted elements are never
[44 lines not shown]
net: phy: aquantia: fix system interface type not updated in forced mode
aqr_gen1_read_status() decodes the MDIO_PHYXS_VEND_IF_STATUS register
to determine which SerDes interface the PHY is currently using on its
system side and stores the result in phydev->interface. phylink relies
on this value to configure the MAC.
The autoneg == AUTONEG_DISABLE check is not correct:
MDIO_PHYXS_VEND_IF_STATUS is set by the PHY firmware based on the
negotiated link speed, not based on whether autoneg was used to reach
it. When the link comes up at 1G in forced mode, the register correctly
reads SGMII, but the early return prevents phydev->interface from being
updated. It stays at whatever value it held before (typically 2500BASE-X
from the initial autoneg run), so phylink configures the MAC for the
wrong interface and the link cannot come up.
Remove the autoneg guard so that the system interface type is always
decoded when the link is up.
[5 lines not shown]
Merge tag 'audit-pr-20260930' of git://git.kernel.org/pub/scm/linux/kernel/git/pcmoore/audit
Pull audit fix from Paul Moore:
"A single audit fix for a potential UAF error in some audit filter
configurations"
* tag 'audit-pr-20260930' of git://git.kernel.org/pub/scm/linux/kernel/git/pcmoore/audit:
audit: fix exe mark UAF in kill_rules()
net: mvneta: clear XDP pfmemalloc flag between frames
mvneta_swbm_add_rx_fragment() sets XDP_FLAGS_FRAGS_PF_MEMALLOC on the
xdp_buff when a fragment page is a pfmemalloc one (page under memory
pressure). The xdp_buff is reused for the next frame, but only the
XDP_FLAGS_HAS_FRAGS bit was cleared at frame start, so the pfmemalloc
bit leaked from one frame into the following ones. mvneta_swbm_build_skb()
propagates the flag to skb->pfmemalloc through xdp_update_skb_frags_info(),
so the skb of a subsequent fragmented frame could be wrongly marked as
pfmemalloc even if none of its pages are under pressure.
Clear all the xdp_buff flags in mvneta_swbm_rx_frame(), which is invoked
for each new frame, instead of just the XDP_FLAGS_HAS_FRAGS bit.
Fixes: ed7a58cb40bd ("net: marvell: rely on xdp_update_skb_shared_info utility routine")
Reviewed-by: Simon Horman <horms at kernel.org>
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi at oss.qualcomm.com>
Reviewed-by: Toke Høiland-Jørgensen <toke at redhat.com>
Link: https://patch.msgid.link/20260929-mvneta-xdp-clear-frag-fix-v4-1-1e63b25eeed8@oss.qualcomm.com
Signed-off-by: Jakub Kicinski <kuba at kernel.org>
ipv6: sr: use skb_get_hash_net() in seg6_make_flowlabel()
Since commit d58e468b1112 ("flow_dissector: implements flow dissector
BPF hook") __skb_flow_dissect() needs a net pointer, either from
skb->dev, skb->sk, or since commit 3cbf4ffba5ee ("net: plumb network
namespace into __skb_flow_dissect") a caller provided pointer.
syzbot was able to reach seg6_make_flowlabel() with an skb having
neither skb->dev nor skb->sk set: a TIPC UDP bearer sends a discovery
message through an IPv4 route using seg6 encap, while
net.ipv6.seg6_flowlabel is set to 1.
seg6_make_flowlabel() already has a net pointer, use skb_get_hash_net().
WARNING: net/core/flow_dissector.c:1131 at __skb_flow_dissect+0x910/0x5368 net/core/flow_dissector.c:1126, CPU#0: syz.0.17/4930
Call trace:
__skb_flow_dissect+0x910/0x5368 net/core/flow_dissector.c:1126 (P)
__skb_get_hash_net+0xe0/0x29c net/core/flow_dissector.c:1903
skb_get_hash include/linux/skbuff.h:1663 [inline]
[26 lines not shown]
net/mlx5e: Fix AF_XDP TX timestamp teardown NULL dereference
During XSK TX queue teardown, outstanding descriptors are completed
without a CQE. If one requested a TX timestamp, the completion path
passes that NULL CQE to mlx5e_xsk_fill_timestamp(), which dereferences
it.
Return zero when no CQE is available, indicating that teardown did not
produce a TX timestamp.
Fixes: ec706a860eba ("net/mlx5e: Implement AF_XDP TX timestamp and checksum offload")
Signed-off-by: Prathamesh Deshpande <prathameshdeshpande7 at gmail.com>
Reviewed-by: Tariq Toukan <tariqt at nvidia.com>
Link: https://patch.msgid.link/20260926165402.5902-1-prathameshdeshpande7@gmail.com
Signed-off-by: Jakub Kicinski <kuba at kernel.org>
r8169: disable EEE on RTL8168h/8111h
Force EEE off on the RTL8168h/8111h (RTL_GIGA_MAC_VER_46) at PHY connect
with phy_disable_eee().
Commit 202fef9bbbf5 ("net: phy: realtek: fix EEE advertisement write on
the internal PHY MMD path") made EEE actually advertise on the generic
Realtek PHY; the write had been a no-op before. On the RTL8168h that
un-masks a latent defect: once EEE negotiates, RX silently stalls after
~7-20 minutes. The carrier stays up, no counter or dmesg moves, and only
"ip link set down/up" recovers it; disabling EEE keeps the link stable.
Root-causing the RTL8168h LPI/RX path needs hardware not available now, so
disable EEE for this version. Use phy_disable_eee() rather than dropping
the version from rtl_supports_eee(): the latter also skips
rtl_enable_tx_lpi(), whose disable branch clears the MAC TX-LPI bits (ERI
0x1b0[1:0]) on link up. rtl_hw_start_8168h_1() does not clear them (the
RTL8402/RTL8106e init does), so a warm reboot from an EEE-active state
could otherwise leave TX-LPI asserted while the PHY no longer negotiates
[10 lines not shown]
octeontx2-pf: Fix RSS indirection table size
The conversion to the dedicated RSS context operations replaced the
pointer to struct otx2_rss_ctx with an inline u32 ind_tbl[] array.
However, otx2_rss_init() still uses sizeof(*rss->ind_tbl) to set rss_size.
This now yields the size of one u32 (4), rather than the 256 entries in
the indirection table.
The RSS initialization and hardware programming loops use rss_size,
so only four entries are initialized and programmed despite the NIX LF
being allocated a 256-entry table. On an OCTEON CN102 with eight RX
queues, the table contained 0, 1, 2, 3 followed by zeros. Traffic was
concentrated on queue 0 and ethtool -X equal 8 did not correct the table.
Use ARRAY_SIZE() to count the indirection table entries. On the same
hardware, this restores the full table repeating queue numbers 0 through
7, and bidirectional multi-flow traffic increments multiple RX queues.
Runtime testing on an Asterfusion ET2500 (OCTEON CN102 A0).
[8 lines not shown]
sctp: check RCV_SHUTDOWN after the sendmsg connect wait
sctp_wait_for_connect() drops the socket lock while it sleeps. An
out-of-the-blue ABORT can then be processed from the socket backlog and
unlink the association. If a concurrent shutdown(fd, SHUT_RD) sets
RCV_SHUTDOWN, the waiter breaks with err == 0 before checking
asoc->base.dead. Its final sctp_association_put() can then free the
association, leaving sctp_sendmsg_to_asoc() to continue with a dangling
pointer.
Check RCV_SHUTDOWN along with the wait error in sctp_sendmsg_to_asoc()
before using the association again. The check only accesses the socket,
so it needs no additional association reference. Return the existing
-ESRCH so that sctp_sendmsg() skips freeing a new association that may
already have been destroyed.
Keep sctp_wait_for_connect() unchanged to preserve its behavior for the
connect() caller.
[7 lines not shown]
net: sparx5: make ports inherit the switch base mac address type
When the switch uses a random base MAC address it does not mark each
port's address as random:
$ dmesg | grep "MAC addr"
sparx5-switch e00c0000.switch: MAC addr was not set, use random MAC
$ cat /sys/class/net/eth0/addr_assign_type
0
0 is NET_ADDR_PERM, so userspace thinks this is a stable address.
With a random MAC reported as NET_ADDR_PERM, systemd's
MACAddressPolicy=persistent [1] leaves it unchanged, so a DHCP client
gets a new IP address every boot and a static reservation can't be used.
Reporting it as NET_ADDR_RANDOM makes systemd replace it with a stable
address derived from the interface name and machine-id.
Fixes: f3cad2611a77 ("net: sparx5: add hostmode with phylink support")
[5 lines not shown]
Merge tag 'wireless-2026-09-30' of https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless
Johannes Berg says:
====================
Still more fixes coming in, notably:
- ath11k: avoid running out of stations on HW restart
- mac80211:
- drop too large fragmented MPDUs
- mesh path handling fixes
- validation improvements
- reject CSA with bad 320 MHz bandwidth
- cfg80211: fix RTS for single radio devices
* tag 'wireless-2026-09-30' of https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless: (27 commits)
wifi: mac80211: fix slab-out-of-bounds read in ieee80211_monitor_select_queue()
wifi: mac80211: reject invalid 320 MHz CSA bandwidth
wifi: mac80211: set info->band for 802.3 encap offload frames
wifi: mac80211: prevent AP VLAN tx from other interfaces
[21 lines not shown]
net: microchip: vcap: stop scanning after deleting key field
The only call outside the KUnit tests is in
sparx5_tc_add_rule_copy(), which passes a fixed key list with no
duplicate entries. Therefore, the potential double-free is not
currently reachable in tree.
However, vcap_filter_rule_keys() is exported and does not require its
key list entries to be unique. If a future or external caller supplies
the same key at another position that the inner loop visits, it calls
list_del() and kfree() a second time on the freed field, leading to a
potential double-free.
Break the inner loop after removing a matching field. The outer safe
list traversal continues to process the remaining rule fields.
Fixes: 465a38a269e9 ("net: microchip: sparx5: Support for copying and modifying rules in the API")
Reported-by: Dan Carpenter <error27 at gmail.com>
Closes: https://lore.kernel.org/all/ahs-lCAXXikkaHky@stanley.mountain/
[3 lines not shown]
netfilter: nft_flow_offload: drop flowtable reference on init error path
nft_flow_offload_init() bumps the flowtable use count with
nft_use_inc() before calling nf_ct_netns_get(). When the latter
fails, the error is returned as-is and the reference is leaked.
The upper layers do not balance it either: nf_tables_newexpr()
clears expr->ops when the expression init callback fails, so the
nft_expr_more() iteration in nft_rule_expr_deactivate() and
nf_tables_rule_destroy() stops right before the failed expression
and its ->destroy callback, which would drop the reference, never
runs.
Each failed rule addition therefore leaks one flowtable reference
and the flowtable can no longer be removed: NFT_MSG_DELFLOWTABLE
keeps reporting -EBUSY even though no rule references it.
Save the nf_ct_netns_get() return value and undo the nft_use_inc()
when it fails, restoring the inc/dec pairing within
[8 lines not shown]
netfilter: bpf: reject invalid NAT manipulation types
As bpf_ct_set_nat_info() is not validating the NAT manipulation type a
wrong value can be passed directly to nf_nat_setup_info(). This triggers
the WARN_ON() at nf_nat_setup_info() and if panic_on_warn isn't set,
then IPS_SRC_NAT_DONE is set without adding nat_bysource and conntrack
cleanup tries to unlink an uninitialized hlist node.
Fix this by checking that NAT manipulation type is correct before
calling nf_nat_setup_info(). In addition, if the WARN_ON is hit, return
NF_DROP instead of continuing with the processing to avoid similar
situations in the future.
Reported-by: VEGA <vega at nebusec.ai>
Fixes: 0fabd2aa199f ("net: netfilter: add bpf_ct_set_nat_info kfunc helper")
Signed-off-by: Fernando Fernandez Mancera <fmancera at suse.de>
Signed-off-by: Pablo Neira Ayuso <pablo at netfilter.org>
netfilter: flowtable: generalize pending status bit
Rename NF_FLOW_HW_PENDING to NF_FLOW_PENDING and use it to inhibit the
flowtable GC worker until pending hw offload work has been completed.
Apparently, nf_flow_offload_stats() can schedule work to retrieve stats
while the flow is being removed by GC.
And this bit can also be used in a follow up patch to disable GC until
the flow has been fully added in both directions.
Revert the reordering done in commit d644b23afe1e ("netfilter:
flowtable: publish HW_DEAD after worker is done") to prevent a race
between GC and hw offload handler.
Fixes: 2c8897953f3b ("netfilter: flowtable: Add pending bit for offload work")
Signed-off-by: Pablo Neira Ayuso <pablo at netfilter.org>
netfilter: flowtable: restore ieee80211 forward path
Before commit 871df5007eda ("netfilter: flowtable: bail out if forward
path cannot be discovered"), there was a fallback to set up a forward
path in case .ndo_fill_forward_path fails or DEV_PATH_MTK_WDMA was used.
Such fallback was used by commit d787a3e38f01 ("mac80211: add support
for .ndo_fill_forward_path").
One possibility is to handle DEV_PATH_MTK_WDMA from the flowtable
forward path discovery. However, this is only used internally by drivers
to retrieve mtk_wdma information to set up hardware offload. Felix
decided to use the .fill_forward_path interface for this purpose due to
the lack of a better interface at that time.
Add a new DEV_PATH_IEEE80211 path which is offered if the new ieee80211
flag is set on in the struct net_device_path_ctx to restore the
flowtable with a ieee80211 netdevice. Handle this new DEV_PATH_IEEE80211
path just like DEV_PATH_ETHERNET and DEV_PATH_DSA, ie. this is the last
netdevice in the stack.
[6 lines not shown]
netfilter: nft_set_rbtree: skip transaction elements during GC
Since nft_set_commit_update() runs set commit callbacks before processing
NEWSETELEM transactions, nft_rbtree_gc_scan() can observe elements added by
the transaction being committed.
The scan records an interval end in rbe_end without checking the element's
transaction state. A later, unrelated expired start then moves both
elements to the expired list. The synchronous GC queue can free the new end
element before the transaction subsequently activates it, causing a
use-after-free.
Only consider elements that are fully active in both generations. This
keeps transaction-state elements out of the GC scan and preserves interval
pairing across skipped elements.
KASAN reports:
BUG: KASAN: slab-use-after-free in nft_setelem_activate
[16 lines not shown]
ipvs: fix missing counter decrement in lblc
LBLC may delete cache entries for destinations that are
removed or overloaded and replace them with available ones.
But ip_vs_lblc_new() forgets to decrement the tbl->entries
counter after calling ip_vs_lblc_del(). This can lead to
increased shrinking of the cache with every new garbage
collection.
Fixes: 2f3d771a35fe ("ipvs: do not use dest after ip_vs_dest_put in LBLC")
Link: https://sashiko.dev/#/patchset/0bdd5abe9968ded7ca2b9cb6844ba83d94cc8d53.1787318053.git.zhilinz%40nebusec.ai
Signed-off-by: Julian Anastasov <ja at ssi.bg>
Signed-off-by: Pablo Neira Ayuso <pablo at netfilter.org>
ipvs: filter some flags received in the backup server
While the IPVS SYNC protocol is not secure by design
we can still protect the backup server from messages that
can wreak havoc.
This commit addresses problems from received connection flags
or their combinations. We now drop messages as follows:
1. the NO_CPORT+TEMPLATE combination allows lookups for normal
connections to hit template which can break in many ways.
While the master does not sync connections with NO_CPORT flag,
i.e. before they are established, we still accept NO_CPORT
without TEMPLATE.
2. ONE_PACKET: it is not sent by master, so we do not
expect it in backup. Before now it was ignored by
IP_VS_CONN_F_BACKUP_MASK for protocol v1 while protocol
v0 created connections that are not hashed and dropped
[6 lines not shown]
ipvs: do not create invisible templates
The IP_VS_CONN_F_ONE_PACKET flag was implemented for normal
connections. When conn template inherits this flag from
dest->conn_flags it will not be hashed. As result, we will
create new template for every new normal connection.
Fix it to allow one template to be used by many normal
connections.
Fixes: 26ec037f9841 ("IPVS: one-packet scheduling")
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260916231652.127456-1-pablo%40netfilter.org
Signed-off-by: Julian Anastasov <ja at ssi.bg>
Signed-off-by: Pablo Neira Ayuso <pablo at netfilter.org>
ipvs: bound LBLCR and LBLC cache growth
ip_vs_lblcr_new() and ip_vs_lblc_new() create cache entries for
every previously unseen destination address. The table max_size only
tells the periodic collector to reclaim entries after the cache has
already exceeded the limit. It does not reclaim entries that the
attacker continues to use.
Reject new cache entries once either table reaches max_size * 3 / 2.
The extra headroom lets the periodic collector catch up while the
existing scheduler fallback continues to use the selected destination
when cache creation fails. New traffic therefore stays serviceable
without growing the tables further.
Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
Cc: stable at vger.kernel.org
Reported-by: Vega <vega at nebusec.ai>
Suggested-by: Julian Anastasov <ja at ssi.bg>
Signed-off-by: Zhiling Zou <zhilinz at nebusec.ai>
[2 lines not shown]
Merge branch 'tcp-old-acks'
tcp: correct timestamp echo for accepted old ACKs
Linux can acknowledge newly received data while echoing an outdated
TCP timestamp. This happens when a reordered packet fills a receive gap
but carries an older acknowledgment for traffic in the other direction.
If the sender uses this echo to measure round-trip time after a long idle
period, the stale timestamp can inflate its estimate and slow its sending.
Changes in v2, following Eric Dumazet's review:
- Remove the redundant SYN_RECV condition and unused flag accumulation.
- Replace the log-parsing wrapper with a plain packetdrill test that checks
the timestamp echo directly, using the existing selftest runner.
- Rebase on net fc6d80eb5044.
v1: https://lore.kernel.org/netdev/20260921222609.50824-4-jeffjo@openai.com/
In this example, S sends the reordered data and R is the Linux receiver
[76 lines not shown]
selftests: net: check timestamp echo after an old ACK
Add a regression test for a gap-filling packet whose acknowledgment has
become old. Linux first sends data to the peer. Deliver two peer data
packets out of order: the later packet acknowledges Linux's data, while
the delayed packet still carries the earlier acknowledgment.
Require the ACK that closes the receive gap to echo the delayed packet's
timestamp, 301000. Without the fix, Linux accepts the data but still echoes
the previously saved timestamp, 1000. Also check that the application can
read all 34 bytes.
Use a large jump in peer timestamps to represent the idle interval, with
no real wait, loss or retransmission. The test directly checks the outgoing
timestamp echo. It requires packetdrill's merged TSecr verification fix
(linked below); older tools incorrectly pass on an unfixed kernel.
Use the existing packetdrill selftest runner for IPv4, IPv6 and
IPv4-mapped IPv6.
[6 lines not shown]
tcp: refresh TS.Recent for accepted old ACKs
A TCP packet can carry new data while acknowledging traffic in the
opposite direction. With overlapping traffic in both directions, a
delayed packet's acknowledgment can be older than one Linux has already
accepted, even when that packet fills a gap in the received data.
Linux accepts the data, but tcp_ack() takes the old_ack path and skips
updating TS.Recent, the timestamp saved for outgoing acknowledgments.
The reply therefore echoes an older timestamp. If the sender uses this
echo to measure round-trip time after a long idle period, its estimate
includes the idle time and can reduce its sending rate.
Update TS.Recent in old_ack using tcp_replace_ts_recent(), before SACK
processing can trigger a transmission. This reuses the existing timestamp
and sequence checks, including PAWS protection against old duplicate
packets. ACK validation already rejects old ACKs in SYN_RECV before this
path, so no additional state check is needed.
[11 lines not shown]
Merge tag 'rtc-7.3-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/abelloni/linux
Pull RTC fixes from Alexandre Belloni:
"Mostly small issues found using AI. The efi change is to avoid a
regression on some platforms
Subsystem:
- fix a possible information leak
Drivers:
- efi: restore alarm support with runtime capability probe"
* tag 'rtc-7.3-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/abelloni/linux:
rtc: spear: initialize IRQ state before requesting alarm IRQ
rtc: mpfs: fix unchecked devm_clk_get() error pointer in probe()
rtc: ac100: Fix clock provider use-after-free on probe failure
rtc: ac100: Assign .num before accessing .hws
rtc: efi: restore alarm support with runtime capability probe
rtc: dev: zero-initialize struct rtc_wkalrm to prevent information leak
net: sparx5: skip ptp deinit if init was skipped
sparx5_ptp_init() returns early on the base lan969x variants
because they do not have the SPX5_FEATURE_PTP flag. It also returns early
when no "ptp" interrupt is described. Unbinding the driver then causes a
NULL pointer dereference when cleaning up uninitialized tx_skbs queues:
Unable to handle kernel NULL pointer dereference at virtual address
0000000000000008
Call trace:
skb_queue_purge_reason+0x68/0x120 (P)
sparx5_ptp_deinit+0x64/0xc0
mchp_sparx5_remove+0x38/0x70
platform_remove+0x20/0x30
device_remove+0x4c/0x80
Before per-port tx_skbs were added, a NULL dereference would have happened
when calling ptp_clock_unregister() on the never-registered PTP clocks.
[7 lines not shown]
net: phy: qcom: at803x: Fix IPQ5018 short-cable DAC values
When "qcom,dac-preset-short-cable" is set, ipq5018_config_init()
programs the MDAC (MMD1 0x8100) and EDAC (debug 0x4380) fields. Both
fields occupy bits 15:8 (IPQ5018_PHY_DAC_MASK), but the value 0x10 is
passed unshifted as the set argument of phy_modify_mmd() and
at803x_debug_reg_mask(). Neither helper shifts or masks that argument,
so both fields are cleared to 0x00 instead of being set to 0x10, and
bit 4 of the low byte, which is outside the field, is set.
Use FIELD_PREP() to place the value in the field. This matches the
vendor SDK, which clears bits 15:8 and ORs in the value shifted left
by 8.
On a Redmi AX5400 board, where the IPQ5018 internal PHY connects to a
QCA8337 switch PHY without a cable, MDAC and EDAC read 0x6868 and
0x7800 before the write. With this change they read back 0x1068 and
0x1000, with the low byte preserved. Without it, the same writes would
leave 0x0078 and 0x0010.
[9 lines not shown]