vm_swapout: Restore handling of RLIMIT_RSS
In commit 13a1129d700c, I removed the mechanism by which the pagedaemon
signals the swapout thread when page reclamation is unable to keep up
with demand. This is because the swapout thread's main action in this
case is to swap out sleeping processes, but we removed this support.
However, it had the secondary effect of causing the swapout thread to
enforce RLIMIT_RSS when racct is not enabled. Without it, if
racct_enabled is false, nothing ever kicks the swapout thread.
Restore the old behaviour of trying to enforce RLIMIT_RSS when the page
daemon is unable to keep up with demand. I'm not at all convinced this
is a good way to implement the limit, but the change wasn't intentional,
so let's restore it for now.
Fixes: 13a1129d700c ("vm: Remove kernel stack swapping support, part 1")
Reviewed by: olce, kib
MFC after: 2 weeks
Differential Revision: https://reviews.freebsd.org/D56142
pf: carry pfra_fback in the netlink pfr_addr encoding
The netlink encoding of struct pfr_addr omits pfra_fback, so no
per-address feedback reaches userspace. In particular,
pfr_get_astats() marks entries of tables without counters with
PFR_FB_NOCOUNT, which pfctl uses to skip their counters, so
"pfctl -v -T show" prints zero counters for such tables.
Add PFR_A_FBACK, emit it from the kernel and decode it in libpfctl.
Add a regression test.
Reviewed by: kp
Approved by: kp (mentor)
Fixes: 08f54dfca197 ("pf: convert DIOCRGETASTATS to netlink")
Sponsored by: Rubicon Communications, LLC ("Netgate")
loader.efi: Retain standalone network configuration
DHCP initmd discovery used net_configure() to open and configure SNP,
then immediately closed the socket and released the protocol.
Selecting net0 as currdev subsequently required another exclusive
SNP attach, which may hang on some EDK II firmware.
Keep the configured socket available for ordinary network consumers
so NFS boot can reuse it. Add an explicit deconfiguration operation
and invoke it only when an initmd URL requires handing SNP back
to the firmware network stack.
Signed-off-by: Krzysztof Galazka <krzysztof.galazka at intel.com>
Reviewed by: imp
Assisted by: Github Copilot (GPT-5.6 Sol)
Sponsored by: Intel Corporation
Differential Revision: https://reviews.freebsd.org/D59999
loader.efi: Apply command-line DHCP overrides earlier
Ability to override DHCP options with command-line arguments
was affected by intoduction of initmd support. Initmd discovery
configures the network before loader arguments
were parsed, so a dhcp.root-path override was unavailable
during the first network configuration. Move parsing
arguments earlier and apply dhcp.root-path even if DHCP response
does not contain option 17. This allows providing a dynamic NFS
root e.g. by chain loading loader.efi from iPXE.
Signed-off-by: Krzysztof Galazka <krzysztof.galazka at intel.com>
Reviewed by: imp
Assisted by: Github Copilot (GPT-5.6 Sol)
Sponsored by: Intel Corporation
Differential Revision: https://reviews.freebsd.org/D59998
igmp: Refresh the ip header pointer after m_pullup()
There is a chance that the m_pullup() call immediately above invalidated
the "ip" pointer, so refresh it as we do with the IGMP header.
Otherwise the test in igmp_input_v3_query() for whether the packet is an
IGMPv3 general query might use a pointer to a freed mbuf.
MFC after: 1 week
Sponsored by: The FreeBSD Foundation
pf: deregister the ifnet_rename_event handler on unload
pfi_initialize() registers pfi_rename_ifnet_event() on
ifnet_rename_event, but pfi_cleanup() never deregisters it. After
"kldunload pf", the next interface rename calls through a stale
pointer into the unloaded module:
kldload pf
kldunload pf
ifconfig epair create
ifconfig epair0a name foo
Reviewed by: kp
Approved by: kp (mentor)
Fixes: 349fcf079ca3 ("net: add ifnet_rename_event EVENTHANDLER(9) for interface renaming")
Sponsored by: Rubicon Communications, LLC ("Netgate")
imgact_aout: Widen the overflow checks in exec_aout_imgact()
I suspect this overflow check isn't needed at all, at least today, since
vm_map_insert() will detect wraparound when it creates segments (and
even a check against UINT_MAX is too loose, since the max user address
is AOUT32_USRSTACK == 0xbfc00000). But if we're going to keep this
check, there doesn't seem to be any downside to applying it on all
platforms, before exec_new_vmspace() tears down the current vmspace.
Reported by: Muhammed Sariyildiz <asiyee994 at gmail.com>
Reviewed by: kib
MFC after: 2 weeks
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D60027
net: Fix handling of sockaddrs in the SIOC{ADD,DEL}MULTI handlers
The SIOCADDMULTI and SIOCDELMULTI handlers add or delete a link-layer
multicast address from an interface's multicast filter list. The
link-layer address is passed using the ifr_addr field of the request
structure.
struct ifreq's ifr_addr field is a struct sockaddr, which is a fair bit
smaller than struct sockaddr_dl (though big enough to hold an ethernet
address). Existing callers set the sockaddr length to
sizeof(struct sockaddr_dl), which is too large, and causes OOB accesses
when if_findmulti() is used to compare the address with others, or when
if_addmulti() makes a copy.
Fix this without breaking compatibility: copy the user-supplied address
into a sockaddr_dl on the stack, and use the latter for the respective
operation.
Also validate the sockaddr_dl internal length fields, suggested by zlei.
[7 lines not shown]
jail_{get,set}: user_ implementations
This will enable compat implementations in the future.
Reviewed by: jamie, jhb
Effort: CHERI upstreaming
Sponsored by: Innovate UK
Differential Revision: https://reviews.freebsd.org/D60026
sys/uio: add updateiov()
This function take a struct uio previously created by copyinuio and and
updates the lengths of the user-space iovec to match those in the uio.
To reduce the risks of pointer leakage and cross-ABI pointer confusion,
lengths are updated individually.
Reviewed by: jamie, jhb
Effort: CHERI upstreaming
Sponsored by: DARPA, AFRL, Innovate UK
Differential Revision: https://reviews.freebsd.org/D60025
openssl: Fix CVE-2026-84782
This is a backport of an upstream commit to fix:
dtls: reset init_off before retransmitting a message
Approved by: so
Security: FreeBSD-SA-26:68.openssl
Security: CVE-2026-84782
vfs: Disallow renameat() with FD_RESOLVE_BENEATH descriptors
The FD_RESOLVE_BENEATH flag was intended to try to resolve bugzilla PR
262179 without entirely disallowing fd passing between jails. However,
one can use renameat() to bypass the restriction: upon receiving a
directory fd with FD_RESOLVE_BENEATH set, a jailed process can still
move its CWD or one of its ancestors to the directory, and just cd
out of its jail root.
So disallow renameat() when either the source or destination directory
fds has FD_RESOLVE_BENEATH set, like we do with fchdir() and fchroot()
to prevent similar escapes.
Approved by: so
Security: FreeBSD-SA-26:66.jail
Security: CVE-2026-101305
PR: 262179
Reported by: firk at cantconnect.ru
Reviewed by: olce, kib
Differential Revision: https://reviews.freebsd.org/D59875
fdescfs: Pass up additional metadata during lookups
When an fdescfs mount has the nodup option set, fdesc_lookup(/dev/fd/n)
returns the vnode referenced by file descriptor n, rather than returning
an fdescfs vnode. This meant that fd metadata attached to fd n was not
preserved when reopening the file, which is contrary to the expected
semantics for capsicum rights and the UF_RESOLVE_BENEATH fd flag. For
regular fdescfs mounts, this metadata is copied via dupfdopen().
Fix the problem by passing up this metadata through the nameidata
structure. Thus, if one opens /dev/fd/n, the returned fd will inherit
UF_RESOLVE_BENEATH and the capability rights of fd n. Add some
regression tests as well.
Approved by: so
Security: FreeBSD-SA-26:66.jail
Security: CVE-2026-101304
Reported by: Jan Bramkamp
Reviewed by: kib
[2 lines not shown]
file: Add a helper function to check whether filecaps are full
In a couple of places we want to know whether someone has limited rights
on an fd. There, we want a predicate which determines whether the set
of rights is smaller than CAP_ALL, and whether there are explicit ioctl
or fcntl lists. Factor this out into a helper function, in preparation
for use elsewhere.
No functional change intended.
Approved by: so
Security: FreeBSD-SA-26:66.jail
Reviewed by: kib
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59884
file: Add filecaps_intersect() and cap_rights_intersect()
These routines let one compute the intersection of two sets of filecaps
or capability rights, just as filecaps_merge() and cap_rights_merge()
compute the union. This will be useful in an upcoming patch.
filecaps_intersect() is complex due to the need to merge sets of ioctls.
For now this is implemented with a dumb nested loop on the basis that
ioctl lists are typically short enough that this is fine. It may be
better to instead sort the two lists first and step through them
together.
No functional change intended.
Approved by: so
Security: FreeBSD-SA-26:66.jail
Reviewed by: kib
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59885
sysvsem: Fix another sequence number wraparound race
semop() may sleep waiting for a semaphore. Upon waking up, it checks to
see if the set's sequence number has changed, indicating that the set
was removed. The sequence number is not wide enough to prevent a false
negative due to wraparound, in which case the subsequent access of
`semakptr->u.__sem_base[sopptr->sem_num]` may be out of bounds. This
race can be leveraged to elevate privileges.
Fix this by introducing a 64-bit sequence number for each semaphore
pool. This is wide enough to make the race impossible to hit. Allocate
a separate array for them, as we cannot really change the layout of
struct semid_kernel since some userspace tools (e.g., ipcrm(1)) embed
the layout.
While here, use semvalid() instead of open-coding its implementation,
convert a couple of flags to be bool, and use a better variable name to
store required permissions.
[8 lines not shown]
vfs: Disallow renameat() with FD_RESOLVE_BENEATH descriptors
The FD_RESOLVE_BENEATH flag was intended to try to resolve bugzilla PR
262179 without entirely disallowing fd passing between jails. However,
one can use renameat() to bypass the restriction: upon receiving a
directory fd with FD_RESOLVE_BENEATH set, a jailed process can still
move its CWD or one of its ancestors to the directory, and just cd
out of its jail root.
So disallow renameat() when either the source or destination directory
fds has FD_RESOLVE_BENEATH set, like we do with fchdir() and fchroot()
to prevent similar escapes.
Approved by: so
Security: FreeBSD-SA-26:66.jail
Security: CVE-2026-101305
PR: 262179
Reported by: firk at cantconnect.ru
Reviewed by: olce, kib
Differential Revision: https://reviews.freebsd.org/D59875
file: Add filecaps_intersect() and cap_rights_intersect()
These routines let one compute the intersection of two sets of filecaps
or capability rights, just as filecaps_merge() and cap_rights_merge()
compute the union. This will be useful in an upcoming patch.
filecaps_intersect() is complex due to the need to merge sets of ioctls.
For now this is implemented with a dumb nested loop on the basis that
ioctl lists are typically short enough that this is fine. It may be
better to instead sort the two lists first and step through them
together.
No functional change intended.
Approved by: so
Security: FreeBSD-SA-26:66.jail
Reviewed by: kib
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59885
sysvsem: Fix another sequence number wraparound race
semop() may sleep waiting for a semaphore. Upon waking up, it checks to
see if the set's sequence number has changed, indicating that the set
was removed. The sequence number is not wide enough to prevent a false
negative due to wraparound, in which case the subsequent access of
`semakptr->u.__sem_base[sopptr->sem_num]` may be out of bounds. This
race can be leveraged to elevate privileges.
Fix this by introducing a 64-bit sequence number for each semaphore
pool. This is wide enough to make the race impossible to hit. Allocate
a separate array for them, as we cannot really change the layout of
struct semid_kernel since some userspace tools (e.g., ipcrm(1)) embed
the layout.
While here, use semvalid() instead of open-coding its implementation,
convert a couple of flags to be bool, and use a better variable name to
store required permissions.
[8 lines not shown]
file: Add a helper function to check whether filecaps are full
In a couple of places we want to know whether someone has limited rights
on an fd. There, we want a predicate which determines whether the set
of rights is smaller than CAP_ALL, and whether there are explicit ioctl
or fcntl lists. Factor this out into a helper function, in preparation
for use elsewhere.
No functional change intended.
Approved by: so
Security: FreeBSD-SA-26:66.jail
Reviewed by: kib
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59884
kqueue: Fix a potential OOB access in kqueue_fork_copy_knote()
Here, fdp points to the new fdtable, copied from that of the parent
process. There is a window after the fdtable is copied, and before
kqueue_fork_copy_knote() runs, where a different thread in the parent
could have grown the parent's fdtable and registered a knote with ident
larger than the size of the child's fdtable. This race can lead to an
out-of-bounds read.
Add a bounds check for this case; skip the knote if it is referencing a
non-existent file.
Approved by: so
Security: FreeBSD-SA-26:65.kqueue
Security: CVE-2026-58100
Reviewed by: kib
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59916
kqueue: Fix handling of marker knotes during fork
Commit d8bdcb08d0eb fixed a problem in kqueue_fork_copy_knote() where we
did not skip over marker knotes when copying. However, that fix was not
sufficient: we bump the influx counter and check for a marker after
dropping the kqueue lock. So, if multiple threads in a process are
forking concurrently, kqueue_fork_copy_list() may mark a marker as
in-flux and drop the lock; if the marker owner then frees the marker,
the first thread will decrement the in-flux counter of a freed knotes.
This use-after-free can be exploited, at least prior to commit
d8bdcb08d0eb, which makes exploitation more challenging.
Approved by: so
Security: FreeBSD-SA-26:65.kqueue
Security: CVE-2026-58099
Reported by: Reo Shiseki
Reviewed by: kib
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59522
fdescfs: Pass up additional metadata during lookups
When an fdescfs mount has the nodup option set, fdesc_lookup(/dev/fd/n)
returns the vnode referenced by file descriptor n, rather than returning
an fdescfs vnode. This meant that fd metadata attached to fd n was not
preserved when reopening the file, which is contrary to the expected
semantics for capsicum rights and the UF_RESOLVE_BENEATH fd flag. For
regular fdescfs mounts, this metadata is copied via dupfdopen().
Fix the problem by passing up this metadata through the nameidata
structure. Thus, if one opens /dev/fd/n, the returned fd will inherit
UF_RESOLVE_BENEATH and the capability rights of fd n. Add some
regression tests as well.
Approved by: so
Security: FreeBSD-SA-26:66.jail
Security: CVE-2026-101304
Reported by: Jan Bramkamp
Reviewed by: kib
[2 lines not shown]