sctp: hold tcp lock when calling sctp_calc_rwnd()
This was done in all code paths except one. Fix this to avoid
race conditions.
Reported by: syzbot+2e2dc35c0e24a74663a7 at syzkaller.appspotmail.com
Reviewed by: tuexen
MFC after: 1 week
MFC to: stable/14
MFC to: stable/15
safe_eval.sh: safe_set_var avoid positional args
Use options to set ssv_{allow,only,extra}
Improve the documentation.
Add a test harness for safe_set_var.
pfctl: remove the prototype of pfr_ina_define()
The function went when DIOCRINADEFINE was converted to netlink; pfctl
calls pfctl_ina_define() from libpfctl instead. Name that one in the
error message that pfctl_load_tables() prints when it fails.
Fixes: 219ce3063840 ("pf: convert DIOCRINADEFINE to netlink")
Sponsored by: Rubicon Communications, LLC ("Netgate")
libpfctl: Fix the counts returned by pfctl_ina_define()
The counts of a table definition were stored only if the caller's
variables already held a non-zero value, and read uninitialised ones
otherwise. pfctl_ina_define() passes the address of two uninitialised
locals and sums them up.
Reviewed by: kp
Fixes: 219ce3063840 ("pf: convert DIOCRINADEFINE to netlink")
Sponsored by: Rubicon Communications, LLC ("Netgate")
Differential Revision: https://reviews.freebsd.org/D60614
LinuxKPI: pci: first round of file cleanup
While the multiple sections of the file are mostly clearly separated
some bits were spread around.
Get the device methods array to the end of the file saving us a lot
of forward function delcarations and making it easier to rename them
to a common lkpi_ prefix instead of linux_ showing that they are file
local internal functions not part of any official LinuxKPI KPI.
Sort the various lkpi_pci_* bus attachment functions together,
followed by the lkpi_pci_iov_ ones. Then add the
pci_[un]register_driver* driver bits. This keeps the group
logically together.
For the PCI functions then group the pdev allocations and free/release
functions--for the two paths--together, giving them more homogeneous
names and add comments to some of them so it is more clear which
ones are shared ("common") and which ones belong to which code path.
[17 lines not shown]
LinuxKPI: pci: make internal variables file local
The only reason the internal lists and lock are shared seems to be
because the SYSINIT code was in linux_compat.c.
Migrate the three variables (pci_drivers, pci_devices, and pci_lock)
to linux_pci.c and make them file local.
While doing so adjust a few things:
(a) mark the various sections of the linux_pci.c file as we add new
SYSINIT code to the middle of the file. The file currently really
contains three parts.
(b) rename the variables by adding a lkpi_ prefix indicating that they
are part of our internal code and not part of the public KPI, and
to clearly distinguish them from any native PCI code upon which
we call as well.
(c) factor the locking code out into macros as we have often done in
various parts of the kernel to make the lock itself more opaque.
[6 lines not shown]
LinuxKPI: pci: cleanup lkpinew_pci_dev_release()
Factor out some cleanup routines from the device detach path so
we can re-use them here too.
Also clarifying that lkpinew_pci_dev_release() should only run
from the (*release) callback function means that we adjust our
other calls to what we had already envisioned in comments, to
go through the official pci_dev_put() path.
Compared to the device detach path, we have little chance to know
if we held the last reference and if the device is really gone or
not but in this case a later release would be fine (if it happens).
In contrast to the device detach path we cannot do the hack of
setting the .release callback to NULL and check upon return from
pci_dev_put() as in this case the pdev is freed.
Sponsored by: The FreeBSD Foundation
MFC after: 3 days
Reviewed by: dumbbell (before rebase for ac9c07a3e908 instead of D57430)
Differential Revision: https://reviews.freebsd.org/D57431
if_gre: Use OID_AUTO for the net.link.gre sysctl node
if_gre(4) registers net.link.gre with the static OID number IFT_TUNNEL.
That number is an interface type that several tunnel drivers share, not
one that belongs to gre(4). if_me(4) used it too, which made loading
both modules panic until 21f2245de219. Nothing addresses the node by
number, so use OID_AUTO as if_me(4) now does.
Suggested by: markj
MFC after: 3 days
Sponsored by: Rubicon Communications, LLC ("Netgate")
safe_eval.sh fix description of safe_set_var returns.
The description was out of sync wrt return values.
1. invalid var
2. invalid val
3. var not in allowed list
safe_eval.sh: add safe_set_var
Get the log message right!
Fix default allowed chars (untabify clobbered TAB)
Add a couple more : messages to help debugging.
safe_eval.sh add safe_var_set
safe_var_set can be used to safely handle var=val args from command line.
eg.
while :
do
case "$1" in
*=*) # eval "$1" can be very dangerous
# safe_var_set will only eval if
# both var and value are sane.
if ! safe_var_set "$1"; then
case "$?" in
1) echo invalid variable name;;
2) echo invalide value;;
3) echo variabel not allowed;;
esac
fi
shift
[5 lines not shown]
nfs_nfsdstate.c: Add an extra safety belt check for the backchannel
I do not think that xp_p2 can be NULL at this point,
but add an extra safety belt, just in case.
(cherry picked from commit a52c50b4b7c2652954ef1bd34710a4b1fc8ef391)
clnt_vc.c: Fix handling of broken TCP connections
After more than, I don't know, maybe 10k operations: mount, copy,
remove, verify and unmount cycles, one cp command hung in
close() / ncl_flush and never recovered. The machine and the mount
continued to work normally through a new connection, but the writes
using the old connection stayed frozen.
I did not understand exactly what happened. I traced what appears to
be the issue in the code. My current understanding is that
clnt_vc_soupcall() saw the EOF and woke the caller waiting for RPC
replies, but one caller remained blocked in sosend().
That thread continued holding a reference to the old client, preventing
it from being completely cleaned up.
The attached patch calls socantsendmore() when EOF is received, which
should wake the blocked sender and let the normal reconnect code replace
the connection.
(cherry picked from commit 49bec8c3dc58cf8e944f9bbbbea07e7e856ad897)
clnt_vc.c: Fix handling of backchannel xprt
When clnt_vc_destroy() is called, it might not be the
current connection. Without this patch, if it is not
the current connection, xp_p2 is set NULL and xprt is released
when it should not be released.
This patch adds a check for "current connection" to fix
the problem. Found during testing to the client RDMA code,
but could happen for TCP as well.
(cherry picked from commit 81a6514689cefdaa5c92b5539ae885ec8c6b3336)
nfs_nfsdstate.c: Add an extra safety belt check for the backchannel
I do not think that xp_p2 can be NULL at this point,
but add an extra safety belt, just in case.
(cherry picked from commit a52c50b4b7c2652954ef1bd34710a4b1fc8ef391)
clnt_vc.c: Fix handling of broken TCP connections
After more than, I don't know, maybe 10k operations: mount, copy,
remove, verify and unmount cycles, one cp command hung in
close() / ncl_flush and never recovered. The machine and the mount
continued to work normally through a new connection, but the writes
using the old connection stayed frozen.
I did not understand exactly what happened. I traced what appears to
be the issue in the code. My current understanding is that
clnt_vc_soupcall() saw the EOF and woke the caller waiting for RPC
replies, but one caller remained blocked in sosend().
That thread continued holding a reference to the old client, preventing
it from being completely cleaned up.
The attached patch calls socantsendmore() when EOF is received, which
should wake the blocked sender and let the normal reconnect code replace
the connection.
(cherry picked from commit 49bec8c3dc58cf8e944f9bbbbea07e7e856ad897)
clnt_vc.c: Fix handling of backchannel xprt
When clnt_vc_destroy() is called, it might not be the
current connection. Without this patch, if it is not
the current connection, xp_p2 is set NULL and xprt is released
when it should not be released.
This patch adds a check for "current connection" to fix
the problem. Found during testing to the client RDMA code,
but could happen for TCP as well.
(cherry picked from commit 81a6514689cefdaa5c92b5539ae885ec8c6b3336)
amd64, powerpc: Enable tpm(4) in supported kernels
tpm(4) was removed from amd64 GENERIC because it broke suspend and
resume. The preceding lifecycle, state-save, interrupt, locality, and
teardown fixes address those failures for both TPM 1.2 and TPM 2.0.
Restore the driver to amd64 GENERIC and MINIMAL, where TPM entropy
harvesting remained enabled. Enable the driver and entropy harvesting
in the MPC85XX and QORIQ64 configurations, which already provide FDT,
spibus, and the platform SPI controller required by FDT-attached TPMs.
Leave the generic AIM and POWER configurations unchanged because they
have no TPM attachment bus.
The TPM 1.2 path completed repeated S3 cycles and command tests on
ThinkPad T430 and T440p systems. The TPM 2.0 path completed repeated
device and full-system suspend/resume cycles on a ThinkPad P51. The
PowerPC configuration matrix was checked to retain tpm(4) only where its
FDT SPI attachment path is present.
[9 lines not shown]
bhyve: tpm: allow the last dword of the CRB command buffer
The bounds check rejected any access ending exactly at the end of the
register block, so a four byte write at offset 0xffc was refused,
returning EINVAL and killing the VM.
PR: 291063
Fixes: 75909086a45d ("bhyve: allow read/write to full CRB buffer")
Sponsored by: Defenso
Signed-off-by: Quentin Thébault <quentin.thebault at defenso.fr>
Reviewed-by: aokblast, kevans, markj
Pull-Request: https://github.com/freebsd/freebsd-src/pull/2362
(cherry picked from commit 4505445ccdb9555efc23262a971a18ca03486969)
bhyve: allow read/write to full CRB buffer
For some reason, we've incorrectly calculated the size of the CRB data buffer
register. There's no need to divide the CRB data buffer size by 4. We should
allow access to the whole buffer instead.
Reviewed by: markj
Sponsored by: Beckhoff Automation GmbH & Co. KG
Pull Request: https://github.com/freebsd/freebsd-src/pull/2169
(cherry picked from commit 75909086a45da3c5aeaff8152728111cf798c6bc)
tpm: crb: check the error bit only after taking locality
tpmcrb_transmit() read CRB_CTRL_STS before requesting locality 0.
An AMD Pluton fTPM (like FrameWork Desktop) using the plain CRB start method
reads the control area as all-ones until locality is assigned, so bit 0 looks
like a stuck tpmSts and every command failed with EIO.
With RANDOM_ENABLE_TPM the harvester retries every 10 seconds, so this printed
"Device has Error bit set" forever.
Reviewed by: kbowling
Approved by: kbowling
Sponsored by: Netflix
Differential Revision: https://reviews.freebsd.org/D59661
(cherry picked from commit 7c375af72ef3f631e89431ae32a855d806e6ec98)
tpm: crb: make the Pluton startmethod more resilient
The original implementation assumed that the start/reply doorbells
lived within the device _CRS space, but that isn't always the case. On
my AMD Ryzen 7640U-based frame.work laptop, device memory runs from
0xc0500000-0xc0500fff while the doorbells are up around 0xc0508000.
Stop sanity checking the addresses and just map them in to work reliably
whether they're within the device range or not.
pluton_wait_reply is cribbed from tpm_wait_for_u32, but rewritten
slightly to read in just one place and to read one last time before
giving up at the end of the timeout, just in case.
Reviewed by: kbowling
Differential Revision: https://reviews.freebsd.org/D59327
(cherry picked from commit 6af3031c951614f2a42dfd67ce2c3a48ee8314ae)
tpm: Do not use timed tsleep() while polling during cold boot
Commit 4e0f283fb97a made tpm_tis12_init() wait for TPM_STS_CMD_READY
after aborting any command. The wait is implemented by the driver's
existing tpm_waitfor_poll() loop, which sleeps with a one-tick tsleep()
between status reads. Until now, that loop only ran from the resume and
command paths after boot. From tpm_attach() it can panic with "timed
sleep before timers are working" when the TPM is attached from ACPI
during cold boot and the chip does not report ready on the first status
read.
Before 4e0f283fb97a, tpm_tis12_init() wrote TPM_STS_CMD_READY and
returned without waiting, so the polling loops only ran after boot.
tpm_request_locality() had the same latent hazard but its fast path
returns before sleeping whenever locality is already active.
Nothing calls wakeup() on the channels used by these polling loops, so
the sleeps are pure delays. Use pause_sig(), which falls back to
DELAY() while the kernel is cold and returns EWOULDBLOCK, a value these
[11 lines not shown]
tpm_tis: Quiesce interrupts before registering a handler
The current interrupt path uses the IRQ resource value directly as the
LPC SIRQ selector in TPM_INT_VECTOR and already restricts it to 1 through
15. This is a driver limitation: a parent interrupt number need not equal
an LPC SIRQ channel, and SPI TPMs can use a separate parallel interrupt.
On the reported system with ACPI IRQ 45, the existing range check runs
after handler registration and returns before disabling firmware interrupt
delivery. This can leave a polling device with a handler on an asserted
source.
Disable and verify interrupt delivery before registering a handler or
starting common TPM services. Preserve the existing range policy, using
polling without registering a handler for routes rejected by that check,
and release their IRQ resources. Keep a failed setup's potentially stale
output cookie out of the device state; the interrupt framework may
already have removed that handler.
[20 lines not shown]
tpm20: Move user copies outside the lifecycle lock
The TPM 2.0 character-device methods held the global device lock
while uiomove() accessed user memory. User page faults could therefore
delay suspend or detach even though the read response was already
buffered.
Add a per-open sleepable lock to serialize operations on each response
buffer. Stage commands under that lock before acquiring the device
lock, and copy them into the response buffer only after the lifecycle
checks succeed. This preserves an unread response when suspend or
detach rejects a write. Release the device lock before copying buffered
responses out. Also advance the response offset by the bytes actually
copied when uiomove() returns after a partial transfer.
Validated on an Intel TPM 2.0 TIS device. PCR reads and GetRandom
passed under 16-process mixed command load. A response was consumed
correctly in 5-byte, 7-byte, and remainder reads. Module unload/reload
recreated the device and entropy source without lock diagnostics.
[7 lines not shown]
tpm20: Correct 32-bit register helpers
OR4() reads only the low byte before writing the complete 32-bit
register. Preserve all register bits by using a matching 32-bit read.
Make BIT() produce an unsigned value so masks containing bit 31 do not
rely on a signed left shift into the sign bit. OpenBSD carries the
same change.
Reviewed by: kevans
Sponsored by: BBOX.io
Differential Revision: https://reviews.freebsd.org/D59244
(cherry picked from commit f613a43c366600e8c388c3200f74c6dd27f8a7a2)
tpm20: Release transport state after command failures
Once a transport acquires locality, several TIS and CRB error paths
return without relinquishing it. They can also leave a partial FIFO
transaction or an active CRB command for the next operation to inherit.
Route post-locality exits through common cleanup. Reset the TIS command
state on every attempt. For CRB, cancel an active failed command when
necessary, request the idle state, and relinquish locality even when the
state transition itself fails.
Successful command handling is unchanged apart from sharing the same
cleanup path.
Reviewed by: kevans
Sponsored by: BBOX.io
Differential Revision: https://reviews.freebsd.org/D59243
(cherry picked from commit 2fe8510b0a06ad577789ece7be8c9319c8ce3d06)
tpm_tis: release TPM resources after reading response
Per TIS 1.3 section 5.6.12, write commandReady to TPM_STS after reading the
response so the TPM can free its ReadFIFO and other internal resources.
The subsequent tpmtis_go_ready() provides the second write the spec describes
and waits for the state transition.
PR: 295103
Reported by: Benoit Sansoni <benoit.sansoni at gmail.com>
Reviewed by: kevans
Approved by: kevans
Differential Revision: https://reviews.freebsd.org/D57841
(cherry picked from commit e4c50d86796d07cf90e6a683234e1c0e9ca5e03c)
bsdinstall: Drop "Technology preview" from package sets
And refer to dist sets as "legacy."
We're planning on turning off dist sets in 16.0 (the code will remain
in the tree, but disabled by default) so we'd like to encourage users
to install 15.2 systems using package sets in order to ease the upgrade
path to 16.0.
Reviewed by: cperciva
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D54156
(cherry picked from commit dac74b20c706b1f73986fb40dac27ed85c1d2850)