hastd: Fix crash on empty message
A HAST message can be empty, in which case ebuf_add_tail() does nothing
and ebuf_data() returns NULL because the size of the ebuf is zero, but
hast_proto_recv_hdr() asserts that the return value is not NULL,
resulting in an immediate crash if hastctl or hastd receive an empty
message. This is trivially reproducable by running `hastctl status` or
`hastctl role init` (as the rc script does prior to stopping hastd).
To avoid this, don't try to grow the ebuf or receive additional data
if the header size is zero.
PR: 298085
MFC after: 3 days
Reviewed by: kevans, gjb
Differential Revision: https://reviews.freebsd.org/D59306
hastd: Clean up the ebuf code
Rename the members of struct ebuf to match their function, replace
bcopy() with memcpy(), add comments explaining what each function does.
Reviewed by: kevans, emaste
Differential Revision: https://reviews.freebsd.org/D59310
fsck_msdosfs: add a test for reconnecting on volumes larger than 4 GiB
Build a 4.5 GiB FAT32 image whose LOST.DIR cluster sits exactly 4 GiB above
the single cluster of a PAYLOAD.BIN, so that truncating the offset of the
former to 32 bits yields the offset of the latter, then inject a lost cluster
chain and let fsck_msdosfs(8) reconnect it.
The test asserts both halves of the bug fixed in the previous commit: that
PAYLOAD.BIN's cluster is unchanged, and that a second pass no longer reports
the chain as lost, which it only stops doing once the directory entry reaches
the real LOST.DIR.
newfs_msdos(8) -C only calls ftruncate(2) and nothing outside the reserved
area, the FATs and a handful of clusters is ever written, so the image stays
sparse and costs about 2 MiB on disk.
The geometry is read back out of the BPB rather than assumed, so
newfs_msdos(8) stays free to lay the file system out differently; the test
fails with a clear message if the volume ever becomes too small to hold a
[3 lines not shown]
fsck_msdosfs: fix 32-bit overflow computing the LOST.DIR offset
reconnect() computed the byte offset of the LOST.DIR cluster in 32-bit
arithmetic and widened the result only on assignment:
lfoff = (lfcl - CLUST_FIRST) * boot->ClusterSize
+ boot->FirstCluster * boot->bpbBytesPerSec;
cl_t is u_int32_t and ClusterSize is u_int, so both products wrap modulo
2**32. Once LOST.DIR's cluster lies past the 4 GiB mark, lfoff aliases
the offset exactly 4 GiB below it, which on such a volume is ordinary
file data.
That offset is used for both the read and the write: reconnect() reads a
cluster of file data, scans it in 32-byte steps for a leading SLOT_EMPTY
or SLOT_DELETED byte, which arbitrary data readily provides, stores the
new directory entry in that slot, and writes the cluster back to the
same wrong place. Thirty-two bytes of an unrelated file are silently
replaced by a directory entry, and since that entry never reaches the
[14 lines not shown]
fsck_msdosfs: add tests for the 32-bit boot block field decoding
Exercise each of the 32-bit BIOS Parameter Block and FSInfo fields that
readboot() decodes, using values whose most significant byte has its high
bit set. Each case checks two things: that fsck_msdosfs(8) reports the
full unsigned 32-bit value back on stdout, and that nothing writes a
sanitizer runtime error to stderr.
The second check is what catches a byte-at-a-time decode. Shifting such
a byte left by 24 is undefined, but every compiler we use wraps it into
the same bit pattern, so the decoded value alone cannot tell a correct
decode from an overflowing one. In a WITH_UBSAN build bsd.sanitizer.mk
compiles with -fsanitize=undefined and -fsanitize-recover=undefined, so
the shift is reported on stderr and execution continues, which the test
can then assert on. Against the byte-at-a-time decode these cases fail
in a WITH_UBSAN build and pass without it.
Note that the stderr check also fails on unrelated undefined behavior
that these images reach anywhere in fsck_msdosfs(8), which is intended.
[2 lines not shown]
fsck_msdosfs: avoid signed integer overflow in readboot()
readboot() decoded the 32-bit little-endian BIOS Parameter Block and
FSInfo fields by shifting the individual bytes of a u_char array into
place. The u_char operands are promoted to signed int, so shifting a
most significant byte of 0x80 or greater left by 24 overflows int, which
is undefined behavior. Use le32dec() from <sys/endian.h> instead, which
is both well defined and easier to read.
No functional change intended.
MFC after: 1 week
Pull Request: https://github.com/freebsd/freebsd-src/pull/2350
examples/jails: Encode ifnames used as derive_mac counters
derive_mac keeps a per-parent branch index in a global named from the
parent interface so the N nibble can increment when the same PHY is
presented more than once. That name must be a POSIX identifier; a
vlan-style parent (em0.20) is not.
Encode the ifname first (alnum unchanged, every other byte as _HH) so
the lookup stays a symbol-table hit and em0.20 does not collide with
em0_20. Same change in jib (9.2) and jng (9.4).
In jng, also address netgraph by node name. ngctl(8) treats `.' and
`:' as control characters, so ng_ether(4) names its node after the
sanitized ifname (vtnet0.20 becomes vtnet0_20). Sanitize the parent
ifname where it enters and use that for every ngctl call; ifconfig(8)
and derive_mac keep the real name. Previously jng failed outright on
such parents where jib did not.
PR: 291143
[4 lines not shown]
fsck_msdosfs: fix memory leaks in checkfilesys()
Invoke releasefat(fat) in checkfilesys() prior to free(fat) on exit paths
so that fatbuf, headbitmap.map, and fat32_cache entries are properly freed.
MFC after: 1 week
Pull Request: https://github.com/freebsd/freebsd-src/pull/2351
tpm20: Move user copies outside the lifecycle lock
The TPM 2.0 character-device methods held the global device lock
while uiomove() accessed user memory. User page faults could therefore
delay suspend or detach even though the read response was already
buffered.
Add a per-open sleepable lock to serialize operations on each response
buffer. Stage commands under that lock before acquiring the device
lock, and copy them into the response buffer only after the lifecycle
checks succeed. This preserves an unread response when suspend or
detach rejects a write. Release the device lock before copying buffered
responses out. Also advance the response offset by the bytes actually
copied when uiomove() returns after a partial transfer.
Validated on an Intel TPM 2.0 TIS device. PCR reads and GetRandom
passed under 16-process mixed command load. A response was consumed
correctly in 5-byte, 7-byte, and remainder reads. Module unload/reload
recreated the device and entropy source without lock diagnostics.
[6 lines not shown]
tpm20: Correct 32-bit register helpers
OR4() reads only the low byte before writing the complete 32-bit
register. Preserve all register bits by using a matching 32-bit read.
Make BIT() produce an unsigned value so masks containing bit 31 do not
rely on a signed left shift into the sign bit. OpenBSD carries the
same change.
Reviewed by: kevans
MFC after: 2 weeks
Sponsored by: BBOX.io
Differential Revision: https://reviews.freebsd.org/D59244
tpm20: Release transport state after command failures
Once a transport acquires locality, several TIS and CRB error paths
return without relinquishing it. They can also leave a partial FIFO
transaction or an active CRB command for the next operation to inherit.
Route post-locality exits through common cleanup. Reset the TIS command
state on every attempt. For CRB, cancel an active failed command when
necessary, request the idle state, and relinquish locality even when the
state transition itself fails.
Successful command handling is unchanged apart from sharing the same
cleanup path.
Reviewed by: kevans
MFC after: 2 weeks
Sponsored by: BBOX.io
Differential Revision: https://reviews.freebsd.org/D59243
tpm_tis: Close interrupt wait races
The TIS interrupt handler can acknowledge and signal an event after
the waiter checks the device status but before it enters tsleep().
Since the handler is MPSAFE, the command lock does not close this
window. A lost wakeup can delay a completed command for its full
timeout, up to 40 seconds for long TPM 2.0 operations.
Publish the expected event under an interrupt mutex and use a generation
counter to record matching interrupts. Recheck the device predicate
without the mutex because register access may sleep on a SPI transport,
then compare the generation before atomically waiting on a condition
variable. This closes the check-to-sleep race without placing sleeping
bus operations under a mutex.
Use an absolute deadline while retrying the predicate after wakeups.
Apply the same scheme to locality acquisition, which had an equivalent
race. Leave the expected event published while polling so the
attach-time test can still prove that an advertised interrupt arrived.
[10 lines not shown]
tpm20: Harden the common device lifecycle
Mark the device as dying before teardown and destroy the character
device before freeing its private state or lock. This prevents cdev
methods from entering with a freed internal buffer or a destroyed sx.
Check the teardown state in command paths, honor failures from the cdev
private data interface, and publish teardown before waiting for the
lifecycle lock. Keep that lock across TPM retry delays so commands
cannot interpose and private state remains pinned, but abort before the
next retry once teardown begins.
Block new cdev operations after a successful Shutdown(STATE). Keep the
suspend gate and the TPM command under the same lock so a userspace
command cannot invalidate the saved state before S3 entry. Clear the
gate only after Startup(STATE) succeeds.
Keep entropy harvesting scheduled after a transient command or suspend
failure, but stop it while suspended or once teardown begins. Queue the
[14 lines not shown]
tpm_tis: Restore validated interrupts after resume
TIS interrupt routing and enable registers may lose their state across
S3, while the driver retains its software indication that interrupts
work. A subsequent locality or command wait can then sleep for an
interrupt that cannot arrive.
Remember whether interrupts worked before suspend and restore the
vector, pending status, and enable mask before TPM2_Startup. Put the
transport in polling mode first; the interrupt handler promotes it back
to interrupt waits only after observing an interrupt from the restored
configuration. If register restoration fails, Startup and subsequent
commands continue using polling.
Preserve the initial interrupt-enable mask, including the firmware's
trigger and polarity selection proven by the attach time interrupt test,
and restore that exact mask rather than accepting post-S3 defaults.
Program the same safe baseline for polling devices during attach and
[16 lines not shown]
tpm20: Initialize common state before testing TIS interrupts
The TIS attach path tested its interrupt by transmitting GetRandom
before tpm20_init() allocated the internal command buffer. A TPM2 FIFO
device with a usable IRQ could therefore dereference a null
internal_priv.
Initialize the common TPM2 state before running the interrupt test.
Make common cleanup safe for partially initialized devices and leave
cleanup to the attachment after tpm20_init() fails, avoiding duplicate
release of the lock, command buffer, and random-source state.
Clear the IRQ resource pointer after releasing it when interrupt handler
setup fails so the later polling-mode detach does not release it twice.
Free the internal command allocation through its object pointer rather
than relying on its embedded buffer being the first structure member.
Reviewed by: kevans
[3 lines not shown]
tpm20: Validate suspend and resume commands
The internal TPM2_Shutdown and TPM2_Startup paths ignored both transport
failures and the TPM response. Suspend could therefore enter S3 without
saved TPM state, while resume could restart entropy harvesting after a
failed state restoration.
Build both commands through one helper, validate their response framing
and TPM return codes, and propagate failures. Retry the standard RETRY
and TESTING responses with bounded exponential backoff. Accept
TPM_RC_INITIALIZE from Startup because firmware may already have started
the TPM during resume.
Do not enter S3 after an unsuccessful state save, and do not restart the
entropy task when TPM state restoration failed. If Shutdown fails after
the entropy task was drained, requeue it before returning so an aborted
suspend does not permanently stop harvesting.
Reviewed by: kevans
[3 lines not shown]
ixl: Route suspend and resume through iflib
Register the iflib device suspend and resume methods. Remove the
direct initialization from the driver resume callback because
iflib_device_resume() performs the datapath restart after the callback
returns.
MFC after: 2 weeks
Sponsored by: BBOX.io
ixv: Wire iflib suspend and resume methods
Register the standard iflib device suspend and resume methods so the
framework reinitializes the VF datapath after a system power
transition.
MFC after: 2 weeks
Sponsored by: BBOX.io
iavf: Wire iflib suspend and resume methods
Register the iflib device suspend and resume methods so the existing
driver callbacks run during system power transitions. This stops
mailbox retry work before suspend and lets iflib reinitialize the
datapath after resume.
MFC after: 2 weeks
Sponsored by: BBOX.io
axgbe: Wire iflib power management methods
Register the standard iflib device methods for shutdown, suspend, and
resume. This gives axgbe the framework managed reinitialization used by
other iflib drivers after a power transition.
MFC after: 2 weeks
Sponsored by: BBOX.io
enic: Route resets through iflib lifecycle
Mark the driver stopped after attach so its first IFDI_STOP() call
does not repeat hardware shutdown.
Defer error interrupt recovery through iflib instead of calling
driver stop and init methods from interrupt context. Let iflib own
the stop and restart around MTU changes as well, avoiding duplicate
lifecycle operations.
MFC after: 2 weeks
Sponsored by: BBOX.io
nfscl: A few more fixes for the NFS over RDMA client glue
A couple of additional fixes for the NFS client side RDMA glue:
- For Readdirplus, the reply needs to be a large chunk, so set
M_PROTO9 instead of M_PROTO8.
- The nfsclrdma.ko module now uses xprt_rdma_unmap_chunk()
instead of xprt_rdma_rekey_chunk().
Hopefully, this is it for the NFS over RDMA client glue changes.
MFC after: 3 months
Fixes: 884ee8d6c9b4 ("nfscl: Add some glue for client side NFS over RDMA")
fsck_msdosfs: add tests for lost cluster chain repair accounting
Add an ATF test suite covering Phase 3 ("Checking for Lost Files")
error accounting. Test images are created using newfs_msdos(8),
and lost cluster chains are injected directly into FAT copies at
offsets derived from the BPB. The LOST.DIR directory required by
reconnect() is constructed similarly: a root directory entry with
ATTR_DIRECTORY set and its first cluster pointing to a zero-filled
cluster containing "." and ".." entries.
The lost_chain_cleared and corrupted_lost_chain_reconnected test
cases provide regression coverage for the preceding commit:
- lost_chain_cleared verifies that clearing a lost chain (the fallback
taken when LOST.DIR is absent) exits with status 0 rather than 8
(unrecovered error).
- corrupted_lost_chain_reconnected verifies that FAT modifications
from a chain truncated by checkchain() prior to reconnection are
written back to disk, requiring "Update FATs? yes" and ensuring a
clean second pass.
[8 lines not shown]
fsck_msdosfs: fix status accounting for lost cluster chains
checklost() scans for lost cluster chains and attempts to
repair each one, first by reconnecting it to LOST.DIR,
and falling back to clearing it if reconnection fails.
However, checklost() incorrectly updates the modification
status flags (mod), which checkfilesys() relies on to
determine whether to write back changes and what exit
status to return.
The current code have three issues:
1. A reconnect() failure immediately sets FSERROR in mod
via "mod |= ret = reconnect(...)". If reconnect() failed
(e.g., because LOST.DIR is missing, or full) but the
fallback clear operation succeeds, clearchain() frees the
chain and sets FSFATMOD. However, the leftover FSERROR
remains in mod: checkfilesys() skips marking the file system
clean and exits with status 8, even though the file system was
[30 lines not shown]
bhyve: Keep passthrough PCI power state virtual
The passthrough Command register is emulated, but PMCSR writes were
sent directly to the physical function. A guest D3hot-to-D0 transition
can perform an internal reset and clear physical Command while its
emulated copy remains enabled.
Cache the Power Management capability and keep the physical D-state
host-owned. Emulate the guest D-state and advertise No_Soft_Reset so
the guest is not promised a function reset by a virtual power cycle.
Restore the assignment-time virtual state after a managed FLR.
Reviewed by: markj
Sponsored by: BBOX.io
(cherry picked from commit 3b90096cf9bcaec70b717e9ff0a9e23d14b600b6)
bhyve: Manage passthrough devices across guest FLR
bhyve emulates the guest PCI Command register so BAR sizing does not
disable physical decoding. However, PCIe Device Control was passed
through. A guest VFIO reset therefore performed a physical FLR, which
cleared physical Command, while the guest restored only its emulated
copy. The device remained assigned with bus mastering disabled and
could not fetch DMA descriptors.
Intercept guest FLR writes and issue a PPT-managed reset. Stop all
vCPUs, verify ownership, quiesce the function, perform only an FLR, and
restore the host-owned PCI configuration, decode, and bus-master state.
Keep the IOMMU domain in place. bhyve removes guest BAR mappings before
this ioctl; a later guest MEMEN write recreates them. Never escalate a
guest FLR to a power reset.
Reset the guest-owned Command, MSI, MSI-X, MSI-X table, INTx, and MRRS
state. PCIe 6.2 section 6.6.2 explicitly preserves MPS across FLR.
Virtualize MPS, MRRS, and Completion Timeout. Keep physical MPS and
[27 lines not shown]
e1000: Report 82571 packet buffer ECC errors
The 82571 PBA_ECC register contains a 12-bit count of packet buffer ECC
detections. The shared code enables single-bit correction, but neither
FreeBSD nor the DPDK base driver consumes the counter.
Sample it with the ordinary statistics timer, accumulate the value under
dev.em.N.memory_errors.detected_packet_buffer, and clear the hardware
counter while preserving correction and reserved register state. Do not
enable its shared interrupt: the register does not distinguish corrected
from uncorrectable events and does not provide a safe fatal recovery
policy.
Validated on a dual port 82571EB. Both functions reported zero after a
clean boot, and a controlled link down/up cycle left the counter at zero
while the management link recovered at 1 Gb/s without issue.
Sponsored by: BBOX.io
(cherry picked from commit aec0f1b85b54d14819747ed3364f366d21e76d88)
iflib: Plumb per-packet RX hardware timestamps to mbufs
Add iri_rcv_tstmp to if_rxd_info so an isc_rxd_pkt_get() driver can
report a hardware RX timestamp. Copy it into m_pkthdr.rcv_tstmp,
reusing the generic mbuf timestamp path.
Widen iri_flags from uint8_t to uint32_t and define the flags drivers
may supply. Mask the flags before copying them into the mbuf so no
other mbuf state can leak through the driver callback.
Place the timestamp next to iri_frags to avoid an alignment hole, and
document its nanoseconds-since-boot representation and validity flags.
Bump __FreeBSD_version because changing if_rxd_info breaks KBI.
Reviewed by: gallatin
Signed-off-by: Sreekanth Reddy <sreekanth.reddy at broadcom.com>
Differential Revision: https://reviews.freebsd.org/D58638