amd64, powerpc: Enable tpm(4) in supported kernels
tpm(4) was removed from amd64 GENERIC because it broke suspend and
resume. The preceding lifecycle, state-save, interrupt, locality, and
teardown fixes address those failures for both TPM 1.2 and TPM 2.0.
Restore the driver to amd64 GENERIC and MINIMAL, where TPM entropy
harvesting remained enabled. Enable the driver and entropy harvesting
in the MPC85XX and QORIQ64 configurations, which already provide FDT,
spibus, and the platform SPI controller required by FDT-attached TPMs.
Leave the generic AIM and POWER configurations unchanged because they
have no TPM attachment bus.
The TPM 1.2 path completed repeated S3 cycles and command tests on
ThinkPad T430 and T440p systems. The TPM 2.0 path completed repeated
device and full-system suspend/resume cycles on a ThinkPad P51. The
PowerPC configuration matrix was checked to retain tpm(4) only where its
FDT SPI attachment path is present.
[9 lines not shown]
bhyve: tpm: allow the last dword of the CRB command buffer
The bounds check rejected any access ending exactly at the end of the
register block, so a four byte write at offset 0xffc was refused,
returning EINVAL and killing the VM.
PR: 291063
Fixes: 75909086a45d ("bhyve: allow read/write to full CRB buffer")
Sponsored by: Defenso
Signed-off-by: Quentin Thébault <quentin.thebault at defenso.fr>
Reviewed-by: aokblast, kevans, markj
Pull-Request: https://github.com/freebsd/freebsd-src/pull/2362
(cherry picked from commit 4505445ccdb9555efc23262a971a18ca03486969)
bhyve: allow read/write to full CRB buffer
For some reason, we've incorrectly calculated the size of the CRB data buffer
register. There's no need to divide the CRB data buffer size by 4. We should
allow access to the whole buffer instead.
Reviewed by: markj
Sponsored by: Beckhoff Automation GmbH & Co. KG
Pull Request: https://github.com/freebsd/freebsd-src/pull/2169
(cherry picked from commit 75909086a45da3c5aeaff8152728111cf798c6bc)
tpm: crb: check the error bit only after taking locality
tpmcrb_transmit() read CRB_CTRL_STS before requesting locality 0.
An AMD Pluton fTPM (like FrameWork Desktop) using the plain CRB start method
reads the control area as all-ones until locality is assigned, so bit 0 looks
like a stuck tpmSts and every command failed with EIO.
With RANDOM_ENABLE_TPM the harvester retries every 10 seconds, so this printed
"Device has Error bit set" forever.
Reviewed by: kbowling
Approved by: kbowling
Sponsored by: Netflix
Differential Revision: https://reviews.freebsd.org/D59661
(cherry picked from commit 7c375af72ef3f631e89431ae32a855d806e6ec98)
tpm: crb: make the Pluton startmethod more resilient
The original implementation assumed that the start/reply doorbells
lived within the device _CRS space, but that isn't always the case. On
my AMD Ryzen 7640U-based frame.work laptop, device memory runs from
0xc0500000-0xc0500fff while the doorbells are up around 0xc0508000.
Stop sanity checking the addresses and just map them in to work reliably
whether they're within the device range or not.
pluton_wait_reply is cribbed from tpm_wait_for_u32, but rewritten
slightly to read in just one place and to read one last time before
giving up at the end of the timeout, just in case.
Reviewed by: kbowling
Differential Revision: https://reviews.freebsd.org/D59327
(cherry picked from commit 6af3031c951614f2a42dfd67ce2c3a48ee8314ae)
tpm: Do not use timed tsleep() while polling during cold boot
Commit 4e0f283fb97a made tpm_tis12_init() wait for TPM_STS_CMD_READY
after aborting any command. The wait is implemented by the driver's
existing tpm_waitfor_poll() loop, which sleeps with a one-tick tsleep()
between status reads. Until now, that loop only ran from the resume and
command paths after boot. From tpm_attach() it can panic with "timed
sleep before timers are working" when the TPM is attached from ACPI
during cold boot and the chip does not report ready on the first status
read.
Before 4e0f283fb97a, tpm_tis12_init() wrote TPM_STS_CMD_READY and
returned without waiting, so the polling loops only ran after boot.
tpm_request_locality() had the same latent hazard but its fast path
returns before sleeping whenever locality is already active.
Nothing calls wakeup() on the channels used by these polling loops, so
the sleeps are pure delays. Use pause_sig(), which falls back to
DELAY() while the kernel is cold and returns EWOULDBLOCK, a value these
[11 lines not shown]
tpm_tis: Quiesce interrupts before registering a handler
The current interrupt path uses the IRQ resource value directly as the
LPC SIRQ selector in TPM_INT_VECTOR and already restricts it to 1 through
15. This is a driver limitation: a parent interrupt number need not equal
an LPC SIRQ channel, and SPI TPMs can use a separate parallel interrupt.
On the reported system with ACPI IRQ 45, the existing range check runs
after handler registration and returns before disabling firmware interrupt
delivery. This can leave a polling device with a handler on an asserted
source.
Disable and verify interrupt delivery before registering a handler or
starting common TPM services. Preserve the existing range policy, using
polling without registering a handler for routes rejected by that check,
and release their IRQ resources. Keep a failed setup's potentially stale
output cookie out of the device state; the interrupt framework may
already have removed that handler.
[20 lines not shown]
tpm20: Move user copies outside the lifecycle lock
The TPM 2.0 character-device methods held the global device lock
while uiomove() accessed user memory. User page faults could therefore
delay suspend or detach even though the read response was already
buffered.
Add a per-open sleepable lock to serialize operations on each response
buffer. Stage commands under that lock before acquiring the device
lock, and copy them into the response buffer only after the lifecycle
checks succeed. This preserves an unread response when suspend or
detach rejects a write. Release the device lock before copying buffered
responses out. Also advance the response offset by the bytes actually
copied when uiomove() returns after a partial transfer.
Validated on an Intel TPM 2.0 TIS device. PCR reads and GetRandom
passed under 16-process mixed command load. A response was consumed
correctly in 5-byte, 7-byte, and remainder reads. Module unload/reload
recreated the device and entropy source without lock diagnostics.
[7 lines not shown]
tpm20: Correct 32-bit register helpers
OR4() reads only the low byte before writing the complete 32-bit
register. Preserve all register bits by using a matching 32-bit read.
Make BIT() produce an unsigned value so masks containing bit 31 do not
rely on a signed left shift into the sign bit. OpenBSD carries the
same change.
Reviewed by: kevans
Sponsored by: BBOX.io
Differential Revision: https://reviews.freebsd.org/D59244
(cherry picked from commit f613a43c366600e8c388c3200f74c6dd27f8a7a2)
tpm20: Release transport state after command failures
Once a transport acquires locality, several TIS and CRB error paths
return without relinquishing it. They can also leave a partial FIFO
transaction or an active CRB command for the next operation to inherit.
Route post-locality exits through common cleanup. Reset the TIS command
state on every attempt. For CRB, cancel an active failed command when
necessary, request the idle state, and relinquish locality even when the
state transition itself fails.
Successful command handling is unchanged apart from sharing the same
cleanup path.
Reviewed by: kevans
Sponsored by: BBOX.io
Differential Revision: https://reviews.freebsd.org/D59243
(cherry picked from commit 2fe8510b0a06ad577789ece7be8c9319c8ce3d06)
tpm_tis: release TPM resources after reading response
Per TIS 1.3 section 5.6.12, write commandReady to TPM_STS after reading the
response so the TPM can free its ReadFIFO and other internal resources.
The subsequent tpmtis_go_ready() provides the second write the spec describes
and waits for the state transition.
PR: 295103
Reported by: Benoit Sansoni <benoit.sansoni at gmail.com>
Reviewed by: kevans
Approved by: kevans
Differential Revision: https://reviews.freebsd.org/D57841
(cherry picked from commit e4c50d86796d07cf90e6a683234e1c0e9ca5e03c)
bsdinstall: Drop "Technology preview" from package sets
And refer to dist sets as "legacy."
We're planning on turning off dist sets in 16.0 (the code will remain
in the tree, but disabled by default) so we'd like to encourage users
to install 15.2 systems using package sets in order to ease the upgrade
path to 16.0.
Reviewed by: cperciva
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D54156
(cherry picked from commit dac74b20c706b1f73986fb40dac27ed85c1d2850)
tpm_tis: Close interrupt wait races
The TIS interrupt handler can acknowledge and signal an event after
the waiter checks the device status but before it enters tsleep().
Since the handler is MPSAFE, the command lock does not close this
window. A lost wakeup can delay a completed command for its full
timeout, up to 40 seconds for long TPM 2.0 operations.
Publish the expected event under an interrupt mutex and use a generation
counter to record matching interrupts. Recheck the device predicate
without the mutex because register access may sleep on a SPI transport,
then compare the generation before atomically waiting on a condition
variable. This closes the check-to-sleep race without placing sleeping
bus operations under a mutex.
Use an absolute deadline while retrying the predicate after wakeups.
Apply the same scheme to locality acquisition, which had an equivalent
race. Leave the expected event published while polling so the
attach-time test can still prove that an advertised interrupt arrived.
[11 lines not shown]
tpm20: Harden the common device lifecycle
Mark the device as dying before teardown and destroy the character
device before freeing its private state or lock. This prevents cdev
methods from entering with a freed internal buffer or a destroyed sx.
Check the teardown state in command paths, honor failures from the cdev
private data interface, and publish teardown before waiting for the
lifecycle lock. Keep that lock across TPM retry delays so commands
cannot interpose and private state remains pinned, but abort before the
next retry once teardown begins.
Block new cdev operations after a successful Shutdown(STATE). Keep the
suspend gate and the TPM command under the same lock so a userspace
command cannot invalidate the saved state before S3 entry. Clear the
gate only after Startup(STATE) succeeds.
Keep entropy harvesting scheduled after a transient command or suspend
failure, but stop it while suspended or once teardown begins. Queue the
[15 lines not shown]
tpm_tis: Restore validated interrupts after resume
TIS interrupt routing and enable registers may lose their state across
S3, while the driver retains its software indication that interrupts
work. A subsequent locality or command wait can then sleep for an
interrupt that cannot arrive.
Remember whether interrupts worked before suspend and restore the
vector, pending status, and enable mask before TPM2_Startup. Put the
transport in polling mode first; the interrupt handler promotes it back
to interrupt waits only after observing an interrupt from the restored
configuration. If register restoration fails, Startup and subsequent
commands continue using polling.
Preserve the initial interrupt-enable mask, including the firmware's
trigger and polarity selection proven by the attach time interrupt test,
and restore that exact mask rather than accepting post-S3 defaults.
Program the same safe baseline for polling devices during attach and
[17 lines not shown]
tpm20: Initialize common state before testing TIS interrupts
The TIS attach path tested its interrupt by transmitting GetRandom
before tpm20_init() allocated the internal command buffer. A TPM2 FIFO
device with a usable IRQ could therefore dereference a null
internal_priv.
Initialize the common TPM2 state before running the interrupt test.
Make common cleanup safe for partially initialized devices and leave
cleanup to the attachment after tpm20_init() fails, avoiding duplicate
release of the lock, command buffer, and random-source state.
Clear the IRQ resource pointer after releasing it when interrupt handler
setup fails so the later polling-mode detach does not release it twice.
Free the internal command allocation through its object pointer rather
than relying on its embedded buffer being the first structure member.
Reviewed by: kevans
[4 lines not shown]
tpm20: Validate suspend and resume commands
The internal TPM2_Shutdown and TPM2_Startup paths ignored both transport
failures and the TPM response. Suspend could therefore enter S3 without
saved TPM state, while resume could restart entropy harvesting after a
failed state restoration.
Build both commands through one helper, validate their response framing
and TPM return codes, and propagate failures. Retry the standard RETRY
and TESTING responses with bounded exponential backoff. Accept
TPM_RC_INITIALIZE from Startup because firmware may already have started
the TPM during resume.
Do not enter S3 after an unsuccessful state save, and do not restart the
entropy task when TPM state restoration failed. If Shutdown fails after
the entropy task was drained, requeue it before returning so an aborted
suspend does not permanently stop harvesting.
Reviewed by: kevans
[4 lines not shown]
tpm: fix multi-threaded access with per-open state
The TPM driver currently has a single buffer per instance to hold the
result of a command, and does not allow subsequent commands to be sent
until the current result is read by the same OS thread that sent the
command, with a timeout to throw away the result after a while if the
result is not read in a timely fashion. This has a couple problems:
- The timeout code has a bug which causes all subsequent commands to
hang forever if a different OS thread tries to read the result
before the OS thread which sent the command, and the OS thread
which sent the command never tries to read the result.
- Even if the first problem is fixed, applications expect to be able
to read the result from a different OS thread than the OS thread
which sent the command. The particular case that we saw was a go
application where the go runtime scheduled the goroutine which read
the result to a different OS thread from one where the goroutine
that sent the command ran, and there's no way to force these to
[13 lines not shown]
e1000: Configure PCH low-power link modes for suspend
The PCH suspend path kept a wake link fully powered and did not restore
the negotiated EEE modes after its stop-time reset. Intel provides the
ULP entry and exit machinery in the shared code, but FreeBSD did not
invoke its Sx policy.
Enter ULP on LPT and newer PCH controllers when wake is armed without
directed-unicast, multicast, or broadcast filters, which ULP cannot
preserve. For a link retained by host wake or management, restore the
100BASE-TX and 1000BASE-T LPI controls selected by the local
advertisement and the cached link-partner ability.
Keep these power reductions best-effort: wake filters and PME are
already configured independently, and a ULP or EEE failure is logged
without converting an optional power optimization into a suspend
failure. The existing PCH resume workaround forcibly exits ULP and
clears automatic Sx LPI state before normal initialization.
[11 lines not shown]
e1000: Rework Wake-on-LAN policy and programming
The driver used the NVM APME default as both the hardware-support
decision and the mutable filter mask. Consequently, an NVM-disabled
but capable port did not advertise wake support, disabling a wake mode
once could keep it disabled across later suspends, and directed-unicast
wake could never be selected.
Require the PCI power management capability to report D3hot PME support
before advertising or arming wake. A PM capability alone does not mean
the function can signal PME from the state used during system sleep.
Separate the board and port capability matrix from the NVM-selected
magic packet default. Read the proper per function NVM word on igb
controllers, cover the newer PCH generations, and retain the documented
legacy, multi-port, and OEM restrictions. Decode the distinct APM
Enable locations used by 82544, 82541EI/82547EI, and the later 8254x
parts. Do not advertise wake on the 82541ER, whose power-management
logic cannot assert PME for wake events. For I210/I211 internal iNVM,
[93 lines not shown]
e1000: Recover from igb(4) controller DEV_RST
CTRL.DEV_RST resets every port on an 82580 and newer igb device.
Hardware reports the event to each affected function through ICR.DRSTA
and requires software to reinitialize the port registers and descriptor
rings. The driver neither enabled nor handled this cause, so a reset
initiated by another function could leave a running interface using
stale state.
In FreeBSD, we do not currently send this, but other OSes including
Linux do, so a PF passed through to such a guest or FreeBSD as a guest
running with a passthrough PF on the same controller can be wedged.
Enable DRSTA for 82580 and newer PFs in both MSI-X and shared MSI/legacy
modes. Bit 30 is reserved on 82575 and is the TCP timer on 82576, so
leave it masked on those parts. Latch the event without programming
port registers from the interrupt filter. Defer the iflib reset request
to admin-task context because the request takes STATE_LOCK. Use IAM to
auto-mask shared interrupts on the first ICR read. Keep admin,
[41 lines not shown]
e1000: Correct igb(4) DMA coalescing register programming
The disabled path wrote the complement of DMAC_EN to DMACR. That
set every other field, including reserved bits, the one-shot EXIT_DC
command, watchdog enables, receive threshold, and PCIe Lx selection.
When DMA coalescing was enabled, the requested watchdog and Lx-delay
values were ORed into their reset values rather than replacing the
fields. PCIEMISC.LX_DECISION was also cleared, preventing
DMACR.DMAC_Lx from controlling PCIe low-power entry.
Disable coalescing by clearing only DMAC_EN and retaining the
documented DMAC_Lx policy and watchdog fields. On enable, replace the
variable fields under their masks, select DMA requirements for PCIe
low-power entry, and program the loopback and BMC watchdog policies
independently of prior state.
Apply this consistently to I350, I354, and I210 while preserving the
I354-specific timer units. Leave I210 reserved fields at their
[13 lines not shown]
Merge commit 681c2ee4dfbf from llvm-project (by Justin King):
asan: refactor interceptor allocation/deallocation functions (#145087)
Do some refactoring to allocation/deallocation interceptors. Expose
explicit per-alloc_type functions and stop accepting explicit AllocType.
This ensures we do not accidentally mix.
NOTE: This change rejects attempts to call `operator new(<some_size>,
static_cast<std::align_val_t>(0))`.
For https://github.com/llvm/llvm-project/issues/144435
Signed-off-by: Justin King <jcking at google.com>
This a prerequisite for adding sanitizer interceptors for free_sized(3)
and free_aligned_sized(3).
PR: 298943
MFC after: 1 week
Merge commit a5fa4dba6e2e from llvm-project (by PiJoules):
[compiler-rt] Add interceptors for free_[aligned_]sized for asan+hwasan (#189109)
This avoids a jemalloc assertion when running sanitized applications
against glib, which uses free_sized(3).
PR: 298943
MFC after: 1 week
compiler-rt: enable use of .preinit_array after ef758b59a44e
After base ef758b59a44e6b5b2d2dc178b97c01cde1b34570 we can enable usage
of .preinit_array in compiler-rt's sanitizers. The comment that stated
"On FreeBSD, .preinit_array functions are called with rtld_bind_lock
writer lock held. It will lead to dead lock ..." can also be removed.
PR: 298943
MFC after: 1 week
man4: Only link mgb.4 to if_mgb.4 where mgb.4 is installed
mgb.4 is only installed on amd64 and i386, move the MLINKS entry into
the amd64/i386 block next to _mgb.4.
Fixes: 05c49e2aa643 ("man: Link mgb.4 to if_mgb.4")
MFC after: 3 days
Sponsored by: The FreeBSD Foundation
contrib/bc: remove vs sub-directory
Some of the files in the vs sub-directory are checked out with CRLF
line breaks via .getattributes. This causes issues when checking out
the source tree with other tools that operate on Git repositories like
"got". Remove the vs sub-directory from the contrib tree, since it is
not required on FreeBSD.