FreeBSD: do not clear dirty bits outside a sub-block write
page_busy() shrinks the written range to DEV_BSIZE boundaries. A write
that starts and ends inside one block leaves nbytes at -DEV_BSIZE, and
vm_page_clear_dirty() then clears that block and every one above it. A
page dirtied through mmap goes clean and the store is never written.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Alexander Motin <alexander.motin at TrueNAS.com>
Signed-off-by: Nick Price <nprice at FreeBSD.org>
Closes #19268
ZTS: compare pre-condense log spacemaps
The flush can write new log spacemaps in a later TXG, so total
smp_length can grow even when older logs were flushed. Compare only
entries present before condense, then sync before reading their
lengths.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19273
ZTS: restore timeout for deadman negative check
The test lowers the deadman timeout to 5 seconds for its positive
check. Restore the default before the short-delay negative check so
a long QEMU scheduling pause does not look like another deadman
failure.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19273
ZTS: retry interrupted ctime test sleeps
The ctime helper assumes sleep(2) always completes. A signal can
end it early, leaving an empty timestamp window. Retry until two
wall-clock seconds have passed before running the operation.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19273
Open only the named pool when zpool checks a pool name
While tracing zpool get on one pool I noticed that it opens every
imported pool before the one named on the command line.
zpool get, zpool set and zpool iostat call is_pool() to tell a pool
name from a vdev name. It opened every imported pool to compare names,
so a command on a small pool also paid for every other pool on the
system. is_pool() now opens only the pool it is asked about.
On a system with three pools, one of them with 1200 disks, zpool get
all on the other two pools went from about 930 ms to 15 ms and 22 ms.
On the 1200 disk pool itself it stayed at about 1.8 seconds, because
the named pool is still opened twice, once for this check and once by
the command.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Ameer Hamza <ameer.hamza at truenas.com>
Reviewed-by: Rob Norris <rob.norris at truenas.com>
Signed-off-by: Caleb St. John <yocalebo at gmail.com>
Closes #19271
Bail out of zil_create() when the pool suspends
zil_create() waits with txg_wait_synced(), a void wrapper around a
plain wait that cannot report a suspend. When the pool suspends while
it waits, the fsync() that got there does not return until the pool
resumes.
The rest of the commit path already handles this. zil_commit_flags()
and zil_commit_writer_stall() wait with TXG_WAIT_SUSPEND, call
zil_crash() on ESHUTDOWN and hand EIO back to the waiters.
zil_create() and zil_commit_activate_saxattr_feature() were left on the
old wait, so a suspend caught in either one still hangs.
Use the suspend-aware wait in all three places and return the failure.
zil_process_commit_list() already has a NULL-lwb path that signals the
nolwb waiters with the error, and its comment already says an ESHUTDOWN
there means zil_crash() was called, so the failure lands somewhere that
expects it.
[11 lines not shown]
Bail out of the objset upgrade when the pool suspends
The two objset upgrade callbacks end with txg_wait_synced(), a void
wrapper around a plain wait that cannot report a suspend. When the
pool suspends during that wait, the upgrade taskq thread stays blocked
until the pool resumes. So does dmu_objset_disown(), which calls
dmu_objset_upgrade_stop(). That waits for a running upgrade task to
finish and then waits for a txg itself.
Give all three waits TXG_WAIT_SUSPEND. The callbacks return EAGAIN,
which dmu_objset_upgrade_task_cb() records in os_upgrade_status.
zfs_ioc_userspace_upgrade() and zfs_ioc_id_quota_upgrade() return that
status, and libzfs reports EAGAIN as a suspended pool.
dmu_objset_upgrade_stop() ignores the result, as it ignored the old
wait.
On master, fstests generic/753 in the eio group hangs with the pool
suspended and z_upgrade blocked here:
[10 lines not shown]
Do not wait forever in spa_vdev_state_exit() on a suspended pool
spa_vdev_state_exit() waits for the txg to sync whenever it is given a
vdev, so that zpool(8) commands are synchronous. If the pool suspends
during that wait, the txg never syncs and the command never returns.
A pool that is already suspended does not get that far.
ZFS_IOC_VDEV_SET_STATE refuses it with EAGAIN before vdev_online()
runs, and zfs_ioc_clear() passes NULL instead of the vdev when
spa_suspended() is true. The hang needs the pool to suspend after
those checks. That happens when the state change's own sync fails,
and when a zpool clear races a new suspend, which generic/753 in the
eio group caught on an encrypted mirror:
zpool D 357s txg_wait_synced <- spa_vdev_state_exit
<- zfs_ioc_clear
txg_sync D 359s
pool SUSPENDED
[24 lines not shown]
Stop DMU_TX_NOWAIT callers spinning on a suspended pool
On a suspended pool, dmu_tx_assign() gives a DMU_TX_WAIT caller EIO
under failmode=continue and blocks it under failmode=wait. A
DMU_TX_NOWAIT caller gets ERESTART in both modes, and every such caller
answers it the same way:
if (error == ERESTART) {
waited = B_TRUE;
dmu_tx_wait(tx);
dmu_tx_abort(tx);
goto top;
}
dmu_tx_assign() sets tx_break_on_suspend for any caller without
DMU_TX_SUSPEND, so this dmu_tx_wait() waits with TXG_WAIT_SUSPEND and
returns as soon as it sees the suspended pool. It returns void, so the
caller retries at once, and the thread spins in the kernel until the
pool resumes. dmu_tx_assign() guards its own retry against this by
[26 lines not shown]
ddt: don't prune every unique entry for a target below one entry
A percentage prune rounds its target down to whole entries. When the
target is zero, for example 1% of a table with fewer than 100 unique
entries, the search for the oldest age bin to keep never runs, the
cutoff stays at the current time, and the walk prunes every unique
entry instead of none.
Return without pruning when the target is zero. dedup_prune_percentage
prunes 1% of three unique entries, which must keep all of them, and
then 100%, which must remove them.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Alexander Motin <alexander.motin at TrueNAS.com>
Signed-off-by: Kamil Monicz <kamil at monicz.dev>
Closes #19264
dsl_scan: resume a scan which suspends at the origin snapshot
dsl_scan_visitbp() checks for suspension before it checks for a hole,
so a scan can suspend even at the root of the empty $ORIGIN snapshot,
which dsl_scan_visit() visits right after the MOS. The caller asserts
that this visit cannot suspend. Debug builds panic in the sync thread.
Without assertions, dsl_scan_visit() goes on to an empty dataset queue:
the suspended visit returned before queueing the snapshots and clones
which descend from the origin. It records that traversal is complete,
and the next TXG finishes the scan without visiting any dataset. A
resilver then retires the new device's missing ranges, and the original
can be detached although the file systems were never copied.
Return with the bookmark intact, as for other dataset visits. The next
TXG resumes at the origin snapshot and queues its descendants.
Fixes: 5815f7ac30e1 ("Fix stalled txg with repeated noop scans")
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Kamil Monicz <kamil at monicz.dev>
Closes #19171
kmodtool: use kABI-baseline naming on RHEL to prevent kmod accumulation
On RHEL, errata kernels within a minor release share a stable kABI.
The existing kmodtool generates a unique kmod package name per exact
kernel version (e.g. kmod-zfs-5.14.0-687.52.1.el9_8), which causes
unbounded kmod package accumulation as new errata kernels are installed
and akmods rebuilds for each one.
Port the kABI-aware naming logic from RPM Fusion's kmodtool:
- Add init_kernel_uname_r_vars() to parse kernel uname -r into
components including the kABI baseline (kernel_uname_r_short).
Example: 5.14.0-687.52.1.el9_8.x86_64 -> 5.14.0-687.el9_8
- On RHEL (%{?rhel}), use kernel_uname_r_short in package names
so all errata kernels within a minor release produce the same
package name (e.g. kmod-zfs-5.14.0-687.el9_8).
- On Fedora (no kABI guarantee), keep the full kernel_uname_r in
[15 lines not shown]
kmodtool: use kABI-baseline naming on RHEL to prevent kmod accumulation
On RHEL, errata kernels within a minor release share a stable kABI.
The existing kmodtool generates a unique kmod package name per exact
kernel version (e.g. kmod-zfs-5.14.0-687.52.1.el9_8), which causes
unbounded kmod package accumulation as new errata kernels are installed
and akmods rebuilds for each one.
Port the kABI-aware naming logic from RPM Fusion's kmodtool:
- Add init_kernel_uname_r_vars() to parse kernel uname -r into
components including the kABI baseline (kernel_uname_r_short).
Example: 5.14.0-687.52.1.el9_8.x86_64 -> 5.14.0-687.el9_8
- On RHEL (%{?rhel}), use kernel_uname_r_short in package names
so all errata kernels within a minor release produce the same
package name (e.g. kmod-zfs-5.14.0-687.el9_8).
- On Fedora (no kABI guarantee), keep the full kernel_uname_r in
[15 lines not shown]
kmodtool: use kABI-baseline naming on RHEL to prevent kmod accumulation
On RHEL, errata kernels within a minor release share a stable kABI.
The existing kmodtool generates a unique kmod package name per exact
kernel version (e.g. kmod-zfs-5.14.0-687.52.1.el9_8), which causes
unbounded kmod package accumulation as new errata kernels are installed
and akmods rebuilds for each one.
Port the kABI-aware naming logic from RPM Fusion's kmodtool:
- Add init_kernel_uname_r_vars() to parse kernel uname -r into
components including the kABI baseline (kernel_uname_r_short).
Example: 5.14.0-687.52.1.el9_8.x86_64 -> 5.14.0-687.el9_8
- On RHEL (%{?rhel}), use kernel_uname_r_short in package names
so all errata kernels within a minor release produce the same
package name (e.g. kmod-zfs-5.14.0-687.el9_8).
- On Fedora (no kABI guarantee), keep the full kernel_uname_r in
[15 lines not shown]
ZIL: avoid deadlock in xattr owner check
zfs_xattr_owner_unlinked() runs with an assigned transaction. A final
zrele() can run inode or vnode cleanup inline. During Linux writeback,
cleanup may wait for the caller's own I_SYNC state. Cleanup may also
open another transaction while the current transaction remains assigned.
The caller cannot return to commit its transaction, so the txg remains
open and cannot sync.
Treat the input znode as borrowed on all platforms. Release only parents
acquired by zfs_zget() with zfs_zrele_async(). This removes the Linux
zhold()/zrele() pair and the platform split. Check the current walk node
in the assertion.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Jaromir Hamala <jaromir.hamala at gmail.com>
Closes #19253
CI: wait for SSH after restarting sshd on FreeBSD
FreeBSD 16-current CI failed while transferring src.txz immediately
after restarting sshd: scp received Connection refused and the build
VM initialization aborted before any tests ran.
Poll SSH readiness for up to thirty attempts before transferring the
archive. Use the existing one-second connection timeout. A server
that remains unavailable still fails at the transfer.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
ZTS: wait for pool suspension and the blocked writer
The ten-poll suspension wait can expire before failed writes suspend
the pool on a busy system. Observing SUSPENDED also does not guarantee
that the background writer has reached dmu_tx_try_assign() yet.
Wait for both pool suspension and an increase in dmu_tx_suspended,
with a shared sixty-poll budget. Sample the counter before starting
the writer so earlier tests cannot satisfy the check. Detect a writer
that exits during the wait and report state and counters on timeout.
Keep the checks that the writer blocks and the pool can be resumed.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
ZTS: use IPv4 loopback explicitly in HTTPS key tests
The Python HTTPS fixture binds an IPv4 socket, while clients resolve
localhost independently and can select IPv6. Use the same explicit
IPv4 loopback address for the server and keylocation URLs. Add an IP
subject alternative name to the test certificate so verification
continues to validate the endpoint.
Debian 13 CI reset every HTTPS key request without logging a request
at the fixture. This change removes address-family ambiguity; the
original reset's precise cause is not established by the CI logs.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
ZTS: retry busy pool export in the zvol replay test
The zvol can remain briefly open for device probing after its ext4
filesystem has been unmounted. Pool export then fails with pool is
busy before the test can import the pool and verify log replay.
Use the existing bounded busy retry for export. Keep the frozen pool,
intent-log inspection, import, and checksum verification unchanged.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
ZTS: sync the resilver finishing txg before checking waiter exit
Pool status reports a finished resilver before its finishing txg has
synced. zpool wait also waits for that txg, whose config and label
writes are subject to the test's injected I/O delay. Starting the
two-second exit grace period at the status change can fail spuriously.
Sync the pool when the activity check observes completion, then check
that the waiter exits within the existing grace period. Keep the
checks for premature exit and nonzero status for both wait commands.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
ZTS: check TRIM rate limits without assuming minimum throughput
Rate limiting caps maximum throughput, but a busy vdev can trim more
slowly. The fixed sleep and cumulative progress checks failed with
32% trimmed when the test expected at least 35%.
Wait for each progress target with a bounded timeout and verify that
it did not arrive faster than the configured limit permits. Check
resumption without specifying a new rate, a higher rate, and complete
trimming at the maximum rate. Retain suspension and completion checks.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
ZTS: wait for each livelist condense before overwriting again
Condensing runs in a background thread and commits in a later txg.
Rapid overwrites can skip condense opportunities while that thread
is busy, leaving seven entries instead of the expected six.
Wait for the expected livelist length after each overwrite, with a
bounded retry and diagnostics on failure. Inspect only the test clone.
Keep the exact entry-count checks and the deactivation checks.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
ZTS: stabilize interior dnode reallocation test
send_realloc_dnode_interior can exhaust all retries without reaching
an interior slot of the freed dnode. In the FreeBSD CI failure, the
freed dnode was object 128 while every retry reused objects 10-17.
Deleting the filler files lets each retry revisit the same lower slots.
Keep filler files between attempts and allow enough allocations to
fill the slots below the freed dnode. Stop at the first interior slot,
then remove the filler files before taking the incremental snapshot.
Preserve the receive and directory comparison checks.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19259
abi: update after umem removal
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Alexander Motin <alexander.motin at TrueNAS.com>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19250
coverity: replace umem with kmem
Manual and untested conversion to match umem removal.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Alexander Motin <alexander.motin at TrueNAS.com>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19250
umem: split out into separate kmem, kmem_cache and vmem
Splits the umem.h inline implementations into separate kmem.[ch],
kmem_cache.[ch] and vmem.h for libspl. Implementations remain the same,
but are now named consistently with kernel SPL and implementation
details hidden.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Alexander Motin <alexander.motin at TrueNAS.com>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19250
umem: convert all calls to kmem
In libspl the kmem_* are simple macro aliases to umem_*, so there's no
functional change here. What we gain is that there's now no ambiguity
over which API to use, and we resolve the assumption baked in a couple
of places that they can be used interchangably.
Changes:
- umem_alloc -> kmem_alloc
- umem_zalloc -> umem_zalloc
- umem_free -> kmem_free
- umem_alloc_aligned -> kmem_alloc_aligned (new kmem_ API)
- umem_free_aligned -> kmem_free_aligned (new kmem_ API)
- UMEM_NOFAIL -> KM_SLEEP
- UMEM_DEFAULT -> KM_NOSLEEP
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Alexander Motin <alexander.motin at TrueNAS.com>
[2 lines not shown]
umem: remove debug, logging and fail callbacks
Never implemented in our umem stubs so nothing has ever called them.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Alexander Motin <alexander.motin at TrueNAS.com>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19250
libzfs: honor literal for vdev fragmentation property
"zpool get -p fragmentation <pool> all-vdevs" printed the value with
a trailing '%', unlike the pool fragmentation property and the vdev
capacity property. Print the bare number when literal is requested.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19222
vdev_prop_get: report real vdev fragmentation
VDEV_PROP_FRAGMENTATION was taken from vd->vdev_stat.vs_fragmentation,
which is never updated: fragmentation is only filled into the copy
returned by vdev_get_stats_ex(). So "zpool get fragmentation <pool>
all-vdevs" always reported 0%.
Report the same values as "zpool list -v": pool fragmentation for the
root vdev, metaslab group fragmentation for top-level concrete vdevs,
and "-" for everything else.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #16459
Closes #19222