OpenZFS/src 4542a31 — include/sys vdev_raidz.h, module/zfs vdev_raidz.c zio.c

zio: prevent DDT extension across RAIDZ expansion

A flat DDT entry can gain DVAs after its initial write, but stores one
physical birth. If a pre-expansion entry receives a post-expansion RAIDZ
DVA, the BP cannot represent both width epochs. RAIDZ then reads the new
copy at the old geometry; repair can map beyond its allocation.

Before extending a live phys, compare each newly allocated RAIDZ top's
logical width at the phys birth with its width at the allocation txg.
Refuse only a mismatched epoch and retry without dedup. This produces a
fresh BP whose DVAs share one birth. Same-epoch extensions, non-RAIDZ
tops, new entries, and traditional DDT tables keep their behavior.

Add raidz_expand_008_pos, which checks the block pointers written:
extension before expansion, refusal across a 3-to-4 disk expansion for
a 128K block and for a 16K block, whose allocation has the same size
at both widths, a clean scrub, and extension of a new entry after the
expansion. The last case distinguishes this exact check from a
permanent pool-wide gate.

    [5 lines not shown]
DeltaFile
+118-0tests/zfs-tests/tests/functional/raidz/raidz_expand_008_pos.ksh
+65-13module/zfs/zio.c
+38-3module/zfs/vdev_raidz.c
+1-1tests/runfiles/common.run
+1-0tests/zfs-tests/tests/Makefile.am
+1-0include/sys/vdev_raidz.h
+224-176 files

OpenZFS/src b7e1ce0 — module/zfs zio.c

zio: consume DDT fallback errors before reporting

DDT extension deliberately returns EAGAIN when adding a copy would need
a mixed gang and non-gang block pointer. The logical write classified
that error as a non-dedup retry only after normal error reporting. It
therefore recorded a persistent data error and posted an ereport for a
hole BP even though the plain-write retry succeeded.

Classify an allocating dedup write's EAGAIN after transforms. Disable
dedup, clear the deliberate error, and request reexecution before vdev
statistics and ereports are processed. The retry terminates because
dedup changes monotonically to false. A backend EAGAIN receives at most
one non-dedup retry before ordinary error handling.

Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Alexander Motin <alexander.motin at TrueNAS.com>
Signed-off-by: Kamil Monicz <kamil at monicz.dev>
Closes #18826
DeltaFile
+17-9module/zfs/zio.c
+17-91 files

OpenZFS/src be9919f — module/zfs zio.c

zio: balance allocation depth on refused DDT extensions

A DDT child that would mix gang and non-gang DVAs is refused at READY.
This occurs before top-vdev children can release its metaslab-group
queue-depth holds. The refusal path decremented those holds even though
only a throttled asynchronous allocation created them. Synchronous and
unthrottled writes could therefore underflow the queue depth.

Gate the per-DVA release on ZIO_FLAG_ALLOC_THROTTLED. Keep it at this
refusal site: these holds use the child's size and tag, unlike gang
header allocations that can reach generic READY error handling with
different accounting keys.

Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Alexander Motin <alexander.motin at TrueNAS.com>
Signed-off-by: Kamil Monicz <kamil at monicz.dev>
Closes #18826
DeltaFile
+7-5module/zfs/zio.c
+7-51 files

OpenZFS/src ddb521c — include/sys ddt.h, module/zfs ddt.c zio.c

zio: fix chained DDT extension rollback

A later write requesting more copies can adopt an in-flight FDT
extension as its child, forming a chain of leads. The chain shared one
rollback snapshot. A completed child copied the live phys after newer
children could already have extended it. Failure could retain an
uncommitted DVA or discard a committed one. The first completion could
also clear a snapshot still needed by a later lead.

Treat the saved phys as a rolling rollback point. Advance it only with
the completing DDT child's private BP. Clear it only after the newest
lead is gone. Failure then restores exactly the DVAs committed by older
children, independent of newer READY completions.

Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Alexander Motin <alexander.motin at TrueNAS.com>
Signed-off-by: Kamil Monicz <kamil at monicz.dev>
Closes #18826
DeltaFile
+8-6module/zfs/zio.c
+5-4include/sys/ddt.h
+2-2module/zfs/ddt.c
+15-123 files

OpenZFS/src 75a7995 — module/zfs vdev_raidz.c, tests/runfiles common.run

vdev_raidz: contain maps that exceed their DVA allocation

RAIDZ selects a block's stripe width from the block pointer's physical
birth. Fast-dedup extension in affected releases could append a DVA
after an expansion while retaining an older birth. This makes the
birth-selected map larger than the DVA's owned allocation. A repair of
that copy can then write into the next allocation.

For dedup BPs read at a historical width, find the dispatched DVA by
top vdev and offset. Compare its owned ASIZE with the map footprint.
Reject an oversized or unmatched direct map with EIO before building
columns, so neither reads nor repair writes can cross the allocation.
Debug builds assert the routing invariant; release builds still fail
closed if it is violated.

The check is limited to dedup BPs because FDT extension is the released
producer of this state. Equal-ASIZE mixed epochs may still fail checksum
verification, but they cannot overrun the allocation. An affected copy
is reported as a RAIDZ read error while other DVA copies remain usable.

    [12 lines not shown]
DeltaFile
+101-8module/zfs/vdev_raidz.c
+95-0tests/zfs-tests/tests/functional/raidz/raidz_dedup_overrun.ksh
+5-0tests/zfs-tests/tests/Makefile.am
+2-1tests/runfiles/common.run
+0-0tests/zfs-tests/tests/functional/raidz/blockfiles/raidz_dedup_overrun-2.dat.bz2
+0-0tests/zfs-tests/tests/functional/raidz/blockfiles/raidz_dedup_overrun-3.dat.bz2
+203-92 files not shown
+203-98 files

OpenZFS/src 7d3e648 — module/zfs dsl_dir.c, tests/runfiles common.run

Allow quota to be changed even when over quota

ZFS allows datasets to use slightly more space than their quota,
typically one write operation totaling less than 1MB. Once in this state
of using more space than the quota, subsequent writes will fail with
ENOSPC. The quota can also be changed to more than the space used (with
`zfs set quota=...`). The quota can not be tightened (decreased) such
that the dataset is in an over-quota state, which is a design decision
dating back to before version 1.

The problem is that when in the state of using more space than the
quota, the quota can not be set to its current value, or relaxed
(increased) to a value that is less than the space used. Since these
operations don't make anything worse in terms of being over-quota, they
should be allowed. This enables `zfs set quota=` to be idempotent.

Original-patch-by: Matthew Ahrens <matt at mahrens.org>
External-issue: https://www.illumos.org/issues/18300
External-issue: https://www.illumos.org/issues/18332

    [6 lines not shown]
DeltaFile
+101-0tests/zfs-tests/tests/functional/quota/quota_007_pos.ksh
+3-0module/zfs/dsl_dir.c
+1-1tests/runfiles/common.run
+1-0tests/zfs-tests/tests/Makefile.am
+106-14 files

OpenZFS/src 98527e0 — module/os/linux/zfs zfs_znode_os.c

Linux: Sleep instead of spinning when zfs_zget() races eviction

When igrab() fails because the VFS is evicting the inode, zfs_zget()
drops its locks, calls cond_resched() and retries.  cond_resched()
only yields when a reschedule is already pending, so every lookup of
an inode that is waiting in a long dispose_list() spins at full speed,
taking and dropping the znode hold locks and allocating a znode_hold_t
on each pass.  With enough of them the thread doing the eviction is
starved of both CPU and those locks, and the eviction stalls.  On a
24-thread NAS serving ~180 rsync clients this turned into a livelock
with zero pool I/O that needed a power cycle.

Sleep for one tick before retrying instead, as xfs_iget() does when it
races the same VFS teardown.  The retry loop is otherwise unchanged and
no locks are held while sleeping.

Co-Authored-By: Claude Opus 5.5 <noreply at anthropic.com>
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Karl Schulze <karl at taniustech.com>

    [2 lines not shown]
DeltaFile
+7-2module/os/linux/zfs/zfs_znode_os.c
+7-21 files

OpenZFS/src e95c552 — module/os/linux/zfs zfs_vfsops.c

Linux: Feed the superblock shrinker in batches from zfs_prune()

zfs_prune() hands the whole ARC prune request to a single
super_cache_scan() call.  When arc_evict() asks for hundreds of
thousands or millions of objects, prune_icache_sb() isolates up to
that many unused inodes, marks every one of them I_FREEING, and only
then evicts them one at a time from dispose_list().  For as long as it
takes to work through that list, a lookup of any of those inodes fails
igrab() in zfs_zget() and has to retry.

The kernel's own reclaim never builds such a list: do_shrink_slab()
calls scan_objects() in chunks of shrinker->batch objects (1024 for
superblocks).  Do the same in zfs_prune(), so that at most one batch
of inodes is I_FREEING at a time.  The total number of objects scanned
per prune request is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply at anthropic.com>
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Karl Schulze <karl at taniustech.com>

    [2 lines not shown]
DeltaFile
+34-4module/os/linux/zfs/zfs_vfsops.c
+34-41 files

OpenZFS/src 623c20f — .github/workflows README.md zfs-precheck.yml, .github/workflows/scripts failfirst-detect.py failfirst-tests.sh

CI: verify "failing test, then fix" PRs

A bug fix is easiest to review and keep fixed when the PR first adds a
test that shows the bug, then fixes it.  Add a workflow that checks
this pattern: for each commit that only changes tests/ and is followed
by a commit changing code, build that commit in a QEMU VM and check its
new or changed tests fail (or crash or hang the kernel), then build the
PR head and check the same tests pass without kernel errors.

The workflow reuses the qemu-* scripts to set up, build and boot one
test VM per build, runs each test on its own with zfs-tests.sh -t, and
restarts the VM after a crash.  PRs without a test commit finish after
the detect job.

The detector gives each job 15 minutes plus 30 per test (at most 240),
and the per-test watchdog in failfirst-tests.sh drops from 40 to 25
minutes: zfs-tests.sh -t stops a test after 600 seconds, and the rest
covers the group's setup and cleanup.  A job with one test times out
after 45 minutes.

    [3 lines not shown]
DeltaFile
+247-0.github/workflows/scripts/failfirst-tests.sh
+201-0.github/workflows/scripts/failfirst-detect.py
+199-0.github/workflows/zfs-precheck.yml
+29-0.github/workflows/README.md
+676-04 files

OpenZFS/src eb011f2 — module/zfs arc.c

arc: fix race between arc_release() and arc_read_done()

Consider the following scenario:

1. arc_release() is called on hdr with one buf, but which has
   IO_IN_PROGRESS (reading more raw data in encypted pool while
   keeping decrypted data in buf).
2. arc_release() moves hdr to anon state and discards its identity.
3. arc_read_done() is called, adds the 2nd buf to hdr, increasing
   b_refcnt to 2.

Now we have hdr in anon state with two bufs and without identity.

Or here's a racing scenario:

1. arc_release() checked that hdr is not in anon state, but before
   taking hash_lock
2. arc_read_done() takes hash_lock, moves hdr to anon state, in
   case of an error.

    [28 lines not shown]
DeltaFile
+41-13module/zfs/arc.c
+41-131 files

OpenZFS/src cece147 — module/zfs arc.c

L2ARC: Reorder header destruction for in-flight L2 writes

With multiple L2ARC devices, headers can be destroyed asynchronously
(e.g., during zpool sync) while L2_WRITING is set. The original code
destroyed L2HDR before L1HDR, causing ABDs to lose their device
association (b_l2hdr.b_dev) when arc_hdr_free_abd() is called.

This caused ABDs to be added to the global free-on-write list without
device information. When any L2ARC device completed its write and
attempted to free these orphaned ABDs, it would panic on
ASSERT(!list_link_active(&abd->abd_gang_link)) because the ABD was
still part of another device's vdev_queue I/O aggregation gang.

Fix by extending l2ad_mtx lock scope to cover L1HDR destruction and
reordering to destroy L1HDR before L2HDR when L2_WRITING is set. This
ensures arc_hdr_free_abd() can access b_l2hdr.b_dev to properly tag
ABDs with their device for deferred cleanup.

Reviewed-by: Alexander Motin <alexander.motin at TrueNAS.com>

    [5 lines not shown]
DeltaFile
+37-18module/zfs/arc.c
+37-181 files

OpenZFS/src acf6f13 — module/zfs arc.c

L2ARC: Preserve L2HDR in arc_release() for in-flight writes

When arc_release() is called on a header with a single buffer and
L2_WRITING set, the L2HDR must be preserved for ABD cleanup (similar
to the arc_hdr_destroy() case). If we destroy the L2HDR here, later
arc_write() will allocate a new ABD and call arc_hdr_free_abd(),
which needs b_l2hdr.b_dev to properly defer ABD cleanup, causing
VERIFY(HDR_HAS_L2HDR(hdr)) to fail.

Allocate a new header for the buffer in the single_buf_l2writing
case (single buffer + L2_WRITING), leaving the original header with
L2HDR intact. The original header becomes an "orphan" (no buffers, no
b_pabd) but retains device association for ABD cleanup when
l2arc_write_done() completes.

The shared buffer case (HDR_SHARED_DATA) is excluded because L2ARC
makes its own transformed copy via l2arc_apply_transforms(), so the
original ABD is not used by the L2 write. The header can be safely
reused without allocating a new one.

    [14 lines not shown]
DeltaFile
+51-37module/zfs/arc.c
+51-371 files

OpenZFS/src 269d147 — module/os/freebsd/zfs zfs_vnops_os.c

FreeBSD: do not clear dirty bits outside a sub-block write

page_busy() shrinks the written range to DEV_BSIZE boundaries.  A write
that starts and ends inside one block leaves nbytes at -DEV_BSIZE, and
vm_page_clear_dirty() then clears that block and every one above it.  A
page dirtied through mmap goes clean and the store is never written.

Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Alexander Motin <alexander.motin at TrueNAS.com>
Signed-off-by: Nick Price <nprice at FreeBSD.org>
Closes #19268
DeltaFile
+1-1module/os/freebsd/zfs/zfs_vnops_os.c
+1-11 files

OpenZFS/src 9b99fe4 — tests/zfs-tests/tests/functional/log_spacemap log_spacemap_flushall.ksh

ZTS: compare pre-condense log spacemaps

The flush can write new log spacemaps in a later TXG, so total
smp_length can grow even when older logs were flushed. Compare only
entries present before condense, then sync before reading their
lengths.

Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19273
DeltaFile
+21-6tests/zfs-tests/tests/functional/log_spacemap/log_spacemap_flushall.ksh
+21-61 files

OpenZFS/src 4c4a65a — tests/zfs-tests/tests/functional/deadman deadman_zio.ksh

ZTS: restore timeout for deadman negative check

The test lowers the deadman timeout to 5 seconds for its positive
check. Restore the default before the short-delay negative check so
a long QEMU scheduling pause does not look like another deadman
failure.

Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19273
DeltaFile
+7-3tests/zfs-tests/tests/functional/deadman/deadman_zio.ksh
+7-31 files

OpenZFS/src 1d73c26 — tests/zfs-tests/cmd ctime.c

ZTS: retry interrupted ctime test sleeps

The ctime helper assumes sleep(2) always completes. A signal can
end it early, leaving an empty timestamp window. Retry until two
wall-clock seconds have passed before running the operation.

Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19273
DeltaFile
+8-2tests/zfs-tests/cmd/ctime.c
+8-21 files

OpenZFS/src e61fe25 — cmd/zpool zpool_main.c

Open only the named pool when zpool checks a pool name

While tracing zpool get on one pool I noticed that it opens every
imported pool before the one named on the command line.

zpool get, zpool set and zpool iostat call is_pool() to tell a pool
name from a vdev name. It opened every imported pool to compare names,
so a command on a small pool also paid for every other pool on the
system. is_pool() now opens only the pool it is asked about.

On a system with three pools, one of them with 1200 disks, zpool get
all on the other two pools went from about 930 ms to 15 ms and 22 ms.
On the 1200 disk pool itself it stayed at about 1.8 seconds, because
the named pool is still opened twice, once for this check and once by
the command.

Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Ameer Hamza <ameer.hamza at truenas.com>
Reviewed-by: Rob Norris <rob.norris at truenas.com>
Signed-off-by: Caleb St. John <yocalebo at gmail.com>
Closes #19271
DeltaFile
+14-12cmd/zpool/zpool_main.c
+14-121 files

OpenZFS/src 5d1ed14 — module/zfs zil.c

Bail out of zil_create() when the pool suspends

zil_create() waits with txg_wait_synced(), a void wrapper around a
plain wait that cannot report a suspend. When the pool suspends while
it waits, the fsync() that got there does not return until the pool
resumes.

The rest of the commit path already handles this. zil_commit_flags()
and zil_commit_writer_stall() wait with TXG_WAIT_SUSPEND, call
zil_crash() on ESHUTDOWN and hand EIO back to the waiters.
zil_create() and zil_commit_activate_saxattr_feature() were left on the
old wait, so a suspend caught in either one still hangs.

Use the suspend-aware wait in all three places and return the failure.
zil_process_commit_list() already has a NULL-lwb path that signals the
nolwb waiters with the error, and its comment already says an ESHUTDOWN
there means zil_crash() was called, so the failure lands somewhere that
expects it.


    [11 lines not shown]
DeltaFile
+34-11module/zfs/zil.c
+34-111 files

OpenZFS/src b38e7a4 — module/zfs dmu_objset.c

Bail out of the objset upgrade when the pool suspends

The two objset upgrade callbacks end with txg_wait_synced(), a void
wrapper around a plain wait that cannot report a suspend. When the
pool suspends during that wait, the upgrade taskq thread stays blocked
until the pool resumes. So does dmu_objset_disown(), which calls
dmu_objset_upgrade_stop(). That waits for a running upgrade task to
finish and then waits for a txg itself.

Give all three waits TXG_WAIT_SUSPEND. The callbacks return EAGAIN,
which dmu_objset_upgrade_task_cb() records in os_upgrade_status.
zfs_ioc_userspace_upgrade() and zfs_ioc_id_quota_upgrade() return that
status, and libzfs reports EAGAIN as a suspended pool.
dmu_objset_upgrade_stop() ignores the result, as it ignored the old
wait.

On master, fstests generic/753 in the eio group hangs with the pool
suspended and z_upgrade blocked here:


    [10 lines not shown]
DeltaFile
+8-3module/zfs/dmu_objset.c
+8-31 files

OpenZFS/src b098c60 — module/zfs spa_misc.c, tests/runfiles common.run

Do not wait forever in spa_vdev_state_exit() on a suspended pool

spa_vdev_state_exit() waits for the txg to sync whenever it is given a
vdev, so that zpool(8) commands are synchronous. If the pool suspends
during that wait, the txg never syncs and the command never returns.

A pool that is already suspended does not get that far.
ZFS_IOC_VDEV_SET_STATE refuses it with EAGAIN before vdev_online()
runs, and zfs_ioc_clear() passes NULL instead of the vdev when
spa_suspended() is true. The hang needs the pool to suspend after
those checks. That happens when the state change's own sync fails,
and when a zpool clear races a new suspend, which generic/753 in the
eio group caught on an encrypted mirror:

    zpool     D  357s   txg_wait_synced <- spa_vdev_state_exit
                        <- zfs_ioc_clear
    txg_sync  D  359s
    pool      SUSPENDED


    [24 lines not shown]
DeltaFile
+103-0tests/zfs-tests/tests/functional/failmode/failmode_vdev_state.ksh
+10-2module/zfs/spa_misc.c
+2-1tests/runfiles/common.run
+1-0tests/zfs-tests/tests/Makefile.am
+116-34 files

OpenZFS/src 2a96397 — module/zfs dmu_tx.c, tests/runfiles linux.run

Stop DMU_TX_NOWAIT callers spinning on a suspended pool

On a suspended pool, dmu_tx_assign() gives a DMU_TX_WAIT caller EIO
under failmode=continue and blocks it under failmode=wait. A
DMU_TX_NOWAIT caller gets ERESTART in both modes, and every such caller
answers it the same way:

        if (error == ERESTART) {
                waited = B_TRUE;
                dmu_tx_wait(tx);
                dmu_tx_abort(tx);
                goto top;
        }

dmu_tx_assign() sets tx_break_on_suspend for any caller without
DMU_TX_SUSPEND, so this dmu_tx_wait() waits with TXG_WAIT_SUSPEND and
returns as soon as it sees the suspended pool. It returns void, so the
caller retries at once, and the thread spins in the kernel until the
pool resumes. dmu_tx_assign() guards its own retry against this by

    [26 lines not shown]
DeltaFile
+79-0tests/zfs-tests/tests/functional/failmode/failmode_nowait_wait.ksh
+14-5module/zfs/dmu_tx.c
+4-0tests/runfiles/linux.run
+1-0tests/zfs-tests/tests/Makefile.am
+98-54 files

OpenZFS/src 7391add — module/zfs ddt.c, tests/runfiles common.run

ddt: don't prune every unique entry for a target below one entry

A percentage prune rounds its target down to whole entries. When the
target is zero, for example 1% of a table with fewer than 100 unique
entries, the search for the oldest age bin to keep never runs, the
cutoff stays at the current time, and the walk prunes every unique
entry instead of none.

Return without pruning when the target is zero. dedup_prune_percentage
prunes 1% of three unique entries, which must keep all of them, and
then 100%, which must remove them.

Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Alexander Motin <alexander.motin at TrueNAS.com>
Signed-off-by: Kamil Monicz <kamil at monicz.dev>
Closes #19264
DeltaFile
+79-0tests/zfs-tests/tests/functional/dedup/dedup_prune_percentage.ksh
+2-1tests/runfiles/common.run
+3-0module/zfs/ddt.c
+1-0tests/zfs-tests/tests/Makefile.am
+85-14 files

OpenZFS/src 3db7dde — module/zfs dsl_scan.c

dsl_scan: resume a scan which suspends at the origin snapshot

dsl_scan_visitbp() checks for suspension before it checks for a hole,
so a scan can suspend even at the root of the empty $ORIGIN snapshot,
which dsl_scan_visit() visits right after the MOS. The caller asserts
that this visit cannot suspend. Debug builds panic in the sync thread.
Without assertions, dsl_scan_visit() goes on to an empty dataset queue:
the suspended visit returned before queueing the snapshots and clones
which descend from the origin. It records that traversal is complete,
and the next TXG finishes the scan without visiting any dataset. A
resilver then retires the new device's missing ranges, and the original
can be detached although the file systems were never copied.

Return with the bookmark intact, as for other dataset visits. The next
TXG resumes at the origin snapshot and queues its descendants.

Fixes: 5815f7ac30e1 ("Fix stalled txg with repeated noop scans")
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Kamil Monicz <kamil at monicz.dev>
Closes #19171
DeltaFile
+2-1module/zfs/dsl_scan.c
+2-11 files

OpenZFS/src 65ea88f — scripts kmodtool

kmodtool: use kABI-baseline naming on RHEL to prevent kmod accumulation

On RHEL, errata kernels within a minor release share a stable kABI.
The existing kmodtool generates a unique kmod package name per exact
kernel version (e.g. kmod-zfs-5.14.0-687.52.1.el9_8), which causes
unbounded kmod package accumulation as new errata kernels are installed
and akmods rebuilds for each one.

Port the kABI-aware naming logic from RPM Fusion's kmodtool:

- Add init_kernel_uname_r_vars() to parse kernel uname -r into
  components including the kABI baseline (kernel_uname_r_short).
  Example: 5.14.0-687.52.1.el9_8.x86_64 -> 5.14.0-687.el9_8

- On RHEL (%{?rhel}), use kernel_uname_r_short in package names
  so all errata kernels within a minor release produce the same
  package name (e.g. kmod-zfs-5.14.0-687.el9_8).

- On Fedora (no kABI guarantee), keep the full kernel_uname_r in

    [15 lines not shown]
DeltaFile
+197-42scripts/kmodtool
+197-421 files

OpenZFS/src 9212b44 — scripts kmodtool

kmodtool: use kABI-baseline naming on RHEL to prevent kmod accumulation

On RHEL, errata kernels within a minor release share a stable kABI.
The existing kmodtool generates a unique kmod package name per exact
kernel version (e.g. kmod-zfs-5.14.0-687.52.1.el9_8), which causes
unbounded kmod package accumulation as new errata kernels are installed
and akmods rebuilds for each one.

Port the kABI-aware naming logic from RPM Fusion's kmodtool:

- Add init_kernel_uname_r_vars() to parse kernel uname -r into
  components including the kABI baseline (kernel_uname_r_short).
  Example: 5.14.0-687.52.1.el9_8.x86_64 -> 5.14.0-687.el9_8

- On RHEL (%{?rhel}), use kernel_uname_r_short in package names
  so all errata kernels within a minor release produce the same
  package name (e.g. kmod-zfs-5.14.0-687.el9_8).

- On Fedora (no kABI guarantee), keep the full kernel_uname_r in

    [15 lines not shown]
DeltaFile
+197-42scripts/kmodtool
+197-421 files

OpenZFS/src 4160224 — scripts kmodtool

kmodtool: use kABI-baseline naming on RHEL to prevent kmod accumulation

On RHEL, errata kernels within a minor release share a stable kABI.
The existing kmodtool generates a unique kmod package name per exact
kernel version (e.g. kmod-zfs-5.14.0-687.52.1.el9_8), which causes
unbounded kmod package accumulation as new errata kernels are installed
and akmods rebuilds for each one.

Port the kABI-aware naming logic from RPM Fusion's kmodtool:

- Add init_kernel_uname_r_vars() to parse kernel uname -r into
  components including the kABI baseline (kernel_uname_r_short).
  Example: 5.14.0-687.52.1.el9_8.x86_64 -> 5.14.0-687.el9_8

- On RHEL (%{?rhel}), use kernel_uname_r_short in package names
  so all errata kernels within a minor release produce the same
  package name (e.g. kmod-zfs-5.14.0-687.el9_8).

- On Fedora (no kABI guarantee), keep the full kernel_uname_r in

    [15 lines not shown]
DeltaFile
+198-42scripts/kmodtool
+198-421 files

OpenZFS/src e9e24ce — module/zfs zfs_log.c

ZIL: avoid deadlock in xattr owner check

zfs_xattr_owner_unlinked() runs with an assigned transaction. A final
zrele() can run inode or vnode cleanup inline. During Linux writeback,
cleanup may wait for the caller's own I_SYNC state. Cleanup may also
open another transaction while the current transaction remains assigned.
The caller cannot return to commit its transaction, so the txg remains
open and cannot sync.

Treat the input znode as borrowed on all platforms. Release only parents
acquired by zfs_zget() with zfs_zrele_async(). This removes the Linux
zhold()/zrele() pair and the platform split. Check the current walk node
in the assertion.

Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Jaromir Hamala <jaromir.hamala at gmail.com>
Closes #19253
DeltaFile
+10-29module/zfs/zfs_log.c
+10-291 files

OpenZFS/src ddba56d — .github/workflows/scripts qemu-2-start.sh

CI: wait for SSH after restarting sshd on FreeBSD

FreeBSD 16-current CI failed while transferring src.txz immediately
after restarting sshd: scp received Connection refused and the build
VM initialization aborted before any tests ran.

Poll SSH readiness for up to thirty attempts before transferring the
archive. Use the existing one-second connection timeout. A server
that remains unavailable still fails at the transfer.

Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
DeltaFile
+5-0.github/workflows/scripts/qemu-2-start.sh
+5-01 files

OpenZFS/src e35cf9b — tests/zfs-tests/tests/functional/failmode failmode_dmu_tx_wait.ksh

ZTS: wait for pool suspension and the blocked writer

The ten-poll suspension wait can expire before failed writes suspend
the pool on a busy system. Observing SUSPENDED also does not guarantee
that the background writer has reached dmu_tx_try_assign() yet.

Wait for both pool suspension and an increase in dmu_tx_suspended,
with a shared sixty-poll budget. Sample the counter before starting
the writer so earlier tests cannot satisfy the check. Detect a writer
that exits during the wait and report state and counters on timeout.
Keep the checks that the writer blocks and the pool can be resumed.

Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
DeltaFile
+14-6tests/zfs-tests/tests/functional/failmode/failmode_dmu_tx_wait.ksh
+14-61 files

OpenZFS/src a1c5513 — tests/zfs-tests/tests/functional/cli_root/zfs_load-key zfs_load-key.cfg zfs_load-key_common.kshlib

ZTS: use IPv4 loopback explicitly in HTTPS key tests

The Python HTTPS fixture binds an IPv4 socket, while clients resolve
localhost independently and can select IPv6. Use the same explicit
IPv4 loopback address for the server and keylocation URLs. Add an IP
subject alternative name to the test certificate so verification
continues to validate the endpoint.

Debian 13 CI reset every HTTPS key request without logging a request
at the fixture. This change removes address-family ambiguity; the
original reset's precise cause is not established by the CI logs.

Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
DeltaFile
+4-1tests/zfs-tests/tests/functional/cli_root/zfs_load-key/zfs_load-key_common.kshlib
+2-1tests/zfs-tests/tests/functional/cli_root/zfs_load-key/zfs_load-key.cfg
+6-22 files