Walk snapshot names without holding the objset in list-next ioctl
ZFS_IOC_SNAPSHOT_LIST_NEXT walks the snapnames ZAP, which lives in the
MOS, so it only needs the dataset. It held the objset instead, and for
an unmounted filesystem or an idle volume that means building a complete
objset_t and tearing it down again on every call, once per snapshot
returned.
Hold the pool and the dataset, and move the ZAP walk into
dsl_dataset_snapshot_list_next(), which dmu_snapshot_list_next() now
wraps for its remaining callers. Listing a dataset's snapshots no
longer fails when the dataset's own objset block cannot be read.
Reviewed-by: Rob Norris <rob.norris at truenas.com>
Reviewed-by: Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Ameer Hamza <ameer.hamza at truenas.com>
Closes #19208
zpool: don't print a huge duration when a scan ends before it started
secs_to_dhms() took an unsigned number of seconds, so when the end time
of a scrub, resilver, expansion or condense was earlier than its start
time, for example because the system clock was set back while it ran,
the negative difference wrapped around and zpool status reported
something like "scrub repaired 0B in 213503982334601 days 06:57:38".
Take a signed value and report a negative duration as zero.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #11597
Closes #19212
libzutil: fix the udev_device_get_is_initialized() config check
Since f040a7b0f ("Fix up FIND_SYSTEM_LIBRARY to work with
cross-compiling") configure detects udev_device_get_is_initialized()
with AC_CHECK_FUNCS, which defines HAVE_UDEV_DEVICE_GET_IS_INITIALIZED,
but udev_device_is_ready() still tested the old
HAVE_LIBUDEV_UDEV_DEVICE_GET_IS_INITIALIZED name. So the fallback that
waits for a DEVLINKS property was always used, and a device which udev
gives no links at all was never considered ready: zpool_disk_wait()
ran into its 30 second timeout for every such device, e.g. importing
a pool on brd ramdisks took ~37s instead of ~0.6s.
Use the name configure actually defines.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19213
Related to #10753
dmu_send: don't send uncompressed ARC data as compressed
issue_data_read() first tries to satisfy compressed (-c) and raw (-w)
sends from the ARC with ARC_FLAG_CACHED_ONLY and ZIO_FLAG_RAW_COMPRESS.
The ARC only honors a request for a compressed buffer when the header
itself is compressed, which is never the case for cached blocks with
zfs_compressed_arc_enabled=0. The returned buffer then holds the
logical data, but the WRITE record is still emitted with the block's
compression type and psize, so the stream carries the first psize
bytes of uncompressed data labeled as compressed. The receiver writes
that payload verbatim with a matching checksum, so scrub reports the
pool as healthy while reading the file fails with EIO.
If the cached buffer is not compressed as the block pointer says,
release it and read the block from disk instead.
Add a ZTS test that sends cached lz4 data with -c and -w while the
compressed ARC is disabled and verifies the received files.
[3 lines not shown]
zloop.sh: don't treat every file as a core when unprivileged
When kernel.core_pattern starts with a pipe or a format specifier,
zloop.sh replaces it with "core". An unprivileged user can't write
core_pattern, so the glob stays "*" and core_file() returns the
first file in the current directory. Every iteration is then
archived as a crash and zloop.sh exits 1.
If core_pattern can't be set, say so and skip core file detection.
Crashes are still caught through the ztest exit status. Only
restore core_pattern at exit if it was changed.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: HongseokChoi <xjx061277 at gmail.com>
Closes #19210
ZTS: drop the FreeBSD inherit_001_pos expected-failure mask
inheritance/inherit_001_pos was added to the FreeBSD 'maybe' list in
583e32054 (#11830) for flaky unmount failures tracked in #11829. The
test itself was fixed later in 8792dd24c (#13686), and #11829 is closed,
but the mask stayed and would hide a real regression.
In the last 40 zfs-qemu CI runs (2026-09-27..29) the test passed all
493 times across every builder, including 139 of 139 runs on FreeBSD
14.4-RELEASE, 15.1-RELEASE, 15.1-STABLE and 16.0-CURRENT, with no
failure, kill or skip.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19214
contrib: dracut: install zfs-load-key hook only with systemd
The zfs-load-key.sh pre-mount hook is only meant for systemd-based
initramfs; without systemd, keys are loaded by mount-zfs.sh in the
mount hook. The script decides this at runtime by checking whether
systemctl exists.
However, the dracut shutdown module installs tools like reboot, halt
and poweroff, which on systemd hosts are symlinks to systemctl, so
systemctl ends up in the initramfs even when the boot is not driven
by systemd. The hook then waits for zfs-import.target forever, and
the boot hangs repeatedly printing:
System has not been booted with systemd as init system (PID 1).
Can't operate.
Failed to connect to bus: Host is down
Install the hook only when the systemd dracut module is included, the
same condition used for the other systemd-specific files.
[4 lines not shown]
CI: Split up tests evenly on runner VMs
Our CI spawns two VMs on each github runner, and runs half the
test suite on each. It naively splits up the tests by count, and
doesn't take into account how long the individual tests groups take
to run. This leads to one VM finishing the test suite before
the other. For example, one recent run on Fedora 44:
vm1 03:16:28
vm2 02:50:54
This commit attempts to balance the tests on the VMs by runtime.
It does this by adding a test completion time database to
zfs-tests.sh which is use to portion out the test groups
equally. The database is just a big associative array that
is generated by the new 'make-testdb.sh' helper script.
Just point make-testdb.sh at a test results tarball and
it will generate the new test times database.
[9 lines not shown]
ZTS: Make redundancy tests faster
The redundancy tests can take a very long time. I was able to reduce
the redundancy tests from 7min 10sec to 4min on a local VM with
these changes:
- Reduce zfs_txg_timeout from 5 -> 1 for all redundancy tests. This
speeds up `zpool replace` operations.
- Reduce the vdev sizes in many of the tests.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Tony Hutter <hutter2 at llnl.gov>
Closes #19140
CI: Enable configure caching by default
Enable configure caching by default in the CI. This doesn't
significantly speed up the build but it does ensure this
functionality is always tested.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Glenn Washburn <development at efficientek.com>
Closes #19173
ZTS: increase send_realloc_dnode_interior attempts
The send_realloc_dnode_interior test depends on being able to
reallocate and interior slot of a freed dnode. The test doesn't
have direct control over the object id assignment so there's a
retry loop which attempts to set this up 5 times then gives up.
Increase that threshhold to 10 times and add some diagnostic
information to log what ranges are being checked on each attempt.
Reviewed-by: George Melikov <mail at gmelikov.ru>
Signed-off-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Closes #19112
spl: remote sys/trace_spl.h
Linux kernel only, not included from anywhere since 801d9b4f96.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/mhd.h
Only in libspl, only included from libefi, and nothing from it gets used.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/priv.h
Empty and unused in userspace.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/uuid.h
Only on FreeBSD, and almost identical to include/sys/uuid.h. However,
only used by sys/efi_partition.h, which is specific to libefi, which is
not used on FreeBSD, and especially not in the kernel. Remove stray
includes of sys/efi_partition.h too and there's no more to see.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/wait.h
Only on Linux, just includes a couple of platform includes that don't
appear to be needed directly anyway.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/mode.h
Only on FreeBSD, empty, and never included.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/poll.h
Only in userspace, as a workaround to get poll.h correctly. However,
there's only one place we include it, in ztest, so its more sensible to
just include poll.h directly there, which is what we want anyway.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/stack.h
Only for userspace, and never included.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/callo.h
Only on Linux. CALLOUT_FLAG_ABSOLUTE is the only thing we need from
here, for cv_timedwait_idle_hires(). FreeBSD and userspace have it in
condvar.h, move it there for Linux.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/inttypes.h
Empty on Linux & FreeBSD. In userspace, only included inttypes.h to get
the "core" types, and then unconditionally defined _INT64_TYPE, which
only gates the 64-bit atomics.
So, remove all versions, and the _INT64_TYPE gate. Finally, move the
inttypes.h include up in sys/types.h, since it is now the only source of
the "core" types, and wasn't included early enough to get them.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/user.h
Only on Linux. Left over from before zfs_file_*, should have been
removed in da92d5cbb3.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/lock.h
Only for FreeBSD, shadowing a system header. The two additional defines
are not used anywhere in OpenZFS.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
Wait for the txg-time database write before setting spa_final_txg
spa_export_common() writes the TXG timestamp database into the MOS via
spa_unload_sync_time_logger(), which assigns the write to whatever txg
is open at the time and returns without waiting for it to sync. The
shutdown fence is then computed as spa_last_synced_txg() +
TXG_DEFER_SIZE + 1, so spa_final_dirty_txg() is spa_last_synced_txg + 1,
on the assumption that nothing is dirty beyond that point.
That assumption only holds while the sync pipeline is idle. With
spa_last_synced_txg at L, txg L+1 syncing and L+2 open, the write lands
in L+2 while the fence permits only L+1. When L+2 syncs, the MOS write
allocates space, the metaslab is dirty for L+2, and metaslab_sync()
trips its own check by exactly one txg:
VERIFY3U(txg, <=, spa_final_dirty_txg(spa)) failed (2751 <= 2750)
That VERIFY3U is compiled into production builds, so a real
'zpool export' can panic the same way ztest does.
[19 lines not shown]
spa: number new txgs above every uberblock on disk
After a successful import of an explicitly requested txg the new txgs
were numbered from that txg onwards, while the uberblocks of the
timeline the import discarded were still in the label ring with higher
txg numbers. If the pool was exported before the new timeline overtook
them, the next plain import selected one of theirs, quietly returning to
the discarded state, or failed with an I/O error once the new timeline
had reused its blocks.
The rewind path does not have this problem: a load that follows a failed
one sets spa_last_ubsync_txg to the txg of the newest uberblock, and
spa_first_txg is taken from it. A load of an explicitly requested txg
succeeds on the first attempt, so that value is zero and the first new
txg comes from spa_last_synced_txg() instead.
Record the newest uberblock seen while loading, whatever txg was asked
for, and start the new timeline above it.
[9 lines not shown]
rpm: Generate initial rpm configure cache file from existing cache
The rpmbuild can not use the existing configure cache generated by the
initial configure run because the build/host/target is likely different.
More generally configure's "precious" variables have different values
or states. However, the rest of the initial cache file can be reused.
This is the bulk of the processing in configure anyway. Also, for the
same reason, caches can not be shared between rpmbuilds either.
Generate the initial rpm configure cache file by taking the initial
configure cache and excluding the precious variables. If the rpm
configure already exists, reuse it but remove the undesired variables.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Glenn Washburn <development at efficientek.com>
Closes #19168
rpm: Rename CONFCACHE_FILE to RPM_CONFCACHE
This configure cache file is specific to rpm builds and should be
labeled appropriately.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Glenn Washburn <development at efficientek.com>
Closes #19168
rpm: Enable configure caching for redhat kmod and zfs user packages
These were missed in 0eff7a30b82e (rpm: Add configure caching if used
in main configure run).
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Glenn Washburn <development at efficientek.com>
Closes #19168
Export zpool_valid_proplist from libzfs
zpool_create checks pool properties with zpool_valid_proplist before it
asks the kernel to create the pool. The function is static, so a caller
that wants to check pool properties ahead of time, such as a dry run of
pool creation, has to copy its rules and keep them in step by hand.
Make the function public and move prop_flags_t, the flags it takes, to
libzfs.h. The function itself is unchanged. Update libzfs.abi to match.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Ameer Hamza <ahamza at ixsystems.com>
Signed-off-by: Caleb St. John <30729806+yocalebo at users.noreply.github.com>
Closes #19190
libzfs_core: block delivery of SIGUSR1 in send_worker thread
This fixes a Linux-specific bug.
3a909fe33 (libzfs, libzfs_core: send: always write to pipe, 2022-02-21)
introduced a subtle bug where zfs send -RPv would not print status
information periodically anymore when the stdout is redirected to
something that isn't a pipe (e.g. >/dev/null).
This is because the send_worker thread introduced by this commit
is created with an unmodified signal mask from libzfs. When
zfs_send_space is called it creates this thread. When we request
verbose status information another thread is created which, in
theory, should periodically receive a SIGUSR1 via a POSIX timer.
The main thread blocks USR1 delivery *after* the creation of the
progress thread, leaving the send_worker thread's signal mask
unmodified. The delivery of SIGUSR1 is now random and for at least
Debian Trixie with kernel 6.12.107+deb13-amd64 this causes the
[12 lines not shown]