CI: Enable configure caching by default
Enable configure caching by default in the CI. This doesn't
significantly speed up the build but it does ensure this
functionality is always tested.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Glenn Washburn <development at efficientek.com>
Closes #19173
ZTS: increase send_realloc_dnode_interior attempts
The send_realloc_dnode_interior test depends on being able to
reallocate and interior slot of a freed dnode. The test doesn't
have direct control over the object id assignment so there's a
retry loop which attempts to set this up 5 times then gives up.
Increase that threshhold to 10 times and add some diagnostic
information to log what ranges are being checked on each attempt.
Reviewed-by: George Melikov <mail at gmelikov.ru>
Signed-off-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Closes #19112
spl: remote sys/trace_spl.h
Linux kernel only, not included from anywhere since 801d9b4f96.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/mhd.h
Only in libspl, only included from libefi, and nothing from it gets used.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/priv.h
Empty and unused in userspace.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/uuid.h
Only on FreeBSD, and almost identical to include/sys/uuid.h. However,
only used by sys/efi_partition.h, which is specific to libefi, which is
not used on FreeBSD, and especially not in the kernel. Remove stray
includes of sys/efi_partition.h too and there's no more to see.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/wait.h
Only on Linux, just includes a couple of platform includes that don't
appear to be needed directly anyway.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/mode.h
Only on FreeBSD, empty, and never included.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/poll.h
Only in userspace, as a workaround to get poll.h correctly. However,
there's only one place we include it, in ztest, so its more sensible to
just include poll.h directly there, which is what we want anyway.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/stack.h
Only for userspace, and never included.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/callo.h
Only on Linux. CALLOUT_FLAG_ABSOLUTE is the only thing we need from
here, for cv_timedwait_idle_hires(). FreeBSD and userspace have it in
condvar.h, move it there for Linux.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/inttypes.h
Empty on Linux & FreeBSD. In userspace, only included inttypes.h to get
the "core" types, and then unconditionally defined _INT64_TYPE, which
only gates the 64-bit atomics.
So, remove all versions, and the _INT64_TYPE gate. Finally, move the
inttypes.h include up in sys/types.h, since it is now the only source of
the "core" types, and wasn't included early enough to get them.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/user.h
Only on Linux. Left over from before zfs_file_*, should have been
removed in da92d5cbb3.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
spl: remove sys/lock.h
Only for FreeBSD, shadowing a system header. The two additional defines
are not used anywhere in OpenZFS.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19206
Wait for the txg-time database write before setting spa_final_txg
spa_export_common() writes the TXG timestamp database into the MOS via
spa_unload_sync_time_logger(), which assigns the write to whatever txg
is open at the time and returns without waiting for it to sync. The
shutdown fence is then computed as spa_last_synced_txg() +
TXG_DEFER_SIZE + 1, so spa_final_dirty_txg() is spa_last_synced_txg + 1,
on the assumption that nothing is dirty beyond that point.
That assumption only holds while the sync pipeline is idle. With
spa_last_synced_txg at L, txg L+1 syncing and L+2 open, the write lands
in L+2 while the fence permits only L+1. When L+2 syncs, the MOS write
allocates space, the metaslab is dirty for L+2, and metaslab_sync()
trips its own check by exactly one txg:
VERIFY3U(txg, <=, spa_final_dirty_txg(spa)) failed (2751 <= 2750)
That VERIFY3U is compiled into production builds, so a real
'zpool export' can panic the same way ztest does.
[19 lines not shown]
spa: number new txgs above every uberblock on disk
After a successful import of an explicitly requested txg the new txgs
were numbered from that txg onwards, while the uberblocks of the
timeline the import discarded were still in the label ring with higher
txg numbers. If the pool was exported before the new timeline overtook
them, the next plain import selected one of theirs, quietly returning to
the discarded state, or failed with an I/O error once the new timeline
had reused its blocks.
The rewind path does not have this problem: a load that follows a failed
one sets spa_last_ubsync_txg to the txg of the newest uberblock, and
spa_first_txg is taken from it. A load of an explicitly requested txg
succeeds on the first attempt, so that value is zero and the first new
txg comes from spa_last_synced_txg() instead.
Record the newest uberblock seen while loading, whatever txg was asked
for, and start the new timeline above it.
[9 lines not shown]
rpm: Generate initial rpm configure cache file from existing cache
The rpmbuild can not use the existing configure cache generated by the
initial configure run because the build/host/target is likely different.
More generally configure's "precious" variables have different values
or states. However, the rest of the initial cache file can be reused.
This is the bulk of the processing in configure anyway. Also, for the
same reason, caches can not be shared between rpmbuilds either.
Generate the initial rpm configure cache file by taking the initial
configure cache and excluding the precious variables. If the rpm
configure already exists, reuse it but remove the undesired variables.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Glenn Washburn <development at efficientek.com>
Closes #19168
rpm: Rename CONFCACHE_FILE to RPM_CONFCACHE
This configure cache file is specific to rpm builds and should be
labeled appropriately.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Glenn Washburn <development at efficientek.com>
Closes #19168
rpm: Enable configure caching for redhat kmod and zfs user packages
These were missed in 0eff7a30b82e (rpm: Add configure caching if used
in main configure run).
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Glenn Washburn <development at efficientek.com>
Closes #19168
Export zpool_valid_proplist from libzfs
zpool_create checks pool properties with zpool_valid_proplist before it
asks the kernel to create the pool. The function is static, so a caller
that wants to check pool properties ahead of time, such as a dry run of
pool creation, has to copy its rules and keep them in step by hand.
Make the function public and move prop_flags_t, the flags it takes, to
libzfs.h. The function itself is unchanged. Update libzfs.abi to match.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Ameer Hamza <ahamza at ixsystems.com>
Signed-off-by: Caleb St. John <30729806+yocalebo at users.noreply.github.com>
Closes #19190
libzfs_core: block delivery of SIGUSR1 in send_worker thread
This fixes a Linux-specific bug.
3a909fe33 (libzfs, libzfs_core: send: always write to pipe, 2022-02-21)
introduced a subtle bug where zfs send -RPv would not print status
information periodically anymore when the stdout is redirected to
something that isn't a pipe (e.g. >/dev/null).
This is because the send_worker thread introduced by this commit
is created with an unmodified signal mask from libzfs. When
zfs_send_space is called it creates this thread. When we request
verbose status information another thread is created which, in
theory, should periodically receive a SIGUSR1 via a POSIX timer.
The main thread blocks USR1 delivery *after* the creation of the
progress thread, leaving the send_worker thread's signal mask
unmodified. The delivery of SIGUSR1 is now random and for at least
Debian Trixie with kernel 6.12.107+deb13-amd64 this causes the
[12 lines not shown]
userspace: remove makedev detection and cleanup headers
makedev() is defined in sys/sysmacros.h on Linux, and sys/types.h on
FreeBSD. However, nothing in any of our userspace code actually uses it,
so all the supporting edifice isn't actually used.
Worse though, the SPL sys/sysmacros.h defines a lot more in userspace
than in kernel, and since sys/types.h would pull it in directly to keep
up the appearance that makedev() is in sys/types.h, it's not always
clear where some of those things in sysmacros.h are coming from, or
_should_ come from.
So clean this up. Remove the the makedev config tests, the empty mkdev.h
header, and the inclusion of sysmacros.h from types.h. For those things
that actually needed some of those definitions, include sysmacros.h
directly.
Sponsored-by: TrueNAS
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Rob Norris <rob.norris at truenas.com>
Closes #19204
zdb: Lower eviction timeouts on prefetched data
The ARC will not evict prefetched data that is too new. When zdb
is doing a lot of reads in a short amount of time, it can lead to
the ARC ballooning past zfs_arc_max, since arc_evict() cannot find
anything to evict (due to the ARC containing mostly prefetched
data). For example, zdb was observed consuming ~1.4GB mem in the
pool_checkpoint tests, even though zfs_arc_max was set to 256MB.
To get around this, allow the prefetched data to be evicted after
only 50ms instead of the normal 1 to 6 seconds. This both
reduces zfs maximum memory usage and allows the zpool_checkpoint
tests to run to completion faster.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Tony Hutter <hutter2 at llnl.gov>
Closes #19191
don't send spill or oversized blocks as WRITE_EMBEDDED
send_do_embed() treated every embedded BP as a DRR_WRITE_EMBEDDED
candidate. Spill blocks (DMU_OT_SA) have to go as DRR_SPILL, and an
embedded block larger than SPA_OLD_MAXBLOCKSIZE cannot be represented
on the receiver unless the stream carries large_blocks: dump_dnode()
clamps the object's block size to 128K in the DRR_OBJECT record, so
the receive then rejects the DRR_WRITE_EMBEDDED.
Skip those cases in send_do_embed() so do_dump() emits a spill record
or splits the block into SPA_OLD_MAXBLOCKSIZE chunks instead.
On receive, accept only BP_EMBEDDED_TYPE_DATA for DRR_WRITE_EMBEDDED.
Add ZTS send_embedded_spill_block (canned pool with an embedded
spill) and send_embedded_large_block (1M zstd block, send with and
without -L).
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Kamil Monicz <kamil at monicz.dev>
[2 lines not shown]
deb: Use release number generated by ZFS for native builds
This allows for a dynamic release part of the version string as is done
for rpm and non-native deb package builds. Importantly, for git builds
this allows having the git commit in the package version string.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Glenn Washburn <development at efficientek.com>
Closes #19186
zpool-prefetch: Fix document description typo
Verb should be imperative mood matching rest of manual.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Alexander Ziaee <ziaee at FreeBSD.org>
Closes #19184
zfs_domount: fix vfs_t double-free on root setup failure
When zfs_root() or d_make_root() fails, zfs_domount() calls
zfs_umount(), which frees zfsvfs->z_vfs via zfsvfs_free(). After
the vfs_t lifetime was matched to fs_context, that vfs_t is still
owned by the caller in fc->fs_private. zpl_get_tree() then returns
without clearing fs_private, and put_fs_context() frees it again.
Detach z_vfs before zfs_umount() so the caller retains ownership,
matching the other zfs_domount() error paths.
This is a follow-up to the vfs_t lifetime change in #18377.
Reviewed-by: Rob Norris <rob.norris at truenas.com>
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Wale Zhang <wale.zhang.ftd at gmail.com>
Closes #19167
Linux: read the snapshot creation time from its bonus buffer
When a '.zfs/snapshot/<name>' entry is looked up by name (stat, open, or
a path walk through it; listing the directory does not do this),
zfsctl_inode_lookup() reads the snapshot's creation time for the
entry's btime through dsl_dataset_hold_obj(). That instantiates the
whole in-core dataset (dsl_dir hold, deadlists, neighbour references,
fsid uniqueness) and tears it down again on release, and it accounts
for most of a dentry-cold lookup. The value is a field of the
snapshot's dsl_dataset_phys_t, so read it from the bonus buffer the way
the fsid is read, through a helper both callers share.
Same value, same locking, same fallback on error. A warm lookup is
unchanged; a cold one loses the dataset instantiation, roughly an order
of magnitude on the systems it was measured on.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Rob Norris <rob.norris at truenas.com>
Signed-off-by: Ameer Hamza <ameer.hamza at truenas.com>
Closes #19174
Linux: report a snapshot's fsid from its .zfs/snapshot entry
statfs() on a control directory inode ('.zfs', '.zfs/snapshot', or an
unmounted snapshot entry reached without the automount) fails with EIO
because zfs_statvfs() verifies the znode's SA handle, which these inodes
lack. Handle them explicitly: '.zfs' and '.zfs/snapshot' report the
containing filesystem, and a '.zfs/snapshot/<name>' entry reports the
fsid of its snapshot, read from the dataset's bonus buffer (the objset id
is encoded in the entry's inode number) without instantiating the
dataset, so a probe costs a few microseconds.
This identifies a snapshot without mounting it: open the entry with
O_PATH, which does not trigger the automount, and fstatfs() it. The
value is what the snapshot's superblock reports once mounted, and path
based statfs() still automounts as before. An NFS server resolving a
file handle for an unmounted snapshot (expired, or imported on another
node) can probe entries instead of mounting every snapshot until one
matches. Add the statfs_nomount helper, the snapdir_statfs_fsid test
and a zfsconcepts(7) note.
[5 lines not shown]
zpool: accept more redundant special and dedup vdevs
The replication check required a special or dedup vdev to tolerate
exactly as many device failures as the normal vdevs in the pool. A
3-way mirror special vdev on a raidz1 pool, or on a 2-way mirror pool,
was rejected as a mismatch, and the only way past it was -f, which
also overrides unrelated checks.
Accept a special or dedup vdev which tolerates at least as many
failures as the normal vdevs of a redundant pool. Special and dedup
vdevs are now compared with the normal vdevs rather than with whichever
vdev happens to precede them, so normal vdevs are still required to
match each other. Less redundant special or dedup vdevs, and redundant
ones added to a non-redundant pool, are still rejected.
As a side effect, a pool that already has a more redundant special
vdev is now considered consistent, so later additions to it are
checked instead of being skipped.
[4 lines not shown]