callout: do not retry a try-lock callout sooner than a tick
When softclock_call_cc() fails to acquire the lock of a CALLOUT_TRYLOCK
callout, it reschedules the callout half its precision after
cc_lastscan and halves the precision. Repeated failures shrink the
delay toward zero, and once the precision reaches 1 the callout is due
immediately: the timer fires again at once and softclock retries the
lock in a tight loop for as long as the lock is held.
If the lock owner runs on the callout's CPU and no other CPU is idle,
the softclock thread preempts it on every attempt, starving the thread
it is waiting on. On an 8-CPU arm64 VM, a test module holding the lock
saw 760,000 attempts per second, each with its own timer interrupt, and
progressed at 38% of its normal rate. On a 4-core amd64 system under
loopback TCP load, a netisr thread holding an inpcb lock made no
progress for 12 minutes while the TCP timer callout was retried 830,000
times per second.
Keep the half-precision retry, but never schedule it less than one
[20 lines not shown]
stand: Load dynamic system attribute offsets
zfs_sa_load looks up all the system attribute offsets and stores them in
the mountpoint.
Sponsored by: Netflix
Differential Revision: https://reviews.freebsd.org/D60264
stand: Implement zfs_dnode_readlink in terms of zfs_dnode_sa_lookup
Get the link offset using the zfs_dnode_sa_lookup helper now.
Recently, the symbolic links we rely on in the boot loader have stopped
working.
Prior to OpenZFS commit e90badec11d3 ("Inherit the project ID for every
object type", Matt Turner, 2026-08-14), symbolic link information was
written at a fixed offset in the SA data. Since that commit, the
inherited PROJIDs mean that all pools with quota enabled have started
writing symbolic links with a new, non-fixed offset. Old symbolic links
remained unchanged, but new ones were written with a different
offset. At work, we have all these things: rewritten BEs, quotas, and a
dependence on symbolic links in our boot path.
This came in on 2026-08-24 OpenZFS merge (22649d4dba73). This was 12
hours after stab week for August, so we didn't hit this until the
September stab week. Since the new kernel has to write links at the new
[4 lines not shown]
stand: update zfs_dnode_stat to use zfs_dnode_sa_lookup
Find the SA values with the zfs_dnode_sa_lookup and read out the
relevant bits for the stat buffer.
Sponsored by: Netflix
Differential Revision: https://reviews.freebsd.org/D60267
stand: Lookup specific SA value in a dnode
zfs_dnode_sa_lookup will look in the bonus part of the dnode for the
requested SA values, and fall back to the spill as if it's not there.
Sponsored by: Netflix
Differential Revision: https://reviews.freebsd.org/D60266
stand: Lookup the offsets for this SA bundle
Compute the offset for the data for this bundle and the requested data
type.
Sponsored by: Netflix
Differential Revision: https://reviews.freebsd.org/D60265
stand: zfs_dnode_stat move from spa to mount argument
When reading dynamic system attributes, we'll need the mount argument
since we can no longer hard-code the offsets. Adjust zfs_dnode_stat to
take it.
Sponsored by: Netflix
Differential Revision: https://reviews.freebsd.org/D60261
stand: zfs_dnode_readlink move from spa to mount argument
For the dynamic system attributes, we'll need the mount argument. Adjust
zfs_dnode_readlink to take that argument.
Sponsored by: Netflix
Differential Revision: https://reviews.freebsd.org/D60262
stand: Remove const from zfs_lookup's zfsmount argument
Dynamic system attribute layout can require modifications to the mount
structure on lookup. Drop the const to allow that.
Sponsored by: Netflix
Differential Revision: https://reviews.freebsd.org/D60260
EC2: Use stream-optimized VMDK format
Once enabled on all of the branches, this will reduce bandwidth
consumption from uploading weekly snapshot builds from ~500 GB to
~100 GB, as well as significantly speeding up the process.
MFC after: 2 months
Sponsored by: Amazon
EC2: Pass --vmdk to bsdec2-image-upload if needed
Starting with version 1.5.0, bsdec2-image-upload supports uploading
stream-optimized VMDK files. If we build images in that format, we
need to upload them appropriately.
MFC after: 2 weeks
Sponsored by: Amazon
mkimg: Add support for stream-optimized VMDK
FreeBSD VM/cloud images tend to be highly compressible: The generic VM
images compress roughly 3.5:1, and cloud images typically even more
since they have significant unused space in their virtual disks. This
generally doesn't matter for users who can download compressed release
images and extract them locally, but for EC2 in particular relying on
uncompressed image formats is painful: We now upload over 500 GB/week
to AWS.
This commit adds the "stream-optimized" version of the VMDK format,
which compresses each 64 kB "grain" individually (and omits grains
which are all zeroes). Compared to uncompressed formats, this can
produce much smaller images; experiments indicate that using this
format for EC2 image uploads will reduce the weekly traffic to under
100 GB/week.
Co-authored-by: Claude Opus 5.5
MFC after: 2 weeks
[2 lines not shown]
__FreeBSD_version: Bump for mkimg -f vmdks
Also, adjust bootstrap tools code in Makefile.inc1 to reflect this; in
addition to rebuilding mkimg if it is out of date, we now need to build
libz, which was not previously needed by mkimg.
Reviewed by: kevans
MFC after: 2 weeks
Sponsored by: Amazon
Differential Revision: https://reviews.freebsd.org/D60153
mkimg: Add tests for vmdks format
This is the 'stream-optimized' version of the VMDK format.
MFC after: 2 weeks
Sponsored by: Amazon
Differential Revision: https://reviews.freebsd.org/D60154
mkimg: Make vmdk's desc_fmt slightly more generic
This will allow it to be reused in upcoming work.
Suggested by: Claude Opus 5.5
MFC after: 2 weeks
Sponsored by: Amazon
Differential Revision: https://reviews.freebsd.org/D60150
mkimg: Add image_buffer_region
This is like image_copyout_region, but copies into a buffer rather than
writing out to a file; it will be used by future support for compressed
images (and potentially other circumstances where image data must be
manipulated before being written out to disk).
Reviewed by: jrm
No objection from: Christos Komis (author)
MFC after: 2 weeks
Sponsored by: Google LLC (GSoC 2025)
Differential Revision: https://reviews.freebsd.org/D60149
mkimg: Avoid leaking a page of mmap
The image_file_map function adjusts the provided file offset to be
page-aligned, with a resulting increase in the size of the mapped
region; the increased size needs to be used when unmapping as well.
Reported by: Claude Opus 5.5
Fixes: baf4abfc39b2 ("Allow building mkimg as cross-tool")
MFC after: 2 weeks
Sponsored by: Amazon
Differential Revision: https://reviews.freebsd.org/D60148
mkimg: Check for realloc failure
If realloc fails, clean up and return ENOMEM; don't just copy data
into NULL.
Reviewed by: emaste
Reported by: Claude Opus 5.5
MFC after: 2 weeks
Sponsored by: Amazon
Differential Revision: https://reviews.freebsd.org/D60147