Add middleware support for LIO ALUA HA
Wire up the middleware side of LIO ALUA high-availability: load
lio_ha.ko with per-node addresses on service start, manage ALUA
state across failover events, clean up STANDBY configfs on pool
export, and add pre-flight validation that targets have static
initiator ACLs before ALUA can be enabled.
For each target, create a portal-less phantom TPG carrying the peer
node's controller group so that a single RTPG response from any
connected port lists both ALUA groups. Write tpgt_N/rtpi explicitly
before enable so that relative target port IDs in RTPG match the
tag formula (portal.tag on Node A, portal.tag + 32000 on Node B)
rather than being auto-assigned sequentially by the kernel.
ALUA group states are driven by role and ha_state:
MASTER + synced local=OPTIMIZED remote=NONOPTIMIZED
MASTER + connected local=OPTIMIZED remote=TRANSITIONING
[4 lines not shown]
Gate SMB block cloning and Veeam shares on SMB_BLOCK_CLONING
This commit adds changes to replace the SMB_FASTPATH and SMB_VEEAM license features with a single SMB_BLOCK_CLONING feature, which gates the ZFS block cloning and integrity streams smb.conf parameters as well as the VEEAM_REPOSITORY_SHARE purpose.
NAS-136915 / 27.0.0-BETA.1 / Consolidate the disk identifier logic and fall back to udev when sysfs has no serial (#19856)
Middleware built a disk's identifier in two places, once over udev's
view of the disk for the disk table and the disk list, and once over
sysfs for everything that goes through the disk object (reporting,
SMART, temperatures, SED, the pool validator), so the two could disagree
about the same disk. Both now go through one implementation. Two private
methods nothing called any more are removed; there is no public API
change.
On the hardware we ship this changes no value. udev and sysfs report the
same serial and lunid for SAS, SATA and NVMe disks, and the disk object
still reads sysfs, so the disk-stats tick costs what it did.
One behavior is new. A disk that reports no serial in sysfs, which is a
usb-storage disk (those expose no VPD pages at all) or a SCSI device
that implements VPD page 0x83 but not page 0x80, now takes the serial
udev resolved for it instead of falling through to its partition table
or its device name. For such a disk:
[11 lines not shown]
NAS-144098 / 26.0.0-RC.1 / Send license issued_at in TNC heartbeat (by sonicaj) (#19888)
This PR adds changes to expose the license's issue date as `issued_at`
on `truenas.license.info` and send it to TNC in the heartbeat payload.
- v2 licenses: the `issued_at` string exactly as recorded in the
license, with no parsing or reformatting.
- Legacy licenses: the support contract start date as `YYYY-MM-DD`.
Legacy blobs carry no issue timestamp, so no time or timezone is
invented.
- System-generated (hardware-only) records, no license, or an invalid
license: `null`.
In the heartbeat, `issued_at` is gated exactly like `license_id` (only
for issued licenses), so TNC can tell a legacy value apart by the
`legacy_` prefix on `license_id`.
Original PR: https://github.com/truenas/middleware/pull/19886
Co-authored-by: Waqar Ahmed <waqarahmedjoyia at live.com>
NAS-144098 / 26.0.0 / Send license issued_at in TNC heartbeat (by sonicaj) (#19887)
This PR adds changes to expose the license's issue date as `issued_at`
on `truenas.license.info` and send it to TNC in the heartbeat payload.
- v2 licenses: the `issued_at` string exactly as recorded in the
license, with no parsing or reformatting.
- Legacy licenses: the support contract start date as `YYYY-MM-DD`.
Legacy blobs carry no issue timestamp, so no time or timezone is
invented.
- System-generated (hardware-only) records, no license, or an invalid
license: `null`.
In the heartbeat, `issued_at` is gated exactly like `license_id` (only
for issued licenses), so TNC can tell a legacy value apart by the
`legacy_` prefix on `license_id`.
Original PR: https://github.com/truenas/middleware/pull/19886
Co-authored-by: Waqar Ahmed <waqarahmedjoyia at live.com>
Send license issued_at in TNC heartbeat
This commit adds changes to expose the license's issue date as `issued_at` on `truenas.license.info` and send it to TNC in the heartbeat. v2 licenses pass the minted timestamp through verbatim, legacy licenses report their contract start date, and system-generated records report null.
NAS-144098 / 27.0.0-BETA.1 / Send license issued_at in TNC heartbeat (#19886)
This PR adds changes to expose the license's issue date as `issued_at`
on `truenas.license.info` and send it to TNC in the heartbeat payload.
- v2 licenses: the `issued_at` string exactly as recorded in the
license, with no parsing or reformatting.
- Legacy licenses: the support contract start date as `YYYY-MM-DD`.
Legacy blobs carry no issue timestamp, so no time or timezone is
invented.
- System-generated (hardware-only) records, no license, or an invalid
license: `null`.
In the heartbeat, `issued_at` is gated exactly like `license_id` (only
for issued licenses), so TNC can tell a legacy value apart by the
`legacy_` prefix on `license_id`.
Send license issued_at in TNC heartbeat
This commit adds changes to expose the license's issue date as `issued_at` on `truenas.license.info` and send it to TNC in the heartbeat. v2 licenses pass the minted timestamp through verbatim, legacy licenses report their contract start date, and system-generated records report null.
Gate SMB fast path on the Veeam entitlement
This commit adds changes to drop the separate SMB_FASTPATH license feature and gate the ZFS block cloning / integrity streams smb.conf parameters on SMB_VEEAM instead, since Veeam Fast Clone depends on them and the two were never meant to be licensed independently.
Pass a ZFSResourceSetArgsData to zfs.resource.set_impl
## Problem
`zfs.resource.set_impl` took a path plus loose `properties`/`user_properties`/`inherit`/`bypass` arguments, unlike `create_impl` which takes a model. Internal callers handed raw dicts straight to libzfs, so a bad property name or value only surfaced as a libzfs error part way through the write.
## Solution
- **Typed input**: `set_impl` takes a `ZFSResourceSetArgsData`, the same model the public `zfs.resource.set` uses, so every caller's values are validated when the model is built, before anything is written. `touched_names` and `changed_fields` read the model too; the event, audit and quota/refquota `none` handling are unchanged.
- **Private fields**: `ZFSResourceSetArgsData` gains a `Private` `bypass` (honoured by both `set` and `set_impl`), and `ZFSResourceSetProperties` gains a `Private` `volthreading` for the iSCSI and NVMe-oF zvol handling. API callers still cannot supply either.
- **Callers**: every internal caller builds the model inline. `pool.dataset.update_impl` keeps its dict arguments since the HA peer calls it by name, and coerces them into the model. The tier special_small_blocks constants are ints now that they go through the typed model.
- **System dataset encryption**: when the pool root is passphrase-encrypted, `setup_datasets` added `encryption=off` to the comparison for existing system datasets too, so an encrypted child would have been sent an encryption change that ZFS refuses and setup would fail. Encryption is left out of the update comparison; creation still sets it.
Fix the checksum choices result, the extent readonly event test and the create mount-path check
## Problem
- **checksum choices**: `pool.dataset.checksum_choices` now returns the checksums `zfs.resource` accepts, which include OFF, but its result model still had the fixed fields master's list had without OFF and forbids extras. Every call failed result validation: integration tests raise, production logs a serialization warning, and older API clients fail outright because the adapter validates against the current model first.
- **extent readonly test**: `test_readonly_on_an_extent_zvol_syncs_the_extent` expected two CHANGED events for one readonly change. The second came from the iSCSI resync writing readonly back onto the zvol, which the resync no longer does, so the test failed.
- **create mount-path check**: create refused a filesystem when `/mnt/<path>` already existed, while the share ACL is applied at the path the filesystem really mounts at. Under an ancestor with a non-default mountpoint the two differ, so an occupied real mount path went unnoticed until the mount failed after the resource was created, and a stray `/mnt/<path>` refused a create that would never mount there.
## Solution
- **checksum choices**: added OFF to the v27 result model. Create and update have always accepted `checksum=OFF`, so the list now matches what they take. Older API versions keep their models, and the adapter drops OFF for those clients as before.
- **extent readonly test**: it now expects the single CHANGED event from the caller's own set.
- **create mount-path check**: it tests the nearest existing ancestor's mountpoint joined with the rest of the path, the same path the ACL step uses. With default mountpoints that is still `/mnt/<path>`.
Measure volume reservation headroom the way master did
## Problem
The reservation headroom checks drifted from master in ways that refused requests ZFS and master accept:
- **set**: headroom was measured as `requested - current refreservation` on every refreservation change. The kernel only charges the part of a reservation above the data the volume already holds, so converting a sparse volume that has data to thick was refused even with `force_size`. For example, an 80G sparse volume with 60G written and 50G free needs 20G for `refreservation=80G`, but was refused as needing 80G.
- **create**: a new volume was measured against the nearest ancestor's `available - usedbyrefreservation`. When that ancestor has a refquota, `available` is already clamped to it, so the refreservation was counted twice. Under a nearly empty parent with refquota=10G and refreservation=10G on a 5T pool, every thick volume create was refused, where master allowed up to 8G.
## Solution
- **set**: the volume headroom check runs only when `volsize` changes, as `pool.dataset.update` did on master. A refreservation-only change is left to ZFS, which refuses what it cannot back.
- **create**: the check measures against the nearest existing ancestor's `available`, as both `pool.dataset.create` and `zfs.resource.create` did on master.
NAS-144061 / 26.0.0 / Add bucket recovery APIs (by anodos325) (#19884)
There are various situations in which we can have orphaned buckets.
Examples include:
* bucket is deleted and admin wants to reinstate.
* the NAS is disaster recovery instance and needs to activate buckets
after switching datasets to read-write.
This is facilited by inserting a .truenas_s3/config_backup.json file
inside each bucket (daemon-owned path) and restoring the DB row / S3
config with some user-provided overrides if required.
Two new API endpoints are added:
* sharing.s3.recoverable_buckets lists mounted datasets containing
orphaned buckets.
* sharing.s3.recover basically takes a list of datasets and rebuilds
[4 lines not shown]
NAS-144061 / 27.0.0-BETA.1 / Add bucket recovery APIs (#19873)
There are various situations in which we can have orphaned buckets.
Examples include:
* bucket is deleted and admin wants to reinstate.
* the NAS is disaster recovery instance and needs to activate buckets
after switching datasets to read-write.
This is facilited by inserting a .truenas_s3/config_backup.json file
inside each bucket (daemon-owned path) and restoring the DB row / S3
config with some user-provided overrides if required.
Two new API endpoints are added:
* sharing.s3.recoverable_buckets lists mounted datasets containing
orphaned buckets.
* sharing.s3.recover basically takes a list of datasets and rebuilds
[3 lines not shown]
NAS-144061 / 27.0.0-BETA.1 / Add bucket recovery APIs (#19873)
There are various situations in which we can have orphaned buckets.
Examples include:
* bucket is deleted and admin wants to reinstate.
* the NAS is disaster recovery instance and needs to activate buckets
after switching datasets to read-write.
This is facilited by inserting a .truenas_s3/config_backup.json file
inside each bucket (daemon-owned path) and restoring the DB row / S3
config with some user-provided overrides if required.
Two new API endpoints are added:
* sharing.s3.recoverable_buckets lists mounted datasets containing
orphaned buckets.
* sharing.s3.recover basically takes a list of datasets and rebuilds
bucket configuration from backups for them.
Read ZFS properties raw in the S3 bucket backup and recovery
`mounted` parses to a bool and a literal `none` to None, so no backup was
ever written and every recovery was refused. The backup test also leaked
its row into the rest of the module.
NAS-144083 / 26.0.0 / remove unconditional call-remote in system update (by anodos325) (#19883)
system.general.update called failover.call_remote unconditionally when
ds_auth or ui_certificate changed, stalling for the connect timeout and
logging a spurious standby warning on non-HA systems.
Original PR: https://github.com/truenas/middleware/pull/19882
Co-authored-by: Andrew Walker <andrew.walker at truenas.com>
Gate standby pam regeneration on failover.licensed
system.general.update called failover.call_remote unconditionally when ds_auth or ui_certificate changed, stalling for the connect timeout and logging a spurious standby warning on non-HA systems.
(cherry picked from commit d31e09427cd917b0b0df4fa8b494d661cb1e2764)
NAS-144083 / 27.0.0-BETA.1 / remove unconditional call-remote in system update (#19882)
system.general.update called failover.call_remote unconditionally when
ds_auth or ui_certificate changed, stalling for the connect timeout and
logging a spurious standby warning on non-HA systems.
Gate standby pam regeneration on failover.licensed
system.general.update called failover.call_remote unconditionally when ds_auth or ui_certificate changed, stalling for the connect timeout and logging a spurious standby warning on non-HA systems.
Keep a created resource whose key cannot be recorded
## Problem
`zfs.resource.create` destroyed the resource it had just created when its encryption key could not be written to the database. That only happens when the middleware database itself is failing, where the next operations fail too, and master never rolled back here. The rollback's destroy also queued a late ZFS event whose handler deletes key rows by name, so a quick retry under the same name could have its freshly written key row removed and lose the dataset's key at the next reboot or failover.
## Solution
Let the key-record failure propagate and keep the resource, as master did. The key is still recorded when the mount fails, so a mount failure no longer loses a generated key, and the docstring and tests are updated to match.