ixgbe: Report SR-IOV VF status
Expose cached VF configuration, policy, and runtime state through the
iflib VF status method. Include access or trunk VLAN mode, the queue
count selected by the current virtualization mode, negotiated mailbox
API, whether traffic is enabled, and the MDD-blocked and quarantine
state.
The query runs under the iflib context lock and does not issue mailbox
requests or read hardware registers.
igb: Report SR-IOV VF status
Expose the cached per-VF configuration through the iflib VF status
method. Report mailbox handshake state, MAC address, access or trunk
VLAN mode, hardware queue count, administrator policy, and MDD blocking
state without issuing mailbox requests or reading hardware registers.
rtnetlink: Report SR-IOV VF status
Honor RTEXT_FILTER_VF on RTM_GETLINK requests and expose the versioned
SR-IOV VF status through typed nested FreeBSD attributes. Report
IFLA_NUM_VF with a successful requested query and preserve per-provider
errors in the status container.
Map the common nvlist schema to native integer, boolean, string, and
binary attributes. Carry namespaced driver extensions as packed
versioned nvlists so adding a driver-specific field does not expand the
common netlink ABI.
Add SNL parsers, parser verification, a constructed nested-status test,
and an RTM_GETLINK test for an interface without SR-IOV support.
Document the query contract and every attribute.
ifconfig: Use nvlist to report SR-IOV VF status
Replace the records with a versioned nvlist transported through struct
ifreq, following SIOCGIFCAPNV. The network stack now packs and copies
results, supports bounded retry for larger results, and handles native
and 32-bit callers centrally. Drivers only populate a kernel nvlist
while their state is locked.
Define optional common fields for identity, configuration and handshake
state, VLAN policy, queue resources, runtime blocks, PF link state, and
namespaced driver extensions. Document the extension and versioning
contract and require providers to omit values they cannot observe.
Improve the ixl provider to track its mailbox handshake and report the
expanded common policy. Render the expanded status as grouped output
under ifconfig -v.
iovctl: Report SR-IOV status
Add -L to query the generic packed-nvlist IOV_GET_STATUS interface.
Report PF enable state and configured and total VF counts. For each VF,
print its PCI address, newbus attachment, bound driver, and ppt state.
Retry size negotiation if the topology changes between ioctls and reject
malformed or incompatible status records.
Keep NIC-specific operational state in ifconfig -v; iovctl owns the
device-neutral PCI topology and applies to any SR-IOV device class.
Relnotes: yes
pci: Add SR-IOV status reporting
Add a generic packed-nvlist status query to each /dev/iov/<PF>
control device. Report the live VF Enable state, configured and total
VF counts, and one record for each configured VF.
Each VF record contains its PF-local index, computed PCI location,
newbus attachment state, attached driver, and ppt binding. Construct
records for hardware VFs whose newbus child is absent so attachment
failures remain visible.
Version the extensible schema in sys/iov.h. Use fixed-width request
fields so the ioctl command and layout are identical for 32-bit callers.
Serialize the topology snapshot with Giant, then pack and copy it after
releasing Giant.
libifconfig: Add an SR-IOV VF status query
Provide a public helper which retrieves, unpacks, and validates the
versioned VF status nvlist. Validate the required VF indices and the
shape and version of driver-specific extension namespaces while allowing
unknown optional fields.
The ioctl argument is not copied back when the command returns EFBIG.
Start with a practical buffer and grow it geometrically rather than
relying on the required length being observable.
Use the helper in ifconfig so other consumers share the same transport
and validation behavior.
ifconfig: Add SR-IOV VF status output
- Adds SR-IOV VF status to the existing ifconfig "-v" output
- Adds ioctl command for reporting VF status info from drivers
- Adds support to iflib for drivers to handle this new ioctl
- Add support for ioctl in ixl(4)
Signed-off-by: Eric Joyner <erj at freebsd.org>
Relnotes: yes
Differential Revision: https://reviews.freebsd.org/D19647
igbv: Do not replay VLANs while stopped
iflib clears IFF_DRV_RUNNING before the driver stop callback but leaves
IFF_DRV_OACTIVE set. Consequently, an already queued admin task can
run after the VF reset. If that task consumes a pending timer sample,
it can retry failed VLAN mailbox operations and restore PF filters for
the stopped VF.
Continue sampling statistics, but only run the VLAN retry worker while
the interface is running.
vfs_mountroot: unmute console in interactive prompt
If boot_mute is set the system appears to hang during the mountroot
prompt. Temporarily unmute the console so the prompt is visible.
Reviewed by: kib
MFC after: 1 week
Differential Revision: https://reviews.freebsd.org/D58549
(cherry picked from commit e96f1cbd690e68594fc8812de634f43c6711aa97)
ixv: Negotiate VF queue-set limits
ixv uses one queue set on 82599 and X540 VFs and assumes two on
X550-family VFs. The PF reports the queues assigned to each VF with
GET_QUEUES after mailbox API 1.1 negotiation.
Query the PF during attach. Bound symmetric iflib queue sets by the PF
grant and available MSI-X data vectors. Retain one queue set per data
vector: ixgbe VFs expose at most three vectors and one is reserved for
the mailbox. The hardware permits each pool to use a subset of its RSS
queues, so a two-queue ceiling is valid when the PF assigns four.
This enables the second data vector on 82599 and X540 while avoiding an
assumed second queue when an X550-family VF is granted only one. Keep
the existing family limits if the mailbox is unavailable or the PF uses
an older API.
MFC after: 2 weeks
ix/ixv: Match Tx writeback thresholds to iflib
PTHRESH controls when the device prefetches transmit descriptors,
HTHRESH controls how many host descriptors must be ready, and WTHRESH
controls completion writeback batching.
iflib places RS on selected descriptors and reclaims through those
checkpoints. The data sheets require WTHRESH to be zero when software
uses RS. Clear WTHRESH while retaining the established PTHRESH 32 and
HTHRESH 1 fetch policy.
This also follows DPDK in pairing sparse RS descriptors with
WTHRESH zero. DPDK defaults to 32/0/0, while Linux ixgbevf uses
32/1/8. The 32/1/0 setting preserves FreeBSD's prefetch policy and the
data-sheet requirement that HTHRESH be nonzero when PTHRESH is used.
MFC after: 2 weeks
igc: Correct descriptor control programming
The transmit-ring setup was copied from the e1000 path. On I225
and I226, bits 22 through 24 are reserved and bit 25 enables the
queue; it is not a legacy low-water threshold. Correct the field
masks, remove the nonapplicable legacy definitions, and program only
defined fields.
Use PTHRESH=8 and HTHRESH=1. Keep WTHRESH at zero so the hardware
honors sparse RS descriptors issued by iflib. Linux and DPDK use a
writeback threshold of 16, but request status on every packet. A
nonzero threshold makes hardware ignore individual RS bits and is
unsuitable for the iflib completion model.
The receive-ring setup likewise used a magic mask that left bit 20
of the five-bit WTHRESH field untouched. Define the receive threshold
fields and replace them exactly before installing the established
PTHRESH=8, HTHRESH=8, WTHRESH=4 policy.
MFC after: 2 weeks
igb: Program Rx descriptor thresholds by family
82576 specification-update erratum 26 says MSI-X EITR expiration can
fail to trigger receive descriptor writeback. A WTHRESH above one can
therefore leave received packets invisible until the threshold fills.
The shared threshold macros selected policy by enum ordering, so an
82576 VF fell into the generic WTHRESH=4 case. VFs always use MSI-X
and require the same WTHRESH=1 workaround as the PF.
Use PTHRESH=8 for 82575 and 82576 PFs and VFs, matching DPDK and the
current Linux PF driver. The legacy FreeBSD PF and Linux igbvf value
of 16 thrashes limited descriptor cache; no specification or erratum
requires it. Retain the i354 PTHRESH=12 exception.
Enumerate every supported igb PF and VF MAC type so each receives its
intended policy. Also clear every threshold bit before installing the
new values. The old mask retained the high WTHRESH bit, and 82575
uses six-bit fields while later controllers use five-bit fields.
[2 lines not shown]
igb: Match Tx descriptor control to iflib
iflib requests transmit completion status only on selected descriptors.
Program a zero writeback threshold so igb hardware honors those sparse
RS bits instead of writing back every descriptor in threshold-sized
batches.
Use the existing family specific prefetch threshold: eight descriptors
on most controllers and 20 on I354, with a host threshold of one. These
values match the Intel-derived Linux and DPDK drivers. Their nonzero
writeback settings are not appropriate here because those drivers set
RS on every packet.
A zero writeback threshold also avoids depending on interrupt timer
flushes affected by 82576 specification update erratum 26. Remove the
old IGB_TX_WTHRESH macro as well. It has had no callers since the iflib
conversion, so its 82575 conditional no longer implements any policy.
MFC after: 2 weeks
e1000: Correct Rx descriptor threshold programming
Jumbo receive tuning on integrated controllers enabled PTHRESH without
a nonzero HTHRESH, contrary to the hardware programming requirements.
It also covered only the integrated MAC generations present when the
workaround was added. Enumerate every jumbo-capable ICH and PCH type
and program PTHRESH=3 with HTHRESH=1. Linux fixed the same HTHRESH
omission in b701cacdbcfb.
The 82574 path combined threshold values with the reset values using
bitwise OR. Requesting WTHRESH=4 while the reset value was one thus
programmed five. Clear the complete threshold fields before installing
the established PTHRESH=32, HTHRESH=4, WTHRESH=4 descriptor-granularity
policy.
MFC after: 2 weeks
e1000: Program Tx descriptor control by family
TXDCTL programming is family dependent. 82543 erratum 35 and
82544 erratum 20 require WTHRESH to remain zero; a nonzero value
can corrupt descriptor writebacks and hang the controller. Leave all
descriptor-control thresholds at their reset values on 82542, 82543,
and 82544.
On the remaining em controllers, retain the established PTHRESH=31,
HTHRESH=1, WTHRESH=1, and descriptor granularity policy. Several
legacy specification updates identify full descriptor writeback as a
workaround for transmit descriptor-queue errata.
TXDCTL bit 22 is also family dependent. It is COUNT_DESC on the
82571 family and 80003ES2LAN. Intel shared initialization explicitly
sets raw bit 22 on both transmit queues of every supported ICH/PCH
generation, although the integrated public documentation marks it
reserved. Preserve that required setting when iflib programs the
thresholds, as DPDK does. Clearing it caused a persistent I219
[12 lines not shown]
ixv: Advertise SCTP checksum offload
The shared ixgbe transmit path already creates SCTP context
descriptors, and the hardware exposes the same checksum capability to
VFs. Advertise it through iflib as the PF driver does.
MFC after: 2 weeks
ixv: Remove unused loader tunables
The flow_control and hdr_split variables have never been read. VF
flow control is controlled by the PF, while implementing header split
would require receive-path support that ixv does not provide.
MFC after: 2 weeks
ixgbe: Reject Flow Director with SR-IOV
The iflib Flow Director path does not assign filters using the
absolute queue and pool identifiers required by SR-IOV. Reject the
combination during preflight validation rather than allowing an
unsupported configuration to alter the PF receive path.
The loader tunable is fixed before VFs can be created, so validation
also prevents the reverse ordering of this combination.
MFC after: 2 weeks
amd_iommu: Honor disabled interrupt remapping
Do not instantiate an interrupt-remapping context for a unit whose IRTE
support is disabled. In that mode the caller must retain the ordinary
interrupt path.
Reviewed by: kib
MFC after: 2 weeks
Differential Revision: https://reviews.freebsd.org/D58725
iflib: Add sysctl stat for TX watchdog reset events
iflib counts resets initiated by its transmit watchdog in 69c3e0de01c1.
Export the counter in the per-device iflib sysctl tree so every
driver provides the diagnostic without a driver callback or duplicate
storage.
A watchdog reset does not establish how many packets failed. It can
recover a hardware stall involving several queued packets or a missed
completion involving no packet loss. Stop adding one output error per
watchdog event in em(4), igb(4), and igc(4).
Remove the redundant driver counters and move the diagnostic to
dev.<driver>.<unit>.iflib.tx_watchdog_events.
MFC after: 1 month
Relnotes: yes
pfsync: handle large MTU pfsync interfaces
pfsync packets were allocated with m_get2(), which can't return packets
larger than MJUMPAGESIZE. As a result 9k MTU pfsync interfaces simply didn't work.
Use m_get3(), which can allocate sufficiently large mbufs.
Extend the pfsync:bulk test case to provoke this problem.
PR: 297307
MFC after: 2 weeks
Sponsored by: Rubicon Communications, LLC ("Netgate")
net: don't panic on ifconfig pfsync0 mtu 9000
pfsync interfaces do not have ifp->if_inet6 set, so when we update the
MTU for those interfaces we panicked.
Add an explicit check for this. This should be temporary, until pfsync
is no longer a struct ifnet (as we've already done for pflog).
Reviewed by: glebius
Sponsored by: Rubicon Communications, LLC ("Netgate")
Differential Revision: https://reviews.freebsd.org/D58701
vmm: Tear down the IOMMU before AMD-Vi detach
Register the vmm module handler after both the bundled device drivers
and SMP. On platforms without EARLY_AP_STARTUP, SI_SUB_SMP follows
SI_SUB_DRIVERS; using the later subsystem preserves the
smp_rendezvous() requirement.
The resulting reverse unload order performs IOMMU cleanup while every
IVHD softc remains valid. Refuse an independent IVHD detach while
translation state remains initialized.
MFC after: 2 weeks
release/Makefile.gce: migrate gsutil usages to gcloud CLI
Google Cloud recommends migrating from gsutil to gcloud storage CLI.
Update gce-do-upload target to use `gcloud storage buckets create` and
`gcloud storage cp` instead of `gsutil mb` and `gsutil cp` commands.
PR: conf/297016
Reviewed by: lwhsu
MFC after: 3 days
Differential Revision: https://reviews.freebsd.org/D58464
ixgbe: Drain events for inactive VFs
The aggregate VF mailbox poll includes only VFs whose driver
configuration completed. A configured VF slot whose vf_add callback
failed can nevertheless report reset, request, or acknowledgement
events. Because the mailbox handler skips inactive entries, such an
event remains latched and can retrigger administrative work
indefinitely.
Build the poll masks from every configured VF index and consume reset,
message, and acknowledgement events for inactive entries without
treating them as usable VFs. Use the index rather than the pool because
early vf_add errors precede pool initialization. Also include E610
PFVFLREC in aggregate reset sampling.
MFC after: 2 weeks
ixgbe: Handle deferred link-status requests
The iflib conversion records link-status interrupts in the
administrative request mask, but the administrative task did not
consume them. Timer polling usually hid the omission; frequent mailbox
interrupts could continually rearm that timer and leave cached link
state down after hardware recovered.
Claim request batches atomically, process link-setup dependencies, and
sample hardware before publishing link state. Bound each invocation to
eight batches and requeue residual work so a continuous producer cannot
monopolize the admin taskqueue.
Queue every link-related request from the legacy interrupt path.
Unlike MSI-X, its threaded continuation services RX and does not enqueue
the admin task. This restores the event-driven behavior of ix-3.4.39.
Fixes: b2c1e8e62049 ("ix(4): Run {mod,msf,mbx,fdir,phy}_task in if_update_admin_status")
MFC after: 2 weeks
ixv: Tolerate temporary PF mailbox unavailability
A PF can be resetting, handling a slow link event, or deliberately
withholding mailbox CTS while its VFs enumerate. Keep the VF attached
when the reset handshake is temporarily unavailable so a later if_init
can retry.
Never leave VF hardware running without a negotiated mailbox API: start
hardware only after reset succeeds, stop it when negotiation fails in
attach or init, and defer later recovery through iflib. This prevents a
tight reset loop while preserving recovery when the PF returns.
MFC after: 2 weeks