ulimit(3): Improve UL_{GET,SET}FSIZE wording
Reword the UL_GETFSIZE and UL_SETFSIZE descriptions to say 512-byte
blocks and make the setter's argument clearer.
While here, improve mdoc markup and pull in other small improvements.
Reviewed by: des
MFC after: 3 days
Obtained from: https://github.com/apple-oss-distributions/libc
Sponsored by: Klara, Inc.
Differential Revision: https://reviews.freebsd.org/D60467
if_bnxt: dcb: serialize HWRM sends and avoid freeing DMA bufs on timeout
Every _hwrm_send_message() call in bnxt_dcb.c ran without BNXT_HWRM_LOCK,
unlike the rest of the driver, letting concurrent HWRM commands race on
the shared MMIO doorbell and response buffer. Route single-shot sends
through the already-locked hwrm_send_message() wrapper, and wrap the
remaining send+response-read sequences in explicit BNXT_HWRM_LOCK/UNLOCK.
Also stop freeing the transient DMA buffers used for structured-data
get/set on ETIMEDOUT: a host-side timeout doesn't guarantee firmware
actually gave up, so a late completion could still DMA into memory
that's since been freed and reused. Leak the buffer instead in that case.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D60179
if_bnxt: don't reserve RoCE MSI-x on chips that don't support RoCE
RoCE is only supported on Thor and onwards chips.
bnxt_init_sctx_variants() unconditionally set isc_admin_intrcnt to
BNXT_ROCE_IRQ_COUNT for every PF device, stealing 9 MSI-x vectors from
older, non-RoCE-capable chips that will never use them.
Gate the reservation on the device ID: RoCE-capable chips keep
BNXT_ROCE_IRQ_COUNT, other chips fall back to 1, the baseline
isc_admin_intrcnt value used before RoCE support (commit 050d28e13cde)
repurposed it for RoCE MSI-x on every device.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D59091
if_bnxt: port workqueue, task-queue, and synchronization kPIs to FreeBSD natives
- Replace Linux workqueue (struct workqueue_struct, bnxt_pf_wq) with a
FreeBSD taskqueue (bnxt_taskq) shared by every attached PF. bnxt_taskq
is created once at module load and destroyed once at module unload
(SYSINIT/SYSUNINIT on SI_SUB_KLD), matching the module-owned-taskqueue
idiom used elsewhere in the tree (e.g. sys/dev/cxgbe/t4_main.c's
reset_tq). This replaces an earlier lazy, per-attach/detach
create-or-reuse scheme guarded by a mutex and a "ready" flag, which
had a real race: bnxt_queue_sp_work()/bnxt_queue_fw_reset_work() and
bnxt_attach_pre() re-read the bnxt_taskq/bnxt_taskq_ready globals
outside the mutex right after the lazy-init call returned, so a
concurrent bnxt_detach() on another PF freeing the shared queue once
bnxt_num_pfs hit 0 could race a still-attached PF into enqueuing onto
a freed taskqueue. Tying the taskqueue's lifetime to the module
instead of to any one PF's attach/detach removes that race outright,
since the module can't unload while a PF is still attached.
- Replace struct work_struct sp_task with struct task sp_task;
[42 lines not shown]
if_bnxt: port endian, memory-barrier, and bitops kAPIs to FreeBSD natives
There are few more Linux kernel APIs left in the code. So,
replaced them with corresponding FreeBSD native APIs.
- Replace cpu_to_le16/32/64/le16_to_cpu/le32_to_cpu/le64_to_cpu with
htole16/32/64/le16toh/le32toh/le64toh throughout bnxt_hwrm.c and
the vf-event producer path in if_bnxt.c.
- Replace rmb()/wmb() interprocessor-ordering uses in bnxt_ulp.c
(publishing/consuming ulp->async_events_bmap and max_async_event_id
across CPUs) with atomic_thread_fence_acq()/atomic_thread_fence_rel();
keep wmb() for the HWRM doorbell/DMA ordering barrier in
_hwrm_send_message() (bnxt_hwrm.c), since atomic_thread_fence_rel()
compiles to a bare compiler barrier on amd64/i386 and does not order
the request-buffer store against the doorbell MMIO write.
- Replace hweight32 with bitcount32, ARRAY_SIZE with nitems, and
DECLARE_BITMAP with a plain unsigned long array (sized via
howmany()) in bnxt_hwrm_func_drv_rgtr(); replace the __set_bit/
test_bit calls there with bit_set()/bit_test() from
[12 lines not shown]
if_bnxt: add support for SIOCGI2CPB page/bank i2c requests
Advertise IFLIB_I2C_PAGE_BANK and forward the requested page/bank to
firmware, so CMIS optics can be read correctly instead of falling
back to bogus legacy SFF-8472 decoding. Also fix bnxt_i2c_req() to
return positive errno values.
Signed-off-by: gallatin
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D59573
if_bnxt: drop epoch_arr[] to detect EPOCH bit toggle
The EPOCH array size is 4096 but the driver does not enforce any limit
on ring size, so setting ring size >4096 corrupts adjacent memory
and leads to undefined behavior.
Toggle epoch_bit directly at every ring-producer wrap point (TX/RX
encap and refill, MPC crypto commands, kTLS presync/replay) instead of
looking it up per-index in the doorbell path, dropping the now-unused
epoch_arr[] snapshot array.
Reported by: gallatin
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D58620
Change-Id: I4673d64670ee0e8ea256b8a8932724bbd73f710d
if_bnxt: rework interrupt coalescing onto the AGGINT_QCAPS scheme
Query firmware-advertised coalescing capabilities and program
per-direction rx/tx coalescing settings against them, replacing the
old hardcoded scheme. Add sysctls for the coalescing mode, budget,
and stats-timer knobs the new scheme exposes.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D58621
if_bnxt: drop redundant PCI disable in bnxt_fw_reset_close()
bnxt_fw_reset_close() unconditionally called pci_disable_device()
after freeing the Rx IRQs, on every firmware reset, not just the
BNXT_STATE_FW_FATAL_COND path which already disables the device via
bnxt_fw_fatal_close(). BNXT_FW_RESET_STATE_ENABLE_DEV in
bnxt_fw_reset_task() already re-enables the device unconditionally,
so this wasn't leaving the device disabled, just adding an
unnecessary PCI disable/re-enable cycle around every non-fatal reset.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D58619
if_bnxt: fix HWRM failures/timeouts after repeated FW resets
Consecutive firmware-initiated reset cycles produced HWRM
failures/timeouts and traffic didn't come back. Fix the FW-reset
recovery path:
- bnxt_fw_reset_close() called bnxt_stop() and
bnxt_hwrm_func_drv_unrgtr(), sending HWRM ring/VNIC/filter free
commands to a firmware that's already mid-reset and unresponsive,
producing the observed timeouts. Both are unnecessary since
firmware comes back with fresh state anyway; drop them along with
the now-redundant iflib_request_reset().
- bnxt_func_reset() now skips bnxt_hwrm_resource_free() entirely
while BNXT_STATE_IN_FW_RESET is set, for the same reason.
- bnxt_open() (used to reopen after a firmware reset) now issues
bnxt_hwrm_func_reset() up front and drives reinit through iflib's
own reset machinery (iflib_request_reset() plus a new
bnxt_iflib_reset_sync() that waits for
IFF_DRV_RUNNING/OACTIVE to flip 1->0->1) instead of calling
[9 lines not shown]
if_bnxt: sriov: replace Linux kPIs with FreeBSD native APIs
Convert bnxt_sriov.c and related SRIOV files from Linux to FreeBSD
native equivalents:
- Replace cpu_to_le16/32/64/le16_to_cpu with htole16/32/64/le16toh.
- Replace Linux is_valid_ether_addr, ether_addr_equal, ether_addr_copy
with local bnxt_eth_addr_valid/bnxt_eth_addr_equal/bnxt_eth_addr_copy
helpers built on ETHER_IS_MULTICAST/ETHER_IS_ZERO.
- Replace kcalloc/kzalloc/kfree with malloc/free (M_DEVBUF,
M_WAITOK|M_ZERO).
- Replace dma_alloc_coherent/dma_free_coherent with iflib_dma_alloc/
iflib_dma_free; add struct iflib_dma_info hwrm_cmd_req_mem[4] to
bnxt_pf_info to hold the DMA handles.
- Replace the Linux unsigned long *vf_event_bmap with FreeBSD
bitstr_t *, allocated via bit_alloc(); bit_ffs_at()/bit_clear()/
bit_set() replace find_next_bit()/clear_bit()/set_bit() on it
(if_bnxt.c's vf-event producer side included).
- Replace DIV_ROUND_UP with howmany; rcu_assign_pointer with a plain
[9 lines not shown]
if_bnxt: sysctl: remove unused linux/delay.h include and dead mutex
<linux/delay.h> and the DEFINE_MUTEX(tmp_mutex) it enabled are both
unused in bnxt_sysctl.c: no msleep/udelay/mdelay call remains, and
tmp_mutex itself was never referenced anywhere.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D59087
if_bnxt: support more Tx rings than Rx rings on P5+
Add an unsupported-config guard rejecting nrxqsets > ntxqsets, and
teach the P5+ NQ datapath to handle the opposite asymmetric case:
extra Tx rings beyond the Rx ring count. softc->nq_rings[i].type now
tags each NQ as SHARED_NQ (paired with an Rx CQ) or BNXT_TX_ONLY_NQ
(dedicated to an extra Tx CQ with no backing Rx ring); bnxt_init()
and bnxt_hwrm_resource_free() gate all Rx ring group/ring/stat-ctx
alloc and free on IS_SHARED_NQ() so the extra Tx-only NQs don't touch
nonexistent Rx ring state. Gate the new type-tagging and IS_SHARED_NQ
checks on BNXT_CHIP_P5_PLUS(), since non-P5+ chips never allocate
softc->nq_rings and would otherwise access a nonexistent array.
BNXT_TX_ONLY_NQ is a distinct value from the existing MPC-private
TX_CP_NQ; the two are unrelated and kept apart to avoid confusion.
Suggested by: gallatin
Reviewed by: gallatin
MFC after: 2 weeks
[2 lines not shown]
if_bnxt: log: fix missing lock coverage in bnxt_log_live()
bnxt_log_live() walked the loggers TAILQ without holding log_lock,
unlike every other bnxt_log.c function, racing against concurrent
bnxt_register_logger()/bnxt_unregister_logger() calls. Take log_lock
around the traversal, and bail out early (with the lock dropped) if
live_msgs_len has already reached max_live_buff_size, which would
otherwise underflow the length passed to bnxt_log_info().
bnxt_start_logging_driver_coredump() now drops log_lock before
invoking logger->log_live_op() (which calls back into
bnxt_log_live()) and re-acquires it afterward, resetting live_msgs
to NULL so a later bnxt_log_live() call (e.g. from a VF async event
handler) can't write into a coredump buffer the caller has already
freed.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D59086
if_bnxt: Collect hw stats every 250 milliseconds
bnxt_if_timer() only scheduled bnxt_update_admin_status(), which
polls HW stats, once per second. Drop the interval to 250ms so stats
collection is timely.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D58610
if_bnxt: avoid FTQM/STQM pg_info alias on reset
Copying the STQM backing-store context to seed FTQM's also copied
STQM's pg_info pointer verbatim. On cold load STQM's pg_info is still
NULL so this is harmless, but on a firmware reset STQM's pg_info is
already non-NULL, making FTQM alias STQM's backing-store pages. The
allocator then skips FTQM since it looks already allocated, and FTQM
is later indexed as its own array, reading out of bounds into STQM's
buffer and dereferencing a bogus DMA address.
Clear ctxm->pg_info after the memcpy so FTQM always gets its own
backing store.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D58615
if_bnxt: ktls: Reject new kTLS sessions once driver is detaching
bnxt_tls_snd_tag_alloc() only gated on BNXT_STATE_OPEN, which is set
once by bnxt_open()/bnxt_attach_pre() but never cleared again, so it
could not reject new sessions once bnxt_detach() started tearing
things down. The snd_tag_alloc_ref-based wait in bnxt_detach(), which
waits for in-flight allocators to exit, was already in place, but
without this gate new sessions could keep being created for the
entire duration of that wait, extending it indefinitely instead of
just draining existing ones.
Reject new sessions once softc->detached is set, and reset it back to
false in bnxt_attach_pre() so it doesn't wrongly persist across a
detach/reattach cycle.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D58612
if_bnxt: Restore Rx CQ doorbell skip-when-idle optimization on P5
The unconditional-doorbell fix for Whp/older NICs (always ring the Rx
CQ doorbell after processing Rx interrupts) had also been applied to
P5+ chips, discarding process_nq()'s work-done signal entirely and
losing the doorbell-skip optimization on P5+ when process_nq() finds
no CQ notification to act on.
Make process_nq() return whether it saw a CQ notification, and only
skip the Rx CQ doorbell on P5+ when it didn't; non-P5+ chips continue
to always ring it, since process_nq() isn't used on that path.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D58613
if_bnxt: Never ARM Tx CQ as they are not interrupt-driven
Tx completion rings are not interrupt-driven, so there is no need to
ring the doorbell with CQ_ARMALL and toggle bits set for them. Stop
arming the Tx CQ doorbell on P5+, which makes tx_cp_rings[]->toggle
dead, so drop the now-unnecessary toggle bookkeeping for it in
process_nq() too.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D58616
if_bnxt: add support for HW generic stats
HW generic stats provide information about PCIe transactions and
out-of-order doorbell writes, useful for debug. Allocate a DMA buffer
for them (disabling BNXT_FW_CAP_GENERIC_STATS if the allocation
fails), fetch them via HWRM_STAT_GENERIC_QSTATS once per admin-status
tick on P5+ chips, and expose them to userspace under
dev.bnxt.<unit>.hwstats.generic_stats.* via sysctl.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D58614
if_bnxt: Pass HWRM cmd response to apps even if FW fails it
The driver was not passing the HWRM command response to management
apps when the firmware failed the command, which is not what apps
expect. Copy the response out whenever firmware populated one
(resp_len != 0), regardless of the command's overall return code, and
stop bailing out of the mgmt ioctl path early on a failed passthrough
so DMA'd indirect data and the response still get copied to
userspace.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D58609
bnxt_re: Fix witness reported calltrace while unloading driver
bnxt_unregister_dev() called synchronize_rcu() while holding
bp->en_ops_lock, which witness flags as a sleeping-while-locked
violation on driver unload. Drop the lock around the RCU grace period
and reacquire it afterward.
Also gate the RoCE extended-stats hex dump in
bnxt_qplib_qext_stat() behind bootverbose instead of printing it
unconditionally.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D58611
if_bnxt: Support backing store v2 for Thor if FW supports it
The driver currently only supports BS v1 for Thor even though the
firmware supports v2, capping max_mr at ~192K when firmware can
support up to ~256K. Enable Backing Store v2 support for the Thor controller.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D58601
if_bnxt: Fix interrupt storm on Thor2
An interrupt storm was observed on Thor2 because the Rx CQ consumer
index was always one less than it should be. Correct the Rx CQ
consumer index to fix the storm.
Reported by: gallatin
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D58602
if_bnxt: Fix HWRM mailbox/DMA teardown race on detach
iflib's generic device-deregister path never drains the admin task
before calling IFDI_DETACH, so bnxt_update_admin_status() (scheduled
once/sec) could still run concurrently with bnxt_detach(), racing on
softc->hwrm_lock and the shared HWRM request/response DMA buffer that
bnxt_detach() destroys. That race could leave the firmware-side HWRM
mailbox inconsistent, surfacing as "Timeout sending HWRM_VER_GET" on
the next module load.
Add a softc->detached guard: set it as the first statement in
bnxt_detach(), and bail out of bnxt_update_admin_status() immediately
when it's set.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D58607
if_bnxt: add bnxt_compat.h version-compat shim
Add bnxt_compat.h, a small __FreeBSD_version-gated compatibility
shim: if_getmtu()/if_name()/if_getsoftc()/if_getflags() fallbacks for
pre-1500000, and a config-task vs. config-gtask abstraction over the
iflib API split at 1403000. On this branch's FreeBSD version it also
unconditionally enables KTLS_IFLIB_SUPPORT/KTLS_REC_SEQ_NUMBER, which
the upcoming kTLS patch depends on.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D58604
if_bnxt: honor firmware min wait before VER_GET on reset
The firmware advertises a minimum wait time it needs before it is
ready. The driver was not honoring it and sent VER_GET too early
during firmware error-recovery, which could return an incomplete
VER_GET buffer. Enforce that minimum wait, measured from when the
RESET_NOTIFY async event arrived, before sending VER_GET.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D58608
f_bnxt: add kTLS (kernel TLS) TX offload
Add kTLS TX crypto offload, layered on MPC (crypto add/delete
commands submitted over its TCE channel) and TX completion coalescing
(shared tx_bd_opaque encoding).
Adds kTLS session lifecycle management, a key-context ID allocator,
HWRM key-context alloc/free with partition-mode support, and the TX
submit path: in-order packets go out on the normal iflib path with
CSUM_SND_TAG set, while out-of-order or retransmitted TLS records get
a presync command plus a replayed copy of the record built from the
driver's own replay ring. Also adds kTLS/MPC completion-time counter
sysctls and the max_ktls_entries tunable.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D58606
if_bnxt: add MPC (Mid-Path Channel) ring infrastructure
Add MPC ring support: a pair of extra, driver-private
TX/completion/notification rings the firmware uses as a side-channel
for offload commands. The only current consumer is kTLS TX crypto
command submission; RCE/CFA channel types are not yet implemented by
any consumer.
Adds MPC ring allocation/teardown, HWRM ring alloc/free, a private
interrupt path for the MPC NQ ring, and the TX submit/completion
path. The RoCE-only IRQ table builder is split into a generic,
exported bnxt_populate_irq(softc, irq_count) that both MPC and RoCE
use, plus an incremental bnxt_populate_irq_roce() that grows the
table on top of whatever is already there, so IRQ rids stay
consistent regardless of attach order.
This patch does not link standalone: bnxt_mpc_cmp() calls into
kTLS's completion handler, added by a later patch.
[4 lines not shown]