stand/powerpc/ofw: do not truncate device tree properties to 1024 bytes
When the OpenFirmware loader flattens the firmware device tree into the
FDT it hands to the kernel (usefdt=1, i.e. on every real-mode OF system
such as pSeries LPARs and QEMU pseries guests), add_node_to_fdt() clamps
every property value to 1024 bytes. Any larger property reaches the
kernel truncated.
On QEMU pseries the PCI host bridge's "interrupt-map" is 3584 bytes
(32 slots x 4 pins x 7 cells), so only the entries for slots 0-8 survive
and the entry for slot 9 is cut in the middle. A PCI device in slot 9
or above therefore gets no INTx routing (irq 0), and with INVARIANTS the
partial trailing entry trips the "ofw_bus_search_intrmap: truncated map"
assertion in ofw_bus_search_intrmap() during PCI attach, panicking the
kernel as soon as such a device is present. "ibm,drc-indexes",
"ibm,drc-names" and "ibm,drc-power-domains" are cut the same way.
Drop the clamp. fdt_setprop() already reports a property that does not
fit into the FDT buffer, so no separate limit is needed.
[4 lines not shown]
zfs: merge openzfs/zfs at 1f380a4f3
Notable upstream pull request merges:
#17864 e903655c5 zpool: Add zpool status -vv error ranges
#18820 -multiple zdb: account pending DDT-log frees in leak detection
#18884 -multiple zio_crypt: establish platform interface; rework common
code to use it
#19010 -multiple zfs_namecheck: reject '.' and '..' before a snapshot or
bookmark
#19031 994fb1703 Fix metaslab count assertion in metaslab_group_alloc()
for small vdevs
#19093 f5b2fc8e2 zfs_ctldir: make .zfs/snapshot/<name> btime the snapshot
creation time
#19097 2dece2a34 zstream: report invalid record context without assertions
#19102 78f49e1dd vdev_disk: simplify alignment checks for linear ABDs
#19108 -multiple Fix permanent errors misfiled into the scrub error log
#19110 fa4bc4dec spa_errlog: don't let one unresolvable entry hide the
whole error log
#19116 b6dde8a17 zio_crypt: free the key unwrap uios when decryption fails
[13 lines not shown]
cxgbe: Report SR-IOV VF status
Retain the PF-accepted MAC and VLAN settings from the per-port t4iov
companion and expose them through the corresponding cxgbe ifnet.
Publish, snapshot, and destroy the cache under the existing adapter
synchronized-operation mechanism so status queries cannot race IOV
configuration or teardown.
Track successful t4iov attachment independently of the active VF count.
Restrict reporting to the port main VI, return an empty status for a
supported but unconfigured PF, and omit status from VF and auxiliary
VIs.
Reviewed by: jhb
Sponsored by: BBOX.io
Differential Revision: https://reviews.freebsd.org/D58741
pci: Reserve bus numbers required by SR-IOV VFs
Some firmware assigns only one bus number to each PCI-PCI bridge. This
prevents later SR-IOV VF enumeration when a VF routing ID falls on a bus
number already allocated to a sibling bridge.
Reserve only the additional bus numbers required by SR-IOV PFs.
Enumerate all directly attached functions before child drivers and
bridges attach, inspect their device_t objects for SR-IOV, and grow the
PCI bus resource through the highest possible VF routing ID.
First VF Offset and VF Stride may change when NumVFs changes. Probe
every valid NumVFs value and preserve the original setting. When the
upstream hierarchy uses ARI, temporarily enable the SR-IOV ARI Hierarchy
control in the lowest-numbered PF while sizing, then restore it. Scope
active-VF detection to each conventional PCI slot; an ARI bus remains
one slot-0 hierarchy. If firmware left VFs enabled on a device, do not
modify it and reserve only its active layout.
[20 lines not shown]
ipfw: guard against NULL deref with clat/plat prefixes without length
Found with: Claude Code Sonnet 5
MFC after: 2 weeks
(cherry picked from commit 47b02ace19c53e72ad2290d974025292ff97464a)
exec: Remove an unneeded capability mode check
The subsequent namei() call is relative to AT_FDCWD, and such lookups
are always disallowed in capability mode.
No functional change intended.
Reviewed by: emaste
Differential Revision: https://reviews.freebsd.org/D59888
sysctl: Return ECAPMODE when trying to access sysctls in capability mode
We have always returned EPERM in this case, but it's incorrect, we
should return ECAPMODE for capability mode violations. Fix the errno
value.
Reviewed by: emaste
MFC after: 2 weeks
Differential Revision: https://reviews.freebsd.org/D59887
pf: do not loop on an address that is cleared twice in pfr_clr_astats()
pfr_clr_astats() looks up each address it is given and inserts the entry
it finds at the head of a work queue. If the same address is given more
than once, the entry is inserted twice and the second insertion makes it
its own successor. pfr_clstats_kentries() then walks the queue forever,
with the rules lock held for writing, so packet processing and every
other pf operation in that vnet stop as well. To reproduce:
pfctl -e
pfctl -t foo -T add 192.0.2.1
pfctl -t foo -T zero 192.0.2.1 192.0.2.1
Do as pfr_del_addrs() does: clear pfrke_mark on the entries named, then
queue an entry only the first time it is seen. An address given more
than once is cleared, and counted, once. Validate all addresses before
any entry is touched.
Add a regression test.
[6 lines not shown]
pf: fix NULL dereference in pfr_set_addrs() with feedback
Since DIOCRSETADDRS was converted to netlink, pf_handle_table_set_addrs()
calls pfr_set_addrs() with a NULL size2, as the netlink interface has no
buffer to return the deleted addresses in. pfr_set_addrs() only checked
size2 for NULL at the end of the function; with PFR_FLAG_FEEDBACK set it
dereferenced it unconditionally first. pfctl sets PFR_FLAG_FEEDBACK
when run with -v, so "pfctl -v -t foo -T replace ..." panicked the
kernel with a NULL pointer dereference. To reproduce:
pfctl -e
pfctl -t foo -T add 192.0.2.1
pfctl -v -t foo -T replace 192.0.2.2
Check size2 for NULL before dereferencing it, as is already done at the
end of the function. The per-address feedback for added and changed
addresses is still copied back as before; only the list of deleted
addresses, which the netlink caller has no room for, is skipped.
[9 lines not shown]
hwpmc_amd: add PerfMonV2 global-control path
Add support for AMD PerfMonV2 (Family 19h+) core counters, which need
both the per-counter EVSEL enable bit and the global GLOBAL_CTL bit set
to count. Detects PerfMonV2 at init and switches to v2-specific
start/stop/interrupt handlers; older CPUs and L3/DF counters keep using
the classic path unchanged.
Adds a read-only sysctl, kern.hwpmc.amd_perfmon_v2, to report which path
is active.
Signed-off-by: Andre Silva <andasilv at amd.com>
Reviewed by: Ali Mashtizadeh <ali at mashtizadeh.com>
Sponsored by: AMD
Differential Revision: https://reviews.freebsd.org/D58256
bhyve: fix boot device ordering
EDK2's QemuBootOrderLib inspects the bootorder file provided
via fw_cfg and requires it to be NUL-terminated. Otherwise,
it rejects the supplied bootorder and falls back to its
default boot order.
Currently, bhyve registers bootorder with qemu_fwcfg_add_file()
using bootorder_len returned by open_memstream(), which excludes
the trailing NUL byte.
Fix that by passing bootorder_len + 1 to qemu_fwcfg_add_file() so
the fw_cfg payload is properly NUL-terminated.
PR: 279720
Reviewed by: markj
Found with: codex (gpt-5.6-sol)
MFC after: 1 week
Sponsored by: The FreeBSD Foundation
[3 lines not shown]
hwpmc: fix IBS fetch and op NMI handling
Service each IBS unit with a valid bit set when fetch and op share an
NMI. Otherwise, op samples can be lost. Treat the extra NMI that follows
as expected (skip it).
Reviewed by: mhorne
Fixes: e51ef8ae490f ("hwpmc: Initial support for AMD IBS")
Differential Revision: https://reviews.freebsd.org/D60004
ipmi: Add some additional diagnostic output on errors
Add some additional diagnostic output for IPMI code,
particularly on error paths. This has been found to be helpful at
$WORK and seems generally useful, so contributing the changes back to
upstream.
Sponsored by: Dell Technologies
Reviewed by: vangyzen@
Differential Revision: https://reviews.freebsd.org/D60091
llvm: add LoongArch target support, not enabled by default
Note there is ongoing work to add LoongArch support to the base system,
but having target support in llvm is an essential component.
This must be explicitly enabled using WITH_LLVM_TARGET_LOONGARCH.
Reviewed by: dim
MFC after: 1 week
Differential Revision: https://reviews.freebsd.org/D59899
bhyve: fix boot device ordering
EDK2's QemuBootOrderLib inspects the bootorder file provided
via fw_cfg and requires it to be NUL-terminated. Otherwise,
it rejects the supplied bootorder and falls back to its
default boot order.
Currently, bhyve registers bootorder with qemu_fwcfg_add_file()
using bootorder_len returned by open_memstream(), which excludes
the trailing NUL byte.
Fix that by passing bootorder_len + 1 to qemu_fwcfg_add_file() so
the fw_cfg payload is properly NUL-terminated.
PR: 279720
Reviewed by: markj
Found with: codex (gpt-5.6-sol)
MFC after: 1 week
Sponsored by: The FreeBSD Foundation
[3 lines not shown]
libkvm: support powerpc64 radix minidumps
PowerPC64 radix minidumps use a 64KB root directory followed by three
levels of 4KB page tables. Add an MMU backend which walks those tables,
handles large-page leaves, and decodes their always-big-endian entries
on both powerpc64 and powerpc64le.
Reject old radix minidumps whose pmap section is empty with a specific
diagnostic. Add a synthetic powerpc64le dump test which reads a page
through both its kernel and direct-map addresses.
PR: 298532
Reviewed by: jhb
Approved by: jhb (mentor)
MFC after: 2 weeks
Sponsored by: FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59743
powerpc/radix: include page tables in minidumps
The radix pmap did not implement the minidump pmap callbacks, so radix
minidumps were emitted with an empty pmap section. Such dumps do not
contain enough information for libkvm to translate kernel virtual
addresses.
Snapshot the 64KB radix root table in the pmap section, add lower-level
page-table pages to the sparse dump, and include pages backing non-DMAP
kernel mappings. Keep the bulk direct map out of the dump while
retaining the relocated kernel image.
PR: 298532
Reviewed by: jhb
Approved by: jhb (mentor)
MFC after: 2 weeks
Sponsored by: FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59743
arm: pad minidump page table
libkvm locates the sparse page array after the page-rounded PTE table
size, but the ARM minidump writer emitted only the unrounded size. When
the table size was not page-aligned, libkvm therefore read every dumped
physical page at the wrong offset.
The mismatch was introduced when libkvm began rounding ptesize. It has
affected ARM minidumps since ffdeef323449 ("libkvm: Improve physical
address lookup scaling."). It is exposed when the dumped KVA span is not
a multiple of 4 MiB.
Zero-pad the final PTE page and include the padding in the dump size.
Reviewed by: jhb
Approved by: jhb (mentor)
Fixes: ffdeef323449 ("libkvm: Improve physical address lookup scaling.")
MFC after: 2 weeks
Sponsored by: FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59716
nd6: Fix regeneration of temp addresses in detached state
When an on-link prefix becomes detached, the kernel keeps
generating new RFC 8981 temporary addresses for that prefix.
Fix it by ignoring the detached addresses in regen_tmpaddr().
While here, change its return type to bool.
PR: 298533
Discussed with: markj
MFC after: 3 days
Differential Revision: https://reviews.freebsd.org/D60051