catopen(3): align returned errors with POSIX.1-2024
Return ENOENT instead of EFTYPE for invalid/empty names, non-existent
catalogue files, and bad catalogue headers. Require the complete header
and reject sizes that exceed mmap() limit.
Also save the errno in handling the open() failure.
Derived-from: FreeBSD (commit 1176390d2d2bbb1e207c840d1f7a66a6ac1096ff)
Reported-by: pmjdebruijn
Bug: https://bugs.dragonflybsd.org/issues/3393
catopen(3): Fix race condition
The current code uses a rwlock to protect the cached list, which
in turn holds a list of catentry objects, and increments reference
count while holding only read lock. Fix this by converting the
reference counter to use atomic operations.
Also improve the cleanups of memory allocations.
Obtained-from: FreeBSD (commit 4188ba1a3b65b4a55fc938b71d9091fafb57967d)
Reported-by: pmjdebruijn
Bug: https://bugs.dragonflybsd.org/issues/3393
drm: Balance dma-buf file references on get and fd export
dma_buf_get() drops its lookup reference before returning, but its callers
release the returned buffer with dma_buf_put(). Repeated imports can
therefore consume references owned by the descriptor and PRIME caches.
Retain an independent file reference before dropping the lookup reference.
dma_buf_fd() transfers the caller's reference to the new descriptor.
fsetfd() acquires another reference, so drop the caller's reference after
successful installation. Keep it on allocation failure.
FreeBSD's dma-buf implementation follows the same ownership rules: retain
the fget() reference on lookup and drop the extra reference taken by
finstall() after successful descriptor installation.
The comments were extracted from the patch bundle by servik and dillon:
https://apollo.backplane.com/DFlyMisc/drm98.patch
Reference: https://github.com/freebsd/drm-kmod/blob/f252a30f27d157d9c763cd408850775096a6263f/drivers/dma-buf/dma-buf.c#L498-L544
Bug: https://bugs.dragonflybsd.org/issues/3428
kdmsg: shut down the transport while waiting for reconnect workers
KILLRX and a wakeup on msg_ctl do not interrupt a reader blocked in
fp_read() on a quiet connection. Reconnect can then wait indefinitely
for the reader and writer to exit. HAMMER2's recluster ioctl holds the
root vnode during this wait, blocking other filesystem operations.
Shut down msg_fp before sleeping in the worker-wait loop. This wakes
blocked transport I/O and leaves the existing state cleanup and file
reference handling in place.
Keep the shutdown inside the loop: lksleep() releases msglk, so another
reconnect may replace the connection while this caller is asleep. A
one-time shutdown before the loop can leave this caller waiting on the
replacement workers.
Bug: https://bugs.dragonflybsd.org/issues/3434
drm: notify userspace of connector changes
Emit the DRM CONNECTOR HOTPLUG devctl event for the primary card. The
empty hotplug handler leaves libudev-devd clients unaware of connector
changes.
Use the event contract implemented in FreeBSD drm-kmod
drivers/gpu/drm/drm_sysfs.c.
Bug: https://bugs.dragonflybsd.org/issues/3440
bpf: add XOR and modulo instructions
libpcap can generate XOR and modulo instructions, but the kernel
interpreter and validator do not support them.
Add constant and register operands for both operations. Reject constant
modulo by zero and return zero for a register zero divisor, as for DIV.
Update the manual to match.
Patch-by: guy
Bug: https://bugs.dragonflybsd.org/issues/3387
drm: reject unload while core teardown is incomplete
Unloading drm.ko reaches ttm_exit(), which waits for device_released.
The callback which sets that flag is compiled out and device unregister
is a stub, so kldunload sleeps indefinitely while holding the linker
lock. The module event handler has already cleared the Linux task and
process cleanup callbacks by then.
DRM also retains worker threads and undrained RCU callouts, so removing
the TTM wait alone would not make unloading safe. Return EBUSY from
MOD_UNLOAD before changing callbacks or entering SYSUNINIT. Keep the
module usable until complete teardown is implemented.
Bug: https://bugs.dragonflybsd.org/issues/3443
vm: use normal COW inheritance for user-wired mappings
Forking an mlock()ed MAP_PRIVATE file mapping can panic with
"vm_fault_copy_wired: page missing". The wired-copy path expects the
page in the front object, but it may be in a backing object or have
been removed after the file was truncated.
Use normal COW inheritance for normal mappings with only a user wire.
The parent stays user-wired and the child remains unwired. The normal
fault path resolves backing pages and handles pager errors. Keep eager
copying for hard-wired and virtual-page-table mappings.
Reviewed-by: dillon
Bug: https://bugs.dragonflybsd.org/issues/3433
vm: allow writes to user-wired COW mappings
A later write to a MAP_PRIVATE mapping that was mlock()ed currently gets
KERN_PROTECTION_FAILURE from vm_map_lookup() because the entry is both
MAP_ENTRY_USER_WIRED and MAP_ENTRY_COW. The normal page-fault path does
not retry with VM_PROT_OVERRIDE_WRITE, so the write never completes and the
process receives SIGSEGV.
Fix the bug by removing the obsolete check so the normal copy-on-write
path runs.
vm_map_user_wiring() already created the shadow object before wiring the
entry, so a later COW is a normal copy into the process's own shadow. The
NEEDS_COPY block below still creates the shadow if it was not created at
wiring time.
Reviewed-by: dillon
Bug: https://bugs.dragonflybsd.org/issues/3431
pc64: balance wired PTE replacement and hard-busy wired refaults
pmap_enter drops the previous wired mapping, so account for the new
wired mapping even if the old PTE was wired. Otherwise protection
changes can underflow the pmap and vm_page wire counts.
Use vm_page_wire_quick for an already-wired same-page replacement, which
may be soft-busied. Skip fictitious pages, which have no physical wiring
count.
The old physical wire is released by pmap_removed_pte, and the old pmap
wire count is dropped separately in pmap_enter. Do not add another unwire.
Also reject FW_WIRED in vm_fault_bypass. An ordinary fault on a wired entry
may need to establish a new wire after its PTE was invalidated, so the
soft-busy shortcut cannot assume an old wired reference still exists.
A shared file mapped twice, mlock on one alias, MADV_INVAL on that alias
and a subsequent read otherwise panics in vm_page_wire. The normal
hard-busy fault path handles this case.
[6 lines not shown]
pc64: Support to apply AMD microcode update
* Implement ucode_load_bsp() that runs in hammer_time() immediately
before identify_cpu(), so the microcode update that's preloaded by the
boot loader can be applied before kernel detecting the CPU features.
The companion ucode_apply() function is called from initializecpu() on
each AP to picks up the microcode update.
Note: Only AMD CPUs are supported now.
* Add "cpu_microcode_load" and "cpu_microcode_name" variables to
loader.conf to load the CPU microcode and document them.
Co-authored-by: Aaron LI <aly at aaronly.me>
GitHub-PR: https://github.com/DragonFlyBSD/DragonFlyBSD/pull/55
boot: Support to preload firmware files
Support to preload the firmware files with the "firmware" type for the
firmware(9) subsystem to register, which is supported in the previous
commit.
Add the "./firmware" directory (which resolves to "/firmware" or
"/boot/firmware") to "module_path" variable for searching for the
firmware files.
Introduce the firmware="dir1/file1.bin dir2/file2.bin ..." variable to
the loader.conf for specifying the firmware files to be preloaded.
Co-authored-by: Aaron LI <aly at aaronly.me>
GitHub-PR: https://github.com/DragonFlyBSD/DragonFlyBSD/pull/55
kern: Support loading firmware from preloaded images or filesystem files
Extend the firmware loading mechanisms to support two more methods:
* preloaded images: load a firmware from an in-memory image preloaded by
the boot loader.
* firmware files: load a firmware by directly reading a firmware file
from the filesystem. A new sysctl variable "hw.firmware_path" and a
tunable of the same name is added to specify the search locations,
which defaults to "/usr/local/lib/firmware;/usr/lib/firmware".
The first method may be used to apply a CPU microcode update, and the
second method is mainly used by modern GPU/WiFi drivers to load the
required firmare directly from binary files (e.g., installed by a
firmware package).
Co-authored-by: Aaron LI <aly at aaronly.me>
GitHub-PR: https://github.com/DragonFlyBSD/DragonFlyBSD/pull/55
kern: Fix write-open "." or ".." to return EISDIR instead of EEXIST
Before the fix, opening "." or ".." for write would return EEXIST, which
was incorrect per the POSIX spec:
https://pubs.opengroup.org/onlinepubs/9699919799/functions/fopen.html
where one would expect one of EISDIR, EINVAL, EACCES.
For example:
```
% sh -c 'echo xxx > /tmp/.'
sh: cannot create /tmp/.: File exists
```
Fix the code to return EISDIR. Note that opening a directory other than
"." or ".." for write already returns EISDIR, e.g.,
[6 lines not shown]
kern: Fix write-open "." or ".." to return EISDIR instead of EEXIST
Before the fix, opening "." or ".." for write would return EEXIST, which
was incorrect per the POSIX spec:
https://pubs.opengroup.org/onlinepubs/9699919799/functions/fopen.html
where one would expect one of EISDIR, EINVAL, EACCES.
For example:
```
% sh -c 'echo xxx > /tmp/.'
sh: cannot create /tmp/.: File exists
```
Fix the code to return EISDIR. Note that opening a directory other than
"." or ".." for write already returns EISDIR, e.g.,
[6 lines not shown]
libnvmm: Improve the page walker
* Implement A/D bit setting.
* Inject #PFs on non-present pages when emulating MMIOs. This is needed
for guests like Windows 2000 that aggressively page out memory that
they then use in MMIO operations.
* Fix the '64bit walk - present' test that failed in the previous
commit.
Obtained-from: NVMM Reference Implementation
libnvmm: Improve the correctness of page table walks
* Fix 32bit PAE: the top-level pdir doesn't have NX.
* Fix 32bit non-PAE: PTE_PS is ignored if CR4.PSE=0.
* Stop the walks early on if a required permission is not met.
* Take CR0.WP and EFER.NXE into account when evaluating permissions.
* Enforce SMEP/SMAP protections.
* Document the semantics of nvmm_gva_to_gpa() more clearly.
* Add unit-tests.
Note that the '64bit walk - present' test is failing, which will be
fixed in a later commit.
Obtained-from: NVMM Reference Implementation
testcases/libnvmm: Fully reset the VCPU after each test
Otherwise if a fault is pending it gets injected in the next test.
Obtained-from: NVMM Reference Implementation