[StandardInstrumentations] Add ExtendedIRContext and trait registration for custom IR types (#153171)
Introduces a new class ExtendedIRType that is used for extending into
types that aren't recognized by llvm so that -print-changed is able to
be used.
RegAllocFast: Don't mark a physreg def dead when a live def covers its units
When scanning an instruction's physical-register defs, the fast allocator marked
a def dead whenever definePhysReg reported that nothing was displaced. This
ignored sibling def operands on the same instruction. e.g,
$sgpr4 = S_MOV_B32 0, implicit-def $sgpr4_sgpr5_sgpr6_sgpr7
the $sgpr4 subreg def was marked dead even though the live implicit-def of the
enclosing tuple keeps $sgpr4's register unit live, producing an inconsistent
dead flag. Try to maintain the invariant that LiveVariables introduces, which is
to not put a dead flag on any defined register which has a live def in any operand.
This will be enforced by a future verifier check.
This works by clearing improperly set dead flags after the fact which I find
distasteful but I don't see a better option without making the state tracking much
more complicated.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
drm: notify userspace of connector changes
Emit the DRM CONNECTOR HOTPLUG devctl event for the primary card. The
empty hotplug handler leaves libudev-devd clients unaware of connector
changes.
Use the event contract implemented in FreeBSD drm-kmod
drivers/gpu/drm/drm_sysfs.c.
Bug: https://bugs.dragonflybsd.org/issues/3440
[clang][NVPTX] Add support for scaled::n1::ue8m0 in FP8, FP6 and FP4 conversions (#227652)
This patch adds support for `scaled::n1::ue8m0` to existing
`f32/f16x2/bf16x2` to `FP8` (`e4m3x2`, `e5m2x2`),
`FP6` (`e2m3x2`, `e3m2x2`) and `FP4` (`e2m1x2`) conversion intrinsics.
Tests have been verified through `ptxas-13.4`.
PTX ISA Reference:
https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#data-movement-and-conversion-instructions-cvt
---------
Signed-off-by: DharuniRAcharya <dharunira at nvidia.com>
[GlobalISel][TableGen] Support for typed G_FCONSTANT Imm in MIR-patttern (#219563)
GlobalISel MIR combine patterns can already build a typed integer
constant directly in an apply pattern, e.g. `(apply (G_CONSTANT $dst,
(GITypeOf<"$dst"> 0)))`. The same was not possible for `G_FCONSTANT`:
writing a typed literal on it silently producing invalid MIR, since the
immediate was routed through the path meant for register operands rather
than being emitted as a floating-point immediate.
This patch extends that support to `G_FCONSTANT`, letting typed
float-constant patterns be written declaratively and type-derived via
`GITypeOf` just like the integer case, instead of falling back to
hand-written C++ combine code or hitting the bad-MIR bug.
This is a prerequisite for migrating existing `wip_match_opcode`
combines that replace a matched value with a float constant of its own
type into pure MIR patterns.
[AMDGPU] Check bundled instructions in canUsePressureDiffs (#227620)
canUsePressureDiffs refuses to use the imprecise cached PressureDiffs
for instructions with subregister defs or physical registers, but it
skips implicit operands, and a BUNDLE header has only implicit operands.
So bundles always used the PressureDiffs, which count a def of one lane
as a def of the whole register. With EXPENSIVE_CHECKS this tripped the
pressure cross-check in GCNSchedStrategy. Check the operands of the
bundled instructions instead.
Co-authored-by: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
CodeGen: Prefer getting the Triple from the Module (#228682)
Continue replacing TargetMachine::getTargetTriple() with the module's
triple at sites where a Module is one hop away through an available
Function, GlobalValue or MachineModuleInfo.
Where the surrounding class already holds a Subtarget, use its triple
rather than routing through the Module.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[LiveDebugVariables] Repair stale SlotIndexes
The analysis keeps its indexes from before the first register allocator
until DBG_VALUEs are emitted, by which point passes in between have
erased some of the instructions they point at. Resolve them at the
start of each allocator run and before emitting.
SlotIndexes can then reclaim the entries of erased instructions without
sparing the ones held here, which would have made generated code depend
on -g. Emitted locations are unchanged, except that intervals resolving
to one position now emit a single DBG_VALUE rather than identical
consecutive ones.
[SlotIndexes] Add queries for stale indexes
An erased instruction leaves its index list entry in place, making the
index indistinguishable from a block boundary entry. Add
isBlockBoundaryIndex() and isStaleIndex() to tell the two apart, and
canonicalizeIndex() to resolve a stale index to the closest preceding
instruction's register slot, or the block start if none survives.
NFC. No caller yet. LiveDebugVariables is next.
bpf: add XOR and modulo instructions
libpcap can generate XOR and modulo instructions, but the kernel
interpreter and validator do not support them.
Add constant and register operands for both operations. Reject constant
modulo by zero and return zero for a register zero divisor, as for DIV.
Update the manual to match.
Patch-by: guy
Bug: https://bugs.dragonflybsd.org/issues/3387
Move dataset encryption to zfs.resource.encryption
## Problem
Dataset encryption still lived in the `pool.dataset` namespace even though `zfs.resource` is now the dataset API, and zr's own create path had to call back into `pool.dataset.insert_or_update_encrypted_record` to record keys.
## Solution
- **New `zfs.resource.encryption` sub-service** owning lock, unlock, unlock_summary, export_key, export_keys, export_replication_keys, change_key and inherit, with zr conventions (`path`, lowercase enums, `Secret` keys, `ZFS_RESOURCE_*` roles, audit). Public methods delegate to private `*_impl` methods for thread local storage, as the rest of zr does, and everything is synchronous except starting attachment delegates, whose API is async.
- **Private methods moved and renamed** (store_key, delete_keys, stored_keys, sync_keys, encryption_roots, encryption_state, replication_keys, encryption_root_mapping, unlock_impl), together with the `storage_encrypteddataset` model. Every internal caller (failover, pool create/import/export, KMIP, replication, zfs events, zr create) now calls zr directly; no `pool.dataset` alias is left.
- **`pool.dataset` encryption methods are thin shims** keeping their models, roles and pipes. They call the zr `*_impl` methods through the namespace with their own job, so no nested jobs are created, and they share zr's job locks so both APIs serialize on the same dataset. Error messages are unchanged; validation attributes now follow zr names.
- `unlock_impl` has a private Secret-typed accepts model so passphrases handed over by failover stay redacted in job listings.
- Hook names and payloads, including the uppercase key formats failover relies on, are unchanged, and nothing renamed is called across HA controllers.
X86: Preserve the dead flag clobber when expanding dynamic allocas (#227702)
The DYN_ALLOCA pseudos clobber EFLAGS, so propagate the pseudo's dead
flag to the stack adjustment they expand to.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
Grey out the WiFi list while a rescan is running
A connection started during a rescan could be overridden when the scan
finished. Rescan re-enabled itself and the status line went back to
"not connected" while the attempt was still running. Disabling the
list until the scan is done stops a connection from starting mid-scan.
[LifetimeSafety] Fix capture_by argument mapping for explicit object params (#228979)
`isInstance()` is also true for explicit object member functions, but
their object argument binds to a real parameter, so arguments and
parameters line up one-to-one. This commit fixes the problem by using
`isImplicitObjectMemberFunction()`.
[TableGen] Avoid crashing on a missing instruction pattern result (#228075)
When an instruction pattern's result name does not match the declared
output operand, TableGen diagnoses the mismatch but continues and
dereferences InstResults.end().
Make this diagnostic fatal, preserving the existing message and pattern
dump while stopping before the invalid access.
Add regression coverage for invalid and valid result names with both
-gen-instr-info and -gen-dag-isel. This test fails without this patch
This crash seems to have been present since:
`635debe85beb7fbb54590f20b9720d99f938cb84`
[MoveAutoInit] Don't move auto-init instructions into EH pad blocks (#222100)
`BasicBlock::getFirstInsertionPt()` skips a leading EH pad, but the
MemorySSA update registers the moved instruction with
`InsertionPlace::Beginning`. A `CatchPadInst` is itself a `MemoryDef`,
so instruction order and access order disagree and `-verify-memoryssa`
asserts. Extend the existing `CatchSwitchInst` guard to all EH pads.
Fixes #221568.
Written with claude-code (Opus 5); reviewed and tested locally.
Mips: Stop setting kill flags on virtual registers before FinalizeISel (#229018)
There is no point in maintaining these before register allocation anymore.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
drm: reject unload while core teardown is incomplete
Unloading drm.ko reaches ttm_exit(), which waits for device_released.
The callback which sets that flag is compiled out and device unregister
is a stub, so kldunload sleeps indefinitely while holding the linker
lock. The module event handler has already cleared the Linux task and
process cleanup callbacks by then.
DRM also retains worker threads and undrained RCU callouts, so removing
the TTM wait alone would not make unloading safe. Return EBUSY from
MOD_UNLOAD before changing callbacks or entering SYSUNINIT. Keep the
module usable until complete teardown is implemented.
Bug: https://bugs.dragonflybsd.org/issues/3443
[VectorCombine] Fold interleave and widen chained operations (#224005)
Generalize the existing single deinterleave-interleave pair fold by starting from the interleave
and walking backwards through its operands.
Starting from the interleave instead exposes the whole expression tree that produces its operands,
allowing the combine to discover multiple deinterleaves participating in the same reconstructed vector.
The walk follows supported element-wise operations and splats backwards until it reaches the originating deinterleaves. Once all interleave operands can be traced back consistently, the operations can be rebuilt
on the original wider vectors and the intermediate deinterleave-interleave operations removed.
This makes the fold handle patterns with multiple deinterleaved inputs.
vm: use normal COW inheritance for user-wired mappings
Forking an mlock()ed MAP_PRIVATE file mapping can panic with
"vm_fault_copy_wired: page missing". The wired-copy path expects the
page in the front object, but it may be in a backing object or have
been removed after the file was truncated.
Use normal COW inheritance for normal mappings with only a user wire.
The parent stays user-wired and the child remains unwired. The normal
fault path resolves backing pages and handles pager errors. Keep eager
copying for hard-wired and virtual-page-table mappings.
Reviewed-by: dillon
Bug: https://bugs.dragonflybsd.org/issues/3433