ipfw: guard against NULL deref with clat/plat prefixes without length
Found with: Claude Code Sonnet 5
MFC after: 2 weeks
(cherry picked from commit 47b02ace19c53e72ad2290d974025292ff97464a)
[AMDGPU] Fold sudot intrinsics with a zero multiplicand
Replace sudot4 and sudot8 with their accumulator when either
multiplicand is zero.
For example:
```
sudot4(sign0, x, sign1, 0, acc, clamp)
->
acc
```
[AMDGPU] Canonicalize constant operands of sudot intrinsics (#226478)
Move a constant multiplicand and its sign flag to the second operand.
For example:
```
sudot4(true, 1, false, x, acc, clamp)
->
sudot4(false, x, true, 1, acc, clamp)
```
exec: Remove an unneeded capability mode check
The subsequent namei() call is relative to AT_FDCWD, and such lookups
are always disallowed in capability mode.
No functional change intended.
Reviewed by: emaste
Differential Revision: https://reviews.freebsd.org/D59888
sysctl: Return ECAPMODE when trying to access sysctls in capability mode
We have always returned EPERM in this case, but it's incorrect, we
should return ECAPMODE for capability mode violations. Fix the errno
value.
Reviewed by: emaste
MFC after: 2 weeks
Differential Revision: https://reviews.freebsd.org/D59887
[CodeGen] Keep wave-profiled spill frequencies positive
SpillPlacement expects positive block weights, but a valid wave profile can
record zero executions for a CFG-reachable block. Giving such a block zero
spill cost can make the allocator choose a very different placement.
Clamp every accepted wave-derived frequency to at least one, as we already
do for nonzero counts that round down to zero. Unmeasured or rejected blocks
still use their existing MBFI frequency. Add a focused MIR test for a valid
zero-wave record.
This pattern arose in a profiled Composable Kernel convolution case. With
the separate spill correctness fixes and partial spilling enabled, the
zero-cost policy failed two CPU-reference checks; the positive floor passed
both. The test checks the cost directly; the application result was checked
separately on gfx950.
[CodeGen] Use validated wave counts for AMDGPU spill costs
Lane-based block frequencies can understate the cost of a spill in
divergent GPU code: a wave still executes a block with only some lanes
active. Use measured block-wave counts to weight SpillPlacement's costs
relative to the original entry-wave count.
Opt in on AMDGPU only. Accept a measured count only when its IR block
maps uniquely to a machine block with matching predecessors and
successors. Keep the existing MBFI cost for unmeasured or rejected
blocks, and leave branch probabilities and general BFI unchanged.
[InstrProf] Replace !PGOFuncName and !PGOName metadata with !guid (#214134)
!PGOFuncName and !PGOName metadata were attached to internal functions
and vtables during profile annotation to record their original
"<file>;<name>" PGO names before ThinLTO promoted and renamed them. In
post-link LTO passes, InstrProfSymtab read that metadata back so profile
records keyed by the original name's hash could still find the renamed
IR object.
Global objects now carry stable !guid metadata assigned before LTO
renaming, which records the MD5 hash of the original PGO name directly.
See: https://discourse.llvm.org/t/rfc-keep-globalvalue-guids-stable/84801
Use !guid instead of maintaining separate PGO name metadata:
- Stop emitting and reading !PGOFuncName and !PGOName in Clang and
PGOInstrumentation, and remove the metadata helper functions
(createPGOFuncNameMetadata, createPGONameMetadata,
getPGOFuncNameMetadata, and the metadata name getters).
[13 lines not shown]
[Transforms] Preserve wave profiles across CFG rewrites
HIP device PGO attaches measured wave counts to IR blocks. Later CFG
rewrites can drop counts from unchanged blocks or leave stale counts
on blocks that now execute differently. Either case makes the profile
unreliable for later optimizations.
Preserve counts through switch lowering, structurization, and loop
rotation only when a block still represents the same executions. Keep
unaffected counts and their IDs even when a loop header's count must be
invalidated. Transfer branch hints only for equivalent decisions, and
avoid assigning switch weights when default traffic cannot be traced
to one edge.
[PGO] Load dense block wave counts from device profiles
Use the profile's dense layout flag to map appended wave-only slots after
the original block/select prefix. Reuse the producer's block selection so
eligible loop and reconvergence blocks retain their directly measured wave
frequencies. Keep select slots out of the block mapping.
Test dense and sparse profiles, measured zeros, generation and metadata
switches, and exact loop block identities after critical-edge splitting.
[PGO] Add a debugging switch for wave profile metadata
Add the hidden pgo-wave-metadata option, enabled by default, for debugging,
performance comparisons and disabling wave annotations when investigating
regressions without turning off ordinary PGO or uniformity hints.
Gate wave metadata emission while retaining the existing clearing of stale
function and block annotations during profile use. Profile collection and
ordinary count reconstruction are unchanged.
Test default/explicit enablement, disabling, retained counts and uniformity
hints, and replacement profiles with missing, mismatched or zero counts.
[PGO] Load GPU wave counts into IR metadata
GPU profiles contain wave counts alongside lane counts, but profile use
does not expose them to optimizations. Wave visits do not obey scalar
flow conservation, so unmeasured blocks cannot use counts reconstructed
from neighboring blocks or ordinary branch weights.
Map wave-counter indices to the blocks selected by PGO instrumentation,
after reproducing its critical-edge splits. Attach measured counts using
wave.profile metadata, retaining measured zeros and marking other blocks
unmeasured. Require a measured entry count for normalization and exclude
select-counter slots from the block mapping.
Validate the wave-counter layout against the accepted lane profile.
Clear old wave metadata when loading a replacement profile, including
when a function has no usable record. Do not emit wave metadata for
previously profiled functions: their branch weights may change counter
placement without changing the CFG hash. Keep ordinary lane-count
reconstruction and branch weights unchanged.
[IR] Define GPU wave-profile metadata
Existing offload GPU profile counters measure lane executions, while
GPU instructions execute at wave granularity under an active-lane
mask. A block visited by every wave can therefore look cold when only
a few lanes are active. Scalar branch weights also cannot represent a
divergent wave visiting both successors before reconverging. These
profiles are a poor fit for optimizations that estimate work performed
by a wave.
Lane counters remain useful for measuring per-lane branch selectivity
and estimating work that scales with the number of active lanes.
Wave counters cannot replace them: a visit with one active lane and a
visit with all lanes active both count as one. The two profiles provide
complementary information about instruction execution and lane activity.
Introduce wave.profile and wave.profile.block metadata to represent
measured wave visits to IR blocks. Each dynamic visit with at least one
active lane contributes one to the count. A measured zero is distinct
[18 lines not shown]
[PGO] Collect dense AMDGPU block wave counts
Wave visits are not additive across divergent control flow. Sparse scalar
counter sites can leave repeated loop blocks unmeasured, so their wave
frequencies cannot be reconstructed from entry and edge counts.
Append zero-step instrumentation for eligible unmeasured AMDGPU blocks,
keeping the existing lane and select counter indices unchanged. Exclude the
appended slots from lane-flow reconstruction and uniformity annotation.
Identify the layout with a profile variant bit, preserve it through raw and
indexed readers/writers, and select it automatically during profile use.
Keep sparse profiles readable and reject incompatible merges, including
concatenated raw profiles. Ignore empty merge-worker contexts.
Enable dense collection for ordinary AMDGPU IR-PGO by default, with the
hidden -pgo-instrument-dense-wave-counts option for debugging. Leave
context-sensitive, coverage, and temporal instrumentation unchanged.
[2 lines not shown]
[mlir][scf] Add unsignedCmp to scf.parallel and use it in the tiling in-bound check (#226130)
This change mirrors `scf.for` where `scf.parallel` gets an `unsignedCmp`
unit attribute, and every pass that rebuilds or lowers a `scf.parallel`
has to respect it:
- Propagate where a loop is rebuilt from another, such as parallel loop
tiling, parallel loop fusion (which also refuses to fuse loops of
different signedness), parallel-to-nested-fors and SCF-to-CF lowering.
- Decline where the bound arithmetic assumes signed values, such as
SCF-to-GPU, SCF-to-OpenMP, async-parallel-for and the
parallel-loop-collapsing test pass.
Fixes #223233
pf: do not loop on an address that is cleared twice in pfr_clr_astats()
pfr_clr_astats() looks up each address it is given and inserts the entry
it finds at the head of a work queue. If the same address is given more
than once, the entry is inserted twice and the second insertion makes it
its own successor. pfr_clstats_kentries() then walks the queue forever,
with the rules lock held for writing, so packet processing and every
other pf operation in that vnet stop as well. To reproduce:
pfctl -e
pfctl -t foo -T add 192.0.2.1
pfctl -t foo -T zero 192.0.2.1 192.0.2.1
Do as pfr_del_addrs() does: clear pfrke_mark on the entries named, then
queue an entry only the first time it is seen. An address given more
than once is cleared, and counted, once. Validate all addresses before
any entry is touched.
Add a regression test.
[6 lines not shown]
ccache.mk: Mark gmake as circular
ccache3 has used gmake to build for a very long time, but it was not
marked as a circular dependency. Resolves build failure under bob
when PKGSRC_COMPILER has ccache.