HBSD: Disable JIT in devel/pcre2
PCRE2 JIT has long been a source of frustration on HardenedBSD because
of our PaX NOEXEC implementation. Now that the PCRE2 JIT implementation
has had a recent security advisory (a buffer overflow, nonetheless), we
probably should not trust PCRE2's JIT.
Signed-off-by: Shawn Webb <shawn.webb at hardenedbsd.org>
See-Also: GHSA-r9hj-j2rw-4q3m
[flang][cuda] Add cuf.on_device op and lower to it (#227089)
Replace the CUFFunctionRewrite pass, which matched fir.call names and
folded them to constants, with a cuf.on_device operation emitted when
the intrinsic is lowered. CUFOpConversion folds the operation once the
code is in its host or device context, and leaves the host copy of an
OpenACC routine unfolded so the device clone is not baked to false. Any
operation that remains is folded by the late CUF conversion, and the
rewrite pass is dropped from the pipeline.
Revert "releand "[clang-repl] Implement IncrementalHIPDeviceParser for HIP device compilation"" (#227123)
Reverts llvm/llvm-project#226930
AMD author has an additional step to perform, that is still needed.
in contact with Author.
[CIR] Support aggregate co_await / co_yield in AggExprEmitter (#225412)
Support evaluating `co_await` and `co_yield` expressions whose result is
an aggregate type in `AggExprEmitter`.
Fixes #225317
[TailCallElim] Add profile annotations to return value selects
Tail call elimination in some cases can create selects on possible
return values conditioned on whether or not execution is currently in
what was a recursive call. That is equal to the probability with which
we recurse, which in turn can be computed from the block frequencies of
blocks that recurse and blocks that directly return.
Reviewers: mtrofin
Reviewed By: mtrofin
Pull Request: https://github.com/llvm/llvm-project/pull/202518
[SelectionDAG] Handle constants in SimplifyMultipleUseDemandedBits
Replace a non-zero constant with zero when none of its set bits are
demanded.
This allows users of `SimplifyMultipleUseDemandedBits` to eliminate
irrelevant constant bits while preserving the convention that a null
SDValue indicates no simplification.
[Clang] Add missing release note entry in #226753 (#227078)
As per the feedback from #226753, we add release note for GH-212211.
Also move the test to new-delete.cpp.
Assisted-by: Claude
[CodeGen] Enable validated AMDGPU wave spill costs by default
Use available validated wave counts for AMDGPU spill placement without an
explicit opt-in. Keep the existing target, mapping and normalization checks
and the ordinary block-frequency fallback for unavailable or rejected data.
Retain -enable-wave-profiled-spill=false for debugging and matched performance
comparisons. Test default behavior, explicit disabling, equivalence to explicit
enabling, fallback cases and the positive frequency floor.
[RISCV][P-ext] Support the missing packed pair (#226992)
This patch supports the packed pair for v2i16.
Some are inconsistent with the spec because existing shuffle-lowering
transforms them into equivalent narrowing shifts or zips.
ipfw: guard against NULL deref with clat/plat prefixes without length
Found with: Claude Code Sonnet 5
MFC after: 2 weeks
(cherry picked from commit 47b02ace19c53e72ad2290d974025292ff97464a)
[AMDGPU] Fold sudot intrinsics with a zero multiplicand
Replace sudot4 and sudot8 with their accumulator when either
multiplicand is zero.
For example:
```
sudot4(sign0, x, sign1, 0, acc, clamp)
->
acc
```
[AMDGPU] Canonicalize constant operands of sudot intrinsics (#226478)
Move a constant multiplicand and its sign flag to the second operand.
For example:
```
sudot4(true, 1, false, x, acc, clamp)
->
sudot4(false, x, true, 1, acc, clamp)
```
exec: Remove an unneeded capability mode check
The subsequent namei() call is relative to AT_FDCWD, and such lookups
are always disallowed in capability mode.
No functional change intended.
Reviewed by: emaste
Differential Revision: https://reviews.freebsd.org/D59888
sysctl: Return ECAPMODE when trying to access sysctls in capability mode
We have always returned EPERM in this case, but it's incorrect, we
should return ECAPMODE for capability mode violations. Fix the errno
value.
Reviewed by: emaste
MFC after: 2 weeks
Differential Revision: https://reviews.freebsd.org/D59887
revert changing some busy waits to tsleep
landry's ThinkPad T470s (Kaby Lake) with external HDMI monitor could
no longer run X, stuck in a loop of:
'modeset(0): hotplug event: connector 106's link-state is BAD'
[CodeGen] Keep wave-profiled spill frequencies positive
SpillPlacement expects positive block weights, but a valid wave profile can
record zero executions for a CFG-reachable block. Giving such a block zero
spill cost can make the allocator choose a very different placement.
Clamp every accepted wave-derived frequency to at least one, as we already
do for nonzero counts that round down to zero. Unmeasured or rejected blocks
still use their existing MBFI frequency. Add a focused MIR test for a valid
zero-wave record.
This pattern arose in a profiled Composable Kernel convolution case. With
the separate spill correctness fixes and partial spilling enabled, the
zero-cost policy failed two CPU-reference checks; the positive floor passed
both. The test checks the cost directly; the application result was checked
separately on gfx950.
[CodeGen] Use validated wave counts for AMDGPU spill costs
Lane-based block frequencies can understate the cost of a spill in
divergent GPU code: a wave still executes a block with only some lanes
active. Use measured block-wave counts to weight SpillPlacement's costs
relative to the original entry-wave count.
Opt in on AMDGPU only. Accept a measured count only when its IR block
maps uniquely to a machine block with matching predecessors and
successors. Keep the existing MBFI cost for unmeasured or rejected
blocks, and leave branch probabilities and general BFI unchanged.
[InstrProf] Replace !PGOFuncName and !PGOName metadata with !guid (#214134)
!PGOFuncName and !PGOName metadata were attached to internal functions
and vtables during profile annotation to record their original
"<file>;<name>" PGO names before ThinLTO promoted and renamed them. In
post-link LTO passes, InstrProfSymtab read that metadata back so profile
records keyed by the original name's hash could still find the renamed
IR object.
Global objects now carry stable !guid metadata assigned before LTO
renaming, which records the MD5 hash of the original PGO name directly.
See: https://discourse.llvm.org/t/rfc-keep-globalvalue-guids-stable/84801
Use !guid instead of maintaining separate PGO name metadata:
- Stop emitting and reading !PGOFuncName and !PGOName in Clang and
PGOInstrumentation, and remove the metadata helper functions
(createPGOFuncNameMetadata, createPGONameMetadata,
getPGOFuncNameMetadata, and the metadata name getters).
[13 lines not shown]
[Transforms] Preserve wave profiles across CFG rewrites
HIP device PGO attaches measured wave counts to IR blocks. Later CFG
rewrites can drop counts from unchanged blocks or leave stale counts
on blocks that now execute differently. Either case makes the profile
unreliable for later optimizations.
Preserve counts through switch lowering, structurization, and loop
rotation only when a block still represents the same executions. Keep
unaffected counts and their IDs even when a loop header's count must be
invalidated. Transfer branch hints only for equivalent decisions, and
avoid assigning switch weights when default traffic cannot be traced
to one edge.
[PGO] Load dense block wave counts from device profiles
Use the profile's dense layout flag to map appended wave-only slots after
the original block/select prefix. Reuse the producer's block selection so
eligible loop and reconvergence blocks retain their directly measured wave
frequencies. Keep select slots out of the block mapping.
Test dense and sparse profiles, measured zeros, generation and metadata
switches, and exact loop block identities after critical-edge splitting.
[PGO] Add a debugging switch for wave profile metadata
Add the hidden pgo-wave-metadata option, enabled by default, for debugging,
performance comparisons and disabling wave annotations when investigating
regressions without turning off ordinary PGO or uniformity hints.
Gate wave metadata emission while retaining the existing clearing of stale
function and block annotations during profile use. Profile collection and
ordinary count reconstruction are unchanged.
Test default/explicit enablement, disabling, retained counts and uniformity
hints, and replacement profiles with missing, mismatched or zero counts.
[PGO] Load GPU wave counts into IR metadata
GPU profiles contain wave counts alongside lane counts, but profile use
does not expose them to optimizations. Wave visits do not obey scalar
flow conservation, so unmeasured blocks cannot use counts reconstructed
from neighboring blocks or ordinary branch weights.
Map wave-counter indices to the blocks selected by PGO instrumentation,
after reproducing its critical-edge splits. Attach measured counts using
wave.profile metadata, retaining measured zeros and marking other blocks
unmeasured. Require a measured entry count for normalization and exclude
select-counter slots from the block mapping.
Validate the wave-counter layout against the accepted lane profile.
Clear old wave metadata when loading a replacement profile, including
when a function has no usable record. Do not emit wave metadata for
previously profiled functions: their branch weights may change counter
placement without changing the CFG hash. Keep ordinary lane-count
reconstruction and branch weights unchanged.
[IR] Define GPU wave-profile metadata
Existing offload GPU profile counters measure lane executions, while
GPU instructions execute at wave granularity under an active-lane
mask. A block visited by every wave can therefore look cold when only
a few lanes are active. Scalar branch weights also cannot represent a
divergent wave visiting both successors before reconverging. These
profiles are a poor fit for optimizations that estimate work performed
by a wave.
Lane counters remain useful for measuring per-lane branch selectivity
and estimating work that scales with the number of active lanes.
Wave counters cannot replace them: a visit with one active lane and a
visit with all lanes active both count as one. The two profiles provide
complementary information about instruction execution and lane activity.
Introduce wave.profile and wave.profile.block metadata to represent
measured wave visits to IR blocks. Each dynamic visit with at least one
active lane contributes one to the count. A measured zero is distinct
[18 lines not shown]
[PGO] Collect dense AMDGPU block wave counts
Wave visits are not additive across divergent control flow. Sparse scalar
counter sites can leave repeated loop blocks unmeasured, so their wave
frequencies cannot be reconstructed from entry and edge counts.
Append zero-step instrumentation for eligible unmeasured AMDGPU blocks,
keeping the existing lane and select counter indices unchanged. Exclude the
appended slots from lane-flow reconstruction and uniformity annotation.
Identify the layout with a profile variant bit, preserve it through raw and
indexed readers/writers, and select it automatically during profile use.
Keep sparse profiles readable and reject incompatible merges, including
concatenated raw profiles. Ignore empty merge-worker contexts.
Enable dense collection for ordinary AMDGPU IR-PGO by default, with the
hidden -pgo-instrument-dense-wave-counts option for debugging. Leave
context-sensitive, coverage, and temporal instrumentation unchanged.
[2 lines not shown]