LLVM/project 14a00a9 — clang/docs ReleaseNotes.md, clang/test/AST/ByteCode new-delete.cpp

fixup! [Clang] Add missing release note entry in GH226753
DeltaFile
+0-29clang/test/SemaCXX/new-nothrow-by-value.cpp
+17-0clang/test/AST/ByteCode/new-delete.cpp
+1-1clang/docs/ReleaseNotes.md
+18-303 files

FreeBSD/ports 956f979 — x11-servers/xlibre-server Makefile Makefile.common

x11-servers/xlibre-server: Fix some dependency paths

- Bump PORTREVISION

With hat:       xlibre
PR:             298682
Reported by:    mirror176 __at_ hotmail.com
DeltaFile
+2-2x11-servers/xlibre-server/Makefile.common
+1-0x11-servers/xlibre-server/Makefile
+3-22 files

LLVM/project 94b07b7 — llvm/lib/CodeGen SpillPlacement.cpp, llvm/test/CodeGen/AMDGPU wave-profile-spill-zero-cost.mir

[CodeGen] Keep wave-profiled spill frequencies positive

SpillPlacement expects positive block weights, but a valid wave profile can
record zero executions for a CFG-reachable block. Giving such a block zero
spill cost can make the allocator choose a very different placement.

Clamp every accepted wave-derived frequency to at least one, as we already
do for nonzero counts that round down to zero. Unmeasured or rejected blocks
still use their existing MBFI frequency. Add a focused MIR test for a valid
zero-wave record.

This pattern arose in a profiled Composable Kernel convolution case. With
the separate spill correctness fixes and partial spilling enabled, the
zero-cost policy failed two CPU-reference checks; the positive floor passed
both. The test checks the cost directly; the application result was checked
separately on gfx950.

DeltaFile
+27-0llvm/test/CodeGen/AMDGPU/wave-profile-spill-zero-cost.mir
+3-2llvm/lib/CodeGen/SpillPlacement.cpp
+30-22 files

LLVM/project 5d86702 — llvm/lib/CodeGen SpillPlacement.cpp, llvm/test/CodeGen/AMDGPU wave-profile-spill.mir

[CodeGen] Use validated wave counts for AMDGPU spill costs

Lane-based block frequencies can understate the cost of a spill in
divergent GPU code: a wave still executes a block with only some lanes
active. Use measured block-wave counts to weight SpillPlacement's costs
relative to the original entry-wave count.

Opt in on AMDGPU only. Accept a measured count only when its IR block
maps uniquely to a machine block with matching predecessors and
successors. Keep the existing MBFI cost for unmeasured or rejected
blocks, and leave branch probabilities and general BFI unchanged.



DeltaFile
+251-0llvm/test/CodeGen/AMDGPU/wave-profile-spill.mir
+97-1llvm/lib/CodeGen/SpillPlacement.cpp
+348-12 files

LLVM/project 9a0c255 — llvm/include/llvm/ProfileData InstrProf.h, llvm/lib/ProfileData InstrProf.cpp

[InstrProf] Replace !PGOFuncName and !PGOName metadata with !guid (#214134)

!PGOFuncName and !PGOName metadata were attached to internal functions
and vtables during profile annotation to record their original
"<file>;<name>" PGO names before ThinLTO promoted and renamed them. In
post-link LTO passes, InstrProfSymtab read that metadata back so profile
records keyed by the original name's hash could still find the renamed
IR object.

Global objects now carry stable !guid metadata assigned before LTO
renaming, which records the MD5 hash of the original PGO name directly.
See: https://discourse.llvm.org/t/rfc-keep-globalvalue-guids-stable/84801

Use !guid instead of maintaining separate PGO name metadata:

- Stop emitting and reading !PGOFuncName and !PGOName in Clang and
  PGOInstrumentation, and remove the metadata helper functions
  (createPGOFuncNameMetadata, createPGONameMetadata,
  getPGOFuncNameMetadata, and the metadata name getters).

    [13 lines not shown]
DeltaFile
+29-83llvm/lib/ProfileData/InstrProf.cpp
+102-9llvm/unittests/ProfileData/InstrProfTest.cpp
+15-30llvm/include/llvm/ProfileData/InstrProf.h
+44-0llvm/test/tools/llvm-profdata/Inputs/guid-roundtrip.proftext
+35-0llvm/test/tools/llvm-profdata/guid-roundtrip.test
+28-0llvm/test/tools/llvm-profdata/Inputs/guid-roundtrip.c
+253-12217 files not shown
+291-17823 files

LLVM/project e521fe9 — llvm/lib/Transforms/Utils LowerSwitch.cpp, llvm/test/Transforms/LowerSwitch wave-profile.ll profile-weights.ll

[Transforms] Preserve wave profiles across CFG rewrites

HIP device PGO attaches measured wave counts to IR blocks. Later CFG
rewrites can drop counts from unchanged blocks or leave stale counts
on blocks that now execute differently. Either case makes the profile
unreliable for later optimizations.

Preserve counts through switch lowering, structurization, and loop
rotation only when a block still represents the same executions. Keep
unaffected counts and their IDs even when a loop header's count must be
invalidated. Transfer branch hints only for equivalent decisions, and
avoid assigning switch weights when default traffic cannot be traced
to one edge.




DeltaFile
+157-0llvm/test/Transforms/LowerSwitch/profile-weights.ll
+113-10llvm/lib/Transforms/Utils/LowerSwitch.cpp
+68-0llvm/unittests/Transforms/Utils/LoopRotationUtilsTest.cpp
+56-0llvm/test/Transforms/LowerSwitch/wave-profile.ll
+53-0llvm/test/Transforms/StructurizeCFG/wave-profile-loop-prefix.ll
+50-0llvm/test/Transforms/StructurizeCFG/wave-profile.ll
+497-108 files not shown
+629-5114 files

LLVM/project 5ef0a66 — llvm/lib/Transforms/Instrumentation PGOInstrumentation.cpp, llvm/test/Transforms/PGOProfile dense-wave-cfg.ll wave-profile-use.ll

[PGO] Load dense block wave counts from device profiles

Use the profile's dense layout flag to map appended wave-only slots after
the original block/select prefix. Reuse the producer's block selection so
eligible loop and reconvergence blocks retain their directly measured wave
frequencies. Keep select slots out of the block mapping.

Test dense and sparse profiles, measured zeros, generation and metadata
switches, and exact loop block identities after critical-edge splitting.
DeltaFile
+49-3llvm/test/Transforms/PGOProfile/wave-profile-use.ll
+15-1llvm/test/Transforms/PGOProfile/dense-wave-cfg.ll
+11-4llvm/lib/Transforms/Instrumentation/PGOInstrumentation.cpp
+75-83 files

LLVM/project dcef0bb — llvm/lib/Transforms/Instrumentation PGOInstrumentation.cpp, llvm/test/Transforms/PGOProfile wave-profile-use.ll

[PGO] Add a debugging switch for wave profile metadata

Add the hidden pgo-wave-metadata option, enabled by default, for debugging,
performance comparisons and disabling wave annotations when investigating
regressions without turning off ordinary PGO or uniformity hints.

Gate wave metadata emission while retaining the existing clearing of stale
function and block annotations during profile use. Profile collection and
ordinary count reconstruction are unchanged.

Test default/explicit enablement, disabling, retained counts and uniformity
hints, and replacement profiles with missing, mismatched or zero counts.

DeltaFile
+34-1llvm/test/Transforms/PGOProfile/wave-profile-use.ll
+6-1llvm/lib/Transforms/Instrumentation/PGOInstrumentation.cpp
+40-22 files

LLVM/project e655a71 — llvm/lib/Transforms/Instrumentation PGOInstrumentation.cpp, llvm/test/Transforms/PGOProfile wave-profile-use.ll

[PGO] Load GPU wave counts into IR metadata

GPU profiles contain wave counts alongside lane counts, but profile use
does not expose them to optimizations. Wave visits do not obey scalar
flow conservation, so unmeasured blocks cannot use counts reconstructed
from neighboring blocks or ordinary branch weights.

Map wave-counter indices to the blocks selected by PGO instrumentation,
after reproducing its critical-edge splits. Attach measured counts using
wave.profile metadata, retaining measured zeros and marking other blocks
unmeasured. Require a measured entry count for normalization and exclude
select-counter slots from the block mapping.

Validate the wave-counter layout against the accepted lane profile.
Clear old wave metadata when loading a replacement profile, including
when a function has no usable record. Do not emit wave metadata for
previously profiled functions: their branch weights may change counter
placement without changing the CFG hash. Keep ordinary lane-count
reconstruction and branch weights unchanged.

DeltaFile
+212-0llvm/test/Transforms/PGOProfile/wave-profile-use.ll
+44-0llvm/lib/Transforms/Instrumentation/PGOInstrumentation.cpp
+256-02 files

LLVM/project 37656e7 — llvm/docs LangRef.md, llvm/include/llvm/IR ProfDataUtils.h

[IR] Define GPU wave-profile metadata

Existing offload GPU profile counters measure lane executions, while
GPU instructions execute at wave granularity under an active-lane
mask. A block visited by every wave can therefore look cold when only
a few lanes are active. Scalar branch weights also cannot represent a
divergent wave visiting both successors before reconverging. These
profiles are a poor fit for optimizations that estimate work performed
by a wave.

Lane counters remain useful for measuring per-lane branch selectivity
and estimating work that scales with the number of active lanes.
Wave counters cannot replace them: a visit with one active lane and a
visit with all lanes active both count as one. The two profiles provide
complementary information about instruction execution and lane activity.

Introduce wave.profile and wave.profile.block metadata to represent
measured wave visits to IR blocks. Each dynamic visit with at least one
active lane contributes one to the count. A measured zero is distinct

    [18 lines not shown]
DeltaFile
+817-0llvm/unittests/IR/ProfDataUtilsTest.cpp
+396-0llvm/lib/IR/ProfDataUtils.cpp
+91-0llvm/test/Verifier/wave-profile.ll
+74-0llvm/include/llvm/IR/ProfDataUtils.h
+47-0llvm/docs/LangRef.md
+40-0llvm/test/Bitcode/wave-profile.ll
+1,465-05 files not shown
+1,525-011 files

LLVM/project 76c7824 — llvm/lib/Transforms/Instrumentation PGOInstrumentation.cpp, llvm/test/Transforms/PGOProfile dense-wave-instrumentation.ll dense-wave-profile-use.ll

[PGO] Collect dense AMDGPU block wave counts

Wave visits are not additive across divergent control flow. Sparse scalar
counter sites can leave repeated loop blocks unmeasured, so their wave
frequencies cannot be reconstructed from entry and edge counts.

Append zero-step instrumentation for eligible unmeasured AMDGPU blocks,
keeping the existing lane and select counter indices unchanged. Exclude the
appended slots from lane-flow reconstruction and uniformity annotation.

Identify the layout with a profile variant bit, preserve it through raw and
indexed readers/writers, and select it automatically during profile use.
Keep sparse profiles readable and reject incompatible merges, including
concatenated raw profiles. Ignore empty merge-worker contexts.

Enable dense collection for ordinary AMDGPU IR-PGO by default, with the
hidden -pgo-instrument-dense-wave-counts option for debugging. Leave
context-sensitive, coverage, and temporal instrumentation unchanged.


    [2 lines not shown]
DeltaFile
+108-0llvm/test/Transforms/PGOProfile/dense-wave-cfg.ll
+82-0llvm/test/Transforms/PGOProfile/dense-wave-profile-use.ll
+60-8llvm/lib/Transforms/Instrumentation/PGOInstrumentation.cpp
+59-0llvm/test/Transforms/PGOProfile/dense-wave-instrumentation.ll
+43-0llvm/test/tools/llvm-profdata/dense-wave-layout.test
+31-0llvm/unittests/ProfileData/InstrProfTest.cpp
+383-88 files not shown
+441-1414 files

LLVM/project 78f2beb — mlir/include/mlir/Dialect/SCF/IR SCFOps.td, mlir/lib/Dialect/SCF/IR SCF.cpp

[mlir][scf] Add unsignedCmp to scf.parallel and use it in the tiling in-bound check (#226130)

This change mirrors `scf.for` where `scf.parallel` gets an `unsignedCmp`
unit attribute, and every pass that rebuilds or lowers a `scf.parallel`
has to respect it:

- Propagate where a loop is rebuilt from another, such as parallel loop
tiling, parallel loop fusion (which also refuses to fuse loops of
different signedness), parallel-to-nested-fors and SCF-to-CF lowering.
- Decline where the bound arithmetic assumes signed values, such as
SCF-to-GPU, SCF-to-OpenMP, async-parallel-for and the
parallel-loop-collapsing test pass.

Fixes #223233
DeltaFile
+44-2mlir/test/Dialect/SCF/parallel-loop-tiling-inbound-check.mlir
+36-0mlir/test/Dialect/SCF/parallel-loop-fusion.mlir
+22-6mlir/include/mlir/Dialect/SCF/IR/SCFOps.td
+15-4mlir/lib/Dialect/SCF/IR/SCF.cpp
+16-0mlir/test/Dialect/Async/async-parallel-for-async-dispatch.mlir
+16-0mlir/test/Conversion/SCFToControlFlow/convert-to-cfg.mlir
+149-1213 files not shown
+243-2119 files

FreeBSD/ports a48d144 — www/qhttpengine Makefile

www/qhttpengine: deprecate

PR:             289548
Reported by:    Daniel Engberg <diizzy at FreeBSD.org>
DeltaFile
+3-0www/qhttpengine/Makefile
+3-01 files

FreeBSD/src c52f99a — sys/amd64/amd64 machdep.c

amd64: use WRMSRNS immediate form to update splitlock control, when available

(cherry picked from commit 20c09e6fece29dbcf4835b4dd90e969f1b9b3e59)
DeltaFile
+34-4sys/amd64/amd64/machdep.c
+34-41 files

FreeBSD/src 0dce4f2 — sys/amd64/amd64 mp_machdep.c machdep.c, sys/amd64/include md_var.h pcpu.h

amd64: cache MSR_MEMORY_CTL in pcpu

(cherry picked from commit 74d325fee2dbcb5f8541819481ce76f87c151a21)
DeltaFile
+11-0sys/amd64/amd64/initcpu.c
+3-4sys/amd64/amd64/machdep.c
+2-1sys/amd64/include/pcpu.h
+2-0sys/amd64/amd64/mp_machdep.c
+1-0sys/amd64/include/md_var.h
+19-55 files

FreeBSD/src d95433e — sys/amd64/amd64 sys_machdep.c, sys/x86/include sysarch.h

amd64: add userspace control for disabling splitlocks

(cherry picked from commit 51cea5416a2421f29d22100bc71ab6486b1dc54a)
DeltaFile
+26-1sys/amd64/amd64/sys_machdep.c
+2-0sys/x86/include/sysarch.h
+28-12 files

FreeBSD/src 6d6a824 — sys/amd64/amd64 vm_machdep.c exec_machdep.c, sys/amd64/include proc.h md_var.h

amd64: support for tracking per-thread 'disable splitlocks' state

(cherry picked from commit e318f24c0f53b52bd280b827509ad752db0497bf)
DeltaFile
+36-0sys/amd64/amd64/machdep.c
+14-0sys/amd64/amd64/pmap.c
+10-0sys/amd64/amd64/exec_machdep.c
+5-0sys/amd64/include/md_var.h
+2-0sys/amd64/include/proc.h
+2-0sys/amd64/amd64/vm_machdep.c
+69-03 files not shown
+72-09 files

FreeBSD/src ca7aa1c — sys/amd64/include proc.h

amd64: add md thread flags word

(cherry picked from commit dc975612213369f597449ced5a3c5edabc2aae6d)
DeltaFile
+1-0sys/amd64/include/proc.h
+1-01 files

FreeBSD/src b4a9afd — sys/amd64/amd64 machdep.c initcpu.c, sys/amd64/include md_var.h

amd64: calculate if hardware supports disabling splitlocks

(cherry picked from commit d8c1abb0e3409314fb89400bc85b5051649e91ab)
DeltaFile
+18-0sys/amd64/amd64/initcpu.c
+11-0sys/amd64/amd64/machdep.c
+3-0sys/amd64/include/md_var.h
+32-03 files

FreeBSD/src 77a99f3 — sys/amd64/amd64 trap.c

amd64: handle #AC in kernel mode

(cherry picked from commit c6d7235ac15d208435a839bf847dd898d9b1d9fe)
DeltaFile
+8-0sys/amd64/amd64/trap.c
+8-01 files

FreeBSD/src 54cb4bf — sys/amd64/include cpufunc.h

amd64: gate WRMSRNS immediate form on compiler support

(cherry picked from commit 53fe018630d06eeef6791cfacaa6dc9acdf4d669)
DeltaFile
+10-0sys/amd64/include/cpufunc.h
+10-01 files

FreeBSD/src 61412c0 — sys/amd64/include cpufunc.h

amd64 cpufunc.h: add WRMSRNS helpers

(cherry picked from commit 4a8acccc2e59b6ac5f4068471d94b40347ebeff0)
DeltaFile
+13-0sys/amd64/include/cpufunc.h
+13-01 files

FreeBSD/src ee05a36 — sys/netpfil/pf pf_table.c, tests/sys/netpfil/pf table.sh

pf: do not loop on an address that is cleared twice in pfr_clr_astats()

pfr_clr_astats() looks up each address it is given and inserts the entry
it finds at the head of a work queue. If the same address is given more
than once, the entry is inserted twice and the second insertion makes it
its own successor. pfr_clstats_kentries() then walks the queue forever,
with the rules lock held for writing, so packet processing and every
other pf operation in that vnet stop as well. To reproduce:

pfctl -e
pfctl -t foo -T add 192.0.2.1
pfctl -t foo -T zero 192.0.2.1 192.0.2.1

Do as pfr_del_addrs() does: clear pfrke_mark on the entries named, then
queue an entry only the first time it is seen. An address given more
than once is cleared, and counted, once. Validate all addresses before
any entry is touched.

Add a regression test.

    [6 lines not shown]
DeltaFile
+34-0tests/sys/netpfil/pf/table.sh
+9-2sys/netpfil/pf/pf_table.c
+43-22 files

NetBSD/pkgsrc N6ySmIO — mk/compiler ccache.mk

   ccache.mk: Mark gmake as circular

   ccache3 has used gmake to build for a very long time, but it was not
   marked as a circular dependency.  Resolves build failure under bob
   when PKGSRC_COMPILER has ccache.
VersionDeltaFile
1.46+2-1mk/compiler/ccache.mk
+2-11 files

FreeBSD/ports 7f0bbeb — math/cvc5 Makefile pkg-plist

math/cvc5: Fix plist when JAVA=OFF

PR:             280044
Reported by:    iron.udjin at gmail.com
DeltaFile
+1-1math/cvc5/pkg-plist
+1-0math/cvc5/Makefile
+2-12 files

LLVM/project 7b5ec23 — llvm/test/tools/llvm-readobj/ELF/RISCV eflags-abi.test eflags.test

[RISCV][llvm-readobj] Generalize the eflags-abi test to cover all of the eflags. NFC (#227152)

Assisted-by: Claude
DeltaFile
+56-0llvm/test/tools/llvm-readobj/ELF/RISCV/eflags.test
+0-51llvm/test/tools/llvm-readobj/ELF/RISCV/eflags-abi.test
+56-512 files

FreeBSD/src f084f28 — sys/netpfil/pf pf_table.c, tests/sys/netpfil/pf table.sh

pf: fix NULL dereference in pfr_set_addrs() with feedback

Since DIOCRSETADDRS was converted to netlink, pf_handle_table_set_addrs()
calls pfr_set_addrs() with a NULL size2, as the netlink interface has no
buffer to return the deleted addresses in. pfr_set_addrs() only checked
size2 for NULL at the end of the function; with PFR_FLAG_FEEDBACK set it
dereferenced it unconditionally first. pfctl sets PFR_FLAG_FEEDBACK
when run with -v, so "pfctl -v -t foo -T replace ..." panicked the
kernel with a NULL pointer dereference. To reproduce:

pfctl -e
pfctl -t foo -T add 192.0.2.1
pfctl -v -t foo -T replace 192.0.2.2

Check size2 for NULL before dereferencing it, as is already done at the
end of the function. The per-address feedback for added and changed
addresses is still copied back as before; only the list of deleted
addresses, which the netlink caller has no room for, is skipped.


    [9 lines not shown]
DeltaFile
+34-0tests/sys/netpfil/pf/table.sh
+2-2sys/netpfil/pf/pf_table.c
+36-22 files

HardenedBSD/src debbf77 — share/man/man4 wsp.4 uep.4, sys/dev/hwpmc hwpmc_amd.h hwpmc_mod.c

Merge remote-tracking branch 'rad/hardened/current/master' into hardened/current/cross-dso-cfi
DeltaFile
+489-18sys/dev/hwpmc/hwpmc_amd.c
+69-0sys/dev/hwpmc/hwpmc_mod.c
+19-18share/man/man4/uep.4
+8-12share/man/man4/wsp.4
+11-0sys/dev/hwpmc/hwpmc_amd.h
+5-0sys/sys/pmc.h
+601-486 files

HardenedBSD/src 17cf76e — share/man/man4 wsp.4 uep.4, sys/dev/hwpmc hwpmc_amd.h hwpmc_mod.c

Merge remote-tracking branch 'rad/hardened/current/master' into hardened/current/pledge
DeltaFile
+489-18sys/dev/hwpmc/hwpmc_amd.c
+69-0sys/dev/hwpmc/hwpmc_mod.c
+19-18share/man/man4/uep.4
+8-12share/man/man4/wsp.4
+11-0sys/dev/hwpmc/hwpmc_amd.h
+5-0sys/sys/pmc.h
+601-486 files

LLVM/project 619028d — llvm/test lit.cfg.py, llvm/utils profcheck-xfail.txt

[ProfCheck] Exclude DirectX (#227157)

I don't think anyone runs profiling on DirectX, so exclude it for now.
Also move AMDGPU to the normal exclusion list given there are efforts
around PGO for AMDGPU currently.
DeltaFile
+12-4llvm/test/lit.cfg.py
+0-1llvm/utils/profcheck-xfail.txt
+12-52 files