OpenZFS/src cf1bc78 — .github/workflows/scripts qemu-4-build-vm.sh

CI: Enable configure caching by default

Enable configure caching by default in the CI. This doesn't 
significantly speed up the build but it does ensure this
functionality is always tested.

Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Glenn Washburn <development at efficientek.com>
Closes #19173
DeltaFile
+16-4.github/workflows/scripts/qemu-4-build-vm.sh
+16-41 files

LLVM/project 531fe44 — mlir/include/mlir/Dialect/XeGPU/uArch uArchBase.h, mlir/lib/Conversion/VectorToXeGPU VectorToXeGPU.cpp

[mlir][xegpu] Check uArch block shapes in VectorToXeGPU transfer lowering (#217179)

vector.transfer_read/transfer_write were lowered to xegpu.load_nd /
xegpu.store_nd whenever the target chip was pvc, bmg or cri and the
transfer had rank >= 2, without asking whether the target's subgroup 2D
block instructions can access the requested tile at all.
Changing it to query the uArch 2D block load/store description up front
and keep the block path only for tiles whose two innermost dims are each
a multiple of a supported block size- the property the later layout
propagation and blocking passes rely on when they split a tile into
hardware-sized blocks. Everything else takes the scattered
load_gather/store_scatter fallback that both patterns already have.
Deriving 2D block support from the uArch instruction registry also
replaces the hardcoded chip list, resolving the TODO left there; the
three chips that registry covers are the same three that were listed.


assisted by claude
DeltaFile
+70-58mlir/test/Conversion/VectorToXeGPU/transfer-read-to-xegpu.mlir
+93-33mlir/lib/Conversion/VectorToXeGPU/VectorToXeGPU.cpp
+44-21mlir/test/Conversion/VectorToXeGPU/transfer-write-to-xegpu.mlir
+17-13mlir/include/mlir/Dialect/XeGPU/uArch/uArchBase.h
+224-1254 files

LLVM/project c2ccc57 — utils/bazel/llvm-project-overlay/mlir BUILD.bazel

[bazel][MLIR] Fix 98331a0a49e6d4a86d2e39ecc71025ce3b85acd9 (#227500)
DeltaFile
+2-0utils/bazel/llvm-project-overlay/mlir/BUILD.bazel
+2-01 files

LLVM/project b8c4e9f — lldb/test/API/tools/lldb-dap/cancel TestDAP_cancel.py, lldb/test/API/tools/lldb-dap/disconnect TestDAP_disconnect.py

[lldb-dap] End the session before sending the disconnect response (#227106)

Previously the `disconnect` request asks the debug adapter to disconnect
from the debuggee, ending the debug session, and then to shut down.
lldb-dap responded right after killing or detaching the process. So the
`exited` event, terminatedCommands and exitedCommands can be sent after
`disconnect` could reaches the client.

`DisconnectRequestHandler` now stores its response in
`DAP::on_session_end`. `DAP::Loop()` sends it once the session ended.
The teardown/cleanup is now:
- The event threads stop, after they report the exit of a killed
process.
- `terminated` Event is sent, if the event thread didn't send it.
- The transport thread stops.
- If there is a pending `configuration_done` we reply.
- The requests queued behind `disconnect` request are cancelled.
- The `disconnect` response is sent last.
- The debugger is destroyed.

    [17 lines not shown]
DeltaFile
+107-42lldb/tools/lldb-dap/DAP.cpp
+83-9lldb/unittests/DAP/Handler/DisconnectTest.cpp
+35-4lldb/test/API/tools/lldb-dap/disconnect/TestDAP_disconnect.py
+36-0lldb/unittests/DAP/TestBase.h
+33-1lldb/test/API/tools/lldb-dap/cancel/TestDAP_cancel.py
+22-5lldb/tools/lldb-dap/Handler/RequestHandler.h
+316-615 files not shown
+367-7411 files

LLVM/project 40edcbb — llvm/test/Transforms/IndirectBrExpand basic.ll

[IndirectBrExpand] Use UTC for basic.ll

To make it easier to modify in the future.

Reviewers: aeubanks, mtrofin

Pull Request: https://github.com/llvm/llvm-project/pull/227467
DeltaFile
+42-21llvm/test/Transforms/IndirectBrExpand/basic.ll
+42-211 files

NetBSD/src xhsPU2S — sys/arch/x86/include specialreg.h

   Include model 7 in CPUID_TO_MODEL() extended model check.

   Accommodates Zhaoxin CPUs using extended model.
   Update the comment accordingly and fix one typo in it.
VersionDeltaFile
1.223+6-4sys/arch/x86/include/specialreg.h
+6-41 files

LLVM/project 8630009 — clang/docs ReleaseNotes.md, clang/lib/AST ItaniumMangle.cpp

[Clang] Do not assume existing substition for template specialization

Mangled symbol is not always available when having template
specialization. For example, a template alias mangled the type when
actually doing substitution. In previous code, B in A<B>::foo() is
always mangled when reaching A<B>. When doing substition insides A<T>,
B is always available. Alias is not the case here, B is mangled when
reaching A<T> inside.

As now, TemplateName::SubstTemplateTemplateParm might needs to be
resolved further as B is not always defined, we remove it from the
switch case and find recursively like the origianl path.

Assisted-by: Claude # Tests, ReleaseNotes
DeltaFile
+57-52clang/lib/AST/ItaniumMangle.cpp
+15-0clang/test/CodeGenCXX/mangle-template.cpp
+4-0clang/docs/ReleaseNotes.md
+76-523 files

LLVM/project 0cf5430 — bolt/lib/Rewrite DWARFRewriter.cpp, bolt/test/AArch64 go_dwarf.test dwarf2-highpc-addr-form.s

[BOLT] Emit DW_AT_high_pc matching the class of its form (#227066)

DW_AT_high_pc holding an offset from DW_AT_low_pc is a DWARF 4 addition;
in DWARF 2 and 3 the attribute is always class address.

updateLowPCHighPC() always wrote "HighPC - LowPC" while reusing whatever
form the input DIE had, so a DWARF 2 CU using DW_FORM_addr ended up
storing a length where an end address is expected.

BOLT should write the end address when the form is DW_FORM_addr, and
the length otherwise. The default form for a new attribute is now
DW_FORM_addr below DWARF 4, and size needs to be widened to 64 bits so
DW_FORM_data8 won't be truncated. The FunctionRanges.empty() code path
no longer treats a raw DW_FORM_addr high_pc as a size.

Assited-by: opus
DeltaFile
+140-0bolt/test/AArch64/dwarf2-highpc-addr-form.s
+25-8bolt/lib/Rewrite/DWARFRewriter.cpp
+6-1bolt/test/AArch64/go_dwarf.test
+171-93 files

LLVM/project f35bd33 — libc/src/string/memory_utils/aarch64 inline_memcpy.h

[libc][aarch64] Use inline_memcpy_aligned_access_64bit under -mstrict-align (#227428)

Under `-mstrict-align`, clang cannot determine both `src` and `dst` are
aligned. The end of `inline_memcpy_aarch64` aligns `src` but it cannot
confirm `dst` is aligned. As a result,
`builtin::Memcpy<64>::loop_and_tail(dst, src, count)` lowers to
`__builtin_memcpy_inline(..., 64)` but clang emits many many single-byte
load/store instructions to ensure correctness. This can be very slow, so
if unaligned accesses are not supported, opt for adjusting both `src`
and `dst` such that we can use 8-byte loads/stores.
`inline_memcpy_aligned_access_64bit` already does this so we can just
reuse that.
DeltaFile
+17-0libc/src/string/memory_utils/aarch64/inline_memcpy.h
+17-01 files

LLVM/project b41ef0f — llvm/test/CodeGen/Hexagon intrinsics-v67.ll

[Hexagon,test] Enable +audio for intrinsics-v67.ll (#227362)

The test uses llvm.hexagon.M7.vdmpy{,.acc}, which lower to
M7_dcmpyrwc{,_acc}. Those instructions require the audio feature, so add
-mattr=+audio to the RUN line to avoid a predicate-check crash in the
asm printer.
DeltaFile
+1-1llvm/test/CodeGen/Hexagon/intrinsics-v67.ll
+1-11 files

LLVM/project b8952a1 — clang/include/clang/Basic AttrDocs.td

[clang][docs] Use doc links to HLSL/ResourceTypes in AttrDocs.td (#227407)

#222504 made the docs build warn about absolute links to pages of the
same Sphinx project, and the docs build turns warnings into errors. The
HLSL resource attribute docs added in #213346 linked to
https://clang.llvm.org/docs/HLSL/ResourceTypes.html, which after #226729
broke the clang docs build (docs-clang-html, docs-clang-man). Use the
{doc} role instead, like the rest of AttrDocs.td.
DeltaFile
+8-8clang/include/clang/Basic/AttrDocs.td
+8-81 files

LLVM/project ba19a43 — lldb/include/lldb/Core Debugger.h, lldb/source/Commands CommandObjectScripting.cpp

[lldb] Add label, title and divider ANSI settings (#226606)

Commands that print structured listings hard-coded their colors, but
terminal themes render ANSI colors differently, so users couldn't adjust
them the way they can with the prompt, progress or autosuggestion
colors.

Add `label-ansi-prefix`, `title-ansi-prefix` and `divider-ansi-prefix`
settings, each with a matching suffix, for the field labels, entry
titles and dividers such listings are made of. Their defaults are the
colors `scripting extension list` uses today, and it now reads them from
these settings.

Signed-off-by: Med Ismail Bennani <ismail at bennani.ma>
DeltaFile
+48-0lldb/test/API/commands/scripting/extension/TestScriptingExtensionListColors.py
+36-0lldb/source/Core/Debugger.cpp
+21-12lldb/source/Commands/CommandObjectScripting.cpp
+24-0lldb/source/Core/CoreProperties.td
+12-0lldb/include/lldb/Core/Debugger.h
+141-125 files

LLVM/project 83826c2 — flang/include/flang/Optimizer/OpenACC Passes.td, flang/lib/Optimizer/OpenACC/Transforms CMakeLists.txt ACCEraseUnusedKernelAllocations.cpp

[flang][openacc] Add ACCEraseUnusedKernelAllocations pass (#227481)

fir.declare carries a debug-memory effect and fir.freemem is a real
free, so an unused fir.allocmem inside acc.compute_region stays live
through ordinary dead-code elimination and is lowered to a checked
device malloc.

Delete that allocation when it has no uses, or when every use is
fir.freemem, a ViewLikeOpInterface view such as fir.convert, or
fir.declare of that storage. A load, store, or other memory use keeps
it. An unused private-recipe allocation is removed the same way as an
unused source array.
DeltaFile
+122-0flang/lib/Optimizer/OpenACC/Transforms/ACCEraseUnusedKernelAllocations.cpp
+51-0flang/test/Fir/OpenACC/acc-erase-unused-kernel-allocations.mlir
+15-0flang/include/flang/Optimizer/OpenACC/Passes.td
+2-0flang/lib/Optimizer/OpenACC/Transforms/CMakeLists.txt
+190-04 files

LLVM/project 280a20b — mlir/include/mlir/Dialect/GPU/Pipelines Passes.h, mlir/lib/Dialect/GPU/Pipelines GPUToXeVMPipeline.cpp

[MLIR][XeGPU][XeVM] Lower to XeVM: support skipping backend. (#227125)

And skip backend invocation for slow tests.
DeltaFile
+4-1mlir/include/mlir/Dialect/GPU/Pipelines/Passes.h
+2-2mlir/test/Integration/Dialect/XeGPU/WG/simple_mxfp_gemm_quantizeA_F4.mlir
+3-1mlir/lib/Dialect/GPU/Pipelines/GPUToXeVMPipeline.cpp
+1-1mlir/test/Integration/Dialect/XeGPU/WG/simple_mxfp_gemm_quantizeA_F8.mlir
+10-54 files

LLVM/project 1179810 — clang/test/CodeGen/RISCV rvp-intrinsics.c, llvm/test/CodeGen/AMDGPU load-constant-i1.ll amdgcn.bitcast.512bit.ll

Merge branch 'main' into users/adams381/cir-callconv-byval-slot-copy
DeltaFile
+42,441-42,426llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+7,527-7,623llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+4,115-8,359clang/test/CodeGen/RISCV/rvp-intrinsics.c
+4,036-4,380llvm/test/CodeGen/RISCV/rvv/clmul-sdnode.ll
+2,544-2,872llvm/test/CodeGen/RISCV/rvv/clmulh-sdnode.ll
+2,603-2,582llvm/test/CodeGen/AMDGPU/load-constant-i1.ll
+63,266-68,2421,637 files not shown
+154,927-141,4821,643 files

LLVM/project 4fec5ba — clang/test/CodeGen/RISCV rvp-intrinsics.c, llvm/test/CodeGen/AMDGPU load-constant-i1.ll amdgcn.bitcast.512bit.ll

Merge branch 'main' into users/adams381/cir-callconv-coerce-return-slot
DeltaFile
+42,441-42,426llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+7,527-7,623llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+4,115-8,359clang/test/CodeGen/RISCV/rvp-intrinsics.c
+4,036-4,380llvm/test/CodeGen/RISCV/rvv/clmul-sdnode.ll
+2,544-2,872llvm/test/CodeGen/RISCV/rvv/clmulh-sdnode.ll
+2,603-2,582llvm/test/CodeGen/AMDGPU/load-constant-i1.ll
+63,266-68,2421,467 files not shown
+148,098-139,5181,473 files

LLVM/project 8a7dedc — clang/test/CodeGen/RISCV rvp-intrinsics.c, llvm/test/CodeGen/AMDGPU load-constant-i1.ll amdgcn.bitcast.512bit.ll

Merge branch 'main' into users/adams381/cir-callconv-indirect-variadic
DeltaFile
+42,441-42,426llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+7,527-7,623llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+4,115-8,359clang/test/CodeGen/RISCV/rvp-intrinsics.c
+4,036-4,380llvm/test/CodeGen/RISCV/rvv/clmul-sdnode.ll
+2,544-2,872llvm/test/CodeGen/RISCV/rvv/clmulh-sdnode.ll
+2,603-2,582llvm/test/CodeGen/AMDGPU/load-constant-i1.ll
+63,266-68,2421,474 files not shown
+148,847-140,3841,480 files

OpenBSD/src hdBBbzy — usr.sbin/rpki-client output-rtrx.c

   correct include order
VersionDeltaFile
1.3+2-2usr.sbin/rpki-client/output-rtrx.c
+2-21 files

LLVM/project 7e0802b — clang/test/CodeGen/AArch64/sve len.c

[clang][CIR] Match CIR return value with OGCG using mem2reg (#227241)

This fixes the test failure introduced by
[226534](https://github.com/llvm/llvm-project/pull/226534)

Currently, the `sve/len.c` test fails because of this difference between
OGCG and CIR:

```
; classic codegen                  ; CIR (-fclangir)
%2 = mul nuw i64 %1, 16            %6 = mul nuw i64 %5, 16
ret i64 %2                         store i64 %6, ptr %3      ; retval slot
                                   %7 = load i64, ptr %3
                                   ret i64 %7
```
So, adding the return value checks as they are breaks the test.

This adds `mem2reg`, so that the CIR output matches OGCG and the same
return value check works for both.
DeltaFile
+8-6clang/test/CodeGen/AArch64/sve/len.c
+8-61 files

OpenBSD/src znFsEwI — usr.sbin/rpki-client output-rtrx.c

   use __packed; ok rcovelli
VersionDeltaFile
1.2+11-13usr.sbin/rpki-client/output-rtrx.c
+11-131 files

OpenBSD/src Mf8hmAy — usr.sbin/rpki-client Makefile extern.h

   Add option -r to output tables directly to the rtrd(8) controller socket.
   OK deraadt@
VersionDeltaFile
1.1+541-0usr.sbin/rpki-client/output-rtrx.c
1.314+11-4usr.sbin/rpki-client/main.c
1.48+12-1usr.sbin/rpki-client/output.c
1.144+7-3usr.sbin/rpki-client/rpki-client.8
1.300+7-1usr.sbin/rpki-client/extern.h
1.42+2-1usr.sbin/rpki-client/Makefile
+580-106 files

LLVM/project 8081af8 — clang/lib/CodeGen/TargetBuiltins RISCV.cpp, clang/lib/Headers riscv_packed_simd.h

[RISCV][P-ext] Add scalar multiply high intrinsics (#225561)

Add intrinsics, Clang builtins and SelectionDAG support for the scalar
multiply high operations.
Support
[multiply-high](https://github.com/riscv/riscv-p-spec/blob/master/P-ext-intrinsics.adoc#multiply-high).

On RV32 the non-rounding forms map to the M-extension mulh/mulhu/mulhsu
instructions and the rounding forms to the P-extension
mulhr/mulhru/mulhrsu instructions.
On RV64 the scalar spellings reuse the packed pmulh.w family on the low
words via the existing PMULH*_W patterns.
DeltaFile
+96-0clang/test/CodeGen/RISCV/rvp-intrinsics.c
+85-0llvm/test/CodeGen/RISCV/rvp-simd-32.ll
+50-0llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+43-0cross-project-tests/intrinsic-header-tests/riscv_packed_simd.c
+32-0clang/lib/CodeGen/TargetBuiltins/RISCV.cpp
+14-0clang/lib/Headers/riscv_packed_simd.h
+320-02 files not shown
+340-08 files

LLVM/project 930a118 — clang/test/CodeGen/AArch64 abi-classify-sve-tuples.c abi-classify-sve-types.c, llvm/include/llvm/ABI Types.h

[LLVMABI][AARCH64] Handle vector types (#225201)

This adds handling for vector types in the AArch64 implementation of the
LLVM ABI library. Legal vectors are passed and returned directly.
Illegal vectors are coerced to an integer or an integer vector, or
passed indirectly if they are too large. Sizeless SVE types are passed
in a register of their own, and fixed-length SVE vectors are coerced to
a scalable vector that occupies the same register. SVE tuples still
report NYI under AAPCS.

Assisted-by: Cursor / claude-opus-5
DeltaFile
+416-10llvm/unittests/ABI/AArch64TargetInfoTest.cpp
+148-14llvm/lib/ABI/Targets/AArch64.cpp
+98-12clang/test/CodeGen/AArch64/abi-classify-arg-types.c
+79-0clang/test/CodeGen/AArch64/abi-classify-sve-types.c
+39-5llvm/include/llvm/ABI/Types.h
+43-0clang/test/CodeGen/AArch64/abi-classify-sve-tuples.c
+823-417 files not shown
+858-4513 files

LLVM/project 2b0cd5f — llvm/lib/Target/AMDGPU SIInstructions.td

[AMDGPU][NFC] De-factorize scalar_to_vector bf16 patterns (#226847)

For Fake16, an SGPR and a VGPR can share the same scalar_to_vector
pattern because there is no 16-bit SGPR at all. The separate uniform
SGPR pattern is therefore only needed for Real True16, where the
divergent case uses a VGPR_16 REG_SEQUENCE.

This de-factorization is to prepare for follow-up PRs to separate SGPR
and VGPR patterns of scalar_to_vector for other types.
DeltaFile
+5-5llvm/lib/Target/AMDGPU/SIInstructions.td
+5-51 files

FreeBSD/ports 102c4ec — irc/quassel Makefile, irc/quassel-core Makefile

irc/quassel: Drop unused dependencies

These deps are not needed for Qt6 based Quassel

PR:     298956
DeltaFile
+3-5irc/quassel/Makefile
+1-0irc/quassel-core/Makefile
+4-52 files

FreeBSD/ports 4418a40 — www/nyxt/files patch-fix-sbcl269+

www/nyxt: Fix build with lang/sbcl 2.6.9
DeltaFile
+15-0www/nyxt/files/patch-fix-sbcl269+
+15-01 files

LLVM/project f4fe034 — llvm/lib/CodeGen/SelectionDAG SelectionDAG.cpp

fixup! Use a better KnownBits syntax
DeltaFile
+1-2llvm/lib/CodeGen/SelectionDAG/SelectionDAG.cpp
+1-21 files

FreeBSD/src afe3e2e — share/man/man4 unionfs.4

unionfs.4: Canonicalize SYNOPSIS + nit SPDX

MFC after:      3 days
DeltaFile
+4-12share/man/man4/unionfs.4
+4-121 files

FreeBSD/src bfe3273 — share/man/man4 ums.4

ums.4: Canonicalize SYNOPSIS + tag SPDX

MFC after:      3 days
DeltaFile
+10-15share/man/man4/ums.4
+10-151 files

LLVM/project 9425f15 — lldb/source/Expression REPL.cpp, lldb/test/Shell/REPL Breakpoint.test

[lldb] Fix deadlock when a REPL expression hits a breakpoint (#227400)

When a REPL expression stops at a breakpoint,
REPL::IOHandlerInputComplete calls RunIOHandlerAsync to drop into the
command interpreter while it still holds the error stream lock.
RunIOHandlerAsync then takes the IOHandler stack mutex. At the same
time, the event handler thread reports the stop through
IOHandlerStack::PrintAsync, which takes the IOHandler stack mutex first
and the output mutex second. The opposite lock order can deadlock,
leaving lldb hung after it prints "Execution stopped at breakpoint."

The lock scope that covers the call was introduced in 5007dd9d0945
(#183600). Keep printing under the lock, but defer RunIOHandlerAsync
until the lock is released.

This patch adds a test that exercises this path with the C REPL. Because
the deadlock is a race, the test only catches a regression
intermittently: without the fix it hung in about 6% of runs under
parallel load, and in every run when a delay was injected inside the

    [2 lines not shown]
DeltaFile
+28-0lldb/test/Shell/REPL/Breakpoint.test
+14-7lldb/source/Expression/REPL.cpp
+42-72 files