LLVM/project 8262a4b — offload/libompaccsupport PluginManager.cpp

[offload][OpenMP] Read the kernel environment from the device image (#229131)

This should be an optimization for #222606. The kernel environment
doesn't actually need to be loaded back from the device since it
shouldn't be different compared to the time where it was transferred to
the device as part of the image.

Claude assisted with this patch.
DeltaFile
+14-15offload/libompaccsupport/PluginManager.cpp
+14-151 files

LLVM/project f8e207f — llvm/lib/Target/AArch64 AArch64FrameLowering.cpp, llvm/test/CodeGen/AArch64 sme-callee-save-restore-pairs.ll sme-vg-to-stack.ll

[AArch64] Fix incorrect frame indexes for ZPR callee-saved pair spills (#228368)

The ZPR pairs callee-saved spills were being assigned frame indexes in
the reverse order. Resulting in CFI reporting swapped memory locations
for the registers in the pairs

Reverse the conditions to find pairs and tweak the reordering algorithm
to maximize this.
Resulting in pairs with correct frame indexes.

Some refactors have also been included in this PR
DeltaFile
+50-50llvm/test/CodeGen/AArch64/sve-callee-save-restore-pairs.ll
+57-28llvm/lib/Target/AArch64/AArch64FrameLowering.cpp
+32-32llvm/test/CodeGen/AArch64/sme-vg-to-stack.ll
+4-4llvm/test/CodeGen/AArch64/sme-callee-save-restore-pairs.ll
+143-1144 files

LLVM/project a5a5929 — mlir/include/mlir/Dialect/XeGPU/Transforms XeGPULayoutImpl.h, mlir/lib/Dialect/XeGPU/Transforms XeGPUPropagateLayout.cpp

[MLIR][XeGPU] Optimize instruction granularity for elemwise (#220986)
DeltaFile
+125-1mlir/lib/Dialect/XeGPU/Transforms/XeGPUPropagateLayout.cpp
+120-0mlir/test/Dialect/XeGPU/sink-elementwise-conversions.mlir
+31-0mlir/test/lib/Dialect/XeGPU/TestXeGPUTransforms.cpp
+0-24mlir/test/Dialect/XeGPU/resolve-layout-conflicts.mlir
+4-0mlir/include/mlir/Dialect/XeGPU/Transforms/XeGPULayoutImpl.h
+280-255 files

LLVM/project 7c09825 — llvm/lib/Target/RISCV RISCVISelLowering.cpp

fixup! Address review comments
DeltaFile
+2-3llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+2-31 files

FreeBSD/ports 2cf063b — x11/xlockmore Makefile distinfo

x11/xlockmore: update to 5.89
DeltaFile
+3-3x11/xlockmore/distinfo
+2-2x11/xlockmore/Makefile
+5-52 files

LLVM/project 809116f — llvm/lib/Transforms/Scalar LICM.cpp

fix formatting
DeltaFile
+3-3llvm/lib/Transforms/Scalar/LICM.cpp
+3-31 files

LLVM/project b5c07a7 — llvm/include/llvm/Analysis AliasSetTracker.h, llvm/lib/Analysis AliasSetTracker.cpp

[LICM] Drop *only* per-iteration AA tags
DeltaFile
+48-27llvm/lib/Transforms/Scalar/LICM.cpp
+6-4llvm/test/Transforms/LICM/scalar-promote-aa-tags.ll
+4-3llvm/lib/Analysis/AliasSetTracker.cpp
+1-1llvm/include/llvm/Analysis/AliasSetTracker.h
+59-354 files

LLVM/project 1b46cad — llvm/lib/Transforms/Scalar LICM.cpp, llvm/test/Transforms/LICM scalar-promote-aa-tags.ll

[LICM] Drop per-iteration AA tags (#223530)

AA tags are scoped inside the loop via
`llvm.experimental.noalias.scope.decl` were being preserved in #222686
to determine promotion, which is incorrect as the alias metadata only
holds per individual loop iteration, and thus cannot be used to infer
alias information of memory locations used between iterations.

This patch checks for loop-local alias declarations and excludes them
from being preserved.
DeltaFile
+50-18llvm/lib/Transforms/Scalar/LICM.cpp
+6-12llvm/test/Transforms/LICM/scalar-promote-aa-tags.ll
+56-302 files

LLVM/project df91acf — llvm/test/Transforms/LICM scalar-promote-aa-tags.ll

[NFC][LICM] Precommit mishandled per-iteration scoped alias metadata (#223529)

Pre-commit tests based on [this comment on
#222686](https://github.com/llvm/llvm-project/pull/222686#issuecomment-5647760268).
DeltaFile
+227-0llvm/test/Transforms/LICM/scalar-promote-aa-tags.ll
+227-01 files

LLVM/project 0fd019b — llvm/include/llvm/Transforms/Scalar LoopUnrollPass.h, llvm/lib/Passes PassBuilderPipelines.cpp

[𝘀𝗽𝗿] initial version

Created using spr 1.3.7
DeltaFile
+14-18llvm/test/Transforms/LoopUnroll/prepare-for-lto.ll
+6-3llvm/lib/Passes/PassBuilderPipelines.cpp
+4-1llvm/lib/Transforms/Scalar/LoopUnrollPass.cpp
+2-2llvm/include/llvm/Transforms/Scalar/LoopUnrollPass.h
+26-244 files

LLVM/project 1c16beb — llvm/lib/Transforms/Scalar LoopUnrollPass.cpp, llvm/test/Transforms/LoopUnroll prepare-for-lto.ll

[𝘀𝗽𝗿] changes to main this commit is based on

Created using spr 1.3.7

[skip ci]
DeltaFile
+14-18llvm/test/Transforms/LoopUnroll/prepare-for-lto.ll
+4-1llvm/lib/Transforms/Scalar/LoopUnrollPass.cpp
+18-192 files

LLVM/project ee8ed63 — llvm/lib/Transforms/Scalar LoopUnrollPass.cpp, llvm/test/Transforms/LoopUnroll prepare-for-lto.ll

[𝘀𝗽𝗿] initial version

Created using spr 1.3.7
DeltaFile
+14-18llvm/test/Transforms/LoopUnroll/prepare-for-lto.ll
+4-1llvm/lib/Transforms/Scalar/LoopUnrollPass.cpp
+18-192 files

LLVM/project 961aae7 — libsycl/test/basic index_space_classes.cpp

[libsycl] Disable werror for deprecated declaration in test (#229145)

Test verifies deprecated SYCL API intentionally.

Signed-off-by: Tikhomirova, Kseniya <kseniya.tikhomirova at intel.com>
DeltaFile
+1-1libsycl/test/basic/index_space_classes.cpp
+1-11 files

LLVM/project 84688fb — llvm/lib/CodeGen/SelectionDAG DAGCombiner.cpp

fixup! Use m_FShL/R instead of m_Node
DeltaFile
+8-8llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp
+8-81 files

LLVM/project 1dafb43 — clang/lib/Driver/ToolChains HIPAMD.cpp, clang/test/Driver spirv-amd-toolchain.c

[NFC][HIP][SPIRV] Rework `-spirv-preserve-auxdata` handling. (#227875)

The way we were using to (unconditionally) enable this has been silently
inert since switching to the new driver / all things going through
`clang-linker-wrapper`. It was also a bit suspect / sneaky in general.
Pivot it to be controlled based on encountering the AMD vendor type.
DeltaFile
+4-1llvm/lib/Target/SPIRV/SPIRVAuxDataHandler.cpp
+2-2clang/lib/Driver/ToolChains/HIPAMD.cpp
+1-1clang/test/Driver/spirv-amd-toolchain.c
+7-43 files

LLVM/project a7f974f — llvm/docs AMDGPUUsage.rst, llvm/lib/Target/AMDGPU/Disassembler AMDGPUDisassembler.cpp

[AMDGPU] Update ENABLE_WAVEFRONT_SIZE32 kernel descriptor definition for gfx125*

Change-Id: I49b29627f9481c7266863c592d4c25caa9b35c50
DeltaFile
+6-8llvm/test/MC/AMDGPU/hsa-gfx1251-v4.s
+6-8llvm/test/MC/AMDGPU/hsa-gfx1250-v4.s
+8-6llvm/test/MC/AMDGPU/hsa-diag-v4.s
+10-2llvm/lib/Target/AMDGPU/Disassembler/AMDGPUDisassembler.cpp
+5-5llvm/docs/AMDGPUUsage.rst
+5-1llvm/test/CodeGen/AMDGPU/hsa-generic-target-features.ll
+40-309 files not shown
+65-3615 files

LLVM/project ffc09b7 —

[SLP]Keep dbg_value of the lanes replaced by extractelement

Redirect the debug values of an erased scalar to the extract emitted for
its external user. A record placed before the extract is cloned right
after it, unless that passes a record of the same variable. Lanes
without external users still lose their debug values.

Part of #45507.

Assisted-by: Cursor

Reviewers: RKSimon, bababuck

Pull Request: https://github.com/llvm/llvm-project/pull/228937
DeltaFile
+0-00 files

FreeBSD/src e89c3ac — sys/dev/acpica/Osd OsdSchedule.c

acpi: Tasks: Document why 'acpi_task_count' is accessed unsynchronized

MFC after:      3 days
Sponsored by:   The FreeBSD Foundation
DeltaFile
+1-0sys/dev/acpica/Osd/OsdSchedule.c
+1-01 files

FreeBSD/src f0825f7 — sys/dev/acpica/Osd OsdSchedule.c

acpi: Tasks: Make OsdSchedule.c whitespace clean

MFC after:      3 days
Sponsored by:   The FreeBSD Foundation
DeltaFile
+1-1sys/dev/acpica/Osd/OsdSchedule.c
+1-11 files

LLVM/project 90137f4 — llvm/lib/Target/AMDGPU AMDGPUAttributor.cpp

Merge branch 'users/zGoldthorpe/wg2wf-nodma/simple-demote' into users/zGoldthorpe/wg2wf-nodma/tablegen
DeltaFile
+3-3llvm/lib/Target/AMDGPU/AMDGPUAttributor.cpp
+3-31 files

FreeBSD/src 2165acc — sys/dev/acpica/Osd OsdSchedule.c

acpi: Tasks: Remove unnecessary includes

MFC after:      3 days
Sponsored by:   The FreeBSD Foundation
DeltaFile
+0-2sys/dev/acpica/Osd/OsdSchedule.c
+0-21 files

LLVM/project e9a93ba — llvm/test/CodeGen/AArch64 vector-compress.ll, llvm/test/CodeGen/NVPTX insertelt-dynamic.ll

[SelectionDAG] Freeze dynamic vector indices lowered through the stack (#225031)

fixes #224200
DeltaFile
+581-521llvm/test/CodeGen/X86/vector-compress.ll
+93-88llvm/test/CodeGen/AArch64/vector-compress.ll
+181-0llvm/test/CodeGen/X86/freeze-vector.ll
+84-45llvm/test/CodeGen/X86/var-permute-128.ll
+53-43llvm/test/CodeGen/NVPTX/insertelt-dynamic.ll
+40-40llvm/test/CodeGen/X86/vector-shuffle-variable-128.ll
+1,032-73710 files not shown
+1,111-77816 files

LLVM/project 0e19665 —

[SLP]Keep dbg_value of the lanes replaced by extractelement

Redirect the debug values of an erased scalar to the extract emitted for
its external user. A record placed before the extract is cloned right
after it, unless that passes a record of the same variable. Lanes
without external users still lose their debug values.

Part of #45507.

Assisted-by: Cursor

Reviewers: RKSimon, bababuck

Pull Request: https://github.com/llvm/llvm-project/pull/228937
DeltaFile
+0-00 files

LLVM/project 21219ac —

Merge branch 'users/zGoldthorpe/wg2wf-nodma/attributor' into users/zGoldthorpe/wg2wf-nodma/simple-demote
DeltaFile
+0-00 files

LLVM/project cbbd494 — llvm/utils/gn/secondary/clang/tools/offload-arch/lib BUILD.gn

[gn build] Port b805a7c2e5ec (#229136)
DeltaFile
+1-0llvm/utils/gn/secondary/clang/tools/offload-arch/lib/BUILD.gn
+1-01 files

LLVM/project 799a346 — llvm/utils/gn/secondary/llvm/lib/Target/SystemZ/AsmParser BUILD.gn, llvm/utils/gn/secondary/llvm/utils/TableGen BUILD.gn

[gn] port 0e7e59dcf4ac (systemz match table) (#229135)
DeltaFile
+8-0llvm/utils/gn/secondary/llvm/lib/Target/SystemZ/AsmParser/BUILD.gn
+1-0llvm/utils/gn/secondary/llvm/utils/TableGen/BUILD.gn
+9-02 files

OpenZFS/src 8de0800 — tests/zfs-tests/tests/functional/removal removal_with_ganging.ksh

ZTS: restore metaslab_force_ganging after removal_with_ganging

The cleanup of removal_with_ganging sets metaslab_force_ganging to
131073 instead of the value it replaced. Every later test on the same
runner then gangs about 3% of its blocks larger than 128K, since
metaslab_force_ganging_pct defaults to 3. A test that needs a plain
1M block fails intermittently; replacement/indirect_reconstruct did on
FreeBSD CI.

Save the tunable before changing it and restore it on exit, as the
other tests that force ganging do. Register cleanup after the save,
before pool setup and the tunable write, so setup failures restore it
too.

Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Kamil Monicz <kamil at monicz.dev>
Closes #19207
DeltaFile
+3-2tests/zfs-tests/tests/functional/removal/removal_with_ganging.ksh
+3-21 files

LLVM/project 74d358f — clang/test CMakeLists.txt

[CIR] Minimize dependencies of check-clang-cir (#227957)

This adds special handling for the CIR test directories so that they
depend only on the components needed to run those tests. The goal of
this is to minimize what needs to be built to verify that MLIR changes
haven't broken CIR. There is no need to run the full set of Clang tests
to verify an MLIR change, since only CIR depends on MLIR.

Assisted-by: Cursor / various models
DeltaFile
+38-0clang/test/CMakeLists.txt
+38-01 files

OpenZFS/src f3ad49d — tests/zfs-tests/include libtest.shlib

ZTS: preserve failures when saving and restoring tunables

save_tunable hides a failed get_tunable behind echo, and
restore_tunable hides a failed set_tunable64 behind rm. A test can
therefore report successful cleanup while leaving its changed tunable
in place and discarding the saved value.

Check the tunable and saved-value reads, and return a failed setter's
status before deleting the saved value. A failed restore then retains
the value for a retry, and log_must reports the cleanup failure.

Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Kamil Monicz <kamil at monicz.dev>
Closes #19207
DeltaFile
+8-3tests/zfs-tests/include/libtest.shlib
+8-31 files

LLVM/project 42a6e0d — llvm/lib/Transforms/InstCombine InstCombineCompares.cpp, llvm/test/Transforms/InstCombine get_active_lane_mask.ll

[InstCombine] Fold canonicalized form of vecreduce_or(get_active_lane_mask) (#220085)

get_active_lane_mask is called with lower and upper bounds. We can just
compare the bounds themselves.

This change is intended to improve vecreduce_or codegen for fixed-length
vectors. vscales deal with vecreduce_or directly.

https://godbolt.org/z/84sjGGdWe

alive2 doesn't seem to support get_active_lane_mask at the moment.
DeltaFile
+76-0llvm/test/Transforms/InstCombine/get_active_lane_mask.ll
+16-0llvm/lib/Transforms/InstCombine/InstCombineCompares.cpp
+92-02 files