LLVM/project 3b8bb27 — libcxx/test/benchmarks/containers/associative associative_container_benchmarks.h

[libc++] Use a union for uninitialized storage in associative container benchmarks (#227068)

This removes the need for reinterpret_cast when accessing the containers
constructed in the scratch space, and fixes the PMR constructor
benchmark passing the wrong pointer to DoNotOptimize.
DeltaFile
+29-24libcxx/test/benchmarks/containers/associative/associative_container_benchmarks.h
+29-241 files

LLVM/project b73d8ab — offload/include PluginManager.h, offload/liboffload/src OffloadImpl.cpp

[offload][omp] Load plugins through liboffload
DeltaFile
+18-20offload/libompaccsupport/PluginManager.cpp
+6-0offload/liboffload/src/OffloadImpl.cpp
+2-1offload/include/PluginManager.h
+26-213 files

LLVM/project 9796f2a — offload/liboffload/API Platform.td, offload/liboffload/src OffloadImpl.cpp

[OFFLOAD] add olIteratePlatforms
DeltaFile
+44-0offload/unittests/OffloadAPI/platform/olIteratePlatforms.cpp
+23-0offload/liboffload/API/Platform.td
+12-2offload/liboffload/src/OffloadImpl.cpp
+79-23 files

LLVM/project 907a809 — libcxx/test/benchmarks/containers/associative associative_container_benchmarks.h

[libc++] Fix out-of-bounds read in the associative container query benchmarks (#227070)

The query benchmarks for associative containers would access the pool of
keys to use in the benchmark out-of-bounds.
DeltaFile
+1-1libcxx/test/benchmarks/containers/associative/associative_container_benchmarks.h
+1-11 files

LLVM/project f51e3f2 — flang/lib/Lower Bridge.cpp

[flang][NFC] Split the OpenACC construct lowering into two lanes

genFIR(OpenACCConstruct) decided twice, in three places, whether the
construct it lowers is structured, and reassigned the evaluation it works
from halfway through: before the descent that evaluation is the construct,
after it the loop the directive absorbs. Everything downstream had to know
which one it was holding.

Give each form its own function and leave genFIR to choose between them.
One lane allocates the exit selector, lowers the evaluations the construct
holds, and emits the jump table; the other reads the collapse clauses,
descends to the absorbed depth, and lowers what is inside it. The prologue
and epilogue are short enough to state in both rather than share.
DeltaFile
+149-101flang/lib/Lower/Bridge.cpp
+149-1011 files

LLVM/project d2d3041 — flang/lib/Lower PFTBuilder.cpp, flang/test/Lower/OpenACC acc-unstructured-combined-construct.f90 acc-directive-loop-bounds.f90

[flang] Let a directive keep the loop it owns when its body branches

A loop whose branching is confined to its body keeps its structured form,
but the construct holding it stayed unstructured. A directive does not
merely contain such a loop, it owns it, and its lowering reads the
construct's own classification to decide whether the loop op carries its
bounds. The directive was left with a bounds-free loop that nothing could
partition, and the loop it owns became a second one nested inside.

Reclassify a directive construct once the loops it holds no longer need it
to stay unstructured, and fold the body of the loop it takes over into a
region, which the DO lowering can no longer do for it.

A construct whose branching leaves it is untouched, as is one holding a
branch of its own.
DeltaFile
+8-124flang/test/Lower/OpenACC/Todo/acc-unstructured-loop-construct.f90
+120-0flang/test/Lower/OpenACC/acc-unstructured-loop-construct.f90
+66-0flang/test/Lower/OpenACC/acc-directive-loop-bounds.f90
+40-0flang/test/Lower/OpenACC/acc-unstructured-combined-construct.f90
+33-4flang/lib/Lower/PFTBuilder.cpp
+34-0flang/test/Lower/OpenMP/wsloop-directive-loop-bounds.f90
+301-1284 files not shown
+333-16810 files

LLVM/project 99ac64c — offload/liboffload/API Program.td, offload/liboffload/src OffloadImpl.cpp

[OFFLOAD] Add olIterateCompatibleDevices API
DeltaFile
+73-0offload/unittests/OffloadAPI/program/olIterateCompatibleDevices.cpp
+25-2offload/liboffload/src/OffloadImpl.cpp
+15-0offload/liboffload/API/Program.td
+1-2offload/unittests/OffloadAPI/platform/olIteratePlatforms.cpp
+1-0offload/unittests/OffloadAPI/CMakeLists.txt
+115-45 files

LLVM/project b3b5e33 — flang/lib/Lower Bridge.cpp, flang/test/Lower do-loop-branch-to-loop-header.f90 do-loop-infinite-body-cycle.f90

[flang] Lower loops whose branching is confined to their body structurally (#225758)

Such a loop was classified separately by a previous change but still
lowered as a raw CFG, so its structured form was lost.

Lower it structurally instead, with its body folded into a region that
can hold the branching. The loop keeps its bounds on the op, so it
remains available to whatever transforms or parallelizes it. Only the
body is folded: the loop control statements are emitted as they are for
any structured loop, since a branch from outside may target either of
them.

Loops an OpenACC or OpenMP directive owns are lowered the same way, so
they keep their form too.
DeltaFile
+20-134flang/test/Lower/do_loop_unstructured.f90
+119-11flang/lib/Lower/Bridge.cpp
+50-28flang/test/Lower/do-loop-infinite-body-cycle.f90
+73-0flang/test/Lower/OpenACC/acc-unstructured-internals.f90
+34-28flang/test/Lower/OpenMP/wsloop-unstructured-cycle.f90
+55-0flang/test/Lower/do-loop-branch-to-loop-header.f90
+351-20111 files not shown
+503-26917 files

LLVM/project 976a04b — llvm/lib/CodeGen TargetLoweringObjectFileImpl.cpp, llvm/lib/MC GOFFObjectWriter.cpp MCSymbolGOFF.cpp

[SystemZ][z/OS] Keep weak references weak (#226840)

Fixes #226835.

An `extern_weak` reference must stay unresolved without an error when
the symbol does not exist. On z/OS two kinds of references were always
strong:

- Taking the address of an external function goes through the indirect
symbol `<name>@indirect` (ADA slot `MO_ADA_INDIRECT_FUNC_DESC`). It
never got the weak attribute of the function symbol. The weak attribute
is already set on the function symbol when the ADA is emitted, because
`AsmPrinter::doFinalization` emits the weak references before
`emitEndOfAsmFile`. So the indirect symbol now becomes a weak reference
too.
- An external data reference is a part reference (PR). `GOFF::PRAttr`
had no binding strength, and the PR constructor in `GOFFObjectWriter`
did not set it. `PRAttr` gets a `BindingStrength` field, and
`defineExtern` passes the strength of the symbol.

    [17 lines not shown]
DeltaFile
+37-0llvm/test/CodeGen/SystemZ/zos-extern-weak.ll
+10-10llvm/lib/MC/MCObjectFileInfo.cpp
+9-9llvm/lib/CodeGen/TargetLoweringObjectFileImpl.cpp
+5-5llvm/lib/MC/MCSymbolGOFF.cpp
+3-2llvm/lib/MC/GOFFObjectWriter.cpp
+4-0llvm/lib/Target/SystemZ/SystemZXPLINKAsmPrinter.cpp
+68-261 files not shown
+69-267 files

LLVM/project e65dba4 — clang/test/Profile cxx-throws.cpp branch-logical-mixed.cpp, llvm/test/CodeGen/AMDGPU amdgpu-sw-lower-lds-multi-static-dynamic-indirect-access.ll amdgpu-sw-lower-lds-multi-static-dynamic-indirect-access-asan.ll

[IRBuilder] Produce canonical constexpr GEPs (#226425)

This switches the IRBuilder to always produce canonical constexpr GEPs
in ptradd form, including in the case where ConstantFolder rather than
TargetFolder is used.

This is done in a slightly hacky way, by passing the DataLayout from the
insertion point into FoldGEP. In the future, when we start
canonicalizing the non-constant GEPs as well, we'll do the
canonicalization directly in IRBuilder and replace FoldGEP with
FoldPtrAdd. (Though it would be even better to make IRBuilder always
require a DataLayout and eliminate the ConstantFolder/TargetFolder
distinction...)
DeltaFile
+107-107clang/test/Profile/c-general.c
+23-23clang/test/Profile/branch-logical-mixed.cpp
+16-16llvm/test/Transforms/LowerIFunc/lower-ifunc.ll
+16-16llvm/test/CodeGen/AMDGPU/amdgpu-sw-lower-lds-multi-static-dynamic-indirect-access.ll
+16-16llvm/test/CodeGen/AMDGPU/amdgpu-sw-lower-lds-multi-static-dynamic-indirect-access-asan.ll
+14-14clang/test/Profile/cxx-throws.cpp
+192-19281 files not shown
+514-51187 files

LLVM/project 5274c3a —

[flang-rt] Use thin I/O in the native GPU builds (#226307)
DeltaFile
+0-00 files

LLVM/project 6cd0f02 — flang-rt/cmake/modules AddFlangRTOffload.cmake, flang-rt/include/flang-rt/runtime io-stmt.h work-queue.h

[flang-rt] Use thin I/O in the native GPU builds (#226307)

The amdgcn libflang_rt.runtime.a has undefined references to the DescriptorIoTicket/DerivedIoTicket methods and to flang_rt_verbose_abort. The tickets live in descriptor-io.cpp, which isn't in gpu_sources, but the work queue still refers to them. Any device code that ends up in the work queue (e.g. ALLOCATE of a derived type with default initialization in a target region) then fails to link.

The CUDA PTX build already avoids this with a thin I/O mode. This renames `RT_CUDA_THIN_IO` to `RT_THIN_IO`, defines it in api-attrs.h for the native GPU builds (`RT_GPU_TARGET` without `RT_DEVICE_COMPILATION`), and keeps the CMake define for the CUDA PTX library under the new name. flang_rt_verbose_abort is defined in stl-overrides.cpp, so that's added to gpu_sources.

Tested on gfx90a. The amdgcn library no longer has those undefined symbols, gains only the flang_rt_verbose_abort definition, and doesn't lose any others. A small reproducer (derived type with default init allocated inside a target region) links and runs, and scalar PRINT from device code gives the same output as before. The host library's symbols are identical to before. I haven't built nvptx or the CUDA offload configuration.

Downstream report: ROCm/llvm-project#3517
DeltaFile
+12-0flang/include/flang/Common/api-attrs.h
+5-5flang-rt/include/flang-rt/runtime/work-queue.h
+2-2flang-rt/include/flang-rt/runtime/io-stmt.h
+1-1flang-rt/lib/runtime/io-api-common.h
+1-1flang-rt/cmake/modules/AddFlangRTOffload.cmake
+1-0flang-rt/lib/runtime/CMakeLists.txt
+22-96 files

LLVM/project 1b5f644 —

[libc++] Encode the standard version in the ABI tag (#218527)

This prevents ODR mismatch issues from biting us across standard
versions. The order of elements in the ABI tag is now: libc++ version,
hardening mode, assertion semantic, exceptions, standard version.

Fixes #218524
DeltaFile
+0-00 files

LLVM/project fc1884a — llvm/test/CodeGen/AMDGPU frem.ll

Update lit merge conflict
DeltaFile
+49-31llvm/test/CodeGen/AMDGPU/frem.ll
+49-311 files

pkgng/pkgng c9b03b4 — external/libecc/include/libecc/words words.h, libpkg pkg_abi.c pkg_elf.c

libpkg: recognize loongarch64 as loongarch:64

Add PKG_ARCH_LOONGARCH64 and map EM_LOONGARCH/ELFCLASS64 to it; no
32-bit ABI exists.  Test it in the frontend suite.

external/libecc: add __loongarch__ to words.h.  Upstream, from
https://github.com/libecc/libecc/pull/15/.
DeltaFile
+8-2tests/frontend/abi.sh
+6-0tests/frontend/test_environment.sh.in
+6-0libpkg/pkg_elf.c
+2-1external/libecc/include/libecc/words/words.h
+3-0libpkg/pkg_abi.c
+1-1tests/frontend/create-parsebin.sh
+26-46 files not shown
+31-412 files

LLVM/project 09cdef0 — libc/src/__support/math CMakeLists.txt cosf_double_eval.h, libc/test/src/math/exhaustive sinf_float_test.cpp

[libc][math] Reorganize sinf and cosf function selection (#226529)

Reorganises sinf and cosf similarly to #224735 such that:
src/__support/math/func.h selects implementation, and
src/__support/math/func_<type>_eval.h implements func with <type> as the
intermediate computational type.
DeltaFile
+18-174libc/src/__support/math/sinf.h
+182-0libc/src/__support/math/sinf_double_eval.h
+17-151libc/src/__support/math/cosf.h
+162-0libc/src/__support/math/cosf_double_eval.h
+56-4libc/src/__support/math/CMakeLists.txt
+0-47libc/test/src/math/exhaustive/sinf_float_test.cpp
+435-37611 files not shown
+621-55117 files

LLVM/project e15e5e6 — llvm/test/CodeGen/RISCV clmul.ll, llvm/test/CodeGen/RISCV/rvv clmulh-sdnode.ll clmul-sdnode.ll

[ISel] Improve `clmul` fallback implementation (#204802)

Generalize the approach from
https://github.com/llvm/llvm-project/pull/203727 to narrower and wider
integers.

We still need the fallback for when multiplication isn't available, and
it turns out that for some widths the fallback emits fewer instructions,
the naive fallback is still used for `i1`, `i3`, `i4` and `i9`. I've
also now enabled wider integers (`i128` and `i256` have uses in
cryptography).

Based on my local experiments, the Karatsuba approach (e.g. as in
https://github.com/rust-lang/rust/pull/152132#discussion_r2778609222) is
not actually better than zero extending the input and using
multiplication with holes on the wider type.

CC https://github.com/llvm/llvm-project/issues/203694
CC: @eisenwave
DeltaFile
+4,036-4,380llvm/test/CodeGen/RISCV/rvv/clmul-sdnode.ll
+2,544-2,872llvm/test/CodeGen/RISCV/rvv/clmulh-sdnode.ll
+1,020-1,412llvm/test/CodeGen/X86/clmul-vector-512.ll
+850-1,332llvm/test/CodeGen/X86/clmul-vector-256.ll
+772-1,269llvm/test/CodeGen/X86/clmul-vector.ll
+399-1,451llvm/test/CodeGen/RISCV/clmul.ll
+9,621-12,71614 files not shown
+11,877-17,47120 files

LLVM/project a537e6e — clang/test/CodeGen/RISCV rvp-intrinsics.c, llvm/test/CodeGen/AMDGPU load-constant-i8.ll load-local-i16.ll

Merge branch 'users/zGoldthorpe/wide-copies/precommit' into users/zGoldthorpe/wide-copies/materialise
DeltaFile
+42,441-42,426llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+7,527-7,623llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+4,022-8,362clang/test/CodeGen/RISCV/rvp-intrinsics.c
+2,603-2,582llvm/test/CodeGen/AMDGPU/load-constant-i1.ll
+1,833-1,839llvm/test/CodeGen/AMDGPU/load-local-i16.ll
+1,790-1,779llvm/test/CodeGen/AMDGPU/load-constant-i8.ll
+60,216-64,6111,268 files not shown
+125,107-116,5321,274 files

LLVM/project 15ec5f2 — llvm/test/CodeGen/AMDGPU bitcast-vector-extract.ll

[AMDGPU] Replace GCN-NOT with autogen checks in test (#227148)

The assembly has many v_mov_b32s, so minor scheduling changes can
trigger failure on the GCN-NOT. It seems the test is designed to show
CSE behavior, which still holds even if minor scheduling changes break
the NOT checks.
DeltaFile
+235-30llvm/test/CodeGen/AMDGPU/bitcast-vector-extract.ll
+235-301 files

FreeBSD/ports e54335c — graphics/osg Makefile

graphics/osg: prepare for CMake4 (+)

PR:             298945
DeltaFile
+1-0graphics/osg/Makefile
+1-01 files

FreeBSD/ports c59a4b9 — net-im/tde2e Makefile distinfo

net-im/tde2e: update to 1.8.76 + unicode fixes snapshot

Approved by:    osa
DeltaFile
+3-3net-im/tde2e/distinfo
+2-2net-im/tde2e/Makefile
+5-52 files

LLVM/project b727f3d — clang/test/CodeGen/RISCV rvp-intrinsics.c, llvm/test/CodeGen/AMDGPU load-constant-i8.ll load-local-i16.ll

Merge branch 'main' into users/zGoldthorpe/wide-copies/precommit
DeltaFile
+42,441-42,426llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+7,527-7,623llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+4,022-8,362clang/test/CodeGen/RISCV/rvp-intrinsics.c
+2,603-2,582llvm/test/CodeGen/AMDGPU/load-constant-i1.ll
+1,833-1,839llvm/test/CodeGen/AMDGPU/load-local-i16.ll
+1,790-1,779llvm/test/CodeGen/AMDGPU/load-constant-i8.ll
+60,216-64,6111,268 files not shown
+125,100-116,5061,274 files

LLVM/project 34ba34d — llvm/lib/Transforms/InstCombine InstCombineCalls.cpp, llvm/test/Transforms/InstCombine bitreverse.ll

[InstCombine] Fix ProfCheck for optimizing bitreverse (#227189)
DeltaFile
+8-3llvm/test/Transforms/InstCombine/bitreverse.ll
+8-2llvm/lib/Transforms/InstCombine/InstCombineCalls.cpp
+0-1llvm/utils/profcheck-xfail.txt
+16-63 files

OpenBSD/src AMEral4 — lib/libcrypto/x509 x509_crld.c

   set_dist_point_name(): tiny tweak to restore previous behavior

   Allocate fnm before allocating *pdp. This way a second call to to
   set_dist_point_name() has a tiny little chance of succeeding.

   ok beck ("I strongly suspect this will never matter anywhere.")
VersionDeltaFile
1.13+4-4lib/libcrypto/x509/x509_crld.c
+4-41 files

LLVM/project 4c79537 — llvm/include/llvm/Analysis AliasSetTracker.h, llvm/lib/Analysis AliasSetTracker.cpp

[LICM] Drop *only* per-iteration AA tags
DeltaFile
+48-27llvm/lib/Transforms/Scalar/LICM.cpp
+6-4llvm/test/Transforms/LICM/scalar-promote-aa-tags.ll
+4-3llvm/lib/Analysis/AliasSetTracker.cpp
+1-1llvm/include/llvm/Analysis/AliasSetTracker.h
+59-354 files

LLVM/project d9e7b3e — llvm/lib/Transforms/Scalar LICM.cpp

fix formatting
DeltaFile
+3-3llvm/lib/Transforms/Scalar/LICM.cpp
+3-31 files

LLVM/project 17352e5 — llvm/lib/Transforms/Scalar LICM.cpp

Collect per-iteration alias scopes in advance
DeltaFile
+37-50llvm/lib/Transforms/Scalar/LICM.cpp
+37-501 files

LLVM/project b614800 — llvm/lib/Transforms/Scalar LICM.cpp, llvm/test/Transforms/LICM scalar-promote-aa-tags.ll

[LICM] Drop per-iteration AA tags
DeltaFile
+46-1llvm/lib/Transforms/Scalar/LICM.cpp
+6-12llvm/test/Transforms/LICM/scalar-promote-aa-tags.ll
+52-132 files

LLVM/project 0df5a52 — llvm/test/Transforms/LICM scalar-promote-aa-tags.ll

Add precommit test for noalias scope declared outside loop
DeltaFile
+58-0llvm/test/Transforms/LICM/scalar-promote-aa-tags.ll
+58-01 files

LLVM/project 9c2a881 — llvm/test/Transforms/LICM scalar-promote-aa-tags.ll

[NFC][LICM] Precommit mishandled per-iteration scoped alias metadata
DeltaFile
+117-0llvm/test/Transforms/LICM/scalar-promote-aa-tags.ll
+117-01 files