[libc++] Use a union for uninitialized storage in associative container benchmarks (#227068)
This removes the need for reinterpret_cast when accessing the containers
constructed in the scratch space, and fixes the PMR constructor
benchmark passing the wrong pointer to DoNotOptimize.
[libc++] Fix out-of-bounds read in the associative container query benchmarks (#227070)
The query benchmarks for associative containers would access the pool of
keys to use in the benchmark out-of-bounds.
[flang][NFC] Split the OpenACC construct lowering into two lanes
genFIR(OpenACCConstruct) decided twice, in three places, whether the
construct it lowers is structured, and reassigned the evaluation it works
from halfway through: before the descent that evaluation is the construct,
after it the loop the directive absorbs. Everything downstream had to know
which one it was holding.
Give each form its own function and leave genFIR to choose between them.
One lane allocates the exit selector, lowers the evaluations the construct
holds, and emits the jump table; the other reads the collapse clauses,
descends to the absorbed depth, and lowers what is inside it. The prologue
and epilogue are short enough to state in both rather than share.
[flang] Let a directive keep the loop it owns when its body branches
A loop whose branching is confined to its body keeps its structured form,
but the construct holding it stayed unstructured. A directive does not
merely contain such a loop, it owns it, and its lowering reads the
construct's own classification to decide whether the loop op carries its
bounds. The directive was left with a bounds-free loop that nothing could
partition, and the loop it owns became a second one nested inside.
Reclassify a directive construct once the loops it holds no longer need it
to stay unstructured, and fold the body of the loop it takes over into a
region, which the DO lowering can no longer do for it.
A construct whose branching leaves it is untouched, as is one holding a
branch of its own.
[flang] Lower loops whose branching is confined to their body structurally (#225758)
Such a loop was classified separately by a previous change but still
lowered as a raw CFG, so its structured form was lost.
Lower it structurally instead, with its body folded into a region that
can hold the branching. The loop keeps its bounds on the op, so it
remains available to whatever transforms or parallelizes it. Only the
body is folded: the loop control statements are emitted as they are for
any structured loop, since a branch from outside may target either of
them.
Loops an OpenACC or OpenMP directive owns are lowered the same way, so
they keep their form too.
[SystemZ][z/OS] Keep weak references weak (#226840)
Fixes #226835.
An `extern_weak` reference must stay unresolved without an error when
the symbol does not exist. On z/OS two kinds of references were always
strong:
- Taking the address of an external function goes through the indirect
symbol `<name>@indirect` (ADA slot `MO_ADA_INDIRECT_FUNC_DESC`). It
never got the weak attribute of the function symbol. The weak attribute
is already set on the function symbol when the ADA is emitted, because
`AsmPrinter::doFinalization` emits the weak references before
`emitEndOfAsmFile`. So the indirect symbol now becomes a weak reference
too.
- An external data reference is a part reference (PR). `GOFF::PRAttr`
had no binding strength, and the PR constructor in `GOFFObjectWriter`
did not set it. `PRAttr` gets a `BindingStrength` field, and
`defineExtern` passes the strength of the symbol.
[17 lines not shown]
[IRBuilder] Produce canonical constexpr GEPs (#226425)
This switches the IRBuilder to always produce canonical constexpr GEPs
in ptradd form, including in the case where ConstantFolder rather than
TargetFolder is used.
This is done in a slightly hacky way, by passing the DataLayout from the
insertion point into FoldGEP. In the future, when we start
canonicalizing the non-constant GEPs as well, we'll do the
canonicalization directly in IRBuilder and replace FoldGEP with
FoldPtrAdd. (Though it would be even better to make IRBuilder always
require a DataLayout and eliminate the ConstantFolder/TargetFolder
distinction...)
[flang-rt] Use thin I/O in the native GPU builds (#226307)
The amdgcn libflang_rt.runtime.a has undefined references to the DescriptorIoTicket/DerivedIoTicket methods and to flang_rt_verbose_abort. The tickets live in descriptor-io.cpp, which isn't in gpu_sources, but the work queue still refers to them. Any device code that ends up in the work queue (e.g. ALLOCATE of a derived type with default initialization in a target region) then fails to link.
The CUDA PTX build already avoids this with a thin I/O mode. This renames `RT_CUDA_THIN_IO` to `RT_THIN_IO`, defines it in api-attrs.h for the native GPU builds (`RT_GPU_TARGET` without `RT_DEVICE_COMPILATION`), and keeps the CMake define for the CUDA PTX library under the new name. flang_rt_verbose_abort is defined in stl-overrides.cpp, so that's added to gpu_sources.
Tested on gfx90a. The amdgcn library no longer has those undefined symbols, gains only the flang_rt_verbose_abort definition, and doesn't lose any others. A small reproducer (derived type with default init allocated inside a target region) links and runs, and scalar PRINT from device code gives the same output as before. The host library's symbols are identical to before. I haven't built nvptx or the CUDA offload configuration.
Downstream report: ROCm/llvm-project#3517
[libc++] Encode the standard version in the ABI tag (#218527)
This prevents ODR mismatch issues from biting us across standard
versions. The order of elements in the ABI tag is now: libc++ version,
hardening mode, assertion semantic, exceptions, standard version.
Fixes #218524
libpkg: recognize loongarch64 as loongarch:64
Add PKG_ARCH_LOONGARCH64 and map EM_LOONGARCH/ELFCLASS64 to it; no
32-bit ABI exists. Test it in the frontend suite.
external/libecc: add __loongarch__ to words.h. Upstream, from
https://github.com/libecc/libecc/pull/15/.
[libc][math] Reorganize sinf and cosf function selection (#226529)
Reorganises sinf and cosf similarly to #224735 such that:
src/__support/math/func.h selects implementation, and
src/__support/math/func_<type>_eval.h implements func with <type> as the
intermediate computational type.
[ISel] Improve `clmul` fallback implementation (#204802)
Generalize the approach from
https://github.com/llvm/llvm-project/pull/203727 to narrower and wider
integers.
We still need the fallback for when multiplication isn't available, and
it turns out that for some widths the fallback emits fewer instructions,
the naive fallback is still used for `i1`, `i3`, `i4` and `i9`. I've
also now enabled wider integers (`i128` and `i256` have uses in
cryptography).
Based on my local experiments, the Karatsuba approach (e.g. as in
https://github.com/rust-lang/rust/pull/152132#discussion_r2778609222) is
not actually better than zero extending the input and using
multiplication with holes on the wider type.
CC https://github.com/llvm/llvm-project/issues/203694
CC: @eisenwave
[AMDGPU] Replace GCN-NOT with autogen checks in test (#227148)
The assembly has many v_mov_b32s, so minor scheduling changes can
trigger failure on the GCN-NOT. It seems the test is designed to show
CSE behavior, which still holds even if minor scheduling changes break
the NOT checks.
set_dist_point_name(): tiny tweak to restore previous behavior
Allocate fnm before allocating *pdp. This way a second call to to
set_dist_point_name() has a tiny little chance of succeeding.
ok beck ("I strongly suspect this will never matter anywhere.")