[libc++] Encode the standard version in the ABI tag (#218527)
This prevents ODR mismatch issues from biting us across standard
versions. The order of elements in the ABI tag is now: libc++ version,
hardening mode, assertion semantic, exceptions, standard version.
Fixes #218524
[libc][math] Reorganize sinf and cosf function selection (#226529)
Reorganises sinf and cosf similarly to #224735 such that:
src/__support/math/func.h selects implementation, and
src/__support/math/func_<type>_eval.h implements func with <type> as the
intermediate computational type.
[ISel] Improve `clmul` fallback implementation (#204802)
Generalize the approach from
https://github.com/llvm/llvm-project/pull/203727 to narrower and wider
integers.
We still need the fallback for when multiplication isn't available, and
it turns out that for some widths the fallback emits fewer instructions,
the naive fallback is still used for `i1`, `i3`, `i4` and `i9`. I've
also now enabled wider integers (`i128` and `i256` have uses in
cryptography).
Based on my local experiments, the Karatsuba approach (e.g. as in
https://github.com/rust-lang/rust/pull/152132#discussion_r2778609222) is
not actually better than zero extending the input and using
multiplication with holes on the wider type.
CC https://github.com/llvm/llvm-project/issues/203694
CC: @eisenwave
[AMDGPU] Replace GCN-NOT with autogen checks in test (#227148)
The assembly has many v_mov_b32s, so minor scheduling changes can
trigger failure on the GCN-NOT. It seems the test is designed to show
CSE behavior, which still holds even if minor scheduling changes break
the NOT checks.
TwoAddressInstructions: Move the rescheduled copy chain back to front (#227289)
rescheduleMIBelowKill sinks an instruction below the kill of its tied
source, along with the run of copies that follows it. With LiveIntervals
the copies are spliced one at a time so handleMove sees a well-formed
block, but they were visited front to back and each inserted before the
previously moved one. This reversed them, and transiently moved a copy
below its use, asserting in handleMoveDown.
Walk them back to front instead, which also preserves the original
order.
Exposed by #225174, which made LiveIntervals available here by default.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[SROA] Remove splitSliceTails loop in presplitLoadsAndStores (#227061)
The asserts in this loop can trigger when we presplit an overlapping
load/store pair, and the extra splits it adds don't appear to have any
benefit: we only get them when we have an unsplittable slice that's
fully enclosed inside a splittable slice, and in the test I've managed
to create for this (no_move_enclosed_unsplittable_store, and there are
no existing tests for this) the final SROA output is the same (except
that the variables end up with different names).
[Hexagon] Add missing -mtriple to live-outs test (#227079)
This test's RUN command did not specify a target triple, so lit could
execute it using the host target.
This causes the test to fail when LLVM is built with only the Hexagon
target enabled. Add -mtriple=hexagon so the test consistently runs as a
Hexagon test.
[flang][Test] Cover the lowering of loops with a non-terminating body
A previous change leaves such a loop unstructured. Check what that produces:
the cycle survives as a block branching to itself, no fir.do_loop is emitted
for the loop control, and a loop that only needs a block of its own still
gets the structured form with its body in a region.
[flang] Lower loops whose branching is confined to their body structurally
Such a loop was classified separately by a previous change but still
lowered as unstructured, so its structured form was lost.
Lower it structurally instead, with its body folded into a region that can
hold the branching. The loop keeps its bounds on the op, so it remains
available to whatever transforms or parallelizes it. Only the body is
folded: the loop control statements are emitted as they are for any
structured loop, since a branch from outside may target either of them.
Loops an OpenACC or OpenMP directive owns are lowered the same way, so
they keep their form too.
[flang] Detect loops whose branching is confined to their body (#225757)
A DO loop is classified as either structured or unstructured, and a
single raw branch anywhere in its body forces the loop -- and every
construct enclosing it -- onto the unstructured path.
That is stronger than necessary. A loop keeps its structured control
flow as long as its branching neither leaves its body nor enters it from
outside. Classify such a loop separately from a fully unstructured one.
This only classifies: lowering is unchanged. PFT dumps mark the new
classification with '~', which is what the tests key on.
[mlir][OpenACC] Match reuse barriers to the reused private scope (#225570)
A region that writes both gang- and worker-private slots was considered
to be a gang scope and a barrier was missed. This change keeps the two
store sets separate and pick the barrier from the scope that is actually
reused. Follow-up for
https://github.com/llvm/llvm-project/pull/224437#discussion_r4066418696.
[clang][OpenMP] Fix crashes on target regions inside namespace-scope lambdas and blocks (#226691)
Fixes #223397
A `target` region inside a lambda or block at namespace scope crashed
clang in two places. In Sema, `isOpenMPCapturedDecl` decides whether a
global must be captured by walking the function scope stack down to the
innermost OpenMP captured region, stopping at an ordinary function
scope. The capture initializers of a directive's outermost region are
built after all of its regions have been popped. Inside a function that
walk ends at the function's scope, but a namespace-scope lambda or block
has nothing underneath it, so the walk ran off the stack and asserted.
This happens for any global reference, such as `int &r = x; auto l = []
{ #pragma omp target r = 1; };`. The self-referential declaration in the
report is incidental. Once past Sema, CodeGen asserted too: it names the
outlined kernel after the region's parent function, and such a region
has none, even without a reference.
In Sema, running out of scopes now means the same as reaching a function
[3 lines not shown]
[mlir][ODS] Copy constant generated interface model prototypes (#226323)
Generate constexpr constructors for interface models whose concept is a
table of callbacks. Construct exact generated models from a static
prototype; keep external and fallback models on their existing path.
In OpenMPDialect.cpp this saves about 0.22B compiler instructions and
10.5KB of text.
Assisted-by: Codex