[SROA] Remove splitSliceTails loop in presplitLoadsAndStores (#227061)
The asserts in this loop can trigger when we presplit an overlapping
load/store pair, and the extra splits it adds don't appear to have any
benefit: we only get them when we have an unsplittable slice that's
fully enclosed inside a splittable slice, and in the test I've managed
to create for this (no_move_enclosed_unsplittable_store, and there are
no existing tests for this) the final SROA output is the same (except
that the variables end up with different names).
[Hexagon] Add missing -mtriple to live-outs test (#227079)
This test's RUN command did not specify a target triple, so lit could
execute it using the host target.
This causes the test to fail when LLVM is built with only the Hexagon
target enabled. Add -mtriple=hexagon so the test consistently runs as a
Hexagon test.
[flang][Test] Cover the lowering of loops with a non-terminating body
A previous change leaves such a loop unstructured. Check what that produces:
the cycle survives as a block branching to itself, no fir.do_loop is emitted
for the loop control, and a loop that only needs a block of its own still
gets the structured form with its body in a region.
[flang] Lower loops whose branching is confined to their body structurally
Such a loop was classified separately by a previous change but still
lowered as unstructured, so its structured form was lost.
Lower it structurally instead, with its body folded into a region that can
hold the branching. The loop keeps its bounds on the op, so it remains
available to whatever transforms or parallelizes it. Only the body is
folded: the loop control statements are emitted as they are for any
structured loop, since a branch from outside may target either of them.
Loops an OpenACC or OpenMP directive owns are lowered the same way, so
they keep their form too.
[flang] Detect loops whose branching is confined to their body (#225757)
A DO loop is classified as either structured or unstructured, and a
single raw branch anywhere in its body forces the loop -- and every
construct enclosing it -- onto the unstructured path.
That is stronger than necessary. A loop keeps its structured control
flow as long as its branching neither leaves its body nor enters it from
outside. Classify such a loop separately from a fully unstructured one.
This only classifies: lowering is unchanged. PFT dumps mark the new
classification with '~', which is what the tests key on.
[mlir][OpenACC] Match reuse barriers to the reused private scope (#225570)
A region that writes both gang- and worker-private slots was considered
to be a gang scope and a barrier was missed. This change keeps the two
store sets separate and pick the barrier from the scope that is actually
reused. Follow-up for
https://github.com/llvm/llvm-project/pull/224437#discussion_r4066418696.
[clang][OpenMP] Fix crashes on target regions inside namespace-scope lambdas and blocks (#226691)
Fixes #223397
A `target` region inside a lambda or block at namespace scope crashed
clang in two places. In Sema, `isOpenMPCapturedDecl` decides whether a
global must be captured by walking the function scope stack down to the
innermost OpenMP captured region, stopping at an ordinary function
scope. The capture initializers of a directive's outermost region are
built after all of its regions have been popped. Inside a function that
walk ends at the function's scope, but a namespace-scope lambda or block
has nothing underneath it, so the walk ran off the stack and asserted.
This happens for any global reference, such as `int &r = x; auto l = []
{ #pragma omp target r = 1; };`. The self-referential declaration in the
report is incidental. Once past Sema, CodeGen asserted too: it names the
outlined kernel after the region's parent function, and such a region
has none, even without a reference.
In Sema, running out of scopes now means the same as reaching a function
[3 lines not shown]
[mlir][ODS] Copy constant generated interface model prototypes (#226323)
Generate constexpr constructors for interface models whose concept is a
table of callbacks. Construct exact generated models from a static
prototype; keep external and fallback models on their existing path.
In OpenMPDialect.cpp this saves about 0.22B compiler instructions and
10.5KB of text.
Assisted-by: Codex
[SSAF] Serialize virtual method summaries and families (#213318)
Per-TU summaries and whole-program results cross process boundaries, and
the JSON layer refuses to write a summary kind it has no format for.
Register both sides so
--ssaf-extract-summaries=VirtualMethod becomes usable and the
family result survives a round trip.
Deserialization tolerates a missing override list, since a root virtual
method legitimately has none.
§3 of rdar://179151603
Assisted-By: claude
security/q-feeds-connector - add db_update action which indexes the additional feed information in sqlite, which will be executed during a regular update after fetch.
sysutils/sftp-backup - verify backups after put, closes https://github.com/opnsense/plugins/issues/5733
There might be valid reasons why we are not allowed to verify our backup, in which case the error message might not be relevant.
[InlineSpiller][AMDGPU] Implement subreg reload during RA spill
Currently, when a virtual register is partially used, the
entire tuple is restored from the spilled location, even if
only a subset of its sub-registers is needed. This patch
introduces support for partial reloads by analyzing actual
register usage and restoring only the required sub-registers.
This improvement enhances register allocation efficiency,
particularly for cases involving tuple virtual registers.
For AMDGPU, this change brings considerable improvements
in workloads that involve matrix operations, large vectors,
and complex control flows.