[SLP]Use tighter runtime alias checks cost bound in vectorized loops
The scalar loop left by the loop vectorizer runs only a few iterations
or the ranges its runtime checks found overlapping, where the versioning
checks, executed on each iteration, are unlikely to pay off. Bound their
cost by 10% of the guarded scalar region cost in such loops, controlled
by -slp-vec-loop-rt-checks-cost-percent.
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/227756
[SSAF] Close unsafe-buffer reachability over override families (#213319)
An unsafe pointer reaching one override's parameter is equally unsafe in
every sibling and base override of that method, because the call site picks
the target dynamically. Without closing over the families, reachability
depended on which override the extractor happened to see the flow through, so a
fix suggested for the base could be contradicted by a derived override.
- Mirroring is level-preserving: families relate slot entities, so a reachable
EPL propagates only to the same pointer level on its family members.
- Mirroring happens inside the pointer-flow search, so flows out of a mirrored
EPL are followed too.
- Type-constrained slots are never mirrored onto, so C3 still holds.
§4 of rdar://179151603
Assisted-By: claude
[mlir][tosa] Lower MAX_POOL2D_ADAPTIVE to Linalg (#225675)
Legalize max_pool2d_adaptive with constant shape operands through the
existing max_pool2d conversion.
Assisted-by: Codex
[flang][OpenMP] Remember decision of allowing past/future clause
When a deprecated or a future clause is used on a directive, and it is
allowed with a warning, remember that decision and consider that clause
allowed on that directive in all subsequent checks.
Introduce an OpenMPKartoffel warning category to guard these warnings
(and the corresponding -Wopenmp-kartoffel option).
[OpenMP] Allow overriding clause-allowed-on-directive in decomposition (#227418)
Construct decomposition needs to know which clause is allowed on which
directive in a given OpenMP spec version. To allow implementations to
conditionally allow clauses from future (or past) versions, require the
helper class to implement the isClauseAllowedOnDirective check.
fixup! [AMDGPU] Measure MFMA read hazards at each producer
Reword the comments of the two tests where an exact srcC write must not
hide an earlier partial one: name the producers, give the nearer one's
window, and keep the farther one's window apart from the padding it
leaves. On gfx940+ the exact write now has a window of its own, since
two different MFMAs sharing an accumulator need wait states.
AI-assisted.
[AMDGPU] Measure MFMA overwrite hazards at each instruction
Apply previously established processing to:
- VALU overwriting an MFMA result
- VALU overwriting a register an MFMA took as srcC
AI-assisted.
[AMDGPU] Measure MFMA read hazards at each producer
Introduce more sophisticated traversal to avoid the following traps:
- order-dependent traversal and discarding seen BBs despite shorter path
- mis-matching distance and window of different producers
Record the best distance per BB instead of a visited flag and sweep the
arrivals in nondecreasing distance (bucket queue). This pairs producers
with their actual distance to a consumer in one go.
Fixed scenarios:
- MFMA reading an MFMA result as srcA, srcB or srcC
- VALU, memory or export instruction reading an MFMA result
AI-assisted.
[libc++] Inline the build-at-commit composite action into the benchmark workflows
Yet another twist in the endless saga of setting up performance benchmarking
infrastructure for libc++. Since the recent switch to Kubernetes runners,
there are now two layers of "pods". First, there's the runner pod which
executes the Github action itself. It's the one that loads the Docker
image specified in the Github workflow and then launches it.
Then, there's the workflow/job pod, which executes the actual Docker
image and the commands described in `steps` in the workflow file.
This causes problems because the workflow pod is the only one that has
access to the checked out monorepo (since it's checked out within the
running Docker container). This means the runner pod does not have access
to composite actions stored in the repository, which leads to errors like:
Can't find 'action.yml', 'action.yaml' or 'Dockerfile' under
'/home/runner/_work/llvm-project/llvm-project/.github/workflows/libcxx/build-at-commit'.
[3 lines not shown]
[MLIR] Emit final-policy remarks in deterministic order (#224616)
`RemarkEmittingPolicyFinal` stores remarks in a `DenseSet` whose hash
covers the location pointer and the process hash seed, so `finalize()`
emitted them in bucket order. That order depends on the build
configuration and on memory layout, which is why
`mlir/test/Pass/remark-final.mlir` used `CHECK-DAG`.
The engine already assigns every remark a `RemarkId` from a monotonic
counter when it is created, and the set already keeps the newer of two
remarks with the same identity. `finalize()` now sorts the drained
remarks by that ID before emitting. Remarks come out in creation order,
a replaced identity takes the position of its last report, and linked
remarks still follow their parent.
Behaviour change: order only. The identity, `DenseMapInfo<Remark>` and
the header are unchanged. Only remarks handed to the policy outside the
engine have no ID; the unit tests that do so assert unordered or single
results.
[11 lines not shown]
[MLIR][LLVMIR] Restore constant folding for global initializer GEPs (#227651)
Before #226904, a GEP in a global initializer went through
IRBuilder::CreateGEP, and MLIR's IRBuilder<TargetFolder> folded the
result
with ConstantFoldConstant. That combines nested GEPs, folds null and
integer
bases, and infers inbounds and nuw when the offset stays within the
global.
Building the constant expression directly skipped the folder, so the
output
lost those flags.
Fold the constant again. This also applies to inrange GEPs, which used
to
bypass the folder; the only visible difference there is that an inbounds
GEP with a non-negative offset now also gets nuw.
CIR's vtable, VTT and constant pointer tests check for the inferred
[3 lines not shown]
[flang] Let a directive keep the loop it owns when its body branches
A loop whose branching is confined to its body keeps its structured
form, but the construct holding it stayed Unstructured. A directive does
not merely contain such a loop, it owns it, and its lowering reads the
construct's own classification to decide whether the loop op carries its
bounds. The directive was left with a bounds-free loop that nothing
could partition, and the loop it owns became a second one nested inside.
Reclassify a directive construct once the loops it holds no longer need
it to stay Unstructured. Children are visited first, so those loops have
already been reclassified by the time the construct is reached. A
construct whose branching leaves it is untouched, as is one holding a
branch of its own.
Taking a loop over also means genFIR(DoConstruct) -- where a plain loop
folds a body whose branching stays inside it into a region -- never runs
for that loop, so fold its body through the same helper. A construct
that takes over no loop, acc data or acc parallel without a loop
[5 lines not shown]
[flang][NFC] Split the OpenACC construct lowering into two lanes
genFIR(OpenACCConstruct) decided twice, in three places, whether the
construct it lowers is structured, and reassigned the evaluation it works
from halfway through: before the descent that evaluation is the construct,
after it the loop the directive absorbs. Everything downstream had to know
which one it was holding.
Give each form its own function and leave genFIR to choose between them.
One lane allocates the exit selector, lowers the evaluations the construct
holds, and emits the jump table; the other reads the collapse clauses,
descends to the absorbed depth, and lowers what is inside it. The prologue
and epilogue are short enough to state in both rather than share.
[libc] Fix str_to_float.h clinger_fast_math for negative exponents. (#227391)
* Correctly handle inequality comparison with `exp10` when it equals
`INT32_MIN` and cannot be negated.
* When rounding mode is non-nearest, add missing handling of negative
exponents as in the nearest-mode fast-path above.
Both of these issues resulted in out-of-bounds access of
`POWERS_OF_TEN_ARRAY`.
While here, rename `exp10` to `exp_10` in str_to_float.h to disambiguate
from the C23 `exp10()` function. Also rename `exp2` to `exp_2` for
consistency.
Co-authored-by: @mxms0
Assisted-by: Automated tooling, human reviewed.
[LV] Reject non-simple outer-loop memory accesses. (#226212)
Reject loops with atomic or volatile access early in
canVectorizeOuterLoops. We cannot safely vectorize those instructions.
And at the VPlan-level, we do not model non-simple loads/stores etc, so
we cannot easily detect them (without reaching to the underlying
instructions).
Reject them at the outset, like in the inner loop path.
Without the checks, the newly added tests crash.
PR: https://github.com/llvm/llvm-project/pull/226212
[InstCombine] Fold comparisons of llvm.usub.sat result with its LHS (#214108)
Fold comparisons between the result of `llvm.usub.sat(X, C)` and the
original
left-hand side when `C` is a nonzero constant.
For `C != 0`:
```llvm
%sat = call iN @llvm.usub.sat.iN(iN %x, iN C)
%cmp = icmp eq iN %sat, %x
```
can be folded to:
```llvm
%cmp = icmp eq iN %x, 0
```
[26 lines not shown]
[SelectionDAG] Clear Known.One sign bit for ISD::FABS in computeKnownBits (#227457)
In `SelectionDAG::computeKnownBits`, `ISD::FABS` previously called
`Known.makeNonNegative()` after computing the known bits of its operand.
If the operand's sign bit was already known to be 1, `makeNonNegative()`
would only set `Known.Zero`'s sign bit without clearing `Known.One`'s
sign bit, resulting in a conflicting `KnownBits` state.
Explicitly set `Known.Zero`'s sign bit and clear `Known.One`'s sign bit
for `ISD::FABS`, and add a unit test.
CodeGen: Move SupportsDebugEntryValues from TargetOptions to TargetMachine
This is a capability reported by the target rather than an option to the
target. It has no command line flag or frontend plumbing to write it.
I'm not sure if we really need this; Triple::supportsDebugEntryValues() exists
although these don't entirely agree today.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
CodeGen: Move SupportsDefaultOutlining from TargetOptions to TargetMachine (#227691)
This flag is not an option passed down to a target, it is a capability reported by
one. It is only written by backend TargetMachine constructors. Move it into
TargetMachine, next to the vaguely similar O0WantsFastISel.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[AMDGPU][NFC] Extract the MFMA read-window calculation
Move the wait states a consumer needs before reading an MFMA result out
of checkMAIHazards90A into getMFMAReadWaitStates, taking the producer as
an argument, so a caller can ask about a specific producer. The caller
passes the producer the walk recorded, so nothing changes.
AI-assisted.
[AMDGPU] Fold sudot intrinsics with a zero multiplicand (#226479)
Replace sudot4 and sudot8 with their accumulator when either
multiplicand is zero.
For example:
```
sudot4(sign0, x, sign1, 0, acc, clamp)
->
acc
```