[InstCombine] Keep branch weights when folding a shift through a select (#227987)
Pulling a binop out of a select and through a constant shift rebuilds
the select, and the new one was losing the original branch weights. Copy
!prof so the weights stay attached to the same condition.
CodeGen: Fix stale live range for undef PHI sources on split edges
SplitCriticalEdge collects the PHI sources coming from the new block so
the trimming loop below does not undo the segment just added for them.
An undef operand gets no segment, but was still added to the set, so a
register that is only an undef PHI operand on the split edge kept the
stale extension of its live range through the new block.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
CodeGen: Check subrange liveness directly when trimming split edges
SplitCriticalEdge trims the stale extension of a live range into the newly
created block. The subrange guard used overlaps(StartIndex, EndIndex), which
happens to work only because insertMBBInMaps creates exactly one fresh index for
the new block, so no pre-existing segment endpoint can fall strictly inside the
range. That only happens to make makes overlap equivalent to containment.
Check liveAt(PrevIndex) instead, mirroring the same condition on the main range.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[Polly] Fix assertion in addUserAssumptions for unreachable blocks (#227311)
ScopBuilder::addUserAssumptions() crashes when an llvm.assume call
resides in a block that has no entry in InvalidDomainMap. This happens
when __builtin_unreachable() is converted to llvm.assume by earlier
passes, but Polly's buildDomainsWithBranchConstraints() skips the block
(e.g. due to an UnreachableInst terminator), leaving its
InvalidDomainMap entry uninitialized.
Skip such assumptions instead of asserting.
Fixes #226718
[lldb][test] Use external debug info on Arm in TestExitDuringStep and TestThreadExit (#228047)
Arm has to care about Arm and Thumb modes, so without the separate debug
info file installed by libc6-dbg on Linux, we cannot backtrace if the
starting point is in a libc function.
This is the case for these tests where a thread may be stopped in a libc
syscall wrapper.
As the backtrace wasn't working, the code added by #227312 did not see
that some threads were part of the main executable, which made both
these tests fail on Arm Linux.
It's possible we should be running more tests on Arm with this option
on, but for now I just want to get the bot back to green.
More details in https://github.com/llvm/llvm-project/issues/228014.
Mips/GlobalISel: Fix adding $gp to calls as a def instead of a use (#227679)
A PIC call needs $gp to point at the GOT for the lazy binding stub.
SelectionDAG adds it as an ordinary argument register. GlobalISel
instead added it as an implicit def, which killed the $gp copy set up
right before the call.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[flang][OpenMP] Remember decision of allowing past/future clause
When a deprecated or a future clause is used on a directive, and it is
allowed with a warning, remember that decision and consider that clause
allowed on that directive in all subsequent checks.
Introduce an OpenMPKartoffel warning category to guard these warnings
(and the corresponding -Wopenmp-kartoffel option).
SLP: Remove artificial shifts from min-bitwidth test (NFC) (#222634)
Replace shifts by zero with zero constants, using a slightly negative
threshold to preserve the test signal. This allows a future update to
the cost of foldable shifts.
[OpenMP] Add default value for Size argument in EnumSet (#227853)
Most of the interesting enums cover a contiguous range of values, and
define members First_ and Last_. By subtracting the underlying values
one can get the number of elements in the enum.
[LV][NFC] Print more debug to account for mismatch in final costs (#227757)
After printing out the costs of recipes we then print out the total
cost, including the cost per lane. Unfortunately, the final cost often
doesn't match the total of all the recipe costs due to extra precomputed
costs and reg spill costs. This PR adds the missing information.
[clang] Fix crash with follow-on diagnostics w/invalid logical operator (#227837)
If the logical operator involves a vector operand, we perform special
vector-specific checks. `CheckVectorOperands()` returns a null QualType
to signal there was an issue, and `CheckVectorLogicalOperands()` was
using that signal to decide to report an "invalid operands to binary
expression" diagnostic. However, `CheckVectorOperands()` also sometimes
modifies the given LHS and RHS values and when that happens, the caller
cannot assume they're still valid on a null QualType return. This was
causing a crash from `CheckVectorLogicalOperands()` because it was
attempting to use those newly invalidated expressions.
This fixes the crash by letting `CheckVectorOperands()` report the
diagnostics directly and removing the fallback logic from
`CheckVectorLogicalOperands()`. This also helpfully removes some
unhelpful follow-on diagnostics in other cases where we would report
"cannot convert operands" followed by "invalid operands to binary
expression".
Fixes #227588
[flang][OpenMP] Sink intervening code unguarded into collapsed loop body (#225159)
### Summary
Lower intervening code in a collapsed imperfect loop nest by sinking it
unguarded into the innermost `omp.loop_nest` body, so it executes once
per collapsed logical iteration. Per OpenMP 6.0 6.4.3.
An earlier implementation of this work guarded intervening code so that
it ran once per enclosing iteration. That required recomputing the inner
loop's bounds inside the region — and in a `target` region, either
accessing or re-evaluating
the enclosing `omp.target`'s `host_eval` bounds (#223475) — which
overcomplicates the solution and is not required by the spec. Sinking
the intervening code removes all bound arithmetic from the region, while
the execution count remains within the permitted range.
### Notes
Fixes #199092
Assisted-by: Copilot
[AMDGPU] Reject 64-bit VOP1 DPP on gfx8/gfx9 (#220834)
64-bit DPP needs FeatureDPALU_DPP (gfx90a+), but HasDPALU_DPP is missing
GCN3Encoding gate that HasDPP has
[LV] Vectorize uncountable early exit store loops with combined conditions (#205109)
Support the case where both the countable and uncountable exit
conditions have been combined by earlier passes.
[ConstraintElim] Skip rows implied by a single existing row. (#227688)
There are a number of cases where we add duplicated rows (e.g. from
transferring facts between the signed and unsigned systems, or from
tightening a non-strict bound using !=).
Before adding rows, check if the system already has a row that the same
variable coefficient and a constant that is <= the current constant.
This helps reduce compile-time, as each row adds extra work during
constraint solving:
stage1-O3: -0.02%
stage1-ReleaseThinLTO: -0.05%
stage1-ReleaseLTO-g: -0.02%
stage1-aarch64-O3: -0.00%
stage2-O3: -0.01%
stage2-clang: -0.00%
[11 lines not shown]
[AMDGPU] Don't merge M0 initializations across calls that clobber M0 (#221963)
hoistAndMergeSGPRInits collected clobbers only from
MRI.def_instructions(M0),
which lists explicit defs. Calls clobber M0 via a regmask and were
therefore
invisible, so M0 inits were merged and hoisted across them and the
required
re-init after a call was dropped. Scan the function using
modifiesRegister.
Issue: https://github.com/llvm/llvm-project/issues/221212
[compiler-rt] Add print_coverage_summary to silence SanitizerCoverage dump logs
Coverage dumps still write .sancov files; only the "PCs written" summary is
optional.
[GVN] Preserve vectorization opportunities when PREing loop loads
Loop load PRE can replace an invariant-address load with a loop-carried
PHI and conditional reload after a may-alias store. That scalar recurrence
can prevent vectorization even when runtime alias checks could disambiguate
the original accesses. Subsequent full unrolling then expands scalar code.
Conservatively preserve the header load in innermost loops whose clobber
is a conditional may-alias store through a varying pointer. Keep existing
PRE behavior for invariant-address clobbers, known aliasing, calls, ordered
memory operations, and loops with vectorization disabled or completed.
This is an opportunity heuristic, not a vectorization legality proof.
[OpenMP] Preserve host capture lifetimes in frontend lowering
Mark captured alloca/global storage nofreeobj when Clang emits a synchronous
host parallel call. This preserves lifetime across outlining without claiming
that escaped capture slots are noalias or immutable, and without extending the
lifetime guarantee to pointers loaded from those slots.
Mark OpenMPIRBuilder's fresh host capture aggregate noalias and nofreeobj.
This covers the lowering path used by Flang and Clang's IRBuilder mode.
Replace callback Attributor seeding with frontend and translation tests,
including escaped captures, heap references, firstprivate pointers, debug
wrappers, serialized regions, and optimized load hoisting with OpenMPOpt
disabled. Refresh the affected parameter-attribute checks.