Fix typo in CHECK line in clang/test/CodeGen/systemz-charset.c (#229762)
Replace `=` with `:` to fix a typo in a CHECK line in
clang/test/CodeGen/systemz-charset.c
[mlir][acc] Report independent loops not parallelized due to unstructured control flow
An independent acc loop whose body has unstructured control flow (for
example a backward GOTO) cannot be represented as a structured loop, so it
is executed sequentially. This happened silently: no remark was emitted for
the loop, and the user only saw a serial kernel without a reason.
Emit a remark for such loops stating that they are not parallelized because
their control flow could not be represented as a structured loop. This lets
users identify loops that are asserted parallel but run sequentially.
Co-Authored-By: Claude
[SDAG] Reuse MMO in `legalizeStoreOps` promotion path
I hit this bug in #229429 when using `i64` for the 64-bit stores.
I ended up using `<2 x i32>` for consistency with other intrinsics but the
bug still remains.
No target hits this so far so I can't test it, but I think it's good to fix it.
Otherwise a i64 store that goes through this path can lose its atomicity for example.
The load path already preserves the MMO I believe.
[AMDGPU][GISel] Honor the nomerge attribute on calls (#229152)
GlobalIsel and SDag never handled the `nomerge` attribute from the IR.
For SDag, this handling is present in other targets. GISsel had no
`NoMerge` handling, and is relevant for branch folding, so it was added.
The primary motivation for this PR is sanitizer reports. Sanitizers
often rely on `__builtin_return_address(0)` from a function call to
symbolize the site of the error. Without `nomerge` it becomes possible
for two report calls to be merged and no longer report the single,
unique source location.
[X86] combineSelect - don't check SimpleValueTypes through bitcast (#230040)
Although unlikely, its always possible that there's a weird type behind the cast
Noticed by inspection
[VPlan] Don't merge allocas across lanes, parts or recipes. (#229929)
Each execution of an alloca creates a new allocation, distinct from the
allocations of all other executions. They must not be merged, either by
CSE or via isUniformAcrossVFsAndUFs.
Update those to explicitly handle Alloca recipes.
Fixes https://github.com/llvm/llvm-project/issues/229389.
PR: https://github.com/llvm/llvm-project/pull/229929
[AArch64] Pre-commit tests for vector SQABS
Cover saturating absolute value expressed as UMIN of ABS and the signed
maximum for Neon and scalable vectors. Include commuted operands, poison,
multiple uses and negative cases.
Record the existing code generation so the SQABS combine can show its
effect as updates to these checks.
[LV] Don't tail-fold the epilogue with EVL-based tail-folding (#228452)
Epilogue tail-folding isn't supported yet with the `DataWithEVL`
tail-folding style. The epilogue plan is selected by duplicating an
existing VPlan, and some recipes used only by EVL-based tail-folding
don't implement clone() yet, so cloning a plan that contains them hits
llvm_unreachable. This affects targets that prefer `DataWithEVL`, such
as RISC-V.
Until those recipes implement clone(), fall back to a normal epilogue
when the preferred (or forced) tail-folding style is DataWithEVL. This
bail-out is based on the preferred style rather than the style that
would actually be chosen, so it also rejects fixed-width epilogue VFs,
where EVL isn't used. It will be refined once the epilogue's chosen
tail-folding style is known when epilogue TF gets supported.
[VPlan] Support tailfolded loops in multi-use-reductions (#214455)
Previously `handleMultiUseReductions()` would bail out for
tailfolded-loops, since the backedge value is no longer the reduction
intrinsic but the predicated select based on the header-mask.
This patch adds pattern matching to handle tailfolded loops accordingly.
I compiled the LLVM testsuite (for RISCV with -march=rv64gcv) and it
only triggers in the corresponding unit-tests in
`SingleSource/UnitTests/Vectorizer/`. I guess that's because there are
still some other limitations when handling more generic multi-use
reduction cases (e.g. differing types in CanonicalIV vs. WideIV).
[ConstraintElim] Bound IVs by compares of the phi in header/latch. (#226297)
Extend addInfoForInductions to add bounds for IVs when the phi is
compared in the header or latch:
In that case, every iteration taking the backedge checked
PN ContinuePred B, which guarantees PN != B for NE and LT predicates.
For an increment by one, PN != B together with StartValue <= B (added
precondition) imply PN <= B.
Alive2 Proofs:
* icmp ult in header: https://alive2.llvm.org/ce/z/x37Uvm
* icmp slt in header: https://alive2.llvm.org/ce/z/ibnn5k
* icmp ne in header: https://alive2.llvm.org/ce/z/wLRfZe
Triggers in a number of additional cases on
https://github.com/dtcxzyw/llvm-opt-benchmark-nightly/pull/1431.
As-is, this comes with a slight compile-time increase:
[9 lines not shown]
[Mips] Lower calls to private functions the same way as internal functions (#229975)
As with 85b4fbfecb65cd1f5a06e707e7ec0fc2de562407, this seems to be
another case where the backend was treating private functions
differently for no apparent reason.
This was particularly problematic for building the Zig compiler with
LLVM linked statically because we'd have enough such calls that linking
failed with thousands of these errors:
relocation R_MIPS_CALL16 out of range: 76048 is not in [-32768, 32767]
RegisterPressure: Detect dead physreg defs from LiveIntervals
This reverts the remainder of #222627, which was partially reverted by
on dead flags. This is a prerequisite to deleting LiveVariables.
When constructing the PressureDiff for an instruction during scheduling
DAG construction, dead defs were only recognized from the dead flag on the
operand. This implicitly relied on preprocessing done by LiveVariables to
fixup inconsistent dead flags with overlapping registers in other operands.
Dead flags have no verifier-enforced rules and are thus unreliable.
Before LiveVariables, consider this example:
dead $eax = MOV32r0 implicit-def dead $eflags, implicit-def $rax
; $rax is never used
$rax is never used, but only the $eax def is dead-flagged and the overlapping
implicit-def $rax is not. The shared $eax register units are covered by the
non-dead $rax def and so are counted as live defs. That shared unit is then
[26 lines not shown]
RegisterCoalescer: Keep remat def dead if it's a copy destination superregister (#230036)
This is a refinement of #226037, which was too strict.
When rematerializing into a physical register that is not exactly the copy's
destination, the def should only stay live if it is a sub-register of the copy
destination, i.e. part of the live value. Checking register unit coverage also
kept the def live when it is a super-register with the same units as the copy
destination, such as $rax for a copy into $eax on x86_64:
dead $rax = MOV64ri32 -11, implicit-def $eax
Only the $eax part is used, so the $rax def is dead. This matches what
LiveVariables produces for a full def with partial uses.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[mlir][tosa] Adding integer data layout operation support to PRO-FP (#229777)
This pull request was made to add integer data layout operation support
to PRO-FP. This is to prevent the need to cast between int and fp to do
these operations, which had the possibility of producing errors or
unwanted behaviour. Partially implements:
https://github.com/arm/tosa-specification/pull/91
Co-authored-by: Luke Hutton <luke.hutton at arm.com>
[OpenMP][DeviceRTL] Report the source location in __kmpc_error diagnostics (#224298)
Follow-up to #220702. Completes #204240.
The device runtime accepted the `ident_t` argument but ignored it, so
`error at(execution)` in a `target` region printed no source location.
This reports it, replicating the host runtime:
```
OMP: error_directive.f90:14:3: Encountered user-directed warning: warning message.
```
When the ident carries no location the result is `unknown:0:0`, same as
the host. flang populates the ident only with `-g`; clang always does.
Assisted-by: Copilot
SelectionDAG: Stop emitting kill flags in InstrEmitter
These is no point to maintaining these before register allocation.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
RuntimeLibcalls: Pass the default calling convention to libraries
Previously a setAvailableLibFuncs_* function computed DefaultCC locally when a
member calling convention referenced it. To do that, the emitter worked
backwards from a library to the SystemRuntimeLibrary records that reference it,
and had to diagnose the cases where that failed, which would be if there is no
referencing system library or several different ones.
The default calling convention belongs to the target, not to a library. The
dispatcher in setTargetRuntimeLibcallSets already computes it, so pass it to
each library function as a parameter. This removes the reverse lookup and both
diagnostics. DefaultCC references are now valid in a library shared by
SystemRuntimeLibrary records with different defaults, and in a library no
SystemRuntimeLibrary references.
This fixes errors when a LibcallLibrary is unused. This will enable defining
the vector math libraries in the future, as well as decoupling the target
specific handling in #229562.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>