[AArch64] Use multi-vector intrinsics for masked load/store users of predicate-as-counter
If the user of the original wide mask is a masked load or store
intrinsic (matching the element size of the predicate-as-counter),
rewrite it directly to a masked multi-vector load/store.
This avoids materializing the vector mask and is easier to handle here
than later (e.g. in SelectionDAG), since we do not need to match the
concatenation of all `pext` segments of the predicate-as-counter.
Assisted-by: Codex
[AArch64] Avoid materializing full masks for extractelement users of predicate-as-counter (#220961)
If the user of the original wide mask is an `extractelement` and the
index is known to be within the first segment of the
predicate-as-counter, replace it with `extractelement(pext(counter,
0))`.
This avoids materializing the vector mask and produces a form that can
be folded into a conditional branch when the predicate-as-counter is
produced by a `whilelo`.
Assisted-by: Codex
[clang][OpenMP] Extend no-loop promotion to the split teams loop
Widen the no-loop promotion to apply not only to the fused target teams
distribute parallel for but include the same region written as separate
directives, other directives promotable to parallel for, and regions
carrying a schedule clause that asks for a mapping a no-loop kernel
already provides.
Use the SPMD_NO_LOOP kernel tag to carry the promotion decision, based
on whether the kernel can get fully promoted to a no-loop kernel. This
eliminates the previous discrepancy where a clause got the tag but not
the promotion.
PPC: Add MIR examples for missed extsw+word-load fold on subregister input
A gprc LWZ/LWZX feeding EXTSW_32_64 folds into a sign-extending LWA/LWAX
load. The equivalent 64-bit zero-extending word loads (LWZ8/LWZX8) whose
sub_32 feeds EXTSW_32_64 are not folded, leaving a redundant lwz+extsw
(or lwzx+extsw) pair. Add MIR examples documenting the missed fold.
Co-authored-by: Claude (Claude Opus 4.8, claude-opus-4-8) <noreply at anthropic.com>
PPC: Fold 64-bit zero-extending word load feeding extsw subregister
A gprc LWZ/LWZX feeding EXTSW_32_64 is rewritten into a sign-extending
LWA/LWAX load. Extend the same fold to the 64-bit zero-extending word
loads LWZ8/LWZX8 when the EXTSW_32_64 reads their sub_32 subregister,
producing a single LWA/LWAX instead of a redundant lwz+extsw pair.
Co-authored-by: Claude (Claude Opus 4.8, claude-opus-4-8) <noreply at anthropic.com>
PPC: Fix extsw elimination when the input reads a subregister
The EXTSW_32_64 sign-extend elimination previously assumed its input
was a full register value. It would then try using that value as the
source of the new (unnecessary) INSERT_SUBREG.
The new test would then hit this verifier error:
```
bb.0:
liveins: $x3
%0:g8rc = COPY killed $x3
%1:g8rc = RLDICL killed %0:g8rc, 0, 33
%3:g8rc = IMPLICIT_DEF
%2:g8rc = INSERT_SUBREG %3:g8rc(tied-def 0), %1:g8rc, %subreg.sub_32
$x3 = COPY killed %2:g8rc
BLR8 implicit $lr8, implicit $rm, implicit killed $x3
*** Bad machine code: INSERT_SUBREG expected inserted value to have equal or lesser size than the subreg it was inserted into ***
[8 lines not shown]
PPC: Fix EXTSW elimination promoting a subregister operand
promoteInstr32To64ForElimEXTSW copies operands from the 32-bit
instruction verbatim into its promoted 64-bit form. When an operand
reads the sub_32 subregister of a 64-bit register, the promoted
instruction (which takes a full register) ended up with an illegal
subregister use and failed the machine verifier.
Drop the sub_32 subregister and use the original full register, which
provides the low 32 bits the promoted instruction operates on. This
avoids verifier error regressions in a future change.
Co-Authored-By: Claude <noreply at anthropic.com> (Claude Opus 4.8, claude-opus-4-8)
[CIR] Destroy non-trivially destructible co_await results after the cir.await (#229585)
Follow-up to #225412. When await_resume() returns a class with a
non-trivial destructor and the result is discarded the destructor was
pushed inside the cir.await resume region. The associated
cir.cleanup.scope captured the region's terminator which led to a "block
with no terminator" verification error.
Push the destructor after the await instead such that the result is
destroyed at the end of the full expression.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[Clang][OpenMP] Mixed signed/unsigned trip counts fix (#226462)
Build that 1 already in the loop-index type (same pattern as the
existing 0).
Fixes #225756
[CIR][CodeGen] Share isStandardLibraryRTTIDescriptor
Deduplicates `isStandardLibraryRTTIDescriptor` between CIR and classic CodeGen
into `ItaniumCXXABIUtils.h`, taking the classic implementation. The two copies
have been equivalent since #227781 filled in the builtin types CIR was missing.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share canUseSingleInheritance
Deduplicates `canUseSingleInheritance` between CIR and classic CodeGen into
`ItaniumCXXABIUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[clang][OpenMP] Emit lastprivate final copies in no-loop kernels
Promotion rejected lastprivate because the no-loop branch leaves the
worksharing path before it privatizes or copies out, so the clause would
have been silently dropped.
Privatize the non-counter variables in the parallel region and copy them
out under the last iteration that applyWorkshareLoop publishes, forcing
the exit barrier the copy reads through. Loop counters stay at the
distribute level and reach EmitOMPSimdFinal.
[CIR][CodeGen] Share isStandardLibraryRTTIDescriptor
Deduplicates `isStandardLibraryRTTIDescriptor` between CIR and classic CodeGen
into `ItaniumCXXABIUtils.h`, taking the classic implementation. The two copies
have been equivalent since #227781 filled in the builtin types CIR was missing.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share canUseSingleInheritance
Deduplicates `canUseSingleInheritance` between CIR and classic CodeGen into
`ItaniumCXXABIUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share requiresAMDGPUProtectedVisibility
Deduplicates `requiresAMDGPUProtectedVisibility` between CIR and classic CodeGen
into `TargetUtils.h`. The shared version takes a bool for "currently hidden" in
place of the `llvm::GlobalValue` and `cir::VisibilityKind` the two callers
passed.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share hasExtraNeonArgument
Deduplicates `hasExtraNeonArgument` between CIR and classic CodeGen into
`TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Arm SME inlinability check
Deduplicates `ArmSMEInlinability` and `getArmSMEInlinability` between CIR and
classic CodeGen into a new `TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Retire the CodeGenUtils.h catch-all header
Moves the last helpers out of `CodeGenUtils.h` into `ClassUtils.h`,
`ModuleUtils.h`, `TargetUtils.h` and `FunctionUtils.h` (`checkTargetFeatures`,
since it came from CodeGenFunction.cpp) and deletes the header. Only moves code
already on main, so it can be dropped on its own.
Assisted-by: Claude Code (Claude Fable 5.1).