[Clang] Fix a regression introduced by #223645. (#227085)
In #223645, I removed a fixme I thought was no longer necessary. But it
was, our test coverage is just patchy.
This just reintroduces it to fix the regression,
I have not investigated whether there was a better solution.
I got the AI to write as much tests as i could think of, however this
revealed regressions introduced in an earlier version of clang, this
weill be fixed separately.
Reported in
https://github.com/llvm/llvm-project/pull/223645#issuecomment-5868853308
Assisted-By: Opus 5.5
[clang] Also disable scalable vectorisation for vectorize(disable) (#227254)
This ensures that #pragma clang loop vectorize(disable) also disables
scalable vectorisation. Only setting the width to 1 could potentially
enable LoopVectorizer to vectorize with VF = vscale x 1. Especially with
the -scalable-vectorization=preferred option.
[flang][Test] Cover the lowering of loops with a non-terminating body
A previous change leaves such a loop unstructured. Check what that produces:
the cycle survives as a block branching to itself, no fir.do_loop is emitted
for the loop control, and a loop that only needs a block of its own still
gets the structured form with its body in a region.
[flang] Lower loops whose branching is confined to their body structurally
Such a loop was classified separately by a previous change but still
lowered as unstructured, so its structured form was lost.
Lower it structurally instead, with its body folded into a region that can
hold the branching. The loop keeps its bounds on the op, so it remains
available to whatever transforms or parallelizes it. Only the body is
folded: the loop control statements are emitted as they are for any
structured loop, since a branch from outside may target either of them.
Loops an OpenACC or OpenMP directive owns are lowered the same way, so
they keep their form too.
[flang] Search every way back to a GO TO when looking for a cycle
A GO TO reaching a label above it closes a cycle, but the way back need not be
a branch: the statement it lands on simply runs on, through the constructs it
meets, until control reaches the GO TO again. Chaining from one GO TO to the
next stopped at the first statement of another kind and missed the cycle,
leaving the body free to run forever inside a region DCE then deleted.
Follow every successor instead, from the GO TO's targets until control returns
to it. An assigned GO TO closes a cycle the same way, and its targets are
already named by its successors.
Two edges are left out of the search. It stays inside the loop body, since
beyond it lies the loop's own iteration edge, which would carry the search back
in and make every branch look cyclic. For the same reason it skips the
iteration edge of any loop nested in that body: a loop reaching its own DO
statement ends on its own control.
The answer is an over-approximation. Some successor leads back, but control
[3 lines not shown]
[flang] Collect ASSIGNed labels before analyzing branches
An assigned GO TO reaches every label ASSIGNed to its variable, wherever the
ASSIGN sits. Branch analysis marked only the labels it had already walked
past, so an ASSIGN written after the GO TO left that target unrecorded. The
successors, and the incoming-branch map built from them, were incomplete for
every reader.
Collect the ASSIGNed labels of a unit before its branches are analyzed. The
recorded successors then name every target, so an assigned GO TO outside a
loop that can enter its body is seen as entering it, and a loop is no longer
reclassified as having self-contained branching when it has not.
With the successors complete, the classification needs no special case for
these statements: a listless GO TO is decided on its targets like any other
branch, rather than being turned away because a label list did not bound
them.
[AMDGPU] Route no-modifier reg-or-inline AsmParser operands through HwMode predicate (#221989)
Follow-up to #221988: routes the no-modifier reg-or-inline operands through the HwMode-split *_AlignTarget register classes, like #221988 did for the with-modifiers ones, so VGPR-tuple alignment is enforced by the operand's register class on subtargets that require aligned VGPRs. NFC on subtargets that don't require aligned VGPRs.
Co-Authored-By: Claude [noreply at anthropic.com](mailto:noreply at anthropic.com)
[AMDGPU] Route no-modifier reg-or-inline AsmParser operands through HwMode predicate (#221989)
Follow-up to #221988: routes the no-modifier reg-or-inline operands through the HwMode-split *_AlignTarget register classes, like #221988 did for the with-modifiers ones, so VGPR-tuple alignment is enforced by the operand's register class on subtargets that require aligned VGPRs. NFC on subtargets that don't require aligned VGPRs.
Co-Authored-By: Claude [noreply at anthropic.com](mailto:noreply at anthropic.com)
[AArch64][GlobalISel] Reorganise shuffle combines. NFC (#227268)
Other combines had been added between shuffle lowerings, make sure they
are adjacent. Also remove show extra whitespace whilst here.
[CIR][CodeGen][NFC] Retire the CodeGenUtils.h catch-all header
Moves the last helpers out of `CodeGenUtils.h` into `ClassUtils.h`,
`ModuleUtils.h`, `TargetUtils.h` and `FunctionUtils.h` (`checkTargetFeatures`,
since it came from CodeGenFunction.cpp) and deletes the header. Only moves code
already on main, so it can be dropped on its own.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share requiresAMDGPUProtectedVisibility
Deduplicates `requiresAMDGPUProtectedVisibility` between CIR and classic CodeGen
into `TargetUtils.h`. The shared version takes a bool for "currently hidden" in
place of the `llvm::GlobalValue` and `cir::VisibilityKind` the two callers
passed.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share hasExtraNeonArgument
Deduplicates `hasExtraNeonArgument` between CIR and classic CodeGen into
`TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Arm SME inlinability check
Deduplicates `ArmSMEInlinability` and `getArmSMEInlinability` between CIR and
classic CodeGen into a new `TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share hasOwnStorage
Deduplicates `hasOwnStorage` between CIR and classic CodeGen into
`RecordLayoutUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the bit-field and vbase layout ABI predicates
Deduplicates `isDiscreteBitFieldABI` and `isOverlappingVBaseABI` between CIR and
classic CodeGen into `RecordLayoutUtils.h`, as free functions taking the
`ASTContext`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share isEmptyFieldForLayout and isEmptyRecordForLayout
Deduplicates `isEmptyFieldForLayout` and `isEmptyRecordForLayout` between CIR
and classic CodeGen into a new `RecordLayoutUtils.h`. `ABIInfoImpl.h` and CIR's
`TargetInfo.h` re-export them with using-declarations, so the ~30 unqualified
callers are untouched.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen] Share isStandardLibraryRTTIDescriptor
Deduplicates `isStandardLibraryRTTIDescriptor` between CIR and classic CodeGen,
taking the classic implementation. CIR's copy hit `llvm_unreachable("NYI")` on
`WasmExternRef` and `HLSLResource`; neither is reachable in CIR today, so no
test changes.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share canUseSingleInheritance
Deduplicates `canUseSingleInheritance` between CIR and classic CodeGen into
`ItaniumCXXABIUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Itanium __vmi_class_type_info flags computation
Deduplicates the `__vmi_class_type_info` and `__base_class_type_info` flags and
`computeVMIClassTypeInfoFlags` between CIR and classic CodeGen into
`ItaniumCXXABIUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Itanium __pbase_type_info flags and predicates
Deduplicates the `__pbase_type_info` flags, `containsIncompleteClassType` and
`extractPBaseFlags` between CIR and classic CodeGen into `ItaniumCXXABIUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen] Share the EH personality selection logic
Deduplicates the EH personality selection (`getEHPersonality`,
`getCXXEHPersonality`) between CIR and classic CodeGen, taking the classic
implementation. CIR's copy lacked the z/OS, Wasm and GNUstep-on-CygMing cases;
none are reachable in CIR today, so no test changes.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Extract the EHPersonality struct into CodeGenUtils
Deduplicates the `EHPersonality` struct and its personality constants between
CIR and classic CodeGen into `EHPersonality.h`. The `EHPersonality::get`
factories become `getEHPersonality` overloads, since the shared struct cannot
name either CodeGenModule.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFCI] Share the COMDAT and common-linkage predicates
Deduplicates `shouldBeInCOMDAT` and `isVarDeclStrongDefinition` between CIR and
classic CodeGen into `ModuleUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Retire the CodeGenUtils.h catch-all header
Moves the last helpers out of `CodeGenUtils.h` into `ClassUtils.h`,
`ModuleUtils.h`, `TargetUtils.h` and `FunctionUtils.h` (`checkTargetFeatures`,
since it came from CodeGenFunction.cpp) and deletes the header. Only moves code
already on main, so it can be dropped on its own.
[CIR][CodeGen][NFC] Share requiresAMDGPUProtectedVisibility
Deduplicates `requiresAMDGPUProtectedVisibility` between CIR and classic CodeGen
into `TargetUtils.h`. The shared version takes a bool for "currently hidden" in
place of the `llvm::GlobalValue` and `cir::VisibilityKind` the two callers
passed.
[CIR][CodeGen][NFC] Share the Arm SME inlinability check
Deduplicates `ArmSMEInlinability` and `getArmSMEInlinability` between CIR and
classic CodeGen into a new `TargetUtils.h`.