[CIR][CodeGen][NFC] Retire the CodeGenUtils.h catch-all header
Moves the last helpers out of `CodeGenUtils.h` into `ClassUtils.h`,
`ModuleUtils.h`, `TargetUtils.h` and `FunctionUtils.h` (`checkTargetFeatures`,
since it came from CodeGenFunction.cpp) and deletes the header. Only moves code
already on main, so it can be dropped on its own.
Assisted-by: Claude Code (Claude Fable 5.1).
AMDGPU: Stop setting kill flags before FinalizeISel
Work on removing all pre-RA flag management. Kill flags should eventually be
removed. Pre-regalloc passes no longer depend on them, LiveIntervals strips them
and VirtRegRewriter re-introduces them.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[CIR][CodeGen][NFC] Share requiresAMDGPUProtectedVisibility
Deduplicates `requiresAMDGPUProtectedVisibility` between CIR and classic CodeGen
into `TargetUtils.h`. The shared version takes a bool for "currently hidden" in
place of the `llvm::GlobalValue` and `cir::VisibilityKind` the two callers
passed.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Arm SME inlinability check
Deduplicates `ArmSMEInlinability` and `getArmSMEInlinability` between CIR and
classic CodeGen into a new `TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share hasExtraNeonArgument
Deduplicates `hasExtraNeonArgument` between CIR and classic CodeGen into
`TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share isEmptyFieldForLayout and isEmptyRecordForLayout
Deduplicates `isEmptyFieldForLayout` and `isEmptyRecordForLayout` between CIR
and classic CodeGen into a new `RecordLayoutUtils.h`. `ABIInfoImpl.h` and CIR's
`TargetInfo.h` re-export them with using-declarations, so the ~30 unqualified
callers are untouched.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share hasOwnStorage
Deduplicates `hasOwnStorage` between CIR and classic CodeGen into
`RecordLayoutUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the bit-field and vbase layout ABI predicates
Deduplicates `isDiscreteBitFieldABI` and `isOverlappingVBaseABI` between CIR and
classic CodeGen into `RecordLayoutUtils.h`, as free functions taking the
`ASTContext`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen] Share isStandardLibraryRTTIDescriptor
Deduplicates `isStandardLibraryRTTIDescriptor` between CIR and classic CodeGen
into `ItaniumCXXABIUtils.h`, taking the classic implementation. The two copies
have been equivalent since #227781 filled in the builtin types CIR was missing.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share canUseSingleInheritance
Deduplicates `canUseSingleInheritance` between CIR and classic CodeGen into
`ItaniumCXXABIUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen] Share the EH personality selection logic
Deduplicates the EH personality selection (`getEHPersonality`,
`getCXXEHPersonality`) between CIR and classic CodeGen, taking the classic
implementation. CIR's copy lacked the z/OS, Wasm and GNUstep-on-CygMing cases;
none are reachable in CIR today, so no test changes.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Itanium __vmi_class_type_info flags computation
Deduplicates the `__vmi_class_type_info` and `__base_class_type_info` flags and
`computeVMIClassTypeInfoFlags` between CIR and classic CodeGen into
`ItaniumCXXABIUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Itanium __pbase_type_info flags and predicates (#228187)
Deduplicates the `__pbase_type_info` flags, `containsIncompleteClassType` and
`extractPBaseFlags` between CIR and classic CodeGen into `ItaniumCXXABIUtils.h`.
Re-landing #223422, which Erich Keane approved; its merge went into the parent
branch of the stack instead of main.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Itanium __pbase_type_info flags and predicates (#228187)
Deduplicates the `__pbase_type_info` flags, `containsIncompleteClassType` and
`extractPBaseFlags` between CIR and classic CodeGen into `ItaniumCXXABIUtils.h`.
Re-landing #223422, which Erich Keane approved; its merge went into the parent
branch of the stack instead of main.
Assisted-by: Claude Code (Claude Fable 5.1).
[BasicAA] Remove TODO in test (NFC) (#228992)
We've received the second AI generated implementation for this
TODO, and both times found that addressing it shows no benefit
to real-world code. As such, remove the TODO.
[clang] Migrate away from PointerUnion::dyn_cast (NFC) (#228991)
Note that PointerUnion::dyn_cast has been soft deprecated in
PointerUnion.h:
// FIXME: Replace the uses of is(), get() and dyn_cast() with
// isa<T>, cast<T> and the llvm::dyn_cast<T>
Literal migration would result in dyn_cast_if_present (see the
definition of PointerUnion::dyn_cast), but this patch uses dyn_cast.
Note that Target here traces back to AnnotationWarningsMap, which is
populated only with nonnull pointers.
Assisted-by: Antigravity
[mlir][x86] Bail out gracefully if accumulator is initialized from block arg (#228080)
`traceToVectorReadLikeParentOperation` may return `nullptr`, so defer
further checks until we know that a suitable op was found.
[CIR][HIP] Use the kernel handle as a kernel's address on the host (#228377)
For HIP, a __global__ function referenced from host code is represented
by its kernel handle. CIR used the address of the device stub instead,
so APIs that take a kernel pointer failed.
Match classic codegen:
- emitFunctionDeclLValue gives the address of the kernel handle.
- Constant initializers, such as tables of kernel pointers, refer to the
kernel handle.
- A launch through a kernel pointer (f<<<...>>>) loads the device stub
from the handle and calls it.
CUDA has no separate kernel handle, instead the address of a kernel is
the device stub. The new test covers both HIP and CUDA.
Assisted-by: Claude Opus 5.5
Signed-off-by: Steffen Holst Larsen <sholstla at amd.com>
[NFC][IR] Use getDataLayout member function for Instruction and BasicBlock (#228989)
Follow up on #96902: remove some uses of the old
`getModule()->getDataLayout()` pattern.
Related PR: #228964
[NFC][IR] Use getDataLayout member function for Function and GlobalValue (#228964)
Follow up on #96919: remove some uses of the old
`getParent()->getDataLayout()` pattern.
[flang] Fix locations of wrapped unstructured constructs and DO loop ends (#228577)
Construct evaluations have no source position, so the scf.execute_region
wrapping an unstructured construct was given the location of the
previously lowered statement. Use the construct's first statement for
the region and its END statement for the scf.yield. Also attribute the
DO loop end code to the END DO statement rather than to the last
statement of the loop body.
This avoids going back to previous lines when stepping in a debugger.
Assisted-by: AI
[mlir][scf] Keep scf.yield locations when lowering scf.execute_region (#228575)
The branches replacing the scf.yield terminators of an
scf.execute_region used the location of the scf.execute_region. Use the
location of the scf.yield they replace instead.
[TailCallElim] Do not mark a call tail when it is handed the frame (#218797)
markTails refuses to mark a call tail if it is passed an alloca or a
byval
argument, but it missed the intrinsics that return an address in the
current
frame, such as llvm.frameaddress(0), llvm.localaddress and
llvm.stacksave. The
frame is torn down before a tail callee runs, so the callee received a
dangling
pointer. gcc.c-torture/execute/frame-address.c aborts because of this.
llvm.stackrestore is no longer treated as an escape, since it does not
capture
its argument. Otherwise calls after a VLA scope would lose their tail
marking.
LangRef now states that a tail callee may not access the caller's stack
frame.
[3 lines not shown]
[VectorCombine] Fold insertelement chains of scalar parts to a bitcast and shuffle (#226224)
An insertelement chain whose elements are all truncated parts of the
same scalar is lowered element by element, unless InstCombine can turn
an in-order pair of halves into a bitcast (`foldTruncInsEltPair`). The
SLP vectorizer produces such chains for the fields of a struct that SROA
loaded as one integer, e.g. when summing two float fields in a loop:
```llvm
%hi = lshr i64 %x, 32
%h = trunc i64 %hi to i32
%l = trunc i64 %x to i32
%v0 = insertelement <2 x i32> poison, i32 %h, i64 0
%v1 = insertelement <2 x i32> %v0, i32 %l, i64 1
```
which X86 lowers to shrq + vmovd + vpinsrd. If TTI says it is cheaper,
replace the chain by a shuffle of the bitcast scalar:
[27 lines not shown]