[AMDGPU] Extract byte lanes of a split vector from the 32-bit source
After a <4 x i8> is split into i16 halves, a byte lane is extended
from an i16 shift of a truncate. Rewrite it on the 32-bit source as an
and of srl or a sign_extend_inreg of srl, which select to a single bit
field extract.
Uniform zero and any extends are left alone, their i16 shift is already
promoted to i32.
PowerPC/GlobalISel: Stop setting kill flags on selected instructions
There is no point in maintaining these before register allocation
anymore.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
RISCV: Stop setting kill flags on virtual registers before FinalizeISel
Kill flags have no remaining use before register allocation and are
stripped by LiveIntervals.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[mlir][tosa] Preserve all-NaN max-pool windows (#225744)
Initialize IGNORE-mode floating-point max pooling with NaN and select
the first finite input. This keeps an all-NaN window NaN instead of
returning the lowest finite value.
Assited-by: Codex
[CIR][CodeGen][NFC] Retire the CodeGenUtils.h catch-all header
Moves the last helpers out of `CodeGenUtils.h` into `ClassUtils.h`,
`ModuleUtils.h`, `TargetUtils.h` and `FunctionUtils.h` (`checkTargetFeatures`,
since it came from CodeGenFunction.cpp) and deletes the header. Only moves code
already on main, so it can be dropped on its own.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share requiresAMDGPUProtectedVisibility
Deduplicates `requiresAMDGPUProtectedVisibility` between CIR and classic CodeGen
into `TargetUtils.h`. The shared version takes a bool for "currently hidden" in
place of the `llvm::GlobalValue` and `cir::VisibilityKind` the two callers
passed.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share hasOwnStorage
Deduplicates `hasOwnStorage` between CIR and classic CodeGen into
`RecordLayoutUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share hasExtraNeonArgument
Deduplicates `hasExtraNeonArgument` between CIR and classic CodeGen into
`TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the bit-field and vbase layout ABI predicates
Deduplicates `isDiscreteBitFieldABI` and `isOverlappingVBaseABI` between CIR and
classic CodeGen into `RecordLayoutUtils.h`, as free functions taking the
`ASTContext`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Arm SME inlinability check
Deduplicates `ArmSMEInlinability` and `getArmSMEInlinability` between CIR and
classic CodeGen into a new `TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share isEmptyFieldForLayout and isEmptyRecordForLayout
Deduplicates `isEmptyFieldForLayout` and `isEmptyRecordForLayout` between CIR
and classic CodeGen into a new `RecordLayoutUtils.h`. The 25 callers now name
them as `CodeGenUtils::isEmptyFieldForLayout` and
`CodeGenUtils::isEmptyRecordForLayout`, like the other shared helpers.
Assisted-by: Claude Code (Claude Fable 5.1).
VE: Stop setting kill flags on virtual registers before FinalizeISel
There is no point in maintaining kill flags before register allocation
anymore.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[clang][DebugInfo] Make fwd decl call site debug info consistent for methods (#222263)
Prior to this patch, using `EmitFunctionDecl` for methods results in
different fields and flags than if `getFunctionDeclaration` (which calls
`CreateCXXMemberFunction`) is used. This can arbitrarily result in
differences depending on the shape of the source code (missing
`scopeLine` or access flags in some cases which are present in others).
[mlir][MemorySlotInterfaces] Rename `elemType` to `valueType` (NFC) (#228466)
The name `elemType` is potentially confusing, as "element type" already
has a well-defined meaning for types such as MemRef and Vector.
As of #211880, `MemorySlotInterfaces` also supports vector ops, and the
type represented by `elemType` can itself be a Vector type. In that case, the
Vector is the value stored in the memory slot, while its element type is
a different type.
This PR renames `elemType` to `valueType` to make this distinction
explicit.
VE: Remove broken nested call frame around dynamic stack allocation
lowerDYNAMIC_STACKALLOC wrapped the __ve_grow_stack call and the
GETSTACKTOP stack-pointer read in a zero-sized CALLSEQ_START/CALLSEQ_END
pair. The call it contains emits its own CALLSEQ, so the outer bracket
only produced a nested ADJCALLSTACKDOWN 0 / ADJCALLSTACKUP 0 around the
inner ADJCALLSTACKDOWN / ADJCALLSTACKUP which is illegal.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
Revert "[clang-repl] Initialized HIP environment for clang-repl (#217582) (#228976)
This reverts commit d220d5e8239db918bd68f404dae589385a8029e8.
The change landed without review from the clang-repl code owners and its
test fails in some build configurations (see the post-commit discussion
on #217582). Revert so the work can go through a proper review, as
agreed with the author.
This is not a pure revert: DeviceOffloadTest.cpp, added in #226975 and
extended in #226977 on top of the reverted commit, is ported back to the
CUDA-specific API (CreateCudaHost, CreateCudaDevice, createWithCUDA).
The CUDA fixes from those two commits are kept.
Supersedes #228088.
VE: Use splitAt in expandExtendStackPseudo
Replace the manual block-splitting in expandExtendStackPseudo with
MachineBasicBlock::splitAt. Reduces boilerplate, but there's some
block renumbering churn in the output.
Co-authored-by: Claude (Claude-Opus-4.8)
VE: Compute live-ins after splitting for EXTEND_STACK expansion
expandExtendStackPseudo splits its block but left the new blocks without
live-in lists, so their uses of registers live across the split are
rejected by -verify-machineinstrs.
Co-authored-by: Claude (Claude-Opus-4.8)
[lldb][docs] Fix the Windows Python build options (#225465)
The Windows section of `build.md` incorrectly marks `PYTHON_HOME` as
required, and the example value `C:\Python35` is below the 3.11 minimum
on Windows.
Furthermore, nothing said that `LLDB_EMBED_PYTHON_HOME` (which is `ON`
by default on Windows) suppresses `LLDB_ENABLE_PYTHON_LIMITED_API`, and
that setting both is a configure error.
VE: Fix ill-typed setjmp result in emitEHSjLjSetJmp (#221549)
Partially fixes machine verifier failures in existing tests;
they still fail due to other issues.
emitEHSjLjSetJmp materialized the 0/1 return values with LEAzii, which
defines an i64 register, into vregs with the i32 result register class.
This ill-typed MIR is rejected by -verify-machineinstrs.
Materialize the values in i64 and copy the low 32 bits (sub_i32) into
the i32 result. NFC on the emitted code.
Co-authored-by: Claude (Claude-Opus-4.8)
[BOLT][DWARF] Fix unit layout with forward DW_FORM_ref_udata references (#226076)
finalizeDIEs() sizes a DW_FORM_ref_udata reference from the referenced
DIE's current offset. For a forward reference, that is still the input
section offset set by allocDIE(), not the unit-relative offset that is
emitted later. If their ULEB128 sizes differ, the unit length and all
following DIE offsets are wrong.
GNU as emits such references in DWARF 5 units for assembly files, e.g.
libgcc's AArch64 outline-atomics helpers.
Emit these references as DW_FORM_ref4 instead, as BOLT already does for
type references in location expressions.
We hit this in Julia's CI, which BOLTs `libLLVM.so` and
`libjulia-internal.so` on aarch64-linux. Those link in libgcc's `lse.S`
units, and symbolizing a backtrace printed
```
[7 lines not shown]
[mlir][tosa] Honor input_unsigned when lowering tosa.cast to linalg (#228233)
The TosaToLinalg lowering ignored the `input_unsigned` attribute of
`tosa.cast` (added in #215838), so unsigned inputs were converted as
signed: an i8 0xFF cast to i32 gave -1 instead of 255.
When the attribute is set, use `arith.uitofp` for int -> float and
`arith.extui` for widening int -> int, since the TOSA 1.1 CAST
pseudocode zero-extends the input. Narrowing and int -> bool are
unchanged.
AI disclosure: I found this bug with a testing harness built with Claude
(Claude Code). The bug were reviewed by me . The fix and the test were
built and tested by me
Assisted-by: Claude
[CIR][CodeGen][NFC] Retire the CodeGenUtils.h catch-all header
Moves the last helpers out of `CodeGenUtils.h` into `ClassUtils.h`,
`ModuleUtils.h`, `TargetUtils.h` and `FunctionUtils.h` (`checkTargetFeatures`,
since it came from CodeGenFunction.cpp) and deletes the header. Only moves code
already on main, so it can be dropped on its own.
Assisted-by: Claude Code (Claude Fable 5.1).
AMDGPU: Stop setting kill flags before FinalizeISel
Work on removing all pre-RA flag management. Kill flags should eventually be
removed. Pre-regalloc passes no longer depend on them, LiveIntervals strips them
and VirtRegRewriter re-introduces them.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>