LLVM/project 9632d63 — llvm/lib/CodeGen RegisterCoalescer.cpp, llvm/test/CodeGen/X86 rematerialize-sub-super-reg-dead-flags.mir

RegisterCoalescer: Keep remat def dead if it's a copy destination superregister (#230036)

This is a refinement of #226037, which was too strict.

When rematerializing into a physical register that is not exactly the copy's
destination, the def should only stay live if it is a sub-register of the copy
destination, i.e. part of the live value. Checking register unit coverage also
kept the def live when it is a super-register with the same units as the copy
destination, such as $rax for a copy into $eax on x86_64:

  dead $rax = MOV64ri32 -11, implicit-def $eax

Only the $eax part is used, so the $rax def is dead. This matches what
LiveVariables produces for a full def with partial uses.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+5-8llvm/lib/CodeGen/RegisterCoalescer.cpp
+3-2llvm/test/CodeGen/X86/rematerialize-sub-super-reg-dead-flags.mir
+8-102 files

LLVM/project 13a8f9c — mlir/include/mlir/Dialect/Tosa/IR TosaComplianceData.h.inc, mlir/test/Dialect/Tosa profile_pro_int_unsupported.mlir tosa-validation-version-1p1-pro-fp-valid.mlir

[mlir][tosa] Adding integer data layout operation support to PRO-FP (#229777)

This pull request was made to add integer data layout operation support
to PRO-FP. This is to prevent the need to cast between int and fp to do
these operations, which had the possibility of producing errors or
unwanted behaviour. Partially implements:
https://github.com/arm/tosa-specification/pull/91

Co-authored-by: Luke Hutton <luke.hutton at arm.com>
DeltaFile
+62-0mlir/test/Dialect/Tosa/tosa-validation-version-1p1-pro-fp-valid.mlir
+28-7mlir/include/mlir/Dialect/Tosa/IR/TosaComplianceData.h.inc
+0-6mlir/test/Dialect/Tosa/profile_pro_int_unsupported.mlir
+90-133 files

LLVM/project fa6cf7a — offload/test/offloading error_directive.c, offload/test/offloading/fortran error_directive.f90

[OpenMP][DeviceRTL] Report the source location in __kmpc_error diagnostics (#224298)

Follow-up to #220702. Completes #204240.

The device runtime accepted the `ident_t` argument but ignored it, so
`error at(execution)` in a `target` region printed no source location.
This reports it, replicating the host runtime:
```
OMP: error_directive.f90:14:3: Encountered user-directed warning: warning message.
```

When the ident carries no location the result is `unknown:0:0`, same as
the host. flang populates the ident only with `-g`; clang always does.

Assisted-by: Copilot
DeltaFile
+53-3openmp/device/src/Misc.cpp
+9-3offload/test/offloading/fortran/error_directive.f90
+4-3offload/test/offloading/error_directive.c
+66-93 files

LLVM/project e794ffd — llvm/test/CodeGen/AArch64 bf16_fast_math.ll, llvm/test/CodeGen/AMDGPU legalize-amdgcn.raw.ptr.buffer.load.ll isel-amdgpu-cs-chain-preserve-cc.ll

SelectionDAG: Stop emitting kill flags in InstrEmitter

These is no point to maintaining these before register allocation.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+298-298llvm/test/CodeGen/AMDGPU/isel-amdgpu-cs-chain-cc.ll
+203-203llvm/test/CodeGen/AMDGPU/carryout-selection.ll
+174-174llvm/test/CodeGen/AMDGPU/llvm.amdgcn.make.buffer.rsrc.ll
+111-111llvm/test/CodeGen/AArch64/bf16_fast_math.ll
+88-88llvm/test/CodeGen/AMDGPU/legalize-amdgcn.raw.ptr.buffer.load.ll
+88-88llvm/test/CodeGen/AMDGPU/isel-amdgpu-cs-chain-preserve-cc.ll
+962-962161 files not shown
+2,717-2,778167 files

LLVM/project aa80d64 — llvm/test/TableGen RuntimeLibcallEmitter-library-ref.td RuntimeLibcallEmitter-library-name-merge.td, llvm/utils/TableGen/Basic RuntimeLibcallsEmitter.cpp

RuntimeLibcalls: Pass the default calling convention to libraries

Previously a setAvailableLibFuncs_* function computed DefaultCC locally when a
member calling convention referenced it. To do that, the emitter worked
backwards from a library to the SystemRuntimeLibrary records that reference it,
and had to diagnose the cases where that failed, which would be if there is no
referencing system library or several different ones.

The default calling convention belongs to the target, not to a library. The
dispatcher in setTargetRuntimeLibcallSets already computes it, so pass it to
each library function as a parameter. This removes the reverse lookup and both
diagnostics. DefaultCC references are now valid in a library shared by
SystemRuntimeLibrary records with different defaults, and in a library no
SystemRuntimeLibrary references.

This fixes errors when a LibcallLibrary is unused. This will enable defining
the vector math libraries in the future, as well as decoupling the target
specific handling in #229562.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+12-81llvm/utils/TableGen/Basic/RuntimeLibcallsEmitter.cpp
+36-0llvm/test/TableGen/RuntimeLibcallEmitter-library-unreferenced.td
+6-8llvm/test/TableGen/RuntimeLibcallEmitter-library-default-cc.td
+3-3llvm/test/TableGen/RuntimeLibcallEmitter-library-isolated.td
+2-2llvm/test/TableGen/RuntimeLibcallEmitter-library-ref.td
+2-2llvm/test/TableGen/RuntimeLibcallEmitter-library-name-merge.td
+61-962 files not shown
+65-1008 files

LLVM/project d424cf2 — mlir/lib/Dialect/Complex/IR ComplexOps.cpp, mlir/test/Dialect/Complex canonicalize.mlir

[mlir][complex] Require fastmath for the add/sub and exp/log folds (#221384)

`complex.add(complex.sub(a, b), b)`, `complex.add(b, complex.sub(a, b))`
and `complex.sub(complex.add(a, b), b)` fold to `a` unconditionally, and
so does `complex.exp(complex.log(a))`:

```mlir
func.func @add_sub(%a: complex<f32>, %b: complex<f32>) -> complex<f32> {
  %sub = complex.sub %a, %b : complex<f32>
  %add = complex.add %sub, %b : complex<f32>
  return %add : complex<f32>
}
// -canonicalize today
func.func @add_sub(%a: complex<f32>, %b: complex<f32>) -> complex<f32> {
  return %a : complex<f32>
}
```

Neither identity holds in floating point. The intermediate result is

    [19 lines not shown]
DeltaFile
+111-26mlir/test/Dialect/Complex/canonicalize.mlir
+35-12mlir/lib/Dialect/Complex/IR/ComplexOps.cpp
+146-382 files

LLVM/project 2ba993c — llvm/lib/Frontend/OpenMP OMPDescriptors.cpp

Address review comments
DeltaFile
+7-2llvm/lib/Frontend/OpenMP/OMPDescriptors.cpp
+7-21 files

LLVM/project b820d26 — offload/plugins-nextgen/common/include PluginInterface.h, offload/plugins-nextgen/common/src PluginInterface.cpp

[Offload] Remove unused entries from Plugin interface (#230078)

These entries in GenericPluginTy are no longer called anywhere so they
can be removed.
DeltaFile
+0-30offload/plugins-nextgen/common/src/PluginInterface.cpp
+0-21offload/plugins-nextgen/common/include/PluginInterface.h
+0-512 files

LLVM/project 05e41e5 — clang/lib/CIR/Dialect/Transforms CallConvLoweringPass.cpp, clang/test/CIR/CodeGen call-conv-lowering-x86_64-vec3.c

[CIR] Lower non-power-of-two vectors in CallConvLowering
DeltaFile
+86-0clang/test/CIR/CodeGen/call-conv-lowering-x86_64-vec3.c
+11-10clang/test/CIR/Transforms/abi-lowering/x86_64-aggregate-nyi.cir
+2-6clang/lib/CIR/Dialect/Transforms/CallConvLoweringPass.cpp
+99-163 files

LLVM/project fb2a526 — flang/test/HLFIR simplify-hlfir-intrinsics-pack.fir

[flang][NFC] Update PACK simplification test for RESHAPE overflow flags (#230076)

Fix test failure from merge conflict between
https://github.com/llvm/llvm-project/pull/220860 and
https://github.com/llvm/llvm-project/pull/229733.
DeltaFile
+9-9flang/test/HLFIR/simplify-hlfir-intrinsics-pack.fir
+9-91 files

LLVM/project 8e5a934 — llvm/test/CodeGen/AArch64 div-i256.ll phi.ll, llvm/test/CodeGen/RISCV abdu.ll idiv_large.ll

CodeGen: Run LiveIntervals before PHIElimination and drop LiveVariables from it (#228618)

Move LiveIntervals to run before PHIElimination in the optimized register allocation 
pipeline, and make PHIElimination maintain LiveIntervals only.

This removes the last explicit use of LiveVariables. The actual analysis is no longer used. 
There are implicit dependencies on the side effects of running the analysis due to 
adjustments of dead flags, so further work is still needed to complete the removal.

This perturbs register allocation in a number of tests. The same codegen result
can be achieved by not preserving the analysis and recomputing fresh. Greedy is
just sensitive to the exact slot index and value numbering with identical MIR.

Measured across every affected test the emitted instruction count goes from
145124 to 145196, +0.050%, with changes in both directions. The largest regression 
is AArch64/phi.ll, where the GlobalISel output gains about 30
instructions and no longer matches the SelectionDAG output; the largest
improvements are ARM/fpclamptosat.ll and PowerPC/common-chain.ll.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+1,648-1,642llvm/test/CodeGen/RISCV/GlobalISel/wide-scalar-shift-by-byte-multiple-legalization.ll
+1,026-1,017llvm/test/CodeGen/X86/i128-udiv.ll
+396-396llvm/test/CodeGen/RISCV/idiv_large.ll
+452-210llvm/test/CodeGen/AArch64/phi.ll
+290-290llvm/test/CodeGen/RISCV/abdu.ll
+242-242llvm/test/CodeGen/AArch64/div-i256.ll
+4,054-3,797126 files not shown
+8,122-8,041132 files

LLVM/project 878a0ce — mlir/include/mlir/Dialect/OpenACC OpenACC.h, mlir/lib/Dialect/OpenACC/Transforms ACCComputeLowering.cpp ACCEmitRemarksLoop.cpp

[mlir][acc] Report independent loops not parallelized due to unstructured control flow

An independent acc loop whose body has unstructured control flow (for
example a backward GOTO) cannot be represented as a structured loop, so it
is executed sequentially. This happened silently: no remark was emitted for
the loop, and the user only saw a serial kernel without a reason.

Emit a remark for such loops stating that they are not parallelized because
their control flow could not be represented as a structured loop. This lets
users identify loops that are asserted parallel but run sequentially.

Co-Authored-By: Claude
DeltaFile
+100-0mlir/test/Dialect/OpenACC/acc-compute-lowering-unstructured.mlir
+22-0mlir/test/Dialect/OpenACC/acc-emit-remarks-loop.mlir
+16-3mlir/lib/Dialect/OpenACC/Transforms/ACCEmitRemarksLoop.cpp
+13-0mlir/lib/Dialect/OpenACC/Transforms/ACCComputeLowering.cpp
+7-0mlir/include/mlir/Dialect/OpenACC/OpenACC.h
+158-35 files

LLVM/project 55ea1f4 — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Analysis/CostModel/AMDGPU fptrunc.ll

[AMDGPU] Price vector f32 to f16 fptrunc by its packing form (#229950)

The base cost scalarizes the conversion and charges 4, 10, 22 and 46
for 2, 4, 8 and 16 lanes. The backend rounds every lane with
v_cvt_f16_f32 and packs the halves in pairs, which takes N + N/2
instructions. A packed conversion rounds a pair per instruction and
true16 writes a lane into either half of a register, which gives
ceil(N/2) and N.

Co-authored-by: Michael Selehov <michael.selehov at amd.com>
Assisted-By: Claude Code Opus 5

---

<sub>Stack created with <a
href="https://github.com/github/gh-stack">GitHub Stacks CLI</a> • <a
href="https://gh.io/stacks-feedback">Give Feedback 💬</a></sub>

Co-authored-by: Michael Selehov <michael.selehov at amd.com>
DeltaFile
+42-49llvm/test/Transforms/SLPVectorizer/AMDGPU/fptrunc-f16-stores.ll
+25-25llvm/test/Analysis/CostModel/AMDGPU/fptrunc.ll
+13-0llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+80-743 files

LLVM/project 2296021 — llvm/test/Analysis/CostModel/AMDGPU fptrunc.ll, llvm/test/Transforms/SLPVectorizer/AMDGPU fptrunc-f16-stores.ll

[NFC][AMDGPU] Add tests for the cost of vector f32 to f16 fptrunc (#229949)

Assisted-By: Claude Code Opus 5

---

<sub>Stack created with <a
href="https://github.com/github/gh-stack">GitHub Stacks CLI</a> • <a
href="https://gh.io/stacks-feedback">Give Feedback 💬</a></sub>
DeltaFile
+115-0llvm/test/Transforms/SLPVectorizer/AMDGPU/fptrunc-f16-stores.ll
+114-0llvm/test/Analysis/CostModel/AMDGPU/fptrunc.ll
+229-02 files

LLVM/project c8fffb7 — llvm/docs AMDGPUUsage.rst AMDGPUDMAOperations.md

eliminate the term "request" ... these are just fancy loads, after all
DeltaFile
+21-23llvm/docs/AMDGPUDMAOperations.md
+1-4llvm/docs/AMDGPUUsage.rst
+22-272 files

LLVM/project 6a3f98e — clang/lib/CIR/CodeGen CIRGenBuiltinAMDGPU.cpp, clang/test/CIR/CodeGenHIP builtins-amdgcn-gfx11.hip

[CIR][AMDGPU] Add support for AMDGCN s_sendmsg_rtn builtins (#223226)

Adds codegen for the following AMDGCN s_sendmsg_rtn builtins:

- __builtin_amdgcn_s_sendmsg_rtn
- __builtin_amdgcn_s_sendmsg_rtnl

These are lowered to the `llvm.amdgcn.s.sendmsg.rtn` intrinsic.

Assisted by: Claude Opus 5

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+64-0clang/test/CIR/CodeGenHIP/builtins-amdgcn-gfx11.hip
+4-6clang/lib/CIR/CodeGen/CIRGenBuiltinAMDGPU.cpp
+68-62 files

LLVM/project 8079bfb — mlir/include/mlir/Dialect/OpenACC OpenACC.h, mlir/lib/Dialect/OpenACC/Transforms ACCComputeLowering.cpp ACCEmitRemarksLoop.cpp

[mlir][acc] Report independent loops not parallelized due to unstructured control flow

An independent acc loop whose body has unstructured control flow (for
example a backward GOTO) cannot be represented as a structured loop, so it
is executed sequentially. This happened silently: no remark was emitted for
the loop, and the user only saw a serial kernel without a reason.

Emit a remark for such loops stating that they are not parallelized because
their control flow could not be represented as a structured loop. This lets
users identify loops that are asserted parallel but run sequentially.

Co-Authored-By: Claude
DeltaFile
+100-0mlir/test/Dialect/OpenACC/acc-compute-lowering-unstructured.mlir
+22-0mlir/test/Dialect/OpenACC/acc-emit-remarks-loop.mlir
+16-3mlir/lib/Dialect/OpenACC/Transforms/ACCEmitRemarksLoop.cpp
+13-0mlir/lib/Dialect/OpenACC/Transforms/ACCComputeLowering.cpp
+7-0mlir/include/mlir/Dialect/OpenACC/OpenACC.h
+158-35 files

LLVM/project c2daf76 — clang/include/clang/AST Expr.h, clang/include/clang/Options Options.td

[Clang] Mark scoped_atomics with !noalias.addrspace(private)

The HIP specification marks atomics on thread private memory as UB.
Scoped atomics used within a HIP context are also considered UB,
unless explicitley specified via a command line argument.
These are now annotated with !noalias.addrspace(5) for amdgpus,
to avoid an expensive runtime check.
DeltaFile
+76-75clang/test/CodeGen/scoped-atomic-ops.c
+22-20clang/test/CodeGenCUDA/atomic-options.hip
+28-0clang/test/CodeGenCUDA/private-atomics-undefined.hip
+5-4clang/lib/CodeGen/Targets/AMDGPU.cpp
+3-4clang/include/clang/AST/Expr.h
+6-0clang/include/clang/Options/Options.td
+140-1032 files not shown
+142-1048 files

LLVM/project 531bd6f — clang/lib/Driver Driver.cpp, clang/test/Driver aix-ld.c preprocessed-input-file-unused.c

[Clang][Driver] Fix inverted diagnostic condition (#225640)

The diagnostic string for warn_drv_preprocessed_input_file_unused and
warn_drv_input_file_unused require %1/%2 to be true iff the causing
option is *un*available. Swap the option. Also use `getSpelling()` to
include the dash in the printed input.

warn_drv_input_file_unused was already part of #218802 which was
reverted.

preprocessed-input-file-unused.c test case generated by AI
DeltaFile
+15-0clang/test/Driver/preprocessed-input-file-unused.c
+4-4clang/lib/Driver/Driver.cpp
+1-1clang/test/Driver/aix-ld.c
+1-0clang/test/Driver/Inputs/preprocessed-input-file-unused.i
+21-54 files

LLVM/project c9e9f62 — llvm/lib/Support UnicodeNameToCodepointGenerated.cpp, llvm/test/CodeGen/AMDGPU atomic_optimizations_local_pointer.ll fcmp.f16.ll

Rebase, fix test checks

Created using spr 1.3.7
DeltaFile
+23,347-23,371llvm/lib/Support/UnicodeNameToCodepointGenerated.cpp
+2,827-5,340llvm/test/CodeGen/AMDGPU/fptrunc.f16.ll
+3,103-3,156llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-private.mir
+1,450-4,799llvm/test/CodeGen/AMDGPU/fcmp.f16.ll
+733-5,435llvm/test/CodeGen/AMDGPU/atomic_optimizations_local_pointer.ll
+3,245-2,647llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-local.mir
+34,705-44,7483,925 files not shown
+222,785-115,8793,931 files

LLVM/project 2812ee7 — llvm/lib/Target/AMDGPU GCNHazardRecognizer.h GCNHazardRecognizer.cpp, llvm/test/CodeGen/AMDGPU wmma-coexecution-valu-hazards.mir

[AMDGPU] Extend new hazard CFG walk for wmma instruction support (#229428)

Replaces `getWaitStatesSinceVALU` with `getMaxVALUWindowDeficit` and
aligns `checkWMMACoexecutionHazards` implementation with
`checkMAIHazards90A`. Now elides the visited set CFG walk and uses the
BFS CFG walk instead.

Fixes ROCM-32079

AI Assisted
DeltaFile
+100-14llvm/test/CodeGen/AMDGPU/wmma-coexecution-valu-hazards.mir
+45-58llvm/lib/Target/AMDGPU/GCNHazardRecognizer.cpp
+1-1llvm/lib/Target/AMDGPU/GCNHazardRecognizer.h
+146-733 files

LLVM/project ddd4d73 — llvm/lib/Frontend/OpenMP OMPIRBuilder.cpp, llvm/unittests/Frontend OpenMPIRBuilderTest.cpp

[OpenMP] Fix debug locations in GPU reductions (#228622)

GPU reduction codegen temporarily switches the IRBuilder to AllocaIP to
create reduction storage. In Clang, AllocaIP points before the unlocated
allocapt marker, which clears the current debug location. Restoring the
insertion point does not restore the debug location, so the following
inlinable runtime calls lack !dbg.

Use InsertPointGuard for temporary alloca insertion regions, preserving
both the caller insertion point and debug location.

Remove redundant saveIP/restoreIP pairs around helper emitters that
already use InsertPointGuard internally to preserve the builder state.

Add an OpenMPIRBuilder unit test that models Clang's allocapt insertion
point and verifies the expected runtime call has a valid debug location.

Assisted by gpt-5.6.


    [2 lines not shown]
DeltaFile
+74-0llvm/unittests/Frontend/OpenMPIRBuilderTest.cpp
+31-31llvm/lib/Frontend/OpenMP/OMPIRBuilder.cpp
+105-312 files

LLVM/project 8c6efc0 — clang/include/clang/Options Options.td

[clang][Flang][SystemZ] Enable -mbackchain on Flang (#229863)

Enable -mbackchain on Flang now that it supports s390x. This is needed
in order to pass some libomp tests:

  libomp :: tasking/omp_untied_taskloop.f90
  libomp :: transform/fuse/do-looprange.f90
  libomp :: transform/fuse/do.f90
  libomp :: transform/tile/do.F90
  libomp :: transform/tile/do_2d.f90
  libomp :: transform/tile/do_2d_varsizes.f90
  libomp :: transform/unroll/heuristic_do.f90
DeltaFile
+1-1clang/include/clang/Options/Options.td
+1-11 files

LLVM/project c5e7515 — clang/test/Sema throw-address-space.cpp

[clang][Sema] Use 64-bit triple in throw-address-space.cpp (#230051)

__ptr32 has no effect on a 32-bit system where addresses are already
32-bit. This means no error and no error means the test fails.

Fixes #224680.
DeltaFile
+1-1clang/test/Sema/throw-address-space.cpp
+1-11 files

LLVM/project a9a144e — llvm/test/CodeGen/AArch64 bf16-imm.ll

[AArch64] Remove unused CHECK lines. NFC (#230058)
DeltaFile
+0-12llvm/test/CodeGen/AArch64/bf16-imm.ll
+0-121 files

LLVM/project 9aecc2e — lldb/source/Core CMakeLists.txt ModuleList.cpp, lldb/source/Plugins/TypeSystem/Clang CMakeLists.txt TypeSystemClang.cpp

[lldb] Set the default clang module cache path in TypeSystemClang
DeltaFile
+7-0lldb/source/Plugins/TypeSystem/Clang/TypeSystemClang.cpp
+0-6lldb/source/Core/ModuleList.cpp
+0-3lldb/source/Core/CMakeLists.txt
+1-0lldb/source/Plugins/TypeSystem/Clang/CMakeLists.txt
+8-94 files

LLVM/project 1301afd — llvm/test/CodeGen/AArch64 sve-streaming-mode-fixed-length-concat.ll vector-ldst-offset.ll

[LLVM][CodeGen][SVE] Improve lowering for v1f32/f64 when NEON is not available. (#229724)

When NEON is not available it is better to scalarise single element
floating-point vectors than widening them to use Streaming-SVE.

Explicitly make v1f64 scalar_to_vector operations always legal, because
we can use scalar instructions and add a combine to avoid "nop" casts.
DeltaFile
+21-30llvm/test/CodeGen/AArch64/vector-ldst-align-float.ll
+0-24llvm/test/CodeGen/AArch64/sve-streaming-mode-fixed-length-fp-rounding.ll
+0-18llvm/test/CodeGen/AArch64/sve-streaming-mode-fixed-length-fp-minmax.ll
+4-10llvm/test/CodeGen/AArch64/vector-ldst-offset.ll
+4-10llvm/test/CodeGen/AArch64/sve-streaming-mode-fixed-length-fp-to-int.ll
+5-8llvm/test/CodeGen/AArch64/sve-streaming-mode-fixed-length-concat.ll
+34-10012 files not shown
+55-13518 files

LLVM/project 24e9116 — llvm/lib/Target/PowerPC PPCMIPeephole.cpp, llvm/test/CodeGen/PowerPC peephole-elim-extsw-subreg-input.mir

PPC: Fold 64-bit zero-extending word load feeding extsw subregister

A gprc LWZ/LWZX feeding EXTSW_32_64 is rewritten into a sign-extending
LWA/LWAX load. Extend the same fold to the 64-bit zero-extending word
loads LWZ8/LWZX8 when the EXTSW_32_64 reads their sub_32 subregister,
producing a single LWA/LWAX instead of a redundant lwz+extsw pair.

Co-authored-by: Claude (Claude Opus 4.8, claude-opus-4-8) <noreply at anthropic.com>
DeltaFile
+12-12llvm/test/CodeGen/PowerPC/peephole-elim-extsw-subreg-input.mir
+13-3llvm/lib/Target/PowerPC/PPCMIPeephole.cpp
+25-152 files

LLVM/project a5f872a — llvm/test/CodeGen/PowerPC peephole-elim-extsw-subreg-input.mir

PPC: Add MIR examples for missed extsw+word-load fold on subregister input

A gprc LWZ/LWZX feeding EXTSW_32_64 folds into a sign-extending LWA/LWAX
load. The equivalent 64-bit zero-extending word loads (LWZ8/LWZX8) whose
sub_32 feeds EXTSW_32_64 are not folded, leaving a redundant lwz+extsw
(or lwzx+extsw) pair. Add MIR examples documenting the missed fold.

Co-authored-by: Claude (Claude Opus 4.8, claude-opus-4-8) <noreply at anthropic.com>
DeltaFile
+152-6llvm/test/CodeGen/PowerPC/peephole-elim-extsw-subreg-input.mir
+152-61 files

LLVM/project 40be1e7 — llvm/lib/Target/PowerPC PPCMIPeephole.cpp, llvm/test/CodeGen/PowerPC peephole-elim-extsw-subreg-input.mir

review comments
DeltaFile
+3-4llvm/lib/Target/PowerPC/PPCMIPeephole.cpp
+1-1llvm/test/CodeGen/PowerPC/peephole-elim-extsw-subreg-input.mir
+4-52 files