RegisterCoalescer: Keep remat def dead if it's a copy destination superregister (#230036)
This is a refinement of #226037, which was too strict.
When rematerializing into a physical register that is not exactly the copy's
destination, the def should only stay live if it is a sub-register of the copy
destination, i.e. part of the live value. Checking register unit coverage also
kept the def live when it is a super-register with the same units as the copy
destination, such as $rax for a copy into $eax on x86_64:
dead $rax = MOV64ri32 -11, implicit-def $eax
Only the $eax part is used, so the $rax def is dead. This matches what
LiveVariables produces for a full def with partial uses.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[mlir][tosa] Adding integer data layout operation support to PRO-FP (#229777)
This pull request was made to add integer data layout operation support
to PRO-FP. This is to prevent the need to cast between int and fp to do
these operations, which had the possibility of producing errors or
unwanted behaviour. Partially implements:
https://github.com/arm/tosa-specification/pull/91
Co-authored-by: Luke Hutton <luke.hutton at arm.com>
[OpenMP][DeviceRTL] Report the source location in __kmpc_error diagnostics (#224298)
Follow-up to #220702. Completes #204240.
The device runtime accepted the `ident_t` argument but ignored it, so
`error at(execution)` in a `target` region printed no source location.
This reports it, replicating the host runtime:
```
OMP: error_directive.f90:14:3: Encountered user-directed warning: warning message.
```
When the ident carries no location the result is `unknown:0:0`, same as
the host. flang populates the ident only with `-g`; clang always does.
Assisted-by: Copilot
SelectionDAG: Stop emitting kill flags in InstrEmitter
These is no point to maintaining these before register allocation.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
RuntimeLibcalls: Pass the default calling convention to libraries
Previously a setAvailableLibFuncs_* function computed DefaultCC locally when a
member calling convention referenced it. To do that, the emitter worked
backwards from a library to the SystemRuntimeLibrary records that reference it,
and had to diagnose the cases where that failed, which would be if there is no
referencing system library or several different ones.
The default calling convention belongs to the target, not to a library. The
dispatcher in setTargetRuntimeLibcallSets already computes it, so pass it to
each library function as a parameter. This removes the reverse lookup and both
diagnostics. DefaultCC references are now valid in a library shared by
SystemRuntimeLibrary records with different defaults, and in a library no
SystemRuntimeLibrary references.
This fixes errors when a LibcallLibrary is unused. This will enable defining
the vector math libraries in the future, as well as decoupling the target
specific handling in #229562.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
CodeGen: Run LiveIntervals before PHIElimination and drop LiveVariables from it (#228618)
Move LiveIntervals to run before PHIElimination in the optimized register allocation
pipeline, and make PHIElimination maintain LiveIntervals only.
This removes the last explicit use of LiveVariables. The actual analysis is no longer used.
There are implicit dependencies on the side effects of running the analysis due to
adjustments of dead flags, so further work is still needed to complete the removal.
This perturbs register allocation in a number of tests. The same codegen result
can be achieved by not preserving the analysis and recomputing fresh. Greedy is
just sensitive to the exact slot index and value numbering with identical MIR.
Measured across every affected test the emitted instruction count goes from
145124 to 145196, +0.050%, with changes in both directions. The largest regression
is AArch64/phi.ll, where the GlobalISel output gains about 30
instructions and no longer matches the SelectionDAG output; the largest
improvements are ARM/fpclamptosat.ll and PowerPC/common-chain.ll.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[mlir][acc] Report independent loops not parallelized due to unstructured control flow
An independent acc loop whose body has unstructured control flow (for
example a backward GOTO) cannot be represented as a structured loop, so it
is executed sequentially. This happened silently: no remark was emitted for
the loop, and the user only saw a serial kernel without a reason.
Emit a remark for such loops stating that they are not parallelized because
their control flow could not be represented as a structured loop. This lets
users identify loops that are asserted parallel but run sequentially.
Co-Authored-By: Claude
[AMDGPU] Price vector f32 to f16 fptrunc by its packing form (#229950)
The base cost scalarizes the conversion and charges 4, 10, 22 and 46
for 2, 4, 8 and 16 lanes. The backend rounds every lane with
v_cvt_f16_f32 and packs the halves in pairs, which takes N + N/2
instructions. A packed conversion rounds a pair per instruction and
true16 writes a lane into either half of a register, which gives
ceil(N/2) and N.
Co-authored-by: Michael Selehov <michael.selehov at amd.com>
Assisted-By: Claude Code Opus 5
---
<sub>Stack created with <a
href="https://github.com/github/gh-stack">GitHub Stacks CLI</a> • <a
href="https://gh.io/stacks-feedback">Give Feedback 💬</a></sub>
Co-authored-by: Michael Selehov <michael.selehov at amd.com>
[NFC][AMDGPU] Add tests for the cost of vector f32 to f16 fptrunc (#229949)
Assisted-By: Claude Code Opus 5
---
<sub>Stack created with <a
href="https://github.com/github/gh-stack">GitHub Stacks CLI</a> • <a
href="https://gh.io/stacks-feedback">Give Feedback 💬</a></sub>
[CIR][AMDGPU] Add support for AMDGCN s_sendmsg_rtn builtins (#223226)
Adds codegen for the following AMDGCN s_sendmsg_rtn builtins:
- __builtin_amdgcn_s_sendmsg_rtn
- __builtin_amdgcn_s_sendmsg_rtnl
These are lowered to the `llvm.amdgcn.s.sendmsg.rtn` intrinsic.
Assisted by: Claude Opus 5
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[mlir][acc] Report independent loops not parallelized due to unstructured control flow
An independent acc loop whose body has unstructured control flow (for
example a backward GOTO) cannot be represented as a structured loop, so it
is executed sequentially. This happened silently: no remark was emitted for
the loop, and the user only saw a serial kernel without a reason.
Emit a remark for such loops stating that they are not parallelized because
their control flow could not be represented as a structured loop. This lets
users identify loops that are asserted parallel but run sequentially.
Co-Authored-By: Claude
[Clang] Mark scoped_atomics with !noalias.addrspace(private)
The HIP specification marks atomics on thread private memory as UB.
Scoped atomics used within a HIP context are also considered UB,
unless explicitley specified via a command line argument.
These are now annotated with !noalias.addrspace(5) for amdgpus,
to avoid an expensive runtime check.
[Clang][Driver] Fix inverted diagnostic condition (#225640)
The diagnostic string for warn_drv_preprocessed_input_file_unused and
warn_drv_input_file_unused require %1/%2 to be true iff the causing
option is *un*available. Swap the option. Also use `getSpelling()` to
include the dash in the printed input.
warn_drv_input_file_unused was already part of #218802 which was
reverted.
preprocessed-input-file-unused.c test case generated by AI
[AMDGPU] Extend new hazard CFG walk for wmma instruction support (#229428)
Replaces `getWaitStatesSinceVALU` with `getMaxVALUWindowDeficit` and
aligns `checkWMMACoexecutionHazards` implementation with
`checkMAIHazards90A`. Now elides the visited set CFG walk and uses the
BFS CFG walk instead.
Fixes ROCM-32079
AI Assisted
[OpenMP] Fix debug locations in GPU reductions (#228622)
GPU reduction codegen temporarily switches the IRBuilder to AllocaIP to
create reduction storage. In Clang, AllocaIP points before the unlocated
allocapt marker, which clears the current debug location. Restoring the
insertion point does not restore the debug location, so the following
inlinable runtime calls lack !dbg.
Use InsertPointGuard for temporary alloca insertion regions, preserving
both the caller insertion point and debug location.
Remove redundant saveIP/restoreIP pairs around helper emitters that
already use InsertPointGuard internally to preserve the builder state.
Add an OpenMPIRBuilder unit test that models Clang's allocapt insertion
point and verifies the expected runtime call has a valid debug location.
Assisted by gpt-5.6.
[2 lines not shown]
[clang][Flang][SystemZ] Enable -mbackchain on Flang (#229863)
Enable -mbackchain on Flang now that it supports s390x. This is needed
in order to pass some libomp tests:
libomp :: tasking/omp_untied_taskloop.f90
libomp :: transform/fuse/do-looprange.f90
libomp :: transform/fuse/do.f90
libomp :: transform/tile/do.F90
libomp :: transform/tile/do_2d.f90
libomp :: transform/tile/do_2d_varsizes.f90
libomp :: transform/unroll/heuristic_do.f90
[clang][Sema] Use 64-bit triple in throw-address-space.cpp (#230051)
__ptr32 has no effect on a 32-bit system where addresses are already
32-bit. This means no error and no error means the test fails.
Fixes #224680.
[LLVM][CodeGen][SVE] Improve lowering for v1f32/f64 when NEON is not available. (#229724)
When NEON is not available it is better to scalarise single element
floating-point vectors than widening them to use Streaming-SVE.
Explicitly make v1f64 scalar_to_vector operations always legal, because
we can use scalar instructions and add a combine to avoid "nop" casts.
PPC: Fold 64-bit zero-extending word load feeding extsw subregister
A gprc LWZ/LWZX feeding EXTSW_32_64 is rewritten into a sign-extending
LWA/LWAX load. Extend the same fold to the 64-bit zero-extending word
loads LWZ8/LWZX8 when the EXTSW_32_64 reads their sub_32 subregister,
producing a single LWA/LWAX instead of a redundant lwz+extsw pair.
Co-authored-by: Claude (Claude Opus 4.8, claude-opus-4-8) <noreply at anthropic.com>
PPC: Add MIR examples for missed extsw+word-load fold on subregister input
A gprc LWZ/LWZX feeding EXTSW_32_64 folds into a sign-extending LWA/LWAX
load. The equivalent 64-bit zero-extending word loads (LWZ8/LWZX8) whose
sub_32 feeds EXTSW_32_64 are not folded, leaving a redundant lwz+extsw
(or lwzx+extsw) pair. Add MIR examples documenting the missed fold.
Co-authored-by: Claude (Claude Opus 4.8, claude-opus-4-8) <noreply at anthropic.com>