[AMDGPU] Skip s_delay_alu for WMMA C-reuse chains (#214101)
Consecutive wmma/swmmac ops accumulating into the same matrix C register
reuse the accumulator in place, so the tied srcC read is omitted and no
delay is needed. AMDGPUInsertDelayAlu did not model this and emitted an
s_delay_alu that stalls the reuse chain.
Detect a C-reuse edge (tied srcC exactly matches the previous wmma/
swmmac dest, with no intervening instruction) and skip the delay for
that operand. This applies on all wmma-capable targets (gfx11+).
[clangd][ParseHLSL] Fix register attribute source range for hover inside arguments (#212881)
Hovering on the slot identifier inside `register(t1)` (e.g. on `t1`)
previously produced no tooltip, only hovering on the `register` keyword
itself worked.
`HLSLResourceBindingAttr`'s `SourceRange` was zero-width: both the start
and end pointed to the start of the `register` keyword.
`ParseHLSLAnnotations` called `Attrs.addNew` with a single
`SourceLocation` instead of a full `SourceRange`. Since clangd's
`SelectionTree` only matches when the cursor falls inside an attribute's
range, a zero-width range never matched positions inside the argument.
Capture the closing `)` location before it's consumed in the
`AT_HLSLResourceBinding` case, and pass a full `SourceRange` (from the
attribute start to the closing paren) to `addNew`.
Fixes #212749
[mlir][bufferization] Bufferize all edges to a repeated successor (#214368)
bufferizeBlockSignature only rewrote the first successor index that
matched the target block. Branch ops such as cf.cond_br can list the
same destination more than once, but the later edges were left as
tensors and broke multi-block bufferization.
Now we simply iterate the block's BlockOperands so each successor edge
is handled once.
[Clang][SPIRV] Add __spirv_event_t builtin type (#207077)
Add a new builtin type __spirv_event_t for SPIR-V targets. It represents
SPIR-V's OpTypeEvent and lowers to the target("spirv.Event") extension
type.
We would like to expose SPIR-V instructions to users via builtins (not
yet
implemented). The builtins return an event type.
Assisted by Claude Opus 4.8 for writing tests.
[llvm-objcopy] Address reviewer feedback on AMDGPU test cleanups
- Remove unused -DMACHINE yaml2obj template variable in cross-arch-headers.test,
hardcode Machine: EM_NONE directly in the YAML instead
- Remove unused Flags: [[FLAGS=<none>]] template variable in cross-arch-headers.test
- Add comment in binary-output-target.test explaining that Arch: unknown is
intentional when converting from binary (e_flags=0, no EF_AMDGPU_MACH set)
[llvm-objcopy] Add AMDGPU case to binary-output-target.test
Add test coverage for converting binary input to elf64-amdgpu format,
verifying the output has the correct format string, arch (amdgpu),
and machine type (EM_AMDGPU 0xE0). Follows the same pattern as all
other architectures in this test file.
[llvm-objcopy] Add elf64-amdgpu to supported formats in command guide
Update the llvm-objcopy command guide's "Supported formats" section to
include elf64-amdgpu, added in the preceding commit.
[llvm-objcopy] Address review feedback for AMDGPU test in cross-arch-headers
Per reviewer feedback, use the existing non-AMDGPU input (%t.o, EM_NONE)
to test conversion to elf64-amdgpu. This properly demonstrates that
--output-format changes the machine type, consistent with all other cases
in this test file.
The output reports Arch: unknown because converting from a non-AMDGPU ELF
produces e_flags=0 (no EF_AMDGPU_MACH set); added a comment explaining
this. Flag control is a separate concern for a follow-on PR.
[llvm-objcopy] Fix AMDGPU arch string in test: amdgpu not amdgcn
llvm-readobj reports 'Arch: amdgpu' for EM_AMDGPU ELF files
(the generic AMDGPU ELF format used by elf64-amdgpu). The test
was incorrectly expecting 'amdgcn', which is the AMDGCN-specific
arch string used by ROCm HSA code objects.
[llvm-objcopy] Fix AMDGPU arch checks in tests
ELFObjectFile.h getArch() for EM_AMDGPU returns Triple::UnknownArch
when e_flags & EF_AMDGPU_MACH is 0 (no GPU target specified). Only
when a MACH flag in the AMDGCN range is present does it return
Triple::amdgpu.
- cross-arch-headers.test: restore EF_AMDGPU_MACH_AMDGCN_GFX900 flag
on the input ELF so that after format conversion the output correctly
reports Arch: amdgpu.
- binary-output-target.test: expect Arch: unknown since converting
from raw binary input (-I binary) produces an ELF with e_flags=0
(no MACH flags), giving UnknownArch. This is correct behavior.
[lldb] Use the standard GDB remote thread for a non-Wasm process (#214380)
CanDebug returns true whenever the plugin is requested by name, and the
architecture is not known until the stub reports it after connecting.
This means that a non-Wasm process can end up with a ThreadWasm whose
register context and unwinder have nothing to operate on.
Create the plain ThreadGDBRemote once the architecture is known, and add
a helper so that the check covers wasm64 as well as wasm32.
rdar://182229301
Reland "[MergeFunctions] Preserve instruction-level profile metadata during merging" (#208009) (#210138)
This relands #208009, which was reverted in #209987 after an ASan
heap-use-after-free surfaced in MergeFunctionsTest.TrueOutputModuleTest.
The failure was caused by MergeFunctionsTest destroying
FunctionAnalysisManager before ModuleAnalysisManager, while MAM holds a
cached proxy result that calls FAM.clear() on destruction. This PR adds
a commit reordering those members, so they are destroyed in the correct
order, fixing the use-after-free.
Original PR: #208009
Revert PR: #209987
[AMDGPU][GlobalISel] Pre-commit tests for readanylane merge regbank combine (NFC) (#214355)
Add regbank-combiner tests covering a copy to vgpr whose source is a
merge or
build_vector of `G_AMDGPU_READANYLANE` results mixed with uniform
values.
These currently keep the round trip through sgprs. The tests also cover
the two
cases where the transform must not fire:
- the sgpr merge has another user, so it has to be kept;
- all merge sources are uniform, so moving the copy to the sources would
not
remove any readanylane.
Pre-commit only, no functional change. The combine that removes the
round trip
is in the stacked PR.
[6 lines not shown]
[SDAG] Handle vector expansion when expanding FABS (#214341)
Before this, vector FABS that were scheduled for expansion would hit the
getSignAsIntValue() case and either assert there or fail later for lock
of instruction selection for a bitcast that's only looking at the scalar
size.
Now, we unroll to scalars as needed.
Test pre-committed in #214288
[llvm-ar][GOFF] Implement symbol attributes for GOFF archives
z/OS archive symbol table entries contain a 32-bit attribute word
alongside each member offset. The low three bits encode:
bit 2 (0x4): 64-bit addressing (AMODE 64)
bit 1 (0x2): XPLink calling convention
bit 0 (0x1): Writable Static Area (WSA)
Previously in e2c8fa0, llvm-ar wrote zero for
these attributes. This patch reads them from GOFF ESD records and stores them
in a SymbolAttrs vector parallel to the existing Symbols vector in
MemberData to emit the correct word per symbol.
These attributes are tested using `llvm-nm --print-armap` implemented in
#212830 within the LIT test.
[llvm-nm][GOFF] Display archive attributes in GOFF archives through --print-armap
GOFF archive symbol table entries contain an attribute word in addition
to the archive member offset. The low three bits describe whether the symbol is
64-bit, uses XPLink, or belongs to the WSA namespace (which was briefly mentioned
in e2c8fa09872cfacba7f73599dcf8557971ebe865).
This patch extends `llvm-nm --print-armap` to print the attribute value (in hex) and
its decoded description beside a symbol and its corresponding member when processing
a GOFF archive. This will functionality will be used to help validate full support for writing
GOFF archives in a subsequent llvm-ar patch.
The output for non-z/OS archives is unchanged.
[flang] Fix incorrect offset for zero-size COMMON block members. (#214174)
Fixes #214171
`ComputeOffsetsHelper::DoSymbol()` in
`flang/lib/Semantics/compute-offsets.cpp` returned early without calling
`symbol.set_offset()` when a symbol had zero size (e.g. CHARACTER*0). As
a result, every zero-size symbol in a COMMON block retained its default
offset of 0 — the block base address — instead of its correct sequential
position.
This incorrect offset caused two observable bugs:
1. **Wrong storage address**: LOC() and lowering always returned the
block base address for zero-size members instead of their actual
sequential position.
2. **False "cannot backward-extend" error**: When a zero-size COMMON
block member appeared in an EQUIVALENCE association, the backward-extend
[18 lines not shown]
[lldb] Split DynamicLoader binary loading into locate and load (NFCI) (#214372)
LoadBinaryWithUUIDAndAddress both searched for a binary and registered
it with the Target. Split it into LocateBinaries, which only searches,
and LoadBinaryInTarget, which mutates the Target, with
LocateAndLoadBinary keeping the single binary case a one-liner.
The eight binary parameters and the results of the search are bundled in
a new BinarySpec struct, and both entry points return an llvm::Expected.
The motivation is a follow-up that runs LocateBinaries in parallel on
the thread pool. NFC, except for some small improvements to the error
handling because we don't write to the async output stream directly (and
fixed the newline).
[mlir][llvm] Add more constrained FP operations (#213745)
This change adds special constrained forms of transcendental operations
for the remaining cases that lower to contrained fp intrinsic calls. It
also adds fast-math flag support to the constrained operations, which is
needed to handle combinations of Clang command-line options such as
"-ffinite-math-only -ftrapping-math".
Assisted-by: Cursor / various models
[CIR][AArch64] Update builtin handlers to use emitNeonCallToOp (#214075)
This is another change to prepare AArch64 builtin handling for the
transition to constrained FP handling. It replaces a number of places
where we were creating CIR operations directly with calls to
emitNeonCallToOp so that we will be able to centralize the constrained
FP handling.
This also updates the vrndns_f32 to eliminate a redundant load of the
operand, which is the only part of this change with a visible difference
in the output.
Assisted-by: Cursor / Grok 4.5