[AMDGPU] Fix missed WMMA C-operand co-exec hazard
The gfx1250 WMMA co-execution hazard check treats only A, B and the
SWMMAC index as registers the in-flight MMA still reads. C (src2 of a
non-SWMMAC WMMA) is missing, so a VALU scheduled into the MMA's shadow
can clobber C and the MMA consumes the new value.
This is latent while C is tied to vdst, since the existing D check then
covers it. It miscompiles where the tie does not hold: for
v_wmma_bf16f32_16x16x32_bf16, whose D is narrower than C, and for the
_threeaddr form of any WMMA.
[AMDGPU] Measure MFMA overwrite hazards at each instruction
Apply previously established processing to:
- VALU overwriting an MFMA result
- VALU overwriting a register an MFMA took as srcC
AI-assisted.
[AMDGPU] Measure MFMA read hazards at each producer
Introduce more sophisticated traversal to avoid the following traps:
- order-dependent traversal and discarding seen BBs despite shorter path
- mis-matching distance and window of different producers
Record the best distance per BB instead of a visited flag and sweep the
arrivals in nondecreasing distance (bucket queue). This pairs producers
with their actual distance to a consumer in one go.
Fixed scenarios:
- MFMA reading an MFMA result as srcA, srcB or srcC
- VALU, memory or export instruction reading an MFMA result
AI-assisted.
fixup! [AMDGPU] Measure MFMA read hazards at each producer
Reword the comments of the two tests where an exact srcC write must not
hide an earlier partial one: name the producers, give the nearer one's
window, and keep the farther one's window apart from the padding it
leaves. On gfx940+ the exact write now has a window of its own, since
two different MFMAs sharing an accumulator need wait states.
AI-assisted.
[AMDGPU][NFC] Prepare the MFMA hazard checks for per-producer windows (#225988)
NFC preparation for the MFMA hazard fix in #225989.
- Extract the MFMA read-window table out of checkMAIHazards90A into
getMFMAReadWaitStates, with the partial srcC overlap half in
getMFMAOverlappedSrcCWaitStates, so a caller can ask what a specific
producer requires.
- Drop the duplicate checkMAIVALUHazards call in PreEmitNoopsCommon,
which
ran the whole check once to test whether its result was positive and
then
again to use it.
No test changes, output is identical.
#218363 extracts the same partial srcC table under the same name.
Will happily rebase my version when it lands.
AI-assisted.
[ASTMatchers] Reduce per-matcher, per-node traversal kind overhead (#228960)
- Give traverseIgnored() a fast path and inline it
- Construct ParentMapContext eagerly and inline its accessor
- Only check if the node is ignored in matchWithFilter() for expressions
clang-tidy with LLVM's .clang-tidy config on five LLVM files (clang's
ASTContext.cpp, CGExpr.cpp, ParseDecl.cpp, SemaOverload.cpp, and
LoopVectorize.cpp), sum of two runs each:
instructions: 436.4e9 => 405.8e9, -7.0%
user time: 32.2 s => 30.2 s, -6.2%
No behavior change.
[libc][math] Fix powf exact cases, improve its performance, and add code-size optimized float-only version. (#226588)
Powf improvement:
- Fix exact case detection
- Improve performance
- Add code-size optimized float-only option
- Reorganize the implementation selection and tests.
Assisted-by: Gemini is used for code-size and performance analysis, and
for refactoring.
[AArch64][PeepholeOpt] Use correct register class for subreg uses from optimizeExtInstr. (#226595)
isCoalescableExtInstr from an instruction like SBFMXri can reuse the def
of the instruction. For some targets this is expected to be a
gpr32->gpr64 extend, but for AArch64 and PPC is gpr64->gpr64 and we are
UseSrcSubIdx and are expecting to generate a subreg copy. The regclass
used was for the full register though, not the subreg, generating an
invalid copy and INSERT_SUBREG operand (in this case).
Fixes #226352
[Support] Replace sprintf with snprintf in z/OS strsignal implementation (#224487)
Use bounded snprintf instead of sprintf when formatting the "Unknown
signal" message, and defensively re-terminate the static buffer in case
the null terminator is clobbered by a concurrent call.
Fixes #224343
[offload-arch] Fix clangOffloadArch link with BUILD_SHARED_LIBS=ON (#229075)
add_llvm_library ignores STATIC, so with BUILD_SHARED_LIBS=ON the
library was built as a shared object and failed to link because Verbose
is defined in the tool. Build it with llvm_add_library STATIC and move
Verbose and OffloadArchCategory into the library.
Fixes build failures after #226552 (e.g.
https://lab.llvm.org/buildbot/#/builders/226/builds/23094).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
Co-authored-by: Joseph Huber <huberjn at outlook.com>
[clang] Deduplicate builtin diagnostic descriptions (#202624)
Generate each distinct builtin diagnostic description once in the
existing `DiagnosticStableIDs.inc` output, emit pointer-free character
storage, and store its 20-bit offset and 12-bit length in
`StaticDiagInfoRec`. Diagtool also stores each diagnostic name once,
uses naturally aligned six-byte `DiagnosticRecord` entries, and replaces
the duplicate ID-sorted record table with a direct `uint16_t` index.
On arm64 macOS Release builds, stripped standalone clang decreases by
49,536 bytes and the LLVM 22 stripped all-tools multicall binary
decreases by 49,544 bytes and linked fixups decrease from 14,604 to 4.
Work towards #202616
AI tool disclosure: Co-authored with OpenAI Codex.
[LiveDebugVariables] Repair stale SlotIndexes
The analysis keeps its indexes from before the first register allocator
until DBG_VALUEs are emitted, by which point passes in between have
erased some of the instructions they point at. Resolve them at the
start of each allocator run and before emitting.
SlotIndexes can then reclaim the entries of erased instructions without
sparing the ones held here, which would have made generated code depend
on -g. Emitted locations are unchanged, except that intervals resolving
to one position now emit a single DBG_VALUE rather than identical
consecutive ones.
[clang] Add a one-entry cache to DiagStateMap::getFile() (#228956)
Once a TU has seen a `#pragma clang diagnostic` (libc++ has them in most
headers), every DiagnosticsEngine::isIgnored() / getDiagnosticSeverity()
call with a location looks up the location's FileID in a
std::map<FileID, File>. The map has an entry for every FileID a state
was ever looked up for, including macro expansions, so it gets large.
For blink's logical_box_fragment.cc, it has 60755 entries, and 92.4% of
the 3.28M lookups are for the same FileID as the previous lookup.
std::map has stable addresses, so remember the result of the last
lookup.
For 60 random Chromium TUs (linux x64, -O2) picked with probability
proportional to their compile time, sum over all TUs:
CPU time: 192.6 s => 191.4 s, -0.65%
instructions: 1903.0e9 => 1888.3e9, -0.77%
[3 lines not shown]
[SelectionDAG] Handle constants in SimplifyMultipleUseDemandedBits
Replace a non-zero constant with zero when none of its set bits are
demanded.
This allows users of `SimplifyMultipleUseDemandedBits` to eliminate
irrelevant constant bits while preserving the convention that a null
SDValue indicates no simplification.
[AMDGPU] Fold mul24 with an operand whose low 24 bits are zero (#224537)
mul24 only reads the low 24 bits of each operand. Try
`SimplifyDemandedBits`
before `SimplifyMultipleUseDemandedBits` so that a single use operand
such as
`(x & 0xff000000)` is simplified to zero, then fold the multiply to zero
when
either operand is a zero constant.
This folds cases such as:
```
mul24(x & 0xff000000, y) -> 0
mul24(x, 0) -> 0
```
[lldb][test] Correct register validity check in TestProcessSaveCoreMinidump.py (#229059)
This API will never return None, instead it will be an invalid SBValue.
I noticed this because I'm checking what can be run on an AArch64 host,
and this part of the test passed due to this mistake.