[SLP] Reland: More accurately cost RISCV scalar splats (#224766)
Originally #213104, reverted due to assertion failure in cases of
reordered gather nodes.
Backends may have a fast path for splatting scalar operands (i.e. rather
than generating the splat vector, the vector instruction may be able to
take a scalar operand), for example RISCV `vfoo.vx` instructions. Pass a
hint to the TTI when costing the insert/shuffle sequence in such cases.
Fixes #212413.
Assisted By: Codex
[DebugInfo] Fix overflow in DWARFDebugLine::SectionParser (#224770)
Adding a DWARF64 unit length can overflow and wrap back into the
section, causing an infinite loop. Saturate the addition so an
overflowing offset fails the bounds check.
rdar://186810393
[NVPTX] Preserve volatile on atomic local loads and stores (#224719)
Local loads and stores discard atomic ordering during instruction
selection, but volatile accesses with acquire, release, or seq_cst
ordering also lose the volatile qualifier. Preserve volatile for all
valid load/store orderings when local volatile instructions are
supported. Follow up to #217764.
[WebKit Checkers] Trace through temporaries in tryToFindPtrOrigin (#224877)
RefPtr checking skips temporaries, reporting a path through any
temporary as unsafe. This is mostly correct, but not always. For
example, the following is a false positive:
// makeKey() returns a temporary
RefCountable* p = condition(makeKey()) ? guardian.ptr() : nullptr;
In the upcoming Borrow checker, it's even more important to trace
through temporaries because not tracing an expression can drop a
`lifetimebound` link, resulting in false **negatives**.
This patch adds tracing through temporaries. The logic is:
* In function call arguments, temporaries are lifetime safe because the
full expression does not end until the call returns
* In ranged for loops, temporaries are lifetime safe because lifetime
[4 lines not shown]
[clang][FreeBSD] Enable KASAN and KMSAN for riscv64 (#202288)
enables KASAN and KMSAN for the riscv64 freebsd target in the Clang
driver.
The corresponding FreeBSD kernel runtime support for riscv64 KASAN is
currently under active review.
Depends on: https://reviews.freebsd.org/D57381
Co-authored-by: aokblast <aokblast at FreeBSD.org>
Reland "[AMDGPU] PromoteAlloca: flatten homogeneous structs to vectors" (#221058)
This relands #217055
The original commit revealed a latent issue in eliminateFrameIndex in
SIRegisterInfo where SCC can be clobbered before reading it on
gfx900/gfx90a. This change itself has no known issues.
[AMDGPU] Update based on review feedback
Replace the lambda with the check inlined at both sites, and report a fatal
error when neither a free SGPR nor FrameReg is available, rather than
silently falling back to the spilling scavenge.
[AMDGPU] Use the scavenger to test whether SCC is live after MI
The register scavenger is stepped backwards to the liveness state
immediately after MI, so RS->isRegUsed(SCC) already answers "is SCC live
after MI" directly. Replace the hand-rolled test with that query.
[AMDGPU] Don't spill an SGPR while SCC is live in frame index lowering
When SCC is live into a scalar frame index user, the scaling path avoids
SALU ops that write SCC by computing the address in a VGPR and reading it
back with V_READFIRSTLANE_B32. If the destination of that readfirstlane is
scavenged with spilling allowed, an AMDGPU SGPR spill writes inactive
lanes, so it flips EXEC with S_NOT_B64 and clobbers SCC. Instead, scavenge
that register with AllowSpill=false.
[NFC][AMDGPU] Add tests for SCC live into a frame index user
Pre-commit tests for the case where SCC is live into a frame index user and
no SGPR is free to hold the V_READFIRSTLANE_B32 result. Scavenging one
emergency-spills an SGPR, and an SGPR spill flips EXEC with S_NOT_B64, so the
EXEC flips land between the S_CMP_EQ_U32 that defines SCC and the read of SCC
that follows, clobbering it in between.
[AMDGPU] Prevent SCC clobber in frame index lowering in scaling path
eliminateFrameIndex has two lowering strategies, but only one has the
proper handling for checking SCC-liveness to prevent clobbering. Unify
them with a helper function to ensure both paths handle the same
[CIR] Regenerate CHECK lines for two callconv opt-out tests
Neither `union.c` nor `paren-list-agg-init.cpp` is blocked by the
calling convention lowering pass any more, and both compile clean with
it running. Remove the `-fno-clangir-call-conv-lowering` opt-out and
update their CHECK lines. The coerced parameters and returns they now
pin match classic.
The bodies still differ, since CIR round-trips the record through a
fresh coerce slot, so those fragments move to the `LLVMCIR` and `OGCG`
prefixes the file already declares. Four `define` lines in
`paren-list-agg-init.cpp` wildcarded their return type and now pin it.
Assisted-by: Cursor / claude-opus-5
WebAssembly: Move target EH passes into backend (#225172)
Currently the wasm-specific EH lowering passes are added in the
generic pass configs, driven by the TargetOptions ExceptionModel.
As preparation for driving this process off the IR flag, move these
pass runs into the target. There should be no change in the relative
pass ordering.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
[clang][CodeGen] Compute the pointer-overflow offset from the index list (#225139)
Follow up to #223446. EmitGEPOffsetInBytes recreated the offset by
walking the GEP value that EmitCheckedInBoundsGEP had just created. That
required a separate path for the case where CreateGEP folded the result,
because a folded GEP need not be a GEP at all: `gep(null, 1)` becomes
`inttoptr(1)`, and `gep(@g, 0)` becomes `@g`.
Pass ElemTy and IdxList instead and walk those directly. This removes
the constant path, the cast to GEPOperator, and two asserts that
restated the caller's own behaviour.
This is not quite NFC. The constant path reported OffsetOverflows as
false unconditionally, so a constant GEP whose byte offset wrapped to
exactly zero satisfied the TotalOffset == Zero early return and emitted
no check. The unified path computes the flag, so such a GEP now emits
one. It is provably valid, since TotalOffset == 0 makes the computed
address equal the base, so this is extra IR at -O0 rather than a change
in behaviour.
[5 lines not shown]
[clang][CodeGen] Compute the pointer-overflow offset from the index list (#225139)
Follow up to #223446. EmitGEPOffsetInBytes recreated the offset by
walking the GEP value that EmitCheckedInBoundsGEP had just created. That
required a separate path for the case where CreateGEP folded the result,
because a folded GEP need not be a GEP at all: `gep(null, 1)` becomes
`inttoptr(1)`, and `gep(@g, 0)` becomes `@g`.
Pass ElemTy and IdxList instead and walk those directly. This removes
the constant path, the cast to GEPOperator, and two asserts that
restated the caller's own behaviour.
This is not quite NFC. The constant path reported OffsetOverflows as
false unconditionally, so a constant GEP whose byte offset wrapped to
exactly zero satisfied the TotalOffset == Zero early return and emitted
no check. The unified path computes the flag, so such a GEP now emits
one. It is provably valid, since TotalOffset == 0 makes the computed
address equal the base, so this is extra IR at -O0 rather than a change
in behaviour.
[5 lines not shown]
[mlir][linalg] Fix exponential walk of tensor.extract indices (#222319)
`isLoopInvariantIdx` showed exponential runtime behavior. Replaced with
a worklist and a visited set.
Encountered in iree-org/iree#24886.
Assisted-by: Claude
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[clang][flang][OpenMP] Fix context selector matching and scoring
Incorrect construct contexts and selector scoring can select the wrong
DECLARE VARIANT function or METADIRECTIVE replacement.
Build construct contexts in source order rather than emitted MLIR order:
Source: teams -> distribute -> parallel -> do
MLIR: teams -> parallel -> distribute -> do
Count executable constructs without selectable traits, omit informational
directives, and start the context at the innermost TARGET.
Compute device weights from the enclosing context depth, choose the
highest-scoring complete ordered construct match, and count each selector's
score once using arbitrary-width arithmetic.
Inside a single PARALLEL region, device={kind(cpu)} previously scored 2,
tying construct={parallel}. Using the enclosing context depth gives the
[24 lines not shown]
[AMDGPU] Fold fpround of fadd and fsub into v_mad/fma_mixlo and mixhi
MadFmaMixFP32Pats turns (fadd x, y) into (fma x, 1.0, y) and (fsub x, y)
into (fma (-y), 1.0, x) so the mix instructions absorb the operation along
with the f16 or bf16 source modifiers. MadFmaMixFP16Pats and
MadFmaMixFP16Pats_t16 only did this for fmul, so a rounded result still
needed a separate convert for a rounding the mix instructions perform
themselves.
Unlike the f32 patterns these do not require an operand to be an fpextend
of an f16, since an fpround on the result always removes the convert. The
rewrite is exact because the mix instructions round the f32 result again
when they write the 16-bit destination, so it stays f32_to_f16(fma(x, 1.0,
y)).
Assisted-by: Claude Code Opus 5
[SandboxVec][LoadStoreVec] Address review feedback for mixed-type constants
Build equivalent constants with ConstantExpr instead of inserting CastInsts, and tighten names, comments, and lit tests.
Co-authored-by: Cursor <cursoragent at cursor.com>
[AMDGPU] Require flushed FP16 denormals for the mad-mix f16 results (#224911)
v_mad_mixlo_f16 and v_mad_mixhi_f16 are the unfused gfx900 forms and flush 16-bit denormals, so a denormal half result is written as zero even when the FP16 mode asks for it to be kept, while the patterns only required the FP32 mode to flush and that is the one a HIP compile turns off on its own.
Assisted-by: Claude Code Opus 5