[Attributor] Skip the dead-internal-function walk for AAs that will not be updated
AAIsDeadFunction::initialize calls isAssumedDeadInternalFunction, which
runs checkForAllCallSites and so creates an AAIsDead for every internal
caller, each of which repeats the walk. Initialization therefore recurses
through the whole cone of internal callers.
When the AA will not be updated (outside of deduction, or for a function
the Attributor is not run on), getOrCreateAAFor fixes it pessimistically
right after initialize returns, discarding the walk's result. Skip the
walk in that case and assume the entry block live.
This is the dominant cost of the CGSCC OpenMPOpt pass, whose cleanupIR
queries the liveness of callers outside the current SCC on every SCC,
giving (#SCCs) x (depth of the internal caller chain). #222226 removed
the callee-seeding half of this cost; this removes the walk itself.
cleanup-no-seeding.ll drops a debug check that only the removed walk
printed.
[2 lines not shown]
[CIR] Support PackIndexingExpr for LValue and Aggregates (#227040)
Support `PackIndexingExpr` in `CIRGenFunction::emitLValue` and
`AggExprEmitter` by visiting its selected expression, matching classic
Clang codegen (`CGExpr.cpp` and `CGExprAgg.cpp`) and completing
`PackIndexingExpr` support across all CIR expression emitters.
Fixes #226928
Assisted by Antigravity and Gemini
[SampleProfileMatcher] Fix direct basename matching for suffixed function names
Direct basename matching (#184409) silently skips any function whose name
carries a suffix, for two reasons:
1. `getDemangledBaseName` demangles the raw name. For names such as
`_ZL3fool.__uniq.123`, `_Z3fool.llvm.7`, `.part.N` or `.cfi` the Itanium
demangler's root node is a DotSuffix, for which `getFunctionBaseName()`
returns null, so the function is never a candidate on either the IR or
the profile side.
2. `UpdateWithSalvagedProfiles` keys `FuncNameToProfNameMap` by the raw IR
name, but `SampleProfileReader::getSamplesFor(const Function &)` looks it
up by `getCanonicalFnName`. A salvaged `.llvm.N` (ThinLTO-promoted),
`.part.N` or `.cfi` function therefore still gets no profile.
Canonicalize the name before demangling (additionally dropping `.__uniq.N`,
which `getCanonicalFnName` keeps when the profile has uniq names), and key
the map by the canonical name. Other suffixes such as coroutine `.resume`
are kept so those clones do not make a basename ambiguous.
[12 lines not shown]
[clang][bytecode] Add `EvalSettings` struct and replace parent `State` (#226175)
We currently always construct a full `EvalInfo` when using the bytecode
interpreter, we then pass it to the `evaluate*` function, which only
copies a few values from it.
Introduce a new `EvalSettings` struct that we pass instead of a `State`.
This is cheaper to construct and we only pass what's needed. There is
one `evaluateAsRValue` overload left that takes a parent `State`, that
will be removed in a subsequent patch.
[CodeGen] Enable validated AMDGPU wave spill costs by default
Use available validated wave counts for AMDGPU spill placement without an
explicit opt-in. Keep the existing target, mapping and normalization checks
and the ordinary block-frequency fallback for unavailable or rejected data.
Retain -enable-wave-profiled-spill=false for debugging and matched performance
comparisons. Test default behavior, explicit disabling, equivalence to explicit
enabling, fallback cases and the positive frequency floor.
[CodeGen] Keep wave-profiled spill frequencies positive
SpillPlacement expects positive block weights, but a valid wave profile can
record zero executions for a CFG-reachable block. Giving such a block zero
spill cost can make the allocator choose a very different placement.
Clamp every accepted wave-derived frequency to at least one, as we already
do for nonzero counts that round down to zero. Unmeasured or rejected blocks
still use their existing MBFI frequency. Add a focused MIR test for a valid
zero-wave record.
This pattern arose in a profiled Composable Kernel convolution case. With
the separate spill correctness fixes and partial spilling enabled, the
zero-cost policy failed two CPU-reference checks; the positive floor passed
both. The test checks the cost directly; the application result was checked
separately on gfx950.
[CodeGen] Use validated wave counts for AMDGPU spill costs
Lane-based block frequencies can understate the cost of a spill in
divergent GPU code: a wave still executes a block with only some lanes
active. Use measured block-wave counts to weight SpillPlacement's costs
relative to the original entry-wave count.
Opt in on AMDGPU only. Accept a measured count only when its IR block
maps uniquely to a machine block with matching predecessors and
successors. Keep the existing MBFI cost for unmeasured or rejected
blocks, and leave branch probabilities and general BFI unchanged.
[Transforms] Preserve wave profiles across CFG rewrites
HIP device PGO attaches measured wave counts to IR blocks. Later CFG
rewrites can drop counts from unchanged blocks or leave stale counts
on blocks that now execute differently. Either case makes the profile
unreliable for later optimizations.
Preserve counts through switch lowering, structurization, and loop
rotation only when a block still represents the same executions. Keep
unaffected counts and their IDs even when a loop header's count must be
invalidated. Transfer branch hints only for equivalent decisions, and
avoid assigning switch weights when default traffic cannot be traced
to one edge.
[PGO] Load dense block wave counts from device profiles
Use the profile's dense layout flag to map appended wave-only slots after
the original block/select prefix. Reuse the producer's block selection so
eligible loop and reconvergence blocks retain their directly measured wave
frequencies. Keep select slots out of the block mapping.
Test dense and sparse profiles, measured zeros, generation and metadata
switches, and exact loop block identities after critical-edge splitting.
[PGO] Add a debugging switch for wave profile metadata
Add the hidden pgo-wave-metadata option, enabled by default, for debugging,
performance comparisons and disabling wave annotations when investigating
regressions without turning off ordinary PGO or uniformity hints.
Gate wave metadata emission while retaining the existing clearing of stale
function and block annotations during profile use. Profile collection and
ordinary count reconstruction are unchanged.
Test default/explicit enablement, disabling, retained counts and uniformity
hints, and replacement profiles with missing, mismatched or zero counts.
[PGO] Load GPU wave counts into IR metadata
GPU profiles contain wave counts alongside lane counts, but profile use
does not expose them to optimizations. Wave visits do not obey scalar
flow conservation, so unmeasured blocks cannot use counts reconstructed
from neighboring blocks or ordinary branch weights.
Map wave-counter indices to the blocks selected by PGO instrumentation,
after reproducing its critical-edge splits. Attach measured counts using
wave.profile metadata, retaining measured zeros and marking other blocks
unmeasured. Require a measured entry count for normalization and exclude
select-counter slots from the block mapping.
Validate the wave-counter layout against the accepted lane profile.
Clear old wave metadata when loading a replacement profile, including
when a function has no usable record. Do not emit wave metadata for
previously profiled functions: their branch weights may change counter
placement without changing the CFG hash. Keep ordinary lane-count
reconstruction and branch weights unchanged.
[IR] Define GPU wave-profile metadata
Existing offload GPU profile counters measure lane executions, while
GPU instructions execute at wave granularity under an active-lane
mask. A block visited by every wave can therefore look cold when only
a few lanes are active. Scalar branch weights also cannot represent a
divergent wave visiting both successors before reconverging. These
profiles are a poor fit for optimizations that estimate work performed
by a wave.
Lane counters remain useful for measuring per-lane branch selectivity
and estimating work that scales with the number of active lanes.
Wave counters cannot replace them: a visit with one active lane and a
visit with all lanes active both count as one. The two profiles provide
complementary information about instruction execution and lane activity.
Introduce wave.profile and wave.profile.block metadata to represent
measured wave visits to IR blocks. Each dynamic visit with at least one
active lane contributes one to the count. A measured zero is distinct
[18 lines not shown]
[Profile] Synchronize the dense-wave profile header in compiler-rt
Keep the compiler-rt copy of InstrProfData.inc in sync with LLVM's
copy by adding the dense-wave variant bit and matching comments.
This fixes the shared-header consistency failure in profile CI.
CI failure: https://github.com/llvm/llvm-project/actions/runs/36506166214/job/109207990906
[mlir][cf] Fix cf.switch round-trip for case values wider than 64 bits (#221490)
The `cf.switch` parser reads case values into `int64_t`, and its printer
uses `APInt::getLimitedValue()`, preventing values wider than 64 bits
from round-tripping. Printing negative narrow values as unsigned can
also trigger an assertion when reparsing.
Parse case values into `APInt` and print them directly, so wide and
negative values round-trip correctly. Accept signed or unsigned literals
that fit the switch operand’s bit width, and report an error for values
that do not.
Add regression tests for wide and negative values, integer-width
boundaries, out-of-range diagnostics, and case selection after parsing.
Fixes #220609
Assisted-by: Codex
HBSD: Disable JIT in devel/pcre2
PCRE2 JIT has long been a source of frustration on HardenedBSD because
of our PaX NOEXEC implementation. Now that the PCRE2 JIT implementation
has had a recent security advisory (a buffer overflow, nonetheless), we
probably should not trust PCRE2's JIT.
Signed-off-by: Shawn Webb <shawn.webb at hardenedbsd.org>
See-Also: GHSA-r9hj-j2rw-4q3m
[flang][cuda] Add cuf.on_device op and lower to it (#227089)
Replace the CUFFunctionRewrite pass, which matched fir.call names and
folded them to constants, with a cuf.on_device operation emitted when
the intrinsic is lowered. CUFOpConversion folds the operation once the
code is in its host or device context, and leaves the host copy of an
OpenACC routine unfolded so the device clone is not baked to false. Any
operation that remains is folded by the late CUF conversion, and the
rewrite pass is dropped from the pipeline.
Revert "releand "[clang-repl] Implement IncrementalHIPDeviceParser for HIP device compilation"" (#227123)
Reverts llvm/llvm-project#226930
AMD author has an additional step to perform, that is still needed.
in contact with Author.
[CIR] Support aggregate co_await / co_yield in AggExprEmitter (#225412)
Support evaluating `co_await` and `co_yield` expressions whose result is
an aggregate type in `AggExprEmitter`.
Fixes #225317
[TailCallElim] Add profile annotations to return value selects
Tail call elimination in some cases can create selects on possible
return values conditioned on whether or not execution is currently in
what was a recursive call. That is equal to the probability with which
we recurse, which in turn can be computed from the block frequencies of
blocks that recurse and blocks that directly return.
Reviewers: mtrofin
Reviewed By: mtrofin
Pull Request: https://github.com/llvm/llvm-project/pull/202518
[SelectionDAG] Handle constants in SimplifyMultipleUseDemandedBits
Replace a non-zero constant with zero when none of its set bits are
demanded.
This allows users of `SimplifyMultipleUseDemandedBits` to eliminate
irrelevant constant bits while preserving the convention that a null
SDValue indicates no simplification.
[Clang] Add missing release note entry in #226753 (#227078)
As per the feedback from #226753, we add release note for GH-212211.
Also move the test to new-delete.cpp.
Assisted-by: Claude
[CodeGen] Enable validated AMDGPU wave spill costs by default
Use available validated wave counts for AMDGPU spill placement without an
explicit opt-in. Keep the existing target, mapping and normalization checks
and the ordinary block-frequency fallback for unavailable or rejected data.
Retain -enable-wave-profiled-spill=false for debugging and matched performance
comparisons. Test default behavior, explicit disabling, equivalence to explicit
enabling, fallback cases and the positive frequency floor.