AMDGPU: Move R600 subtarget features to R600Features.td
Split the R600 SubtargetFeature definitions and
R600FrontendVisibleFeatures out of R600Processors.td, mirroring
GCNFeatures.td.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
AMDGPU: Move GCN subtarget features to GCNFeatures.td
Move the GCN SubtargetFeature definitions, the FeatureISAVersion lists
and AMDGPUFrontendVisibleFeatures out of AMDGPU.td, so they can be used
without parsing the instruction, register and intrinsic definitions.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
AMDGPU: Move scheduling model defs to AMDGPUSchedModels.td
Split the SchedMachineModel definitions out of SISchedule.td so the
processor definitions can be parsed without the InstRW and other
instruction-dependent scheduling information.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[InstCombine] Fix InstCombine pass to sink llvm.assume calls (#229658)
When trying to sink a GEP instruction into the loop preheader the
InstCombine pass ignores and deletes an llvm.assume call that sets the
alignment for the pointer of the GEP instruction which later passes
depend on. This patch sinks such llvm.assume calls while sinking a
dependent instruction.
Fixes https://github.com/llvm/llvm-project/issues/226698.
[alpha.webkit.UnretainedLocalVarsChecker] Treat the collection of a fast enumeration as the origin of its element (#230818)
alpha.webkit.UnretainedLocalVarsChecker reported every element variable
of an Objective-C fast enumeration such as "for (T *x in collection)"
since the variable has no initializer, even when the collection was kept
alive by a RetainPtr local variable.
The element of a fast enumeration is kept alive by its collection, which
can't be mutated during the enumeration. So treat the collection as the
initial value of the element, as in "T *x = collection;". Because the
collection is evaluated before the loop body is entered, a guardian
local variable only needs to outlive the loop body; it can be declared
in the same scope as the loop as long as the loop body doesn't mutate
it.
The collection is a full-expression of its own so a temporary smart
pointer in the collection is destroyed before the loop body is entered
and isn't considered safe.
Authored with Claude Code.
[alpha.webkit.UncountedLocalVarsChecker] Don't treat a const operator call on a guardian as a mutation (#230900)
GuardianVisitor treated `guardian->method()` as a mutation of the
guardian because the implicit object argument of a member operator call
had no corresponding parameter, and the check fell back to the non-const
type of the guardian itself. Skip the implicit object argument when the
operator is a const member function, and only apply the argument offset
for implicit object member operators so that arguments of non-member
operators map to the right parameters.
Also treat passing a guardian to a const reference parameter as
non-mutating by checking the constness of the referenced type instead of
the reference type.
Coded with Claude code.
[AMDGPU][SplitModule] Add entry points for unreachable call cycles
A function in a call cycle always has an incoming direct call, so
`SplitGraph` never makes it an entry point. If no kernel and no other
entry point reaches the cycle, the cycle is assigned to no partition.
This can hit an assertion in `verifyGraph`.
In this PR, we visit the unreached nodes in reverse post-order of the
direct call edges, and make each node that is still unreached an entry
point. Only the outermost unreached cycles get one, so their callees
are not copied into other partitions.
[NFC][AMDGPU] Add a test showing unreachable call cycles in module splitting
If no entry point reaches a call cycle, `AMDGPUSplitModule` creates no
entry point for it. In assertion builds, `verifyGraph()` fails. In
release builds, no partition defines the functions in the cycle.
Add a test that shows the current crash.
[docs] Update obsolete Phabricator references for GitHub workflows (#229542)
Use GitHub pull requests in the contribution and code-review guidance,
remove obsolete Phabricator contact handles and infrastructure listings,
and update source links to their current locations.
Assisted-by: Codex
[AMDGPU][SplitModule] Add entry points for unreachable call cycles
A function in a call cycle always has an incoming direct call, so
`SplitGraph` never makes it an entry point. If no kernel and no other
entry point reaches the cycle, the cycle is assigned to no partition.
This can hit an assertion in `verifyGraph`.
In this PR, we visit the unreached nodes in reverse post-order of the
direct call edges, and make each node that is still unreached an entry
point. Only the outermost unreached cycles get one, so their callees
are not copied into other partitions.
[NFC][AMDGPU] Add a test showing unreachable call cycles in module splitting
If no entry point reaches a call cycle, `AMDGPUSplitModule` creates no
entry point for it. In assertion builds, `verifyGraph()` fails. In
release builds, no partition defines the functions in the cycle.
Add a test that shows the current crash.
[clang-tidy][docs] Fix ExplicitConstructorCheck example links (#229932)
Point the contribution guide at the current
misc/ExplicitConstructorCheck header and implementation. Replace the
obsolete Phabricator source-browser link with the GitHub source link.
Assisted-by: Codex
[MLGO][RegAlloc] Update test expectations after #228618 (#230973)
Test fixes after #228618: SlotIndex distances and resulting LiveInterval
sizes changed.
[CodeGenPrepare] Handle negative GEP offsets in splitLargeGEPOffsets (#227627)
CodeGenPrepare::splitLargeGEPOffsets collects GEP candidates with large
constant offsets and rebases them to a shared common base, but only for
positive offsets (`ConstantOffset > 0`). Negative far-offset accesses
are left with independent `SUBXri` base materializations per access,
giving them different base registers and preventing the
`LoadStoreOptimizer` from pairing them into `LDP/STP`.
Relax the gate to `ConstantOffset != 0`, and extend
`AArch64TargetLowering::getPreferredLargeGEPBaseOffset` to handle
negative offsets using the same `HighPart = MinOffset & ~0xfff` rebase
as positive ones. For negative `MinOffset`, `HighPart` is the next lower
4096-aligned address, so all residuals (`Offset - HighPart`) are
non-negative and fit `LDR/STR`'s 12-bit unsigned scaled immediate. When
the residual also falls within `LDP/STP`'s 7-bit pairing range, `LSO`
pairs directly; otherwise `LSO`'s base-adjust (#223684) folds in an
extra `ADDXri` (or merges it into a preceding `SUBXri`/`ADDXri` when
adjacent), recovering the same instruction count as a direct rebase to
[42 lines not shown]
[Alignment] Add a static function to construct an `Align` from power-of-2 values (#230920)
Many call sites of the `Align` constructor uses it as `Align(1ULL << n)`, and
the constructors takes the `Log2` of the value. Add a new static function to
allow `Align::fromLog2(n)` and update old call sites. This avoids the
unnecessary `Log2(1ULL << n)` computations when constructing `Align`s.
[SLP]Fix miscompile of absorbing lanes in flattened chains
The operand columns, peeled while flattening the associative chains,
modeled the lane with the absorbing constant as op(C, poison). Such
lanes belong to the real instructions of the flattened node, so the
poison operand is not frozen there and the lane becomes poison. Keep the
identity as the other operand for the peeled columns.
Follow-up to #228872.
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/230959
[clang] Do not compute typo-correction suggestions for disabled diagnostics (#209694)
While benchmarking I noticed that Boost.MPL compiles with ~1.2% more
instructions
after #140629.
clang spends a fair amount of time computing typo-correction
suggestions for diagnostics that are never emitted, for example Boost
doesn't contain any
typos, so all of this work produces nothing.
We can easily avoid this overhead by checking `isIgnored()` before
computing
suggestions, and by not treating known non-conditional directives as
typos.
You can see the improvement here:
https://llvm-compile-time-tracker.com/compare.php?from=49de424f45389cb757c3cc8c50daf38d024e2314&to=89a68cd24f9fabf15897d7b20b77bb5b0bfb9c16&stat=instructions%3Au
[MISched] Unify SUnit formatting among users (#229756)
This patch migrates diverging SUnit formats, that form a minority in the
codebase, to a single unified format dictated by the new SUnit's
operator<<.
[llubi] Don't ignore address taken by assume-like intrinsics (#230938)
The function pointer used by `llvm.assume` is still evaluated.
The test is generated by DeepSeek-V4.1-Flash.
[LAA] Compute MaxBTC * MaxStride in a wide enough type. (#230860)
isSafeDependenceDistance proves independence if |Dist| > MaxBTC *
MaxStride. The original code did not account for MaxBTC * MaxStride
wrapping, e.g. if either has a narrow type.
Fix by computing the product in the wider of the types of the distance
and MaxBTC, zero-extending MaxBTC. The distance is sign-extended as
before.
The product can still wrap if MaxBTC is as wide as the distance, e.g. an
unbounded i64 MaxBTC. Requiring it not to wrap regresses existing proofs
for symbolic distances (e.g. PR31098), so that case is left as is for
now.
[SCEV] Remove unused code from ScalarEvolutionExpressions.h (NFC) (#230859)
Remove the non-static
SCEVSequentialMinMaxExpr::getEquivalentNonSequentialSCEVType overload,
which has no callers, and SCEVLoopAddRecRewriter together with the
LoopToScevMapT alias, which have no users.