[libc++] Implement LWG4290: Missing Mandates clauses on is_sufficiently_aligned (#229765)
LWG4290 adds a Mandates clause that requires `Alignment` to be a power
of two for `std::is_sufficiently_aligned`. This patch adds the power of
two check as a `static_assert` in the internal helper of
`std::is_sufficiently_aligned`. The change mirrors
`std::assume_aligned`. I added a new libc++ verify test.
Fixes #189823.
## AI Disclosure
AI was used to review the changes before putting this PR up.
[BOLT][AArch64] Preserve tail-call annotation on local trampolines (#228105)
Propagate the tail-call annotation so cluster relaxation can relax the
trampoline's outgoing branch with a long thunk.
[orc-rt] Add SymbolLookupSet-based lookup to SimpleSymbolTable (#230761)
SimpleSymbolTable::lookup(const SymbolLookupSet&) looks up the given
linker-level names in the table and returns a SymbolLookupResult with
one entry per element of the set, in the same order. Missing
weakly-referenced symbols are reported as null addresses; missing
required symbols are reported as empty optionals, matching
NativeDylibManager::lookup.
Reapply "[clang-repl] Fix crashing on unusable top-level declarations" (#230054) (#230742)
Drop the previously failing test lines:
```c
printf("not crashed\n");
// CHECK-NEXT: not crashed
```
A bare `return;` is sufficient for detecting the regression.
Reverts #230734.
[MLIR][CAPI][Python] Add C API and Python bindings for the remark engine (#229094)
Expose the optimization remark engine (`mlir/IR/Remarks.h`) through the
C API and the Python bindings, so clients outside C++ can enable it,
receive the reported remarks and emit their own.
### C API (`mlir-c/Remarks.h`)
- `mlirContextEnableOptimizationRemarks{,ToFile,WithCallback}` select
the categories (one regex per kind plus `all`), the emitting policy
(`all` or `final`) and the sink: MLIR remark diagnostics, a
YAML/bitstream file, or a callback invoked with an opaque `MlirRemark`.
- `mlirContextFinalizeOptimizationRemarks` flushes postponed remarks,
writes the file and drops the engine; `mlirContextHasRemarkEngine`
queries it.
- `MlirRemark` accessors (kind, remark/category/function names,
location, id, key/value arguments, print), valid for the duration of the
callback.
- `mlirEmitOptimizationRemark` emits a remark at a location through the
engine of its context.
[24 lines not shown]
[LAA] Add tests for a set of miscompiles. (NFC) (#230773)
Add tests for the following issues:
* interleaved-accesses-retry-runtime-checks.ll: when LAA retries with
runtime checks, the store of an interleave group is moved past a load
from the same address.
* is-safe-dep-distance-narrow-btc.ll: accesses with a real backward
dependence are reported as safe if the byte stride does not fit in a
narrow backedge-taken count type or MaxBTC * MaxStride wraps in it.
* is-safe-dep-distance-partial-overlap.ll: accesses are reported as
independent if |Dist| > MaxBTC * MaxStride, even if the distance is not
a multiple of the access size and the accesses partially overlap.
* different-access-types-rt-checks.ll: the runtime check bounds of a
pointer loaded as i32 and stored as i8 only account for the i8 store, so
the tail of the loaded range is not checked.
* runtime-check-grouping-wrapping-bounds.ll: pointers whose bounds can
wrap relative to each other are grouped, so the group's Low can be above
its High and the runtime checks never detect a conflict.
[VPlan] Re-use LCSSA phis when expanding AddRecs of other loops. (#230772)
Don't re-use IR values defined in loops that do not contain the plan's
entry in VPSCEVExpander::tryToReuseIRValue, as such uses outside the
loop break LCSSA. Instead, match SCEVExpander behavior and re-use an
existing LCSSA phi, reusing SCEVExpander::findReusableLCSSAPhi.
RuntimeLibcalls: Generate an enum for runtime library names
Emit an RTLIB::RuntimeLibrary enumerator for each distinct LibraryName,
and use it instead of a string for isLibraryAvailable.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
fix(x86): reuse shifted masks in SSE ANDNP
Demanded-bit simplification can bypass a shared arithmetic shift when
ANDNP only needs sign bits. Reuse the existing shift to avoid keeping
its input live across a two-address SSE instruction.
Restore the original SSE2 vector select checks.
Refs #230700
AMDGPU: Fix true16 build_vector (0, x) pattern using a 16-bit shift operand
The real true16 pattern for (build_vector 0, VGPR_16:$x) fed the 16-bit
register directly to V_LSHLREV_B32, which takes a 32-bit operand. Widen
it with a REG_SEQUENCE first. This avoids redundant 16-bit moves in
SelectionDAG, and fixes a GlobalISel selection failure when the 16-bit
input is a G_TRUNC of a 32-bit value, as the shift's operand class
constrained the trunc result to vgpr_32.
I also don't know why this pattern is overcomplicating this. I would expect
true16 to literally translate build_vector to reg_sequence plus a materialize
of the 0.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
RuntimeLibcalls: Remove the Default* libcall lists
Every system library now lists only LibcallLibrary members, so nothing
references DefaultRuntimeLibcallImpls. arm64ec's '#'-prefixed set was
still derived from the whole default list minus the width slices and
the Windows exclusions. Build it from the compiler-rt, libm and libc
slices instead, still omitting the calls the Windows runtime lacks.
Take AArch64's fp128 long double calls from the libm slice, and delete
the remaining default and Windows base lists.
The generated tables are unchanged.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
fix(X86): restore CCMP through boolean NOT
Demanded-bits simplification can turn XOR with 1 into NOT, hiding
SETCC operands from CCMP formation. Invert the condition under AND
when the other SETCC masks high bits and the NOT has one use.
Restore the original ccmp.ll checks for all four configurations.
Refs #230700
AMDGPU: Restore kernel tests in early-if-convert-cost.ll (#230764)
59911f438c47 converted these test kernels into functions using vgpr
function arguments, but this perturbed the code too much. This just
happened to run into a preexisting bug in later passes which appears
in expensive checks builds.
[AMDGPU] Don't select scalar loads for LDS/scratch (#227423)
A uniform load from LDS marked `!invariant.load` gets selected as
`s_load_dword`. SMEM can't access LDS, so it ends up reading global
memory at the LDS offset (address 0x0 + offset). On MI350 this faults.
```llvm
%v = load i32, ptr addrspace(3) %p, align 4, !invariant.load !0
```
```
s_mov_b32 s1, 0
s_load_dword s0, s[0:1], 0x0 ; LDS offset used as a global address
```
I hit this while marking LDS reads as `!invariant.load` from FlyDSL
(MLIR), to let LLVM rematerialize them instead of spilling.
Fix: only allow the SMRD load patterns for global, constant, and 32-bit
constant address spaces. GlobalISel already excluded LDS/scratch, so
[2 lines not shown]
[AMDGPU] Extract byte lanes of a split vector from the 32-bit source
After a <4 x i8> is split into i16 halves, a byte lane is extended
from an i16 shift of a truncate. Rewrite it on the 32-bit source as an
and of srl or a sign_extend_inreg of srl, which select to a single bit
field extract.
Uniform zero and any extends are left alone, their i16 shift is already
promoted to i32.
AMDGPU: Restore kernel tests in early-if-convert-cost.ll
59911f438c47 converted these test kernels into functions using vgpr
function arguments, but this perturbed the code too much. This just
happened to run into a preexisting bug in later passes which appears
in expensive checks builds.
clang: Add convergence control bundles when creating calls
Instead of creating a call and then recreating it with a convergencectrl
operand bundle, add the bundle when the call is created. Recreating the
call left EmitCall's callOrInvoke pointing at the erased call, and lost
any metadata already attached to it, such as !alloc_token and
!callee_type.
For calls from EmitCall, decide whether a token is needed from the call
site attributes derived from the callee declaration, rather than from
the IR.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[Analysis][CodeGen] Remove unused haveFastClmul (NFC) (#230748)
The last caller of TargetTransformInfo::haveFastClmul was removed on
July 31, 2026 in commit 28d9aec513c63beffd2631412adfae327763559e,
leaving TargetTransformInfoImplBase::haveFastClmul and
BasicTTIImplBase::haveFastClmul unused as well.
Assisted-by: Antigravity
[Coroutines] Use SmallMapVector for AllocaInfo::Aliases (#230585) (#230749)
`coro::AllocaInfo::Aliases` and `AllocaUseVisitor::AliasOffetMap` were
previously declared as `DenseMap<Instruction *, std::optional<APInt>>`.
Iterating over `Alloca.Aliases` in `insertSpills` visited instructions
in pointer-hash order, causing non-deterministic IR and codegen in
`CoroSplitPass` across compiler invocations.
Replace `DenseMap` with `SmallMapVector<Instruction *,
std::optional<APInt>, 4>` to preserve deterministic insertion order.
Fixes #230585
[clang] Replace PointerUnion::dyn_cast with llvm::dyn_cast (NFC) (#230747)
PointerUnion::dyn_cast has been soft-deprecated in favor of
llvm::dyn_cast and llvm::dyn_cast_if_present. This patch replaces the
former with llvm::dyn_cast where the operand is guaranteed to be
nonnull.
{Class,Var}TemplateSpecializationDecl::getSpecializedTemplateOrPartial()
always returns a nonnull pointer by construction regardless of whether
the specialization is from a primary template or a partial
specialization.
Assisted-by: Antigravity
[orc-rt] Use const void* for SymbolLookupResults. (#230753)
Lookup is performed for symbol resolution, not for direct access to
symbols. Using 'const void*' ensures that the underlying memory is not
accidentally modified through the lookup result.