[RuntimeDyld] Implement more relocations for 32-bit PowerPC (#229933)
RuntimeDyldELF handled only R_PPC_ADDR16_{LO,HI,HA} on 32-bit PowerPC
and aborted on anything else. Any object with a call (R_PPC_REL24) or an
.eh_frame section (R_PPC_REL32) hit that, so MCJIT could not run even
trivial modules.
Add R_PPC_ADDR32, R_PPC_REL32, R_PPC_REL16_{LO,HA}, R_PPC_REL24 and
R_PPC_PLTREL24. The addend of R_PPC_PLTREL24 selects the GOT pointer of
a PLT call stub and is not an offset from the symbol, so it is ignored.
Calls to external symbols or out-of-range targets go through a stub that
loads the address into r12 and branches through CTR, as on 64-bit
PowerPC.
Also give getGOTEntrySize() an answer for 32-bit PowerPC. It is reached
when the memory manager reserves allocation space, as remote MCJIT does.
This lets Mesa's llvmpipe run on 32-bit PowerPC.
Assisted-by: Claude Code
[GlobalISel] Preserve demanded lanes in G_CTLS KnownBits queries (#230096)
Pass DemandedElts to the recursive sign-bit query so selected vector
lanes retain their count range.
Assisted by: GPT 6.1 Sol
[memprof] Remove MemProf format Version 2 (#229262)
It's been a couple of years since we introduced Version 3. Since our
active development has moved on to Version 4 these days, this patch
removes the old version.
[llubi] Add support for `bitinsert` and `bitextract`
Add interpreter support for the `bitinsert` and `bitextract` instructions.
The bits are read and written at any bit offset, through the existing
`Context::fromBytes`/`Context::toBytes` overloads. They keep poison bits and
pointer provenance tags per bit.
An out-of-range or `poison` offset returns `poison`.
[llubi] Make sure the allocated address can be represented (#230656)
`deriveFromMemoryObject` assumes that the address always fits in the
address space. Check this in `Context::allocate` to avoid crashes.
The test is generated by DeepSeek-V4.1-Flash.
[clang][docs] Check in attribute reference Markdown
Tracking issue: #227907
Implements stage 3 of #227907 by checking in AttributeReference.md and
AttributeReference/*.md, and by removing the docs build rules that
generated and split those Markdown sources.
This also adds the docs-build check for alphabetically sorted H3
attribute headings now that those headings are hand-authored in the
checked-in Markdown files.
The checked-in Markdown was populated with:
mkdir -p clang/docs/AttributeReference
cp build/tools/clang/docs/AttributeReference.md clang/docs/AttributeReference.md
cp build/tools/clang/docs/AttributeReference/*.md clang/docs/AttributeReference/
The AttrDocs TableGen backend is retained as a deterministic migration
[3 lines not shown]
[clang][docs] Render attribute syntaxes from a Sphinx role
Implements stage 2 of https://github.com/llvm/llvm-project/issues/227907 by generating a JSON syntax database from Attr.td and rendering supported syntaxes through the clang-attr-syntaxes Sphinx role.
[clang][docs] Generate split attribute reference docs
Implements stage 1 of https://github.com/llvm/llvm-project/issues/227907 by teaching the attribute docs generator to emit split-file-formatted Markdown and wiring the docs build to split it for Sphinx.
[compiler-rt][ROCm] Collect Linux profiles from resident images (#229518)
On Linux, profile collection looks up every registered HIP image.
Those lookups can load images the program never used just to collect
zero counters.
Collect profiles from images already loaded by the HSA runtime. If
the loader inspection APIs are unavailable, keep the existing image
lookup path.
[libc++] Implement LWG4290: Missing Mandates clauses on is_sufficiently_aligned (#229765)
LWG4290 adds a Mandates clause that requires `Alignment` to be a power
of two for `std::is_sufficiently_aligned`. This patch adds the power of
two check as a `static_assert` in the internal helper of
`std::is_sufficiently_aligned`. The change mirrors
`std::assume_aligned`. I added a new libc++ verify test.
Fixes #189823.
## AI Disclosure
AI was used to review the changes before putting this PR up.
[BOLT][AArch64] Preserve tail-call annotation on local trampolines (#228105)
Propagate the tail-call annotation so cluster relaxation can relax the
trampoline's outgoing branch with a long thunk.
[orc-rt] Add SymbolLookupSet-based lookup to SimpleSymbolTable (#230761)
SimpleSymbolTable::lookup(const SymbolLookupSet&) looks up the given
linker-level names in the table and returns a SymbolLookupResult with
one entry per element of the set, in the same order. Missing
weakly-referenced symbols are reported as null addresses; missing
required symbols are reported as empty optionals, matching
NativeDylibManager::lookup.
Reapply "[clang-repl] Fix crashing on unusable top-level declarations" (#230054) (#230742)
Drop the previously failing test lines:
```c
printf("not crashed\n");
// CHECK-NEXT: not crashed
```
A bare `return;` is sufficient for detecting the regression.
Reverts #230734.
[MLIR][CAPI][Python] Add C API and Python bindings for the remark engine (#229094)
Expose the optimization remark engine (`mlir/IR/Remarks.h`) through the
C API and the Python bindings, so clients outside C++ can enable it,
receive the reported remarks and emit their own.
### C API (`mlir-c/Remarks.h`)
- `mlirContextEnableOptimizationRemarks{,ToFile,WithCallback}` select
the categories (one regex per kind plus `all`), the emitting policy
(`all` or `final`) and the sink: MLIR remark diagnostics, a
YAML/bitstream file, or a callback invoked with an opaque `MlirRemark`.
- `mlirContextFinalizeOptimizationRemarks` flushes postponed remarks,
writes the file and drops the engine; `mlirContextHasRemarkEngine`
queries it.
- `MlirRemark` accessors (kind, remark/category/function names,
location, id, key/value arguments, print), valid for the duration of the
callback.
- `mlirEmitOptimizationRemark` emits a remark at a location through the
engine of its context.
[24 lines not shown]
[LAA] Add tests for a set of miscompiles. (NFC) (#230773)
Add tests for the following issues:
* interleaved-accesses-retry-runtime-checks.ll: when LAA retries with
runtime checks, the store of an interleave group is moved past a load
from the same address.
* is-safe-dep-distance-narrow-btc.ll: accesses with a real backward
dependence are reported as safe if the byte stride does not fit in a
narrow backedge-taken count type or MaxBTC * MaxStride wraps in it.
* is-safe-dep-distance-partial-overlap.ll: accesses are reported as
independent if |Dist| > MaxBTC * MaxStride, even if the distance is not
a multiple of the access size and the accesses partially overlap.
* different-access-types-rt-checks.ll: the runtime check bounds of a
pointer loaded as i32 and stored as i8 only account for the i8 store, so
the tail of the loaded range is not checked.
* runtime-check-grouping-wrapping-bounds.ll: pointers whose bounds can
wrap relative to each other are grouped, so the group's Low can be above
its High and the runtime checks never detect a conflict.
[VPlan] Re-use LCSSA phis when expanding AddRecs of other loops. (#230772)
Don't re-use IR values defined in loops that do not contain the plan's
entry in VPSCEVExpander::tryToReuseIRValue, as such uses outside the
loop break LCSSA. Instead, match SCEVExpander behavior and re-use an
existing LCSSA phi, reusing SCEVExpander::findReusableLCSSAPhi.
RuntimeLibcalls: Generate an enum for runtime library names
Emit an RTLIB::RuntimeLibrary enumerator for each distinct LibraryName,
and use it instead of a string for isLibraryAvailable.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
fix(x86): reuse shifted masks in SSE ANDNP
Demanded-bit simplification can bypass a shared arithmetic shift when
ANDNP only needs sign bits. Reuse the existing shift to avoid keeping
its input live across a two-address SSE instruction.
Restore the original SSE2 vector select checks.
Refs #230700
AMDGPU: Fix true16 build_vector (0, x) pattern using a 16-bit shift operand
The real true16 pattern for (build_vector 0, VGPR_16:$x) fed the 16-bit
register directly to V_LSHLREV_B32, which takes a 32-bit operand. Widen
it with a REG_SEQUENCE first. This avoids redundant 16-bit moves in
SelectionDAG, and fixes a GlobalISel selection failure when the 16-bit
input is a G_TRUNC of a 32-bit value, as the shift's operand class
constrained the trunc result to vgpr_32.
I also don't know why this pattern is overcomplicating this. I would expect
true16 to literally translate build_vector to reg_sequence plus a materialize
of the 0.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
RuntimeLibcalls: Remove the Default* libcall lists
Every system library now lists only LibcallLibrary members, so nothing
references DefaultRuntimeLibcallImpls. arm64ec's '#'-prefixed set was
still derived from the whole default list minus the width slices and
the Windows exclusions. Build it from the compiler-rt, libm and libc
slices instead, still omitting the calls the Windows runtime lacks.
Take AArch64's fp128 long double calls from the libm slice, and delete
the remaining default and Windows base lists.
The generated tables are unchanged.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
fix(X86): restore CCMP through boolean NOT
Demanded-bits simplification can turn XOR with 1 into NOT, hiding
SETCC operands from CCMP formation. Invert the condition under AND
when the other SETCC masks high bits and the NOT has one use.
Restore the original ccmp.ll checks for all four configurations.
Refs #230700