[BOLT] Keep ambiguous references next to function boundaries valid
A reference into code without a relocation that names its target, e.g. a
RIP-relative LEA whose relocation was against a section symbol, cannot be told
apart from "Next - Delta" and "Prev + Offset" when it lands right before or
after a function start. HHVM built with LLVM 23 has such a reference to
"sqlite3RCStrUnref - 1", which lands in padding and makes BOLT fail with
-strict. Without padding, BOLT silently kept such references relative to the
preceding function. See bolt/test/X86/unanchored-code-reference.s.
BOLT now collects these references from code and, when they are within
--boundary-ref-distance bytes (default 2) of a function start, keeps the
functions around them in place in lite mode. When all functions are processed,
it emits those functions unoptimized and back-to-back with the original bytes
between them, and checks after linking that they kept their size and distance.
Absolute references right before a function start are now relative to that
function. This replaces the workaround for "fptr - 1" in de-virtualized member
function pointer calls, which dropped the relocation and left the input
address in the code.
[BuiltinsX86] Use long long for the __int64 MS intrinsics (#154946)
Microsoft declares _Interlocked*64, _xgetbv and _xsetbv with __int64,
which is
long long on every target. BuiltinsX86.td used int64_t and uint64_t,
which are
long on LP64, so declaring them as Microsoft does fails there:
error: conflicting types for '_InterlockedIncrement64'
Use long long int and unsigned long long int, as __emul does.
[docs] Enable absolute documentation link checks
Configure the LLVM documentation URL prefixes so the Sphinx build rejects
absolute links to documents in the same project. Clang is already configured.
Part of #214861
[lldb][docs] Use project-local documentation links
Replace same-project absolute URLs with relative Markdown links and Sphinx
cross-references so local and archived documentation stays self-contained.
Update generated Python API docstrings at their header and SWIG inputs,
and repair stale Python and frame-recognizer destinations.
Validation: docs-lldb-html and docs-lldb-man, followed by fresh Sphinx
rebuilds (-E), on a merge of the three independent self-link fixes with
the checker enabled and warnings-as-errors disabled. No self-link warnings
or new warning messages compared with the audit baseline.
C/C++ formatting: git-clang-format --diff against main for the three
modified API headers reported no changes.
Part of #214861
Assisted-by: Codex
[SimplifyLibCalls] Mark profiles unknown in optimizeMemRChr (#230289)
The selects created here are purely synthesized so we have no way of
preserving weights in the general case. Thus, mark the weights as
unknown.
[libc++][docs] Use project-local links in release notes (#228624)
Replace absolute libc++ homepage links with Sphinx document references
in
release notes 20 through 24 and the release-note template.
I would leave these historical documents alone, but I have to fix these,
or the doc build will fail with warnings when I enable the absolute
self-link Sphinx doc build warning.
Validation: docs-libcxx-html, followed by a fresh Sphinx rebuild (-E),
on
a merge of the three independent self-link fixes with the checker
enabled
and warnings-as-errors disabled. No self-link warnings or new warning
messages compared with the audit baseline.
Part of #214861
Assisted-by: Codex
[Clang] Fix cluster_dims thread block limit checks
The limit is on the total number of thread blocks, 8 for NVPTX and 16
for AMDGPU, but two checks got in the way. A 4-bit cap on each dimension
rejected cluster_dims(16, 1, 1) while accepting the same 16 blocks as
cluster_dims(4, 4, 1), and an omitted dimension contributed 0 rather
than 1, zeroing the product so cluster_dims(15, 15) compiled cleanly.
Bound each dimension only enough to keep the product from overflowing
and let the existing total size diagnostic enforce the limit. Code that
relied on the unenforced limit now fails to compile.
[InstCombine] Fix profile propagation in visitMul (#230304)
Within visitMul there is a fold that creates a new select with the same
condition as an existing select, which lets us directly propagate the
metadata.
[CIR] Don't double-init global variables anymore. (#230279)
My last patch, #228586, did some additional work to make sure we
properly initialized globals even if they had multiple blocks in their
initializer rather than assuming there was one.
As a result, the dynamic initialization of a variable referenced
multiple times regressed, as we started failing verification because we
are double-emitting the initializer.
Note we ALWAYS double-emitted the initializer, but we replaced it every
time. The change above caused us to re-use the same set of blocks to
better tolerate multiple blocks, and as a side effect, the second block
initialization added instead of replacing.
This patch fixes this by making sure we only define a global variable if
it isn't yet defined.
[docs] Repair remaining absolute self-documentation links (#228625)
Replace same-project absolute URLs with relative source links or Sphinx
cross-references in LLVM, Flang, libc, clang-tools-extra, and OpenMP.
Use the explicit LangRef label for atomic.ignore.denormal.mode metadata.
This lets Sphinx validate the targets and keeps local and archived
documentation self-contained.
Validated by building all docs-*-html targets for these subprojects with
the pending Sphinx build warning merged in.
Part of #214861
Assisted-by: Codex
[flang][semantics] Reject DATA-style initializer on EXTERNAL/INTRINSIC (#222256)
An `entity-decl` carrying a legacy `/initialization/` (an extension) was silently
accepted, and the initializer dropped, when the name had already been
declared `EXTERNAL` or `INTRINSIC`:
```fortran
subroutine s
external foo
integer foo /1/ ! accepted, no initialization emitted
end subroutine
```
`DataChecker::Leave(const parser::EntityDecl &)` passed the value list to
`AccumulateDataInitializations` without checking the symbol's class, so the
initialization never reached the emitted code and no diagnostic was produced.
Reversing the two statements already errors, so the behaviour was also
order-dependent. For an intrinsic name that is not an unrestricted specific
function (`sum`), the accumulated value instead tripped
[9 lines not shown]
[orc-rt] Mark LockedAccess's constructor as noexcept. (#230311)
The ORC runtime does not use exceptions for errors internally, so mark
this noexcept. In practice the mutex and lock types we use (STL mutexes
and locks) don't throw unless corrupted or misconfigured. If we ever
want to support locks whose acquisition can legitimately fail we'll need
an alternative Error-based representation of failure.
[AMDGPU] Use vector memory types for 64/128-bit cooperative atomics (#229501)
The 16x8B and 8x16B cooperative atomic intrinsics had integer memory
types (`i64`/`i128`) that didn't match their vector values, which broke
value tracking on the loaded value and could crash the compiler. This
patch uses the value type as the memory type and updates the selection
patterns to match. Codegen for existing cooperative atomics is
unchanged.
Don't let the ThreadPlanSingleStepTimeout interrupt an already stopped process (#227898)
When the ThreadPlanSingleThreadTimeout timeout fires, first check
whether the process is stopped before sending the interrupt.
This is a simpler way to address the problem that was identified in:
https://github.com/llvm/llvm-project/pull/224272
The problem solved there is that if the stop processing is still going
on when the timer fires, we send the interrupt request event which gets
enqueued and then handled after the stop event processing has restarted
the inferior to continue the thread plan work.
That patch involved trying to figure out, when you receive the event,
whether it was stale or not. It was harder to reason about - and not all
the way right, though that probably could be fixed.
But when you get a stop event, we first set the private state to stopped
[6 lines not shown]
RuntimeLibcalls: Require system library members to be libraries
Every SystemRuntimeLibrary now lists only LibcallLibrary and LibraryRef
members, so the inline path that expanded unhomed RuntimeLibcallImpl members
directly into the system's block, including its SystemAvailableImpls
bitset, is dead. Remove it, and error on any member that is not a
library. The generated RuntimeLibcalls.inc is unchanged.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[Analysis] Don't treat code after analyzer_noreturn calls as unreachable (#229577)
Since #150952, the CFG treats calls to functions attributed
'analyzer_noreturn' like calls to 'noreturn' functions, so
-Wunreachable-code reported the code following such a call as never
executed even though the call can return.
Record in CFGBlock whether a noreturn block ends in a real 'noreturn'
call or an 'analyzer_noreturn' call. The CFG keeps the code after an
'analyzer_noreturn' call as the alternate successor of the exit edge and
the reachability scan used by -Wunreachable-code follows it. All other
clients still treat the two kinds identically.
rdar://188741139
[SLP]Fix crash on fused alternate node with reuse shuffles
Size the vector type by the unique scalars, same as the opcode mask;
the reuse-shuffle vectorization factor made them mismatch.
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/230303
[CIR] Pass x87 long double vectors on x86_64
Vectors of x87 long double now go through x86_64 calling-convention
lowering instead of hitting NYI. The ABI library sizes their elements at
128 bits like clang does, so the signatures match classic codegen.
Unions are still moved as a value of their storage type, and a long
double stores only 10 of its 16 bytes. So a union holding an x87 value
next to another member stays NYI unless it's a plain long double and the
other members fit in those 10 bytes. That also stops a silent miscompile
of unions like `union { long double ld; char c[16]; }`.
Assisted-by: Cursor / Claude Opus 5.5