[flang][Driver] Enable bare -O flag alias when specified with -flto (#228483)
Currently, `O_flag` (`-O` as an alias for `-O1`) in Options.td is not
visible to FlangOption. As a result, invoking flang with a bare `-O`
(without a trailing digit) is treated by the joined `-O` definition as
`-O""`, causing an error when forwarded to the linker.
Add `FlangOption` to `O_flag`'s Visibility so that `flang -O` correctly
aliases to `-O1`.
Assisted-by: IBM Bob
Resolve https://github.com/llvm/llvm-project/issues/227474
[scudo] Validate list endpoints and links before removing a node (#229245)
`DoublyLinkedList::remove()` currently modifies the predecessor's link
before validating the successor's reciprocal link, and checks endpoint
consistency only in debug builds. A corrupted successor can therefore be
rejected after a list write has already occurred, while inconsistent
null links can bypass endpoint checks in release builds.
Check that the list is nonempty, enforce endpoint consistency in release
builds, and validate both reciprocal links before modifying any list
state. The production change is confined to `remove()`. This strengthens
consistency checks; it does not authenticate list membership or protect
against mutually consistent forged interior nodes.
Tests cover all 24 removal orders of four nodes, plus empty-list
removal, both directions of endpoint inconsistency, and corrupted
reciprocal links, using pointer and index links.
Local validation on Linux x86_64 in WSL, Clang 21.1.8:
[23 lines not shown]
[AMDGPU] Require wave ID masks to be constants (#229917)
PR #177713 added some cases to the wave ID recognizer, such
as (ThreadID & Mask) >> log2(WaveSize) but, after it landed, Claude
noticed a bug. `Mask` in those types of expressions was an arbitrary
vale, which could be divergent, meaning that the "uniform" value of
that expression wouldn't actually be uniform.
The quick fix that preserves most of the cases we're worried about in
practice is to restrict these patterns to constant masks.
AI disclosure: Claude found the issue and created the patch, I wrote
this message.
Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply at anthropic.com>
[AMDGPU][NFC] Pre-commit tests to not match non-constant wave ID masks (#229916)
Non-constant masks can cause divergences even if we're shifting the
divergent part of a thread ID, and we didn't account for that in the
pattern matches.
AI disclosure: Claude found these and wrote the tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply at anthropic.com>
Use block identity to find escapes in createIteratorLoop
createIteratorLoop decided that a branch target was outside the loop body
by comparing block numbers with a snapshot taken before the body generator
ran. Block numbers can change: renumbering tripped an assertion, and a
pre-existing block moved within the function got a new number and was
treated as part of the body.
Record the blocks that exist before the body generator runs instead, and
reject branches to any of them other than the body and the latch. Add tests
for renumbering, a moved block, and a callback-created block that branches
to the loop exit.
[Mips][FastISel] Also add low part when materializing private functions (#229717)
I have no idea why the code was explicitly written to exclude private
functions; it seems obviously incorrect to me.
[Clang][AArch64] Command-line options for A-profile's Return Address Authentication Hardening (#176171)
This patch introduces a new command-line option to enable the AArch64
A-profile's Return Address Authentication Hardening. It also introduces
a new function attribute with the same naming as the new command-line
option.
At the time of this patch, this new option enables the hardening against
the PACMAN attack [1] using a load from the return address [2].
The new option, -mharden-pac-ret, can take one of two values:
- none: disable hardening. (The default if the option is absent)
- load-return-address: enables hardening using the mitigation based on
load from return address.
The corresponding function attribute takes the option and its possible
values using the same naming. Also, the function attributes take
precedence over the command-line options.
[6 lines not shown]
[lldb] Update formatter_bytecode.py to ABI version 2 (#229923)
Update the assembler and the Python to bytecode compiler to target the
version 2 formatter ABI, in which `self` is a Dictionary owned by LLDB
and passed as the first argument to every synthetic method.
The primary feature of this change is the separation of storage, locals
are on the stack, attributes are stored in the `self` Dictionary. Since
locals and attrs no longer intermingle in the stack, local variables are
now allowed in `__init__` and `update`, and attributes can be
reassigned.
In the compiler:
* `self.x = expr` is compiled to `0 pick "x" <expr> dict_set`
* `self.x` is compiled to `0 pick "x" dict_get`
* Arguments and local variables are on the stack above `self`
Assisted-by: claude
[OpenMP][OMPIRBuilder] Allow multi-block bodies in createIteratorLoop
createIteratorLoop requires the body generator to leave a single block that
falls through or branches to the loop latch. A body generator that lowers
expressions with their own control flow cannot meet that requirement. The
upcoming user is omp.iterator translation, once its region can contain
several blocks (#227454).
Allow the body to span several blocks. If exactly one block is left without
a terminator, it is branched to the latch; otherwise some block must already
branch there. Bodies that leave more than one block unterminated, or that
branch out of the loop, are rejected with an error.
Assisted with Copilot and Claude Opus 5.
[InstCombine] Fold fdiv by splat of pow/exp/powi into fmul (#227238)
Extend foldFDivPowDivisor to look through a one-use splat divisor. The
exponent is negated on the scalar and the result is splatted again:
$$\frac{Z}{\mathrm{splat}(x^{y})} \to Z \cdot \mathrm{splat}(x^{-y})$$
$$\frac{Z}{\mathrm{splat}(e^{y})} \to Z \cdot \mathrm{splat}(e^{-y})$$
$$\frac{Z}{\mathrm{splat}(2^{y})} \to Z \cdot \mathrm{splat}(2^{-y})$$
$$\frac{Z}{\mathrm{splat}(x^{n})} \to Z \cdot \mathrm{splat}(x^{-n}),
\quad n \in \mathbb{Z}\ (\mathrm{powi})$$
Same FMF requirements as the scalar fold: reassoc and arcp, plus ninf
for powi. This removes the reciprocal, e.g. v_rcp on AMDGPU.
AMDGPU example: https://godbolt.org/z/nvWc38GTx
[orc-rt] Add SymbolLookupFlags and SymbolLookupSet utilities (#229958)
Move NativeDylibManager's nested LookupFlags enum and SymbolLookupSet
typedef out into reusable support utilities:
- SymbolLookupFlags.h: enum class SymbolLookupFlags, with the existing
RequiredSymbol and WeaklyReferencedSymbol values.
- SymbolLookupSet.h: SymbolLookupSet and SymbolLookupResult, thin
wrappers around std::vector<std::pair<std::string, SymbolLookupFlags>>
and std::vector<std::optional<void*>> respectively.
- sps/SPSSymbolLookupSet.h: SPS serialization for all three. The flags
serialization (previously private to NativeDylibManagerSPSCI.cpp, and
duplicated in its unit test) is unchanged: RequiredSymbol serializes as
true, matching llvm::orc::RemoteSymbolLookupSetElement's 'Required'
field.
NativeDylibManager::lookup now takes a SymbolLookupSet and reports a
SymbolLookupResult, and sys::lookupLibrarySymbols takes the
SymbolLookupSet directly, saving lookup a copy of the names.
Assisted-by: Claude
[CodeGen] Drop dead SlotIndexes before allocation
SlotIndexes keeps the index list entry of an erased instruction and only
clears its instruction pointer. Live range sizes are measured in slot
indexes and greedy ranks ranges by size, so the leftovers inflate some
ranges more than others and reorder allocation, spilling heavily on
register-starved functions.
Add SlotIndexes::compactIndexes() to erase them. Erased entries are
unlinked, so LiveIntervals first reports the indexes it holds via
appendReferencedIndexes().
Off by default behind -greedy-compact-slot-indexes, since it changes
allocation across much of the test suite.
[lldb] Don't fail initialization when Python can't be loaded (#229891)
SystemInitializerFull::Initialize returns early when the Python runtime
loader fails, preventing remaining plugins (like the platform) to be
initialized.
Both with the statically and dynamically linked plugins, not finding
Python at that point is non-fatal. The error can be more appropriately
handled later, when initializing the plugin as all the subsystems (i.e.
logging) have been initialized.
rdar://188858954
[BoundsSafety] Handle the counted_by family as a type attribute
counted_by / sized_by (and their _or_null variants) were handled as a
declaration-position attribute through handleCountedByAttrField.
Build the type during type construction instead and properly handle as a
type-position attribute. Building the node in type position also allows
the new way of late-parsing for type attributes: creates a type node
in place and fills it in once late parsing is done, which will be
enabled in the follow-up commits.
Assisted-by: Opus 4.8
[CIR] Classify reference-to-pointer catch parameters in CIRGen (#229579)
When a handler catches a reference to a pointer, __cxa_begin_catch
returns the caught pointer by value instead of the address of the
exception object, and how the reference is bound depends on whether the
pointer points to a class. The CIR type of the catch parameter cannot
always tell whether it is a reference to a pointer, or whether the
pointee is a class. CIRGen now records both facts from the AST in two
new InitCatchKind values, reference_to_pointer and
reference_to_record_pointer. The Itanium EH lowering switches on the
kind alone, the way classic codegen does.
Assisted-by: Claude Code / Claude Opus 5.5
[libcxx][test-support] Improve thread_unsafe_shared_ptr (#195932)
Adapting test support class `thread_unsafe_shared_ptr` for future use
with `fancy_pointer_allocator`, in particular, in constant evaluation.
* Implementing the rule of 5;
* Annotating `noexcept` methods
[AArch64] Preserve store width for ptr32 address spaces (#229181)
When handling stores for `ptr32` address spaces, we didn't check if they
were already truncating stores and emitted an ordinary store. This
widened `i8` and `i16` writes to 32 bits and could overwrite adjacent
fields.
Fix is to build a truncating store that preserves the memory type.
Fixes #228782
Assisted-by: GPT-6 Sol (via VS Code)
[lldb] Cache the symbol table for memory-read modules by default (#229305)
When lldb debugs a process on a remote system/device, there may be
binaries loaded in the process that it cannot find on the debug host. It
will read the object file header / load commands (mach-o terminology)
and symbol table out of memory at ObjectFile/Module creation time. The
symbol table may involve many separate reads to complete, and can be a
serious performance issue. We've traditionally worked hard to always
have binaries for the remote target present on the host computer,
because it was such a poor user experience, but in practice it's never
perfect.
This PR adds a new setting,
`symbols.enable-lldb-index-cache-memory-modules`, which is defaulted to
on. With this change, when lldb reads a remote binary out of memory, it
will create a DataFileCache symtab serialization of its symbol table on
the host computer.
I'm open to discussions about changing the expiration-days default. I
[35 lines not shown]
[libc++][test] Remove spurious `REQUIRES: has-fblocks` (#227972)
As drive-by, also update the lengthy `UNSUPPORTED` lit comment to
`REQUIRES: std-at-least-c++26`.
[HLSL] Allow Interlocked original values with different types (#228262)
Fix HLSL `Interlocked*` overloads to treat `original_value` as an `out`
parameter. This permits writeback conversions, including signedness
changes.
Add semantic and code generation regression tests.
Fixes #224151.
Assisted by: Github Copilot
---------
Copilot-Session: f6814fd7-47e6-4edf-af68-4a69514a2212
[AMDGPU] Price vector f32 to f16 fptrunc by its packing form
The base cost scalarizes the conversion and charges 4, 10, 22 and 46
for 2, 4, 8 and 16 lanes. The backend rounds every lane with
v_cvt_f16_f32 and packs the halves in pairs, which takes N + N/2
instructions. A packed conversion rounds a pair per instruction and
true16 writes a lane into either half of a register, which gives
ceil(N/2) and N.
Co-authored-by: Michael Selehov <michael.selehov at amd.com>
Assisted-By: Claude Code Opus 5