[InstCombine] Fold fdiv by splat of pow/exp/powi into fmul (#227238)
Extend foldFDivPowDivisor to look through a one-use splat divisor. The
exponent is negated on the scalar and the result is splatted again:
$$\frac{Z}{\mathrm{splat}(x^{y})} \to Z \cdot \mathrm{splat}(x^{-y})$$
$$\frac{Z}{\mathrm{splat}(e^{y})} \to Z \cdot \mathrm{splat}(e^{-y})$$
$$\frac{Z}{\mathrm{splat}(2^{y})} \to Z \cdot \mathrm{splat}(2^{-y})$$
$$\frac{Z}{\mathrm{splat}(x^{n})} \to Z \cdot \mathrm{splat}(x^{-n}),
\quad n \in \mathbb{Z}\ (\mathrm{powi})$$
Same FMF requirements as the scalar fold: reassoc and arcp, plus ninf
for powi. This removes the reciprocal, e.g. v_rcp on AMDGPU.
AMDGPU example: https://godbolt.org/z/nvWc38GTx
[orc-rt] Add SymbolLookupFlags and SymbolLookupSet utilities (#229958)
Move NativeDylibManager's nested LookupFlags enum and SymbolLookupSet
typedef out into reusable support utilities:
- SymbolLookupFlags.h: enum class SymbolLookupFlags, with the existing
RequiredSymbol and WeaklyReferencedSymbol values.
- SymbolLookupSet.h: SymbolLookupSet and SymbolLookupResult, thin
wrappers around std::vector<std::pair<std::string, SymbolLookupFlags>>
and std::vector<std::optional<void*>> respectively.
- sps/SPSSymbolLookupSet.h: SPS serialization for all three. The flags
serialization (previously private to NativeDylibManagerSPSCI.cpp, and
duplicated in its unit test) is unchanged: RequiredSymbol serializes as
true, matching llvm::orc::RemoteSymbolLookupSetElement's 'Required'
field.
NativeDylibManager::lookup now takes a SymbolLookupSet and reports a
SymbolLookupResult, and sys::lookupLibrarySymbols takes the
SymbolLookupSet directly, saving lookup a copy of the names.
Assisted-by: Claude
[CodeGen] Drop dead SlotIndexes before allocation
SlotIndexes keeps the index list entry of an erased instruction and only
clears its instruction pointer. Live range sizes are measured in slot
indexes and greedy ranks ranges by size, so the leftovers inflate some
ranges more than others and reorder allocation, spilling heavily on
register-starved functions.
Add SlotIndexes::compactIndexes() to erase them. Erased entries are
unlinked, so LiveIntervals first reports the indexes it holds via
appendReferencedIndexes().
Off by default behind -greedy-compact-slot-indexes, since it changes
allocation across much of the test suite.
[lldb] Don't fail initialization when Python can't be loaded (#229891)
SystemInitializerFull::Initialize returns early when the Python runtime
loader fails, preventing remaining plugins (like the platform) to be
initialized.
Both with the statically and dynamically linked plugins, not finding
Python at that point is non-fatal. The error can be more appropriately
handled later, when initializing the plugin as all the subsystems (i.e.
logging) have been initialized.
rdar://188858954
[BoundsSafety] Handle the counted_by family as a type attribute
counted_by / sized_by (and their _or_null variants) were handled as a
declaration-position attribute through handleCountedByAttrField.
Build the type during type construction instead and properly handle as a
type-position attribute. Building the node in type position also allows
the new way of late-parsing for type attributes: creates a type node
in place and fills it in once late parsing is done, which will be
enabled in the follow-up commits.
Assisted-by: Opus 4.8
[CIR] Classify reference-to-pointer catch parameters in CIRGen (#229579)
When a handler catches a reference to a pointer, __cxa_begin_catch
returns the caught pointer by value instead of the address of the
exception object, and how the reference is bound depends on whether the
pointer points to a class. The CIR type of the catch parameter cannot
always tell whether it is a reference to a pointer, or whether the
pointee is a class. CIRGen now records both facts from the AST in two
new InitCatchKind values, reference_to_pointer and
reference_to_record_pointer. The Itanium EH lowering switches on the
kind alone, the way classic codegen does.
Assisted-by: Claude Code / Claude Opus 5.5
[libcxx][test-support] Improve thread_unsafe_shared_ptr (#195932)
Adapting test support class `thread_unsafe_shared_ptr` for future use
with `fancy_pointer_allocator`, in particular, in constant evaluation.
* Implementing the rule of 5;
* Annotating `noexcept` methods
[AArch64] Preserve store width for ptr32 address spaces (#229181)
When handling stores for `ptr32` address spaces, we didn't check if they
were already truncating stores and emitted an ordinary store. This
widened `i8` and `i16` writes to 32 bits and could overwrite adjacent
fields.
Fix is to build a truncating store that preserves the memory type.
Fixes #228782
Assisted-by: GPT-6 Sol (via VS Code)
[lldb] Cache the symbol table for memory-read modules by default (#229305)
When lldb debugs a process on a remote system/device, there may be
binaries loaded in the process that it cannot find on the debug host. It
will read the object file header / load commands (mach-o terminology)
and symbol table out of memory at ObjectFile/Module creation time. The
symbol table may involve many separate reads to complete, and can be a
serious performance issue. We've traditionally worked hard to always
have binaries for the remote target present on the host computer,
because it was such a poor user experience, but in practice it's never
perfect.
This PR adds a new setting,
`symbols.enable-lldb-index-cache-memory-modules`, which is defaulted to
on. With this change, when lldb reads a remote binary out of memory, it
will create a DataFileCache symtab serialization of its symbol table on
the host computer.
I'm open to discussions about changing the expiration-days default. I
[35 lines not shown]
[libc++][test] Remove spurious `REQUIRES: has-fblocks` (#227972)
As drive-by, also update the lengthy `UNSUPPORTED` lit comment to
`REQUIRES: std-at-least-c++26`.
[HLSL] Allow Interlocked original values with different types (#228262)
Fix HLSL `Interlocked*` overloads to treat `original_value` as an `out`
parameter. This permits writeback conversions, including signedness
changes.
Add semantic and code generation regression tests.
Fixes #224151.
Assisted by: Github Copilot
---------
Copilot-Session: f6814fd7-47e6-4edf-af68-4a69514a2212
[AMDGPU] Price vector f32 to f16 fptrunc by its packing form
The base cost scalarizes the conversion and charges 4, 10, 22 and 46
for 2, 4, 8 and 16 lanes. The backend rounds every lane with
v_cvt_f16_f32 and packs the halves in pairs, which takes N + N/2
instructions. A packed conversion rounds a pair per instruction and
true16 writes a lane into either half of a register, which gives
ceil(N/2) and N.
Co-authored-by: Michael Selehov <michael.selehov at amd.com>
Assisted-By: Claude Code Opus 5
[BoundsSafety] Handle the counted_by family as a type attribute
counted_by / sized_by (and their _or_null variants) were handled as a
declaration-position attribute through handleCountedByAttrField.
Build the type during type construction instead and properly handle as a
type-position attribute. Building the node in type position also allows
the new way of late-parsing for type attributes: creates a type node
in place and fills it in once late parsing is done, which will be
enabled in the follow-up commits.
Assisted-by: Opus 4.8
[LAA] Avoid stray predicates from replaceSymStrides (#216350)
To avoid stray predicates being added from the use of
replaceSymbolicStridesSCEV, simply add an optional Predicates argument
to be filled in when the caller doesn't want fresh predicates to be
added to PSE directly. This allows us to drop some stray predicates when
vectorizing.
TableGen: Allow IsTruncStore predicates on atomic PatFrags (#229687)
Atomic stores can be truncating in the same way as regular stores. Allow
IsTruncStore and IsNonTruncStore to be set on atomic store PatFrags so
patterns can check the store is not truncating without needing to specify
a fixed memory size.
I still find the hierarchy of load/store PatFrags frustrating. This would be
easier if we fixed the legacy mistake of treating atomic store as an
"atomic" rather than a store.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[clangd] Extract to function: Do not reject unconditionally in C files
If the parameters can be passed by value, extraction in C files works
the same way as for C++. Otherwise, due to the lack of references, a
pointer parameter is used instead, with the call site taking its address
and every use inside the extracted body rewritten into a dereference.
Assisted-by: Claude Code
Closes https://github.com/clangd/clangd/issues/1810
[lldb] Avoid UB in ansi::TrimAtWordBoundary() (#229324)
`UtilityTest` fails on Windows in the debug configuration when
`AnsiTerminal.TrimAtWordBoundary()` calls `ansi::TrimAtWordBoundary()`
with Unicode strings containing bytes with the high bit set, which leads
to `std::isspace()` being called with a negative value. C11 §7.4
paragraph 1 requires the argument to be representable as unsigned char
or equal to EOF. Passing any other value is UB.
[clang][OpenMP] Fix assertion in vtable registration for incomplete types (#229521)
Assertion failure, "queried property of class with no definition", when
mapping a pointer to incomplete type. Function 'emitAndRegisterVTable()'
assumed types would always be complete when mapped. When it's not (as in
it has no definition), there is no vtable to emit/register, so it should
just be skipped rather than assert fail.
[libcxx] Fix indirect includes for nolocale (#229601)
The indirect include tests for several headers were failing when turning
on filesystem for LLVM-libc. This turned out to be because the nolocale
header was leaking ctime and other headers when building outside of
library mode. This PR just moves all the includes under the library
guard.
[InstCombine] Move sext(trunc(lshr(...))) combine before same type combine (#227322)
The same type sext(trunc(...)) combine will fire on the lshr case, but
the lshr produces one less shift. Move it afterwards to give the lshr
combine a chance to fire first.
Doing this fold early improves codegen for this pattern seen in x264:
```c
int f(short x) {
x = (x + 32) >> 6;
return x;
}
```
- RISC-V: https://godbolt.org/z/djWza678o
- AArch64: https://godbolt.org/z/adbv49Yx5
- X86: https://godbolt.org/z/azsWKW9bq
[2 lines not shown]