[lld-macho] Remove symbol name assumptions from category merging (#217276)
The category merger required every __objc_catlist entry to point to a
symbol named with the `__OBJC_$_CATEGORY_` or `__CATEGORY_` prefix and
hit llvm_unreachable otherwise. Such names cannot be relied upon: `ld
-r` rewrites the names of category body symbols to generated names like
`l002`, and linking its output crashes lld.
The merger also used symbol names to predict the layout of protocol
lists, which is fragile even for conventionally named inputs. The repro
https://github.com/llvm/llvm-project/pull/95124#issuecomment-4267900795
fired the "Protocol list does not match expected size" assertion.
Remove the category symbol name requirement, and drop the layout
assertion together with the SourceLanguage machinery.
[BOLT][RISCV] Implement indirect PLT calls (#219184)
This patch implements `MCPlusBuilder::createIndirectPLTCall` for RISC-V,
enabling BOLT's `--plt=hot` and `--plt=all` optimizations for RISC-V
binaries.
The PLT call pass replaces direct calls and tail calls to PLT entries
with indirect calls through the corresponding resolved GOT slot. The
generated sequence is:
auipc t3, %pcrel_hi(target at GOT)
l[dw] t3, %pcrel_lo(.Lpcrel_hi)(t3)
jalr ra, t3, 0
[WebAssembly] Expand v8f16 SELECT_CC (#218922)
Follow up for #213280 (read
https://github.com/llvm/llvm-project/pull/213280#discussion_r3797603654)
Mark `v8f16 SELECT_CC` for expansion so scalar comparison-based selects
lower through the existing comparison and `v128.select` patterns
[orc-rt] Split headers into support/ and bedrock/ layers. NFC. (#219374)
Follow-up to 8c7563a40a5b, which nested the runtime's headers under
include/orc-rt/bedrock/ and noted that library-neutral headers would
later be split back out.
The split names a layer -- who may include whom. support/ holds
vocabulary and utilities that depend on nothing else in orc-rt; bedrock/
holds the runtime components (Session, Service, the memory map, the
dylib manager, the SPS controller interfaces) and may include support/.
SPIRE will be able to include both. orc-rt-c/ gains the same layering.
Note that support/ is a layer inside the bedrock library, not a separate
one: Error.cpp and RTTI.cpp still compile into orc-rt-bedrock.
Also folded in: bedrock/sps-ci/ -> bedrock/sps/ in both include/ and
lib/; include guards derived from each header's path, as LLVM does
(ORC_RT_SUPPORT_ERROR_H); test/unit/ mirrored onto the new layout, with
cross-layer test helpers left at its root; test-target FOLDER properties
[3 lines not shown]
[mlir][SPIRV] Fix `StorageBuffer` access conversion for emulated i16 (#218693)
Follows up on commit 202ece6. In the absence of `Int16` and
`StorageBuffer16BitAccess` in the target, `i16` isn't any different from
byte & sub-byte types. As exposed by downstream smoke tests of the IREE
project, an edge case where this causes issues is a 0/1-rank memref.
Semantically:
```
memref<i16> -> ptr<struct<array<1 x i32>>>
```
Since the array lengths are the same in the absence of actual packing,
just the index bounds check doesn't catch this and `InBoundsAccessChain`
still gets chosen. In the end, the memref op fails to lower through the
same restriction in `MemRefToSPIRV` that the original change apparently
had to work around - only `AccessChain` is expected there.
As a more general criterion, the change just compares array the element
types and picks `AccessChain` upon mismatch.
[6 lines not shown]
[X86] Emit adox instead of adc for overflow add (#216609)
ADOX is like ADC but with OF instead of the CF and can only be encoded
with a pair of 32 or 64 bit regs.
Basically this applies in cases where the overflow flag is being added.
[WebAssembly] Select lane stores for floating-point vectors (#219186)
This extends the existing integer vector lane-store patterns to the
equivalent floating-point vector types. The underlying WebAssembly
instructions are type-agnostic lane stores.
That being said I think something like `STORE_LANE_I32x4_A32` can be
misleading when dealing with floating-point vectors. (should there be a
rename or something ?)
[WebAssembly] Fold offsets into extending SIMD loads (#219144)
I saw this TODO and realized that instead of lowering to
```
local.get 0
i32.const 8
i32.add
v128.load64_zero 0
f64x2.promote_low_f32x4
```
We could choose
```
local.get 0
v128.load64_zero 8
f64x2.promote_low_f32x4
```
So I used the existing WebAssembly address operand patterns when
lowering v2f32-to-v2f64 extending loads.
[3 lines not shown]
[Clang-Tidy] Improve `bugprone-implicit-widening-of-multiplication-result`. (#214501)
Implicit integer promotions make it a bit difficult to deduce the
correct type in the following expression:
```
std::uint64_t calc_array_size(std::uint16_t width, std::uint16_t height) {
return width * height;
}
```
Originally, Clang-Tidy suggested to use the following code:
```
return static_cast<long long>(width) * height;
```
It is fully correct according to the C++ rules, but it makes it a bit
harder to reason for people. This change adds a more readable "FixIt"
taking into account the source type and avoid intermediate
representations.
Co-authored-by: Dmitrii Kuragin <dkuragin at adobe.com>
[clang][OffloadBundler] Fix uninitialized iterator in BinaryFileHandler (#219346)
This patch initializes NextBundleInfo at the top of ReadHeader to
prevent an uninitialized iterator comparison.
ReadHeader has several early return points where it exits without
reading any bundles. Upon an early return, NextBundleInfo never reaches
the assignment at the bottom of ReadHeader:
NextBundleInfo = BundlesInfo.begin();
leaving NextBundleInfo default-constructed. A subsequent call to
ReadBundleStart then attempts an invalid iterator comparison:
if (NextBundleInfo == BundlesInfo.end())
where NextBundleInfo is still default-constructed.
This bug was discovered with tightened epoch checks in
[2 lines not shown]
[flang][OpenMP] Support omx/ompx extension sentinels (#218475)
This adds support for the OpenMP 5.2 extension sentinels: !$omx, c$omx,
*$omx in fixed form and !$ompx in free form. Known directives after
these sentinels are handled just like !$omp, and unknown ones are
ignored with a warning so code using vendor extensions stays portable.
Added lit tests covering fixed form, free form, and the
ignore-with-warning behavior.
Assisted-by: Claude Opus 4.6
---------
Co-authored-by: Chandra Ghale <ghale at pe34genoa.hpc.amslabs.hpecorp.net>
Co-authored-by: Krzysztof Parzyszek <Krzysztof.Parzyszek at amd.com>
[ScalarizeMaskedMemIntrin][ProfCheck] Correctly annotate branch weights (part 2) (#219286)
https://github.com/llvm/llvm-project/pull/218753 broke LLVM CI because
it added a new test in `ScalarizeMaskedMemIntrin` that was not opted out
of during profcheck. Profcheck failed because this pass creates new
branches that did not attach branch weight metadata. We don't have any
information on the distribution of masks at runtime, so we have to mark
branch weights as explicitly unknown.
This basically extends https://github.com/llvm/llvm-project/pull/181568,
Aiden am I missing something for why you didn't add the branch weight
metadata for all branch creation before?
Tested the `ScalarizeMaskedMemIntrin` tests with profcheck locally and
they all pass.
[ADT] Remove unused IDHash parameter from Equals (NFC) (#219313)
This patch removes the unused IDHash parameter from several functions.
Now that FoldingSetTrait<SDVTListNode>::Equals no longer checks IDHash,
no implementation of Equals uses this parameter.
Assisted-by: Antigravity
[VPlan] Append recipes created via builder to worklist
The previous PR appended the top most created recipe to the worklist, and this PR extends it to any other nested recipes that were created, similar to InstCombine.
This removes the header mask in a good few more places on RISC-V as measured on SPEC CPU 2017, e.g. for the following loop:
```c
long f(const int *p, const int *q, long n) {
long a = 0, b = 0;
for (long i = 0;; i++) {
if (p[i] && q[i]) { a += i; b += i; }
if (i + 1 == n) break;
}
return a + b;
}
```
Before:
[49 lines not shown]
[VPlan] Process simplifyRecipes in a worklist
This brings simplifyRecipes further in line with InstCombine, and asides from unlocking more simplifications it also helps avoid spurious test churn whenever passes are moved around simplifyRecipes.
For now just push the new recipe onto the worklist, not its users.
This uses a post order traversal so we maintain the same simplification order as before.
I've gone through and checked every simplification we do is a canonicalisation that converges, and I checked on llvm-test-suite + SPEC CPU 2017 in various configurations that we don't hit any cycles.
[flang] Speed up large CHARACTER DATA initializers (#218813)
[flang] Speed up large CHARACTER DATA initializers
Repeated CHARACTER(KIND=1) array constants were lowered as one
fir.insert_value per element. Converting those chains to LLVM IR is
quadratic and can make compilation take tens of minutes.
Lower consecutive equal KIND=1 character elements with
fir.insert_on_range
and emit full-range initializers as a single flattened [N x i8] LLVM
global
string, keeping Fortran blank padding.
A 160000-element character DATA statement now compiles in well under a
second and before was more than 10 minutes.
[AST] Make err_struct_too_large check target-aware (#218749)
ASTContext::getASTRecordLayout used a fixed 1ULL << 60 threshold for
err_struct_too_large, regardless of the target's size_t width.
Scale the threshold to the target's size_t width instead, so it is below
(1 << 32) on 32-bit architectures. Diagnosing the overflow in Sema
avoids the crash in codegen.
rdar://183351516
[libc] Add stubs for POSIX netdb.h and getaddrinfo (#219337)
* Add the `<netdb.h>` POSIX header and declare `struct addrinfo` and
`freeaddrinfo` and `getaddrinfo` methods
as defined in
https://pubs.opengroup.org/onlinepubs/9799919799/functions/getaddrinfo.html
;
* Provide Linux-specific definitions for `EAI_` macro family;
* Add header/entrypoints to the list of "experimental" (i.e. WIP)
entrypoints on Linux systems;
* Create the proxy header harness for types / Linux-specific macro.
* Provide stub implementations - no-op `freeaddrinfo` and `getaddrinfo`
that returns `EAI_SYSTEM` and sets errno to `ENOSYS`. Validate this
behavior in unit tests.
Assisted by automated tooling, human-reviewed