[BOLT] Match mmap events by build id (#229596)
Profiling on a host running two versions of the target binary with `perf
-a` might result in a profile containing mmap events from the desired
target and the second copy of that binary. In this case, perf2bolt will
attribute samples from both processes to the target binary. Match by
build id to filter samples.
Co-authored-by: Alexey Moksyakov <moksyakov.alexey at huawei.com>
[flang][cuda][semantics] Reject unguarded host-only and device-only calls in CUDA HOST,DEVICE procedures (#228176)
Previously, semantic checking allowed host-only and device-only calls
throughout a HOST,DEVICE procedure. This change requires unguarded
callees to support both its host and device versions.
A condition that checks the intrinsic ON_DEVICE permits host-only and
device-only calls in its guarded arms. The allowance includes one-line
IF statements, nested branches, and USE aliases. In an IF/ELSE IF chain,
a guard also applies to subsequent ELSE IF and ELSE arms. Guards
established inside an arm remain local to that arm.
When an ON_DEVICE arm, or its associated plain ELSE arm, contains an
unconditional RETURN, the allowance extends to the continuation after
the IF. Unconditional BLOCK constructs propagate these guards. A RETURN
in an unrelated IF/ELSE IF arm does not establish a continuation guard.
Unrelated branches restore their incoming checking context while
retaining an enclosing guard.
[37 lines not shown]
[Flang] Generalize intrinsic CUBLAS USE association (#220024)
This PR generalizes the work done in
[PR#217455](https://github.com/llvm/llvm-project/pull/217455) to work
for the entire cublas library.
[flang] Diagnose and optionally repair missing MODULE procedure prefixes (#220783)
This change diagnoses submodule procedure definitions that match a
separate module procedure interface but omit the required `MODULE`
prefix.
Without the extension, Flang leaves an unprefixed definition as a local
procedure. `-Wportability` or `-pedantic` diagnoses a likely missing
prefix without changing the program.
The new `-fimplicit-module-prefix` option provides an opt-in
compatibility mode that binds a matching definition in the current
source to its ancestor module interface. Definitions read from module
files retain the interpretation chosen when those files were produced.
`-Wimplicit-module-prefix` or `-pedantic` reports when the repair is
applied; the diagnostic can be suppressed without disabling the repair.
Since this extension cannot distinguish an omitted prefix from an
intentionally local procedure with the same name as an ancestor
[8 lines not shown]
[clang-repl] Handle mllvm args for clang-repl before calling executeAction (#197133)
Fixes #178139
This work is built on top of discussions done in
https://github.com/llvm/llvm-project/pull/132670
Making a fresh PR cause the branch is a a year old and a bit outdated.
### Problem Statement : This is a 3 part problem statement. Imagine our
compiler builder args look like this
```
CB.SetCompilerArgs({
"-v",
"-xc++",
"-std=c++23",
"-Xclang", "-iwithsysroot/include/compat",
"-fwasm-exceptions",
"-mllvm", "-wasm-enable-sjlj"
});
[118 lines not shown]
[X86] Make EXTRQI/INSERTQI SDTypeProfile accurate and enable verifyTargetNode (#229655)
The type profile says the type needs to be v2i64, but it is created with
v16i8, v8i16, or v2i64 type.
Relax the type profile to allow this and add extra isel patterns to make
isel continue to work.
I judged that this was easier than adding bitcasts.
[OpenMPOpt][PGO] Set entry count on merged parallel wrapper (#221654)
When OpenMPOpt merges adjacent parallel regions, it creates a new
wrapper (*..omp_par) that calls the original outlined callbacks through
a single __kmpc_fork_call.
Preserve profiling information for the merged wrapper by assigning it
the
largest function_entry_count from the original callbacks. Also preserve
callsite counts for calls to the original callbacks, since SamplePGO may
otherwise treat them as cold once the wrapper itself has profile data.
[flang] Reject MODULE prefix in abstract interface body (#212980)
Issue: Flang accepts a MODULE prefix on a function or subroutine inside
an ABSTRACT INTERFACE. C1547 allows MODULE only on a module subprogram,
or on a nonabstract interface body in a module or submodule.
Fixes #208200
Root cause
- BeginSubprogram() only enforced the module/submodule half of C1547. An
abstract interface inside a module or submodule still has that scope, so
the check passed.
- In PushSubprogramScope(), an abstract interface body takes the
isAbstract() path, sets ABSTRACT, and never looks at hasModulePrefix, so
the prefix was ignored with no diagnostic.
```
if (isAbstract()) {
SetExplicitAttr(*symbol, Attr::ABSTRACT);
} else if (hasModulePrefix) {
[21 lines not shown]
[Flang][HLFIR] Lower PACK(array, .TRUE.) to hlfir.reshape (#220860)
When the PACK mask is the compile-time scalar .TRUE., the result is
equivalent to RESHAPE(array, [SIZE(array)]). Lower all PACK calls to
hlfir.pack, then simplify pack(array, .TRUE.) to hlfir.reshape with an
index-typed SHAPE temporary in SimplifyHLFIRIntrinsics. Variable or
array masks, VECTOR, and non-trivial or polymorphic operands continue to
use the runtime PACK path.
Add regression tests for character and real(k) PACK with scalar .TRUE.
Relanding the patch #213603 after fixing the regression
---------
Co-authored-by: Cursor <cursoragent at cursor.com>
Add various defaulted postfix operator overloads with restrictions and error checks
Introduce new structures demonstrating defaulted postfix increment and decrement operators with various qualifiers such as `__restrict__`, `volatile`, and `const`. Include error checks for invalid return types and parameter types in defaulted operators. Enhance the handling of placeholder types and clarify the rules around defaulted operators in the context of C++ concepts and templates.
[tsan][go] Relax alignment for 128-bit Go atomic args buffer access (#213123)
The Go 128-bit atomic helpers added by #196833 read and write the
128-bit value through an a128* cast of the u8* args buffer (e.g.
*(a128*)(a + 8)). Clang treats a128* as requiring 16-byte alignment and
emits MOVAPS or MOVDQA on x86_64. The Go runtime only guarantees 8-byte
alignment of the args buffer, so a+8 may not be 16-byte aligned. The
misaligned MOVAPS raises GP and causes a SIGSEGV.
Use an 8-byte-aligned typedef (`using a128_u64 ALIGNED(8) = a128;`) for
args buffer accesses so Clang emits unaligned 128-bit moves (MOVDQU).
The target address of the atomic operation itself (`*(a128**)a`) is
always 16-byte aligned from the Go side.
Refactor defaulted function handling by introducing a helper to define functions with synthesized bodies. Update exception specification computation to utilize this new helper for better code clarity and maintainability.
[NFC][clang][Modules] Add regression test for implicit lambda capture from merged module copies (#222020)
Adds a regression test for
https://github.com/llvm/llvm-project/issues/221548.
When the same header is textually included into the global module
fragments of two different module interface units, and the header
defines an out-of-line member template whose body contains a lambda
capturing a local variable, instantiating that template for a type named
through the second copy must keep the capture. On affected compilers the
reproducer miscompiles (a `static_assert` on the same value fails with
the bogus `unimplemented constexpr lambda feature: captures not
currently allowed` note); on current main it compiles cleanly.
Verification performed locally:
- Reproduced both symptoms from the issue with the system clang 18
(runtime exit value 21 instead of 42; bogus constexpr note), and
confirmed the closure is missing the capture in IR.
- Built current main with assertions: the reproducer passes (constexpr
[15 lines not shown]
[WebKit Checkers] Allow a view into a const CanBorrow object without a Borrow<T> (#229843)
Example:
const Vector<char> vector = copyToVector(x);
use(vector.span());
A const object can never be mutated, so it can never invalidate a view
into
its interior.
This applies to a const-declared object only, not a const reference or
const
pointer -- which might refer to a non-const object.
A const object also makes the values it contains const, so a view into a
CanBorrow object held in a const std::optional qualifies too:
const std::optional<Vector<char>> vector = copyToVector(x);
[17 lines not shown]
[clang][Sema] Reject oversized arrays deduced from initializer lists (#226983)
When an array's size is deduced from its initializer list, Sema skips
the array size check.
On x86_64,
```cpp
typedef char B[1ULL << 60];
B e[16]; // error: array is too large
B d[] = { {0}, {0}, /* ... 16 elements */ }; // accepted, sizeof(d) == 0
```
On i386,
```cpp
struct B { char c[1 << 30]; };
struct B e[5]; // error: array is too large
struct B d[] = { {{0}}, {{0}}, {{0}}, {{0}}, {{0}} }; // accepted, sizeof(d) == 1 GiB
[12 lines not shown]
[Mips] Track ISA mode and distinguish ELF data labels (#228883)
Represent standard MIPS, microMIPS and MIPS16 with a shared ISA mode in
the ELF target streamer. Initialize it from the subtarget, save and
restore it across .set push/pop, and reset it for .set mips0. Disabling
one compressed ISA must not clear the other ISA's active mode. Keep the
assembler's feature bits consistent when switching between them.
Use the mode to mark function and pending instruction labels, propagate
it to aliases, and set the MIPS16 ELF flag for command-line MIPS16 mode.
Clear pending code labels when emitting byte strings or fill directives,
as already done for integer data, so data does not inherit the following
instruction's mode.
Extend the existing label and .ent tests with command-line microMIPS
mode, nested mode changes, initial-mode resets, and data emission.
Assisted-by: OpenAI Codex # Testcases
[flang][Driver] Enable bare -O flag alias when specified with -flto (#228483)
Currently, `O_flag` (`-O` as an alias for `-O1`) in Options.td is not
visible to FlangOption. As a result, invoking flang with a bare `-O`
(without a trailing digit) is treated by the joined `-O` definition as
`-O""`, causing an error when forwarded to the linker.
Add `FlangOption` to `O_flag`'s Visibility so that `flang -O` correctly
aliases to `-O1`.
Assisted-by: IBM Bob
Resolve https://github.com/llvm/llvm-project/issues/227474
[scudo] Validate list endpoints and links before removing a node (#229245)
`DoublyLinkedList::remove()` currently modifies the predecessor's link
before validating the successor's reciprocal link, and checks endpoint
consistency only in debug builds. A corrupted successor can therefore be
rejected after a list write has already occurred, while inconsistent
null links can bypass endpoint checks in release builds.
Check that the list is nonempty, enforce endpoint consistency in release
builds, and validate both reciprocal links before modifying any list
state. The production change is confined to `remove()`. This strengthens
consistency checks; it does not authenticate list membership or protect
against mutually consistent forged interior nodes.
Tests cover all 24 removal orders of four nodes, plus empty-list
removal, both directions of endpoint inconsistency, and corrupted
reciprocal links, using pointer and index links.
Local validation on Linux x86_64 in WSL, Clang 21.1.8:
[23 lines not shown]
[AMDGPU] Require wave ID masks to be constants (#229917)
PR #177713 added some cases to the wave ID recognizer, such
as (ThreadID & Mask) >> log2(WaveSize) but, after it landed, Claude
noticed a bug. `Mask` in those types of expressions was an arbitrary
vale, which could be divergent, meaning that the "uniform" value of
that expression wouldn't actually be uniform.
The quick fix that preserves most of the cases we're worried about in
practice is to restrict these patterns to constant masks.
AI disclosure: Claude found the issue and created the patch, I wrote
this message.
Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply at anthropic.com>
[AMDGPU][NFC] Pre-commit tests to not match non-constant wave ID masks (#229916)
Non-constant masks can cause divergences even if we're shifting the
divergent part of a thread ID, and we didn't account for that in the
pattern matches.
AI disclosure: Claude found these and wrote the tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply at anthropic.com>