[libc++][string] Improve constexpr performance
Adding `if (__libcpp_is_constant_evaluated()) return` allows
significantly increase complexity of extression.
In case of Asan it changed from 1000 to 6000.
Pull Request: https://github.com/llvm/llvm-project/pull/184724
Wyles/update libclctests (#228681)
Using libclc with different versions of clang emit different output.
computeConstantRange now pushes ranges through zext and sext so we now
can prove noundef. This test returns AMDGPU workgroup ID which is zero
extended, so we can now prove it needs noundef.
At some point we may need to update the CI targets to look at libclc?
[mlir][wasmssa] Fix if/else round-trip and verify return types (#227828)
Fixes two WasmSSA dialect bugs:
**#226072: `wasmssa.if` does not round-trip.** The else region was
printed with `printKeywordOrString("else ")`. Because of the trailing
space, `"else "` is not a valid bare keyword, so it was printed as a
quoted string (`} "else "{`), which `parseElseRegion` does not accept.
The else keyword is now printed directly, giving `} else {`.
**#226073: `wasmssa.return` is not checked against the function's
results.** `wasmssa.return` had no verifier, so a function could return
a different number or type of values than its signature declares. This
adds a `ReturnOp` verifier that checks the operands against the
enclosing `wasmssa.func`. Checking in the return op rather than in
`FuncOp::verifyBody` also covers returns nested inside
`block`/`loop`/`if` regions.
Tests:
[16 lines not shown]
[NFC][TSan] Allocate ScopedReport as a stack variable
Now that ScopedReport is constructed before acquiring ThreadRegistryLock
or slot locks across all reporting functions, it no longer needs to be
constructed inside the lock scope via placement new on __builtin_alloca
storage.
Declare ScopedReport as a normal stack variable before the lock scope
and remove the manual destructor calls.
Assisted-by: Gemini
Pull Request: https://github.com/llvm/llvm-project/pull/228637
[TSan] Lock ScopedErrorReportLock before slot and thread_registry locks
OutputReport runs while ScopedErrorReportLock is held after slot_mtx and
thread_registry have been unlocked. Because code executed during
OutputReport (symbolizer, callbacks, or signal handlers) can acquire
slot_mtx or thread_registry, ScopedErrorReportLock must precede slot and
thread_registry locks in the lock hierarchy to avoid AB-BA deadlocks
between concurrent reports or fork().
- Move ScopedErrorReportLock::Lock() before slot.mtx, thread_registry,
and slot_mtx in ForkBefore (and unlock in reverse order in ForkAfter).
- Replace ctx->thread_registry.CheckLocked() in ScopedReportBase's
constructor with CheckedMutex::CheckNoLocks(), and add CheckLocked() to
AddThread(const ThreadContext *) and CheckNoLocks() to OutputReport.
- Construct ScopedReport before acquiring ThreadRegistryLock across all
reporting functions, and close the RestoreStack lock scope before
constructing ScopedReport in ReportRace.
Assisted-by: Gemini
[2 lines not shown]
[NFC][TSan] Move ObtainCurrentStack out of ThreadRegistryLock scope
ObtainCurrentStack only reads the current thread's shadow stack and
allocates a VarSizeStackTrace buffer, which does not require
ThreadRegistryLock (and already runs outside ThreadRegistryLock in
ReportRace, SignalUnsafeCall, and ReportErrnoSpoiling).
Move ObtainCurrentStack (and dummy_pc in ReportDeadlock) before
ScopedReport in ReportMutexHeldWrongContext, ReportMutexMisuse,
ReportDeadlock, and ReportDestroyLocked so the stack trace buffers also
outlive ScopedReport and OutputReport.
Assisted-by: Gemini
Pull Request: https://github.com/llvm/llvm-project/pull/228794
[TSan] Defer symbolization in AddStack and AddSleep to SymbolizeStackElems
PR #151495 delayed symbolization for memory accesses, locations,
threads, and mutexes until OutputReport (after ThreadRegistryLock is
released), but missed ScopedReport::AddStack and ScopedReport::AddSleep.
As a result, ReportRace (via AddSleep) and ReportMutexMisuse,
ReportDeadlock, and ReportDestroyLocked (via AddStack) still invoked the
symbolizer while holding ThreadRegistryLock.
Store the unsymbolized stack traces and sleep stack ID in ReportDesc and
symbolize them in ScopedReport::SymbolizeStackElems().
Assisted-by: Gemini
Pull Request: https://github.com/llvm/llvm-project/pull/228795
[NFC][TSan] Use in-class member initializers in tsan_report.h
Use in-class member initializers for all structs and classes in
tsan_report.h and default constructors and destructor in
tsan_report.cpp.
Assisted-by: Gemini
Pull Request: https://github.com/llvm/llvm-project/pull/228780
[CIR] Preserve address spaces in emitPointerWithAlignment casts (#228652)
emitPointerWithAlignment handled CK_AddressSpaceConversion like
CK_BitCast and never emitted the address-space cast, so the result kept
the source address space. When the address escaped (returned reference,
stored pointer), the store path bitcast the destination slot to the
wrong address space. This miscompiled sycl::multi_ptr::operator[] on
SPIR-V: a local-memory offset was used as a generic address. Emit the
cast after the element bitcast, as classic CodeGen does.
createElementBitCast also built the new pointer type in the default
address space, so changing the element type of a non-default-AS pointer
produced an AS-changing bitcast that the verifier rejects (for example
((int *)p)[i] or __builtin_stdc_memreverse8 on an AS3 pointer). Keep the
source address space, which matches classic's withElementType.
Co-authored-by: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[Clang][Smea] Accept weak reference after declaration
GCC accepts a weak reference even after there is a declaration. It is
fine for us to just append a weak attribute in the previous
declaration directly.
[mlir][XeGPU] Add an SLM round-trip fallback for sg-to-lane convert_layout (#227884)
This PR adds a general fallback lowering for xegpu.convert_layout in the
sg-to-lane distribution pass, which round-trips the value through shared
local memory.
The existing lane-level lowerings are all special cases: the layouts
fold into each other, or they differ in a way a xegpu.lane_shuffle or a
gpu.shuffle can express. Anything else failed to legalize and stopped
the pipeline — for instance a conversion that only moves the lane_data
of a dimension distributed over part of the subgroup:
```mlir
xegpu.convert_layout %src
<{input_layout = #xegpu.layout<lane_layout = [1, 2, 8], lane_data = [1, 1, 4]>,
target_layout = #xegpu.layout<lane_layout = [1, 2, 8], lane_data = [1, 1, 1]>}>
: vector<1x2x32xbf16>
```
[17 lines not shown]