[CIR] Find conditional cleanups in implicit code (#229829)
`ConditionalEvaluationFinder` in `CIRGenCleanup` skipped implicit code
because that's the default for `RecursiveASTVisitor`. This lead to
default arguments and default member initializers being skipped.
This patch enables the traversal of implicit code with the exception of
the implicit call to `await_resume()` in `co_await` and `co_yield`
expressions. This requires cleanup scopes for await full-expressions
which don't exist yet.
---------
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[LLVM][CodeGen][SVE] Add lowering for bfloat strict-fp cast operators. (#223709)
The majority of the changes are just a case of ensuring the chain is
routed correctly and the matching STRICT passthrough node is used.
NOTE: At present full strict-fp support has a minimum requirement of
+sve2+bf16, otherwise we lack the necessary cast instructions. Of these
+bf16 is fundamental whereas +sve2 is only required for double->bfloat.
[SLP]Fix poison shift after MinBW demotion
Lanes converted from mul-by-power-of-2 are emitted as shl by the
exponent; check the node shift amounts, not the scalar operands.
Fixes #230392
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/230464
[flang][openacc] Diagnose a construct name on the END DO of an ACC labeled DO (#230308)
When an OpenACC loop or combined construct is associated with a labeled DO
loop, AccNonBlockDoConstruct turns a terminating END DO statement into a
labeled CONTINUE statement, which silently dropped any construct name on the
END DO. The DO statement of an unnamed labeled DO loop has no construct
name, so a name on its END DO is not allowed (C1135), and it is already
diagnosed when the loop is not associated with a directive:
```fortran
!$acc parallel loop
do 10 i = 1, n
a(i) = 0
10 end do foo
```
Report "Unexpected DO construct name" at the name before it is dropped.
Also add tests for labeled END DO forms that are accepted: a branch to the
END DO from inside the loop, END DO followed by an end directive, the kernels
[3 lines not shown]
[Clang] Fix oversized bit-field layout on big-endian targets (#225494)
Fixes #225361.
This patch fixes two issues related to bit-fields:
- **Big-endian CodeGen:** Clang incorrectly places the value bits of
oversized bit-fields after the padding bits, contrary to the Itanium C++
ABI (§2.4). Fix the layout so that value bits precede padding bits.
- **`__builtin_clear_padding` (LE and BE):** Correct the occupied-bit
calculation for bit-fields, including `bool` and `_BitInt`, by using
`min(declared width, type size)`. This preserves bits that should not be
treated as padding.
[SLP]Fix wrap flags for lanes emitted in negated form
sub C, x emitted as add x, -C (or a swapped add/sub lane) negates the
value; nsw/nuw of the original do not cover the negated overflow.
Drop poison-generating flags on the emitted vector add/sub then.
Fixes #230391
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/230453
[libc] Implement tcgetwinsize in termios (#228435)
Implement the standard POSIX.1-2024 function `tcgetwinsize` in
`<termios.h>`.
Fixes #228379
Part of #228378
Implementation was assisted by Antigravity by analysing other functions
in header and reviewed by Aman Maurya.
[lldb][Windows] Register the Python runtime's directory for DLL dependency resolution (#225815)
Follow-up to the discussion on llvm/llvm-project#206585: @mstorsjo found
that pointing LLDB at a specific Python install via
`LLDB_PYTHON_DLL_RELATIVE_PATH` breaks with the limited API DLL.
`python3.dll` forwards to the version numbered DLL and the forwarder
only resolves if that directory is already on `PATH`.
This patch applies @Nerixyz's suggesetion of reinstating
`SetDllDirectory` handling that was dropped in ff65d81. This uses
`AddDllDirectory` instead, registering the runtime's directory so lazy
forwarder resolution finds it regardless of PATH. More details here:
https://learn.microsoft.com/en-us/windows/win32/dlls/dynamic-link-library-security.
> Use the `LOAD_LIBRARY_SEARCH` flags with the `LoadLibraryEx` function,
or use these flags with the `SetDefaultDllDirectories` function to
establish a DLL search order for a process and then use the
`AddDllDirectory` or `SetDllDirectory` functions to modify the list.
[5 lines not shown]
[lldb][Windows] Ignore stale traps delivered after a resume (#228440)
A thread can hit a trap at the same time as the thread whose exception
caused the stop. Windows only delivers one debug event at a time, so the
second thread's event is not sent until the process is resumed. By then
the trap it reports can be stale:
1. A
[`STATUS_SINGLE_STEP`](https://learn.microsoft.com/en-us/windows-hardware/drivers/debugger/specific-exceptions)
for a thread that the current resume did not step.
`NativeProcessWindows::HandleSingleStepException` reports it as a
`eStopReasonTrace` stop. `NativeThreadWindows` now records whether its
last `DoResume()` stepped it, and masks the trap which arrives too late.
2. A
[`STATUS_BREAKPOINT`](https://learn.microsoft.com/en-us/windows-hardware/drivers/debugger/specific-exceptions)
for a breakpoint the client removed during the stop. The original byte
is back in memory, so `HandleBreakpointException` does not find a
breakpoint site and reports an exception stop, with the program counter
[8 lines not shown]
[BOLT][AArch64] Add support for conditional tail calls in cold code (#227693)
BOLT handles AArch64 conditional tail calls in cold code using
the workaround implemented in #140669. This leaves the
conditional branches unchanged and patches the target function’s
original entry to redirect execution to the relocated function.
This patch adds support for R_AARCH64_CONDBR19 and
R_AARCH64_TSTBR14 relocations so that BOLT can now
update B.cond, CB(N)Z, and TB(N)Z to directly reach relocated
functions when they are within range. Branches out-of-range
are left unchanged and target the original patched function.
Assisted-by: Codex
[AArch64][llvm] Fix incorrect diagnostic (index/immediate in simm9)
Improve the error message in the simm9 diagnostics and use the word
'immediate' instead of 'index'. This was noticed when creating the
CFLT instructions for Armv9.8-A, but would have caused a lot of
unrelated churn if merged as part of that change.
Co-authored-by: Martin Wehking <martin.wehking at arm.com>
[X86] Remove mul-constant-optimization flag (#229376)
The flags doesn't give us complete control as there's a lot of crossover with other folds,
some generic, some in X86DAGToDAGISel and X86FixupLEAs.
Without it we have partial control based off minsize builds and the Slow
LEA tuning flags, but tbh I'd prefer this fold wasn't in the DAG at all,
but scheduler driven in MachineCombiner or X86FixupLEAs.
Test recipe execution for vector types
This ensures that execute() and get(VPV, LaneIdx) give expected
results for "vector lanes".
This requires a slight adjustment to type asserts in
VPInstruction::execute to be safe for REVEC.
And ensure a value is correctly detected as "widened", even when
VF == VFScaleFactor
[LV][REVEC][AArch64] Proof of concept for re-vectorisation
This shows the changes required to enable basic re-vectorisation support in LoopVectorizer. Most of the diff comes from the added tests, the changes to LoopVectorizer files are rather minimal. This proof-of-concept has obvious limitations and only represents the first building block.
My hope is that this helps discussions and complements the RFC at https://discourse.llvm.org/t/rfc-re-vectorisation-to-wider-vectors-in-loopvectorizer/91071.
Support for re-vectorisation is hidden behind a -vectorize-vector-loops flag and LV will bail out if it encounters constructs that are not yet supported. For example:
- shufflevectors
- gather/scatter and interleaved accesses
- target intrinsics
- reductions
- if-conversion or tail folding
[lldb][test] Don't strip debug info twice with SPLIT_DEBUG_SYMBOLS" (#228118)
With `SPLIT_DEBUG_SYMBOLS`, the Makefile rule writes the debug file
first and then rewrites the executable, so the executable ends up newer
than the debug file. The next make in the same build directory therefore
runs the rule again, on an executable that has already been stripped,
and the debug file is left with no debug info.
On Windows, `make` compares file times to the second, so this only
happens when the two steps cross a second boundary. In swiftlang, this
makes `TestSwiftSplitDebug` flaky on CI, where `check-lldb-swift`
rebuilds in `check-lldb`'s build directory.
The rule now touches the debug file last, so a second make does nothing.
rdar://188918828
[Object] Move the createBinary/createObjectFile factories into their own file (#229795)
`Binary.cpp`, `SymbolicFile.cpp` and `ObjectFile.cpp` each define both
base class members and a factory that dispatches to every object file
format. Anything that uses one format, references the base class
constructors and vtables, so the linker loads all three members from the
static archive, and with them the factories' references to the ELF,
Mach-O, Wasm, XCOFF, TAPI and IR readers.
`--gc-sections` and `/OPT:REF` cannot help here: the newly loaded
members (BitcodeReader, the LLVMCore files) have dynamic initializers
for their `cl::opt` globals, which are GC roots.
This patch moves `createBinary`, `ObjectFile::createObjectFile` and
`SymbolicFile::createSymbolicFile` to `ObjectFactories.cpp`. No
functional change.
Together with https://github.com/llvm/llvm-project/pull/229798 and
https://github.com/llvm/llvm-project/pull/229817, this takes lldb-server
on Windows from 8.4 MB to 3.7 MB.
[DAG] Expand mask_beforefirst during promotion (#230163)
Fixed length mask vector types (v16i1 etc.) are promoted on NEON, so
mask_beforefirst @ v16i1 becomes mask_beforefirst @ v16i8, which in turn
gets expanded to get_active_lane_mask + cttz_elts @ v16i8, and finally
undergoes generic expansion.
If we expand it earlier to get_active_lane_mask + cttz_elts @ v16i1,
then AArch64 can custom lower it to AArch64ISD::CTTZ_ELTS and we get the
good optimizeBrk lowering.
CodeGen: Only visit live-out vregs when splitting a critical edge with LIS
To update LiveIntervals, SplitCriticalEdge checked every virtual
register in the function for liveness at the end of the split block.
That made PHIElimination's edge splitting O(splits * vregs).
PHIElimination now computes the set of virtual registers live out of
each block before the first split and passes it through
SplitCriticalEdgeAnalyses. Only those registers are visited, and the set
for the new block is added once the intervals are updated.
This essentially resurrects the per-block sparse register sets used to
update LiveVariables, which were removed in #228618 and #230147, applied
to the LiveIntervals update instead.
Instructions retired in phi-node-elimination, x86_64 -O3, on a generated
chain of N compare blocks branching to a shared PHI block:
N before after after/before
[12 lines not shown]