LLVM/project 9368462 — llvm/lib/Target/AArch64 AArch64TargetTransformInfo.cpp, llvm/test/Transforms/LoopVectorize/AArch64 replicating-load-store-costs-apple.ll transform-narrow-interleave-to-widen-memory-factor-gt-vf.ll

[AArch64] Use VectorInstrContext in getScalarizationOverhead. (#177201)

Use VectorInstrContext to return more accurate scalarization overhead
costs when inserts/extracts can be folded into ld1/st1 and CPUs where
ld1/st1 are fast (same perf as regular loads).

Depends on https://github.com/llvm/llvm-project/pull/175982

PR: https://github.com/llvm/llvm-project/pull/177201
DeltaFile
+1,535-215llvm/test/Transforms/LoopVectorize/AArch64/transform-narrow-interleave-to-widen-memory-factor-gt-vf.ll
+259-52llvm/test/Transforms/LoopVectorize/AArch64/replicating-load-store-costs-apple.ll
+9-0llvm/lib/Target/AArch64/AArch64TargetTransformInfo.cpp
+1,803-2673 files

LLVM/project 4362ba8 — clang/include/clang/CIR MissingFeatures.h, clang/lib/CIR/CodeGen CIRGenCleanup.cpp

[CIR] Find conditional cleanups in implicit code (#229829)

`ConditionalEvaluationFinder` in `CIRGenCleanup` skipped implicit code
because that's the default for `RecursiveASTVisitor`. This lead to
default arguments and default member initializers being skipped.

This patch enables the traversal of implicit code with the exception of
the implicit call to `await_resume()` in `co_await` and `co_yield`
expressions. This requires cleanup scopes for await full-expressions
which don't exist yet.

---------

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+191-3clang/test/CIR/CodeGen/cleanup-conditional.cpp
+27-2clang/lib/CIR/CodeGen/CIRGenCleanup.cpp
+1-0clang/include/clang/CIR/MissingFeatures.h
+219-53 files

LLVM/project d1d0fed — llvm/lib/CodeGen/SelectionDAG LegalizeVectorOps.cpp, llvm/lib/Target/AArch64 AArch64ISelLowering.cpp

[LLVM][CodeGen][SVE] Add lowering for bfloat strict-fp cast operators. (#223709)

The majority of the changes are just a case of ensuring the chain is
routed correctly and the matching STRICT passthrough node is used.

NOTE: At present full strict-fp support has a minimum requirement of
+sve2+bf16, otherwise we lack the necessary cast instructions. Of these
+bf16 is fundamental whereas +sve2 is only required for double->bfloat.
DeltaFile
+409-0llvm/test/CodeGen/AArch64/sve-bf-constrained-intrinsics.ll
+46-30llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
+6-0llvm/lib/CodeGen/SelectionDAG/LegalizeVectorOps.cpp
+461-303 files

LLVM/project 97d2c0c — llvm/lib/Transforms/Vectorize SLPVectorizer.cpp, llvm/test/Transforms/SLPVectorizer/X86 minbw-interchangeable-mul-shl.ll

[SLP]Fix poison shift after MinBW demotion

Lanes converted from mul-by-power-of-2 are emitted as shl by the
exponent; check the node shift amounts, not the scalar operands.

Fixes #230392

Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/230464
DeltaFile
+9-7llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+4-2llvm/test/Transforms/SLPVectorizer/X86/minbw-interchangeable-mul-shl.ll
+13-92 files

LLVM/project 538cec9 — llvm/test/Transforms/SLPVectorizer/X86 minbw-interchangeable-mul-shl.ll

[SLP][NFC]Add a test with incorrect optimization, NFC



Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/230462
DeltaFile
+56-0llvm/test/Transforms/SLPVectorizer/X86/minbw-interchangeable-mul-shl.ll
+56-01 files

LLVM/project 6511483 — flang/lib/Parser openacc-parsers.cpp, flang/test/Parser acc-label-do-end-name.f90

[flang][openacc] Diagnose a construct name on the END DO of an ACC labeled DO (#230308)

When an OpenACC loop or combined construct is associated with a labeled DO
loop, AccNonBlockDoConstruct turns a terminating END DO statement into a
labeled CONTINUE statement, which silently dropped any construct name on the
END DO. The DO statement of an unnamed labeled DO loop has no construct
name, so a name on its END DO is not allowed (C1135), and it is already
diagnosed when the loop is not associated with a directive:

```fortran
!$acc parallel loop
do 10 i = 1, n
  a(i) = 0
10 end do foo
```

Report "Unexpected DO construct name" at the name before it is dropped.
Also add tests for labeled END DO forms that are accepted: a branch to the
END DO from inside the loop, END DO followed by an end directive, the kernels

    [3 lines not shown]
DeltaFile
+53-0flang/test/Parser/acc-label-do-end-name.f90
+49-0flang/test/Semantics/OpenACC/acc-label-do.f90
+8-1flang/lib/Parser/openacc-parsers.cpp
+110-13 files

LLVM/project f3c93ac — clang/lib/AST ASTContext.cpp, clang/lib/CodeGen CGRecordLayoutBuilder.cpp

[Clang] Fix oversized bit-field layout on big-endian targets (#225494)

Fixes #225361.

This patch fixes two issues related to bit-fields:

- **Big-endian CodeGen:** Clang incorrectly places the value bits of
oversized bit-fields after the padding bits, contrary to the Itanium C++
ABI (§2.4). Fix the layout so that value bits precede padding bits.
- **`__builtin_clear_padding` (LE and BE):** Correct the occupied-bit
calculation for bit-fields, including `bool` and `_BitInt`, by using
`min(declared width, type size)`. This preserves bits that should not be
treated as padding.
DeltaFile
+263-270clang/test/CodeGenCXX/builtin-clear-padding-codegen.cpp
+83-0clang/test/CodeGenCXX/big-endian-oversized-bitfield.cpp
+10-14clang/lib/AST/ASTContext.cpp
+5-0clang/lib/CodeGen/CGRecordLayoutBuilder.cpp
+361-2844 files

LLVM/project 1c10209 — llvm/lib/Transforms/Vectorize SLPVectorizer.cpp, llvm/test/Transforms/SLPVectorizer/X86 negated-lane-wrap-flags.ll

[SLP]Fix wrap flags for lanes emitted in negated form

sub C, x emitted as add x, -C (or a swapped add/sub lane) negates the
value; nsw/nuw of the original do not cover the negated overflow.
Drop poison-generating flags on the emitted vector add/sub then.

Fixes #230391

Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/230453
DeltaFile
+35-0llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+2-2llvm/test/Transforms/SLPVectorizer/X86/negated-lane-wrap-flags.ll
+37-22 files

LLVM/project 3f098e0 — libc/include termios.yaml, libc/src/termios tcgetwinsize.h

[libc] Implement tcgetwinsize in termios (#228435)

Implement the standard POSIX.1-2024 function `tcgetwinsize` in
`<termios.h>`.

Fixes #228379
Part of #228378

Implementation was assisted by Antigravity by analysing other functions
in header and reviewed by Aman Maurya.
DeltaFile
+91-0libc/test/src/termios/tcgetwinsize_test.cpp
+34-0libc/src/termios/linux/tcgetwinsize.cpp
+26-0libc/src/termios/tcgetwinsize.h
+20-0libc/test/src/termios/CMakeLists.txt
+15-0libc/src/termios/linux/CMakeLists.txt
+9-0libc/include/termios.yaml
+195-05 files not shown
+207-011 files

LLVM/project 29707f4 — lldb/source/Host/common PythonRuntimeLoader.cpp

[lldb][Windows] Register the Python runtime's directory for DLL dependency resolution (#225815)

Follow-up to the discussion on llvm/llvm-project#206585: @mstorsjo found
that pointing LLDB at a specific Python install via
`LLDB_PYTHON_DLL_RELATIVE_PATH` breaks with the limited API DLL.

`python3.dll` forwards to the version numbered DLL and the forwarder
only resolves if that directory is already on `PATH`.

This patch applies @Nerixyz's suggesetion of reinstating
`SetDllDirectory` handling that was dropped in ff65d81. This uses
`AddDllDirectory` instead, registering the runtime's directory so lazy
forwarder resolution finds it regardless of PATH. More details here:
https://learn.microsoft.com/en-us/windows/win32/dlls/dynamic-link-library-security.

> Use the `LOAD_LIBRARY_SEARCH` flags with the `LoadLibraryEx` function,
or use these flags with the `SetDefaultDllDirectories` function to
establish a DLL search order for a process and then use the
`AddDllDirectory` or `SetDllDirectory` functions to modify the list.

    [5 lines not shown]
DeltaFile
+40-0lldb/source/Host/common/PythonRuntimeLoader.cpp
+40-01 files

LLVM/project 4a019b7 — lldb/source/Plugins/Process/Windows/Common NativeThreadWindows.cpp NativeThreadWindows.h

[lldb][Windows] Ignore stale traps delivered after a resume (#228440)

A thread can hit a trap at the same time as the thread whose exception
caused the stop. Windows only delivers one debug event at a time, so the
second thread's event is not sent until the process is resumed. By then
the trap it reports can be stale:

1. A
[`STATUS_SINGLE_STEP`](https://learn.microsoft.com/en-us/windows-hardware/drivers/debugger/specific-exceptions)
for a thread that the current resume did not step.
`NativeProcessWindows::HandleSingleStepException` reports it as a
`eStopReasonTrace` stop. `NativeThreadWindows` now records whether its
last `DoResume()` stepped it, and masks the trap which arrives too late.

2. A
[`STATUS_BREAKPOINT`](https://learn.microsoft.com/en-us/windows-hardware/drivers/debugger/specific-exceptions)
for a breakpoint the client removed during the stop. The original byte
is back in memory, so `HandleBreakpointException` does not find a
breakpoint site and reports an exception stop, with the program counter

    [8 lines not shown]
DeltaFile
+42-1lldb/source/Plugins/Process/Windows/Common/NativeProcessWindows.cpp
+7-0lldb/source/Plugins/Process/Windows/Common/NativeProcessWindows.h
+6-0lldb/source/Plugins/Process/Windows/Common/NativeThreadWindows.h
+1-0lldb/source/Plugins/Process/Windows/Common/NativeThreadWindows.cpp
+56-14 files

LLVM/project 3820435 — llvm/test/Transforms/SLPVectorizer/X86 negated-lane-wrap-flags.ll

[SLP][NFC]Add a test with incorrect flag propagation, NFC



Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/230448
DeltaFile
+130-0llvm/test/Transforms/SLPVectorizer/X86/negated-lane-wrap-flags.ll
+130-01 files

LLVM/project 73c1fe1 — llvm/lib/Transforms/Vectorize VPlanAnalysis.cpp, llvm/test/Transforms/LoopVectorize revec-reg-usage.ll

[LV][REVEC] Correctly compute register usage

For REVEC, the initial types might already be vectors, so make sure the
right register class is picked.
DeltaFile
+39-0llvm/test/Transforms/LoopVectorize/revec-reg-usage.ll
+17-12llvm/lib/Transforms/Vectorize/VPlanAnalysis.cpp
+56-122 files

LLVM/project a09df97 — bolt/lib/Core BinaryFunction.cpp Relocation.cpp, bolt/lib/Target/AArch64 AArch64MCPlusBuilder.cpp

[BOLT][AArch64] Add support for conditional tail calls in cold code (#227693)

BOLT handles AArch64 conditional tail calls in cold code using
the workaround implemented in #140669. This leaves the
conditional branches unchanged and patches the target function’s
original entry to redirect execution to the relocated function.

This patch adds support for R_AARCH64_CONDBR19 and
R_AARCH64_TSTBR14 relocations so that BOLT can now
update B.cond, CB(N)Z, and TB(N)Z to directly reach relocated
functions when they are within range. Branches out-of-range
are left unchanged and target the original patched function.

Assisted-by: Codex
DeltaFile
+90-0bolt/unittests/Core/Relocation.cpp
+78-0bolt/test/AArch64/condbr19-tailcall.s
+76-0bolt/test/AArch64/tstbr14-tailcall.s
+26-3bolt/lib/Core/Relocation.cpp
+9-18bolt/lib/Core/BinaryFunction.cpp
+18-7bolt/lib/Target/AArch64/AArch64MCPlusBuilder.cpp
+297-284 files not shown
+315-4010 files

LLVM/project 8ffd3ce — llvm/include/llvm/DebugInfo/DWARF DWARFExpressionPrinter.h

fixup! [DebugInfo] Print DWARF expressions without a DWARFUnit dependency
DeltaFile
+1-1llvm/include/llvm/DebugInfo/DWARF/DWARFExpressionPrinter.h
+1-11 files

LLVM/project 07b499b — clang/lib/AST/ByteCode InterpBuiltin.cpp, clang/test/AST/ByteCode c.c

[clang][bytecode] Reject strcmp with non-primitive element types (#230396)
DeltaFile
+6-0clang/test/AST/ByteCode/c.c
+2-1clang/lib/AST/ByteCode/InterpBuiltin.cpp
+8-12 files

LLVM/project 955285d — llvm/lib/Target/AArch64/AsmParser AArch64AsmParser.cpp, llvm/test/MC/AArch64 armv8.5a-mte-error.s basic-a64-diagnostics.s

fixup! And fix all the other cases
DeltaFile
+69-69llvm/test/MC/AArch64/basic-a64-diagnostics.s
+53-53llvm/test/MC/AArch64/armv8.5a-mte-error.s
+37-29llvm/lib/Target/AArch64/AsmParser/AArch64AsmParser.cpp
+16-16llvm/test/MC/AArch64/SVE/st1h-diagnostics.s
+16-16llvm/test/MC/AArch64/SVE/ld1h-diagnostics.s
+14-14llvm/test/MC/AArch64/SVE/st1w-diagnostics.s
+205-197113 files not shown
+667-659119 files

LLVM/project 9bb6bc1 — llvm/test/MC/AArch64 armv9.8a-cflt-diagnostics.s arm64-diags.s, llvm/test/MC/AArch64/SVE str-diagnostics.s ldr-diagnostics.s

[AArch64][llvm] Fix incorrect diagnostic (index/immediate in simm9)

Improve the error message in the simm9 diagnostics and use the word
'immediate' instead of 'index'. This was noticed when creating the
CFLT instructions for Armv9.8-A, but would have caused a lot of
unrelated churn if merged as part of that change.

Co-authored-by: Martin Wehking <martin.wehking at arm.com>
DeltaFile
+82-82llvm/test/MC/AArch64/basic-a64-diagnostics.s
+39-39llvm/test/MC/AArch64/armv8.4a-ldst-error.s
+5-5llvm/test/MC/AArch64/arm64-diags.s
+4-4llvm/test/MC/AArch64/SVE/str-diagnostics.s
+4-4llvm/test/MC/AArch64/SVE/ldr-diagnostics.s
+2-2llvm/test/MC/AArch64/armv9.8a-cflt-diagnostics.s
+136-1361 files not shown
+137-1377 files

LLVM/project 20ea4a8 — llvm/lib/Target/X86 X86Options.td X86ISelLowering.cpp, llvm/test/CodeGen/X86 mul-constant-i32.ll mul-constant-i64.ll

[X86] Remove mul-constant-optimization flag (#229376)

The flags doesn't give us complete control as there's a lot of crossover with other folds,
some generic, some in X86DAGToDAGISel and X86FixupLEAs.

Without it we have partial control based off minsize builds and the Slow
LEA tuning flags, but tbh I'd prefer this fold wasn't in the DAG at all,
but scheduler driven in MachineCombiner or X86FixupLEAs.
DeltaFile
+679-1,139llvm/test/CodeGen/X86/mul-constant-i64.ll
+624-991llvm/test/CodeGen/X86/mul-constant-i32.ll
+2-5llvm/lib/Target/X86/X86ISelLowering.cpp
+0-2llvm/lib/Target/X86/X86Options.td
+1,305-2,1374 files

LLVM/project ff3baf3 — llvm/test/Transforms/SLPVectorizer/X86 negated-lane-wrap-flags.ll

[𝘀𝗽𝗿] initial version

Created using spr 1.3.7
DeltaFile
+130-0llvm/test/Transforms/SLPVectorizer/X86/negated-lane-wrap-flags.ll
+130-01 files

LLVM/project a93e321 — llvm/lib/Transforms/Vectorize VPlan.h VPlanRecipes.cpp, llvm/test/Transforms/LoopVectorize revec-liveout.ll

Test recipe execution for vector types

This ensures that execute() and get(VPV, LaneIdx) give expected
results for "vector lanes".

This requires a slight adjustment to type asserts in
VPInstruction::execute to be safe for REVEC.

And ensure a value is correctly detected as "widened", even when
VF == VFScaleFactor
DeltaFile
+76-0llvm/test/Transforms/LoopVectorize/revec-liveout.ll
+50-0llvm/unittests/Transforms/Vectorize/VPlanTest.cpp
+13-2llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
+2-1llvm/lib/Transforms/Vectorize/VPlan.h
+141-34 files

LLVM/project e9da10d — llvm/test/Transforms/LoopVectorize revec-memory-contiguous.ll revec-memory-interleaved.ll

[LV][REVEC][AArch64] Proof of concept for re-vectorisation

This shows the changes required to enable basic re-vectorisation support in LoopVectorizer. Most of the diff comes from the added tests, the changes to LoopVectorizer files are rather minimal. This proof-of-concept has obvious limitations and only represents the first building block.

My hope is that this helps discussions and complements the RFC at https://discourse.llvm.org/t/rfc-re-vectorisation-to-wider-vectors-in-loopvectorizer/91071.

Support for re-vectorisation is hidden behind a -vectorize-vector-loops flag and LV will bail out if it encounters constructs that are not yet supported. For example:
 - shufflevectors
 - gather/scatter and interleaved accesses
 - target intrinsics
 - reductions
 - if-conversion or tail folding
DeltaFile
+293-0llvm/test/Transforms/LoopVectorize/revec-select.ll
+119-0llvm/test/Transforms/LoopVectorize/revec-memory-gather-scatter.ll
+103-0llvm/test/Transforms/LoopVectorize/revec-predication.ll
+99-0llvm/test/Transforms/LoopVectorize/revec-unroll.ll
+90-0llvm/test/Transforms/LoopVectorize/revec-memory-contiguous.ll
+90-0llvm/test/Transforms/LoopVectorize/revec-memory-interleaved.ll
+794-016 files not shown
+1,316-4322 files

LLVM/project c1d2834 — llvm/lib/Target/X86 X86ISelLowering.cpp, llvm/test/CodeGen/X86 vector-interleaved-store-i16-stride-8.ll vector-interleaved-store-i32-stride-8.ll

[X86] concatSubVectors helper - use POISON as INSERT_SUBVECTOR base instead of UNDEF (#230427)

First tentative attempt at UNDEF cleanup
DeltaFile
+32-32llvm/test/CodeGen/X86/vector-interleaved-store-i32-stride-8.ll
+16-16llvm/test/CodeGen/X86/vector-interleaved-store-i16-stride-8.ll
+3-3llvm/lib/Target/X86/X86ISelLowering.cpp
+51-513 files

LLVM/project 8d4fc8f — lldb/packages/Python/lldbsuite/test/make Makefile.rules

[lldb][test] Don't strip debug info twice with SPLIT_DEBUG_SYMBOLS" (#228118)

With `SPLIT_DEBUG_SYMBOLS`, the Makefile rule writes the debug file
first and then rewrites the executable, so the executable ends up newer
than the debug file. The next make in the same build directory therefore
runs the rule again, on an executable that has already been stripped,
and the debug file is left with no debug info.

On Windows, `make` compares file times to the second, so this only
happens when the two steps cross a second boundary. In swiftlang, this
makes `TestSwiftSplitDebug` flaky on CI, where `check-lldb-swift`
rebuilds in `check-lldb`'s build directory.

The rule now touches the debug file last, so a second make does nothing.

rdar://188918828
DeltaFile
+8-0lldb/packages/Python/lldbsuite/test/make/Makefile.rules
+8-01 files

LLVM/project 95346de — lldb/source/Expression DWARFExpression.cpp DWARFExpressionList.cpp, lldb/source/Symbol UnwindPlan.cpp

[lldb] Don't link DWARFUnit to print DWARF expressions
DeltaFile
+2-3lldb/source/Expression/DWARFExpressionList.cpp
+1-1lldb/source/Symbol/UnwindPlan.cpp
+1-1lldb/source/Expression/DWARFExpression.cpp
+4-53 files

LLVM/project 05af87a — lldb/source/Plugins/ObjectFile/PECOFF ObjectFilePECOFF.cpp

[lldb] Create COFFObjectFile directly in ObjectFilePECOFF
DeltaFile
+4-10lldb/source/Plugins/ObjectFile/PECOFF/ObjectFilePECOFF.cpp
+4-101 files

LLVM/project b549a8f — llvm/include/llvm/DebugInfo/DWARF DWARFExpressionPrinter.h, llvm/lib/DebugInfo/DWARF DWARFExpressionPrinterImpl.h DWARFExpressionPrinterUnit.cpp

[DebugInfo] Print DWARF expressions without a DWARFUnit dependency
DeltaFile
+31-39llvm/lib/DebugInfo/DWARF/DWARFExpressionPrinter.cpp
+57-0llvm/lib/DebugInfo/DWARF/DWARFExpressionPrinterUnit.cpp
+40-0llvm/lib/DebugInfo/DWARF/DWARFExpressionPrinterImpl.h
+26-0llvm/unittests/DebugInfo/DWARF/DWARFExpressionCompactPrinterTest.cpp
+6-0llvm/include/llvm/DebugInfo/DWARF/DWARFExpressionPrinter.h
+1-0llvm/utils/gn/secondary/llvm/lib/DebugInfo/DWARF/BUILD.gn
+161-391 files not shown
+162-397 files

LLVM/project 3632e3c — llvm/lib/Object CMakeLists.txt SymbolicFile.cpp, llvm/utils/gn/secondary/llvm/lib/Object BUILD.gn

[Object] Move the createBinary/createObjectFile factories into their own file (#229795)

`Binary.cpp`, `SymbolicFile.cpp` and `ObjectFile.cpp` each define both
base class members and a factory that dispatches to every object file
format. Anything that uses one format, references the base class
constructors and vtables, so the linker loads all three members from the
static archive, and with them the factories' references to the ELF,
Mach-O, Wasm, XCOFF, TAPI and IR readers.

`--gc-sections` and `/OPT:REF` cannot help here: the newly loaded
members (BitcodeReader, the LLVMCore files) have dynamic initializers
for their `cl::opt` globals, which are GC roots.

This patch moves `createBinary`, `ObjectFile::createObjectFile` and
`SymbolicFile::createSymbolicFile` to `ObjectFactories.cpp`. No
functional change.

Together with https://github.com/llvm/llvm-project/pull/229798 and
https://github.com/llvm/llvm-project/pull/229817, this takes lldb-server
on Windows from 8.4 MB to 3.7 MB.
DeltaFile
+252-0llvm/lib/Object/ObjectFactory.cpp
+0-99llvm/lib/Object/Binary.cpp
+0-93llvm/lib/Object/ObjectFile.cpp
+0-76llvm/lib/Object/SymbolicFile.cpp
+1-0llvm/utils/gn/secondary/llvm/lib/Object/BUILD.gn
+1-0llvm/lib/Object/CMakeLists.txt
+254-2686 files

LLVM/project 001b1ab — llvm/include/llvm/CodeGen TargetLowering.h, llvm/lib/CodeGen/SelectionDAG LegalizeIntegerTypes.cpp LegalizeVectorOps.cpp

[DAG] Expand mask_beforefirst during promotion (#230163)

Fixed length mask vector types (v16i1 etc.) are promoted on NEON, so
mask_beforefirst @ v16i1 becomes mask_beforefirst @ v16i8, which in turn
gets expanded to get_active_lane_mask + cttz_elts @ v16i8, and finally
undergoes generic expansion.

If we expand it earlier to get_active_lane_mask + cttz_elts @ v16i1,
then AArch64 can custom lower it to AArch64ISD::CTTZ_ELTS and we get the
good optimizeBrk lowering.
DeltaFile
+16-39llvm/test/CodeGen/AArch64/mask-beforefirst-sve.ll
+12-0llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
+1-8llvm/lib/CodeGen/SelectionDAG/LegalizeVectorOps.cpp
+7-0llvm/lib/CodeGen/SelectionDAG/LegalizeIntegerTypes.cpp
+5-0llvm/include/llvm/CodeGen/TargetLowering.h
+41-475 files

LLVM/project d51e79e — llvm/include/llvm/CodeGen MachineBasicBlock.h, llvm/lib/CodeGen MachineBasicBlock.cpp PHIElimination.cpp

CodeGen: Only visit live-out vregs when splitting a critical edge with LIS

To update LiveIntervals, SplitCriticalEdge checked every virtual
register in the function for liveness at the end of the split block.
That made PHIElimination's edge splitting O(splits * vregs).

PHIElimination now computes the set of virtual registers live out of
each block before the first split and passes it through
SplitCriticalEdgeAnalyses. Only those registers are visited, and the set
for the new block is added once the intervals are updated.

This essentially resurrects the per-block sparse register sets used to
update LiveVariables, which were removed in #228618 and #230147, applied
to the LiveIntervals update instead.

Instructions retired in phi-node-elimination, x86_64 -O3, on a generated
chain of N compare blocks branching to a shared PHI block:

  N     before          after           after/before

    [12 lines not shown]
DeltaFile
+26-10llvm/lib/CodeGen/PHIElimination.cpp
+22-5llvm/lib/CodeGen/MachineBasicBlock.cpp
+7-0llvm/include/llvm/CodeGen/MachineBasicBlock.h
+55-153 files