[LLVM][CodeGen][SVE] Improve lowering for v1f32/f64 when NEON is not available. (#229724)
When NEON is not available it is better to scalarise single element
floating-point vectors than widening them to use Streaming-SVE.
Explicitly make v1f64 scalar_to_vector operations always legal, because
we can use scalar instructions and add a combine to avoid "nop" casts.
[CIR][CodeGen][NFC] Retire the CodeGenUtils.h catch-all header
Moves the last helpers out of `CodeGenUtils.h` into `ClassUtils.h`,
`ModuleUtils.h`, `TargetUtils.h` and `FunctionUtils.h` (`checkTargetFeatures`,
since it came from CodeGenFunction.cpp) and deletes the header. Only moves code
already on main, so it can be dropped on its own.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share requiresAMDGPUProtectedVisibility
Deduplicates `requiresAMDGPUProtectedVisibility` between CIR and classic CodeGen
into `TargetUtils.h`. The shared version takes a bool for "currently hidden" in
place of the `llvm::GlobalValue` and `cir::VisibilityKind` the two callers
passed.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen] Share the EH personality selection logic
Deduplicates the EH personality selection (`getEHPersonality`,
`getCXXEHPersonality`) between CIR and classic CodeGen, taking the classic
implementation. CIR's copy lacked the z/OS, Wasm and GNUstep-on-CygMing cases;
none are reachable in CIR today, so no test changes.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen] Share isStandardLibraryRTTIDescriptor
Deduplicates `isStandardLibraryRTTIDescriptor` between CIR and classic CodeGen
into `ItaniumCXXABIUtils.h`, taking the classic implementation. The two copies
have been equivalent since #227781 filled in the builtin types CIR was missing.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share hasExtraNeonArgument (#227260)
Deduplicates `hasExtraNeonArgument` between CIR and classic CodeGen into
`TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
Transforms: Use changeToCall in LowerInvoke
The pass predates changeToCall, which does the same rewrite and is what other
contexts already use. The only difference is the new branch takes the invoke's
DebugLoc.
[SLP]Fix masked gather GEPs cost, skip GEP index seeds (#228918)
Pass the loaded type and the real base sharing to the chain cost of the
masked gather GEPs. Do not use the single-use GEP index chains as seeds:
they were vectorized with all lanes extracted for the scalar addresses.
Part of #227393.
Assisted-by: Cursor
[DAG] Explicitly truncate splat operand in get_active_lane_mask expansion
x86 can't handle the implicit truncation of build_vector operands so the whilewr_8 test case crashes after expanding get_active_lane_mask (stemming from llvm.loop.dependence.war.mask). Fix it by explicitly truncating it.
[lldb][Windows] Report the DebugBreakProcess halt as SIGSTOP (#229770)
`NativeProcessWindows` reports signal 19 `SIGSTOP` on Linux but
`SIGCONT` for Windows targets, so each interrupt lldb sent to issue a
packet prints:
```
Process 26748 stopped and restarted: thread 2 received signal: SIGCONT
```
This makes typing stdin to a debuggee tedious when debugging.
This patch uses
`UnixSignals::CreateForHost()->GetSignalNumberFromName("SIGSTOP")` to
use the proper signal and adds a regression test.
[lldb][Windows] End the debug session before destroying an lldb-server process (#229433)
When the debuggee exits, lldb-server sends the exit to the client and
then destroys its process object, while the thread that reported the
exit is still running code that uses that object. lldb-server then
sometimes crashes on its way out.
The destructor now ends the debug session first: it stops the debugger
thread and waits until that thread is done with the object.
A process that exits before its first stop never started, so its exit is
no longer reported to the server, which never had that process. Launch
failures are already reported as an error, and the extra exit would
otherwise arrive as the reply to the launch request (This occurs in the
TestMissingDll.py test).
In a stress run on Windows, 194 of 320 debug sessions crashed before
this change and none after it.
[BasicAA] Fix miscompilation with setjmp/longjmp due to missing longjmp re-entry paths in alias analysis (#212297)
Fixes [#198967](https://github.com/llvm/llvm-project/issues/198967).
`EarliestEscapeAnalysis::getCapturesBefore` (used by GVN, DSE,
MemCpyOpt) determines whether an object is captured before a given
instruction using `isPotentiallyReachable()/isNotInCycle()`, both of
which only see the forward CFG. In a function containing a
`returns_twice` call (e.g. `setjmp`), a `longjmp` can re-enter at that
call site — a back-edge invisible to both checks.
This causes `BasicAA` to incorrectly return `NoAlias` for a local alloca
whose address was captured on a branch that only runs before the
`longjmp` re-entry. Downstream passes (GVN, DSE) then miscompile the
function by eliminating stores that are still live on the re-entry path.
Fix: `EarliestEscapeAnalysis` now caches whether its function contains a
call that may return twice (`callsReturnsTwiceFn()`, backed by the
existing `Function::callsFunctionThatReturnsTwice()` scan, memoized per
instance). If so, `getCapturesBefore` conservatively treats the object
[23 lines not shown]
[NVVM][NVPTX] Support spcompress and spdecompress intrinsics (#221751)
This change supports `spcompress` and `spdecompress` intrinsics and
their lowering to the NVPTX backend.
[AArch64] Implement CRC32 const folding (#228985)
This implements constant folding for `__crc32` intrinsics when both the
arguments are compile time known constants.
This is a follow up of https://github.com/llvm/llvm-project/pull/219452
which does the same for X86.
[PassManager] Factor out some common code to a utility function (NFC) (#229432)
There are two places containing common code where module inlining is
added to the ModulePassManager, with a third one coming up.
This commit factors out the common code into a utility function to
reduce duplication.
[InstCombine] Remove target specific sext(trunc(...)) combine tests. NFC
Follow up to https://github.com/llvm/llvm-project/pull/227321#issuecomment-6055905388
t0 already exists in the generic test file, t1 was moved over as "ashr_signbits". A new RUN line has been added for a datalayout with a native i32 type to exercise the change in #227321
CodeGen: Run LiveIntervals before PHIElimination and drop LiveVariables from it
Move LiveIntervals to run before PHIElimination in the optimized register
allocation pipeline, and make PHIElimination maintain LiveIntervals only.
This removes the last explicit use of LiveVariables. The actual analysis is no
longer used. There are implicit dependencies on the side effects of running the
analysis due to adjustments of dead flags, so further work is still needed to
complete the removal.
This perturbs register allocation in a number of tests. The same codegen result
can be achieved by not preserving the analysis and recomputing fresh. Greedy is
just sensitive to the exact slot index and value numbering with identical MIR.
Measured across every affected test the emitted instruction count goes from
145124 to 145196, +0.050%, with changes in both directions. The largest
regression is AArch64/phi.ll, where the GlobalISel output gains about 30
instructions and no longer matches the SelectionDAG output; the largest
improvements are ARM/fpclamptosat.ll and PowerPC/common-chain.ll.
[2 lines not shown]
MC: Error if target did not register null streamer
Without this the AsmPrinter would see a null pointer
for the target streamer. More than likely this is just
going to crash. Only 5 of the in tree targets properly
registered a null streamer before I started looking into this.
The rest had randomly, inconsistent null checks for the streamer.
There are still a few targets not registering this, which should
update (some of which seem to not make use of the TargetStreamer).
There's no reason to leave this as a silently hazardous edge case.
[DAGCombiner] Add tests for zero-extended low-bits masks (NFC) (#229981)
A low-bits mask built in a narrow type and zero-extended into a wider
and, (and X, (zext (not (shl -1, Y)))), with and without other uses of
the mask, on AArch64, X86 and RISC-V (with and without Zbb).
Assisted-by: Claude Code
[cmake] Properly link against LLVM components (#224624)
The proper way to reference LLVM components is to pass them via the
LLVM_LINK_COMPONENTS variable. This is needed when building LLVM as a
dylib, so the proper dependency (the LLVM library) is passed.
This does not apply to plain add_executable() targets, which should
manually select to link against LLVM or individual LLVM components.
The effort to build LLVM as a dylib is tracked in #109483.
[flang][OpenMP] Do not set nuw on canonical loop trip count span (#229776)
lb <= ub is a signed comparison and does not imply that ub - lb does not
wrap as unsigned (e.g. when lb is negative and ub is not), so the nuw
flag on the span computation could produce poison.
Fixes https://github.com/llvm/llvm-project/issues/229739