LLVM/project a986c09mlir/lib/IR PatternMatch.cpp, mlir/test/Transforms remove-dead-values-return-segments.mlir remove-dead-values-llvm-call-segments.mlir

[MLIR] Preserve operand segment sizes when removing dead inputs (#223154)

`RemoveDeadRegionBranchOpSuccessorInputs` was calling `eraseOperands`
and ignoring potential `AttrSizedOperandSegments`. Trying to use
`populateRegionBranchOpInterfaceCanonicalizationPatterns` on downstream
control flow op with it resulted in broken IR. Update the segments sizes
if present.

Also fix LLVM dialect `CallOp` and `InvokeOp` `getArgOperandsMutable`
which were dropping segment sizes and a parser bug uncovered by the
tests.

Code was vibed and then edited, tests were vibed and not particularly
edited.
DeltaFile
+125-0mlir/test/Transforms/region-branch-operand-segments.mlir
+93-0mlir/test/Transforms/remove-dead-values-llvm-call-segments.mlir
+77-0mlir/test/Transforms/remove-dead-values-return-segments.mlir
+42-0mlir/test/lib/Dialect/Test/TestOpDefs.cpp
+31-0mlir/test/lib/Dialect/Test/TestOps.td
+31-0mlir/lib/IR/PatternMatch.cpp
+399-05 files not shown
+438-1011 files

LLVM/project 3dcd66dmlir/lib/Dialect/SCF/Transforms LoopSpecialization.cpp, mlir/test/Dialect/Linalg transform-op-peel-and-vectorize.mlir transform-op-peel-and-vectorize-conv.mlir

[mlir][scf] Do not peel an iteration out of an empty loop (#223257)
DeltaFile
+20-10mlir/lib/Dialect/SCF/Transforms/LoopSpecialization.cpp
+9-9mlir/test/Dialect/SCF/for-loop-peeling.mlir
+7-4mlir/test/Dialect/SCF/for-loop-peeling-front.mlir
+2-2mlir/test/Dialect/SparseTensor/sparse_vector_peeled.mlir
+2-2mlir/test/Dialect/Linalg/transform-op-peel-and-vectorize.mlir
+2-2mlir/test/Dialect/Linalg/transform-op-peel-and-vectorize-conv.mlir
+42-296 files

LLVM/project a9bd24allvm/include/llvm/CodeGen BasicTTIImpl.h, llvm/test/Analysis/CostModel/AArch64 sve-mulh.ll mulh.ll

[TTI] Add basic costs for llvm.[su]mulh. (#224002)

Defaults to the expansion (trunc ((ext(A) * ext(B)) >> BW)).
DeltaFile
+414-0llvm/test/Analysis/CostModel/RISCV/mulh.ll
+160-160llvm/test/Analysis/CostModel/X86/arith-mulh.ll
+181-0llvm/test/Analysis/CostModel/AArch64/mulh.ll
+36-0llvm/test/Analysis/CostModel/AArch64/sve-mulh.ll
+25-0llvm/include/llvm/CodeGen/BasicTTIImpl.h
+816-1605 files

LLVM/project 483b4dellvm/test/CodeGen/ARM vlldm-vlstm-uops.mir

ARM: Fix mixed dead and not-dead LR operands in vlldm-vlstm-uops.mir

This operand list had LR listed twice, once from its implicit-defs on
the instruction definition, and another in the variadic argument list.
One had a dead flag, and the other didn't which should be a verifier error
in the future, so remove the redundant operand.

Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
DeltaFile
+1-1llvm/test/CodeGen/ARM/vlldm-vlstm-uops.mir
+1-11 files

LLVM/project 1415226llvm/lib/Target/LoongArch LoongArchExpandPseudoInsts.cpp, llvm/test/CodeGen/LoongArch test_bl_fixupkind.mir

LoongArch: Simplify handling of call pseudo operands when expanding.

copyImplicitOperands isn't really intended for the call instruction case
where there are also variadic operands. This avoids duplicating the R1
implicit-def operand when expanding calls. This avoids having mixed dead
and not dead flags on the same register, which will fail a future verifier
check.

Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
DeltaFile
+22-35llvm/lib/Target/LoongArch/LoongArchExpandPseudoInsts.cpp
+2-2llvm/test/CodeGen/LoongArch/test_bl_fixupkind.mir
+24-372 files

LLVM/project 2bb259allvm/lib/CodeGen/SelectionDAG SelectionDAG.cpp, llvm/test/CodeGen/AArch64 partial-reduction-sub.ll

[SelectionDAG] Allow PARTIAL_REDUCE_SUMLA in getPartialReduceMLS (#220886)

foldPartialReduceMLAMulOp turns

    partial_reduce_*mla(acc, neg(mul(sext(a), zext(b))), splat(1))

into a PARTIAL_REDUCE_SUMLA node and, because of the negation, builds it
through SelectionDAG::getPartialReduceMLS. That helper only accepted
UMLA
and SMLA, so the mixed-sign case hit its "Unexpected opcode" assertion
whenever the target marks SUMLA as legal or custom for the transformed
types. The loop vectorizer produces exactly this shape for a reduction
chain that mixes adds and subs, e.g.

    acc += (int)s8_a[i] * (int)u8_b[i];
    acc -= (int)s8_c[i] * (int)u8_d[i];

so this crashed at -O2 in builds with assertions enabled on AArch64 with
+dotprod, and since #205373 also on X86 with AVX512-VNNI (release builds

    [9 lines not shown]
DeltaFile
+466-0llvm/test/CodeGen/X86/vector-partial-reduce-sumla.ll
+69-0llvm/test/CodeGen/AArch64/partial-reduction-sub.ll
+2-1llvm/lib/CodeGen/SelectionDAG/SelectionDAG.cpp
+537-13 files

LLVM/project 991d80dllvm/include/llvm/Option OptTable.h, llvm/lib/Option OptTable.cpp

[Option] Store the StringTable by value. NFC (#224818)

Suggested by
https://github.com/llvm/llvm-project/pull/224805#discussion_r4052408942

A StringTable is just a StringRef. Hold it by value and build it from
the storage array in optionTables(), so the generated OptionStrTable
object (16 bytes of .data.rel.ro and a relocation per table) is
unreferenced and discarded, and string lookups skip a load.

Aided by Opus 5
DeltaFile
+14-14llvm/lib/Option/OptTable.cpp
+11-11llvm/include/llvm/Option/OptTable.h
+3-2llvm/utils/TableGen/OptionParserEmitter.cpp
+28-273 files

LLVM/project 7364128llvm/lib/Target/Mips Mips.td MipsSubtarget.cpp

[Mips] Remove experimental warning bits in a few places (#223230)

- Remove the warning message "warning: MIPS-I support is experimental" from
  MipsSubtarget.cpp. The warning can interfere with build systems that do not
  expect diagnostic output on stderr.
- Remove experimental / highly experimental from a few MIPS processors.
  MIPS I is on track and the rest I wouldn't consider experimental anymore.
DeltaFile
+0-7llvm/lib/Target/Mips/MipsSubtarget.cpp
+3-3llvm/lib/Target/Mips/Mips.td
+3-102 files

LLVM/project 8a371caorc-rt/include/orc-rt/bedrock/sps SimpleRemoteCA.h, orc-rt/lib/bedrock/sps SimpleRemoteCA.cpp

[orc-rt] In SimpleRemoteCA, carry tag field as uint64_t (#224814)

A message's tag field is a 64-bit value on the wire, so carry it as a
uint64_t inside SimpleRemoteCA and only convert to an ExecutorAddr or
ResultKind only where needed, with validation.
DeltaFile
+31-9orc-rt/lib/bedrock/sps/SimpleRemoteCA.cpp
+4-5orc-rt/include/orc-rt/bedrock/sps/SimpleRemoteCA.h
+2-2orc-rt/test/unit/bedrock/sps/SimpleRemoteCATest.cpp
+37-163 files

LLVM/project b3ef5fellvm/include/llvm/Option Option.h OptTable.h, llvm/lib/Option OptTable.cpp

[Option] Shrink Info from 40 to 24 bytes (#224807)

Few options set MetaVar, AliasArgs, Values, help text variants, or
subcommands (12%, 4%, 4%, 0.1%, 0.06% of 6955 options), yet every entry
carries all five. Move them to a deduplicated InfoExtra side table
reached by a 16-bit offset; row 0 serves options that set none. Narrow
Visibility to 16 bits; clang uses seven.

clang's table shrinks from 155 KB to 93 KB plus 4 KB of extras.

Aided by Opus 5
DeltaFile
+43-35llvm/include/llvm/Option/OptTable.h
+36-34llvm/utils/TableGen/OptionParserEmitter.cpp
+2-7llvm/include/llvm/Option/Option.h
+3-3llvm/lib/Option/OptTable.cpp
+84-794 files

LLVM/project 130c6b1llvm/lib/Target/AMDGPU AMDGPUISelLowering.cpp, llvm/test/CodeGen/AMDGPU llvm.amdgcn.mul.i24.ll

[AMDGPU] Fold 24 bit multiply with zero low bits

Fold `MUL_I24` and `MUL_U24` to zero when either operand has known zero
low 24 bits.

For example:
```
  llvm.amdgcn.mul.i24(x, 0x01000000) -> 0
```
DeltaFile
+21-33llvm/test/CodeGen/AMDGPU/llvm.amdgcn.mul.i24.ll
+5-0llvm/lib/Target/AMDGPU/AMDGPUISelLowering.cpp
+26-332 files

LLVM/project 8b4c363llvm/test/CodeGen/AMDGPU llvm.amdgcn.mul.i24.ll

[AMDGPU] Add tests for zero MUL_I24 operands
DeltaFile
+155-2llvm/test/CodeGen/AMDGPU/llvm.amdgcn.mul.i24.ll
+155-21 files

LLVM/project 88c7a2allvm/lib/Target/AMDGPU AMDGPUCodeGenPrepare.cpp, llvm/test/CodeGen/AMDGPU wave32.ll spill-scavenge-offset.ll

[AMDGPU] Always optimize out llvm.amdgcn.mbcnt.hi in wave32 mode (#224573)

The instruction v_mbcnt_hi_u32_b32 does nothing useful in wave32 mode so
we can always optimize it regardless of the workgroup size.
DeltaFile
+3-16llvm/test/CodeGen/AMDGPU/wqm.ll
+9-10llvm/test/CodeGen/AMDGPU/dual-source-blend-export.ll
+3-9llvm/lib/Target/AMDGPU/AMDGPUCodeGenPrepare.cpp
+1-2llvm/test/CodeGen/AMDGPU/GlobalISel/divergence-divergent-i1-phis-no-lane-mask-merging.ll
+0-2llvm/test/CodeGen/AMDGPU/wave32.ll
+0-2llvm/test/CodeGen/AMDGPU/spill-scavenge-offset.ll
+16-411 files not shown
+16-427 files

LLVM/project 7d091d5llvm/include/llvm/Option OptTable.h, llvm/lib/Option OptTable.cpp

[Option] Derive the prefix union in the constructor. NFC (#224813)

The union is a function of the prefix table, so compute it in the
OptTable constructor instead of emitting OptionPrefixesUnion and
carrying it in Tables.

LLM-aided, extracted from #224805
DeltaFile
+4-24llvm/utils/TableGen/OptionParserEmitter.cpp
+12-6llvm/lib/Option/OptTable.cpp
+0-1llvm/include/llvm/Option/OptTable.h
+16-313 files

LLVM/project 9109452libcxx/test/std/atomics/atomics.types.operations/atomics.types.operations.wait lost_wakeup.pass.cpp

[libcxx] Avoid a busy-loop in atomic's lost_wakeup.pass.cpp (#211242)

Fixes #207763
On single-core or heavily loaded systems, using this_thread::yield()
prevents the notifier from burning all its allocated time on the
busy-loop.

The measurements were done on my Linux box, the time is the one reported
by lit. To isolate the test on a single core I used the following
commands:

export LIT_FILTER="wait/lost_wakeup.pass.cpp"
taskset -c 0 ninja -C build check-cxx

| Configuration  | Linux futex |             | Fallback |             |
| -------------- | ----------- | ----------- | -------- | ----------- |
|                | Normal      | Pinned core | Normal   | Pinned core |
| Before         | 1.50s       | 8.44s       | 1.57s    | 301.02s     |
| With yield     | 1.73s       | 7.36s       | 1.64s    | 8.88s       |

Both configurations are with the `run < 10` change.
DeltaFile
+4-1libcxx/test/std/atomics/atomics.types.operations/atomics.types.operations.wait/lost_wakeup.pass.cpp
+4-11 files

LLVM/project 3efbad6llvm/tools/llvm-ifs llvm-ifs.cpp, llvm/tools/llvm-objcopy ObjcopyOptions.cpp

[Option] Replace the OptionTables object with a function. NFC (#224805)

The generated `Tables` holds seven addresses, so needs dynamic
relocations in PIE and shared library builds:

```
static constexpr llvm::opt::OptTable::Tables OptionTables = {
    OptionStrTable, OptionPrefixesTable, OptionPrefixesUnion, ...};
...
  FooOptTable() : OptTable(OptionTables) {}
```

Emit a function instead. The caller builds the aggregate on the stack
from PC-relative addresses, which the linker resolves:

```
static constexpr llvm::opt::OptTable::Tables optionTables() {
  return {OptionStrTable, OptionPrefixesTable, OptionInfoTable, ...};
}

    [7 lines not shown]
DeltaFile
+27-25llvm/utils/TableGen/OptionParserEmitter.cpp
+6-5llvm/tools/llvm-objcopy/ObjcopyOptions.cpp
+2-2llvm/tools/llvm-rc/llvm-rc.cpp
+2-2llvm/tools/llvm-objdump/llvm-objdump.cpp
+3-1llvm/tools/llvm-readtapi/llvm-readtapi.cpp
+3-1llvm/tools/llvm-ifs/llvm-ifs.cpp
+43-3642 files not shown
+89-7948 files

LLVM/project 6f516d1flang/lib/Optimizer/Transforms/CUDA CUFAddConstructor.cpp, flang/test/Fir/CUDA cuda-constructor-2.f90

[flang][cuda] Emit cuda sentinel only when compiling main (#224507)

The sentinel must be added only when the `PROGRAM` is build in the TU.
DeltaFile
+20-6flang/test/Fir/CUDA/cuda-constructor-2.f90
+6-2flang/lib/Optimizer/Transforms/CUDA/CUFAddConstructor.cpp
+26-82 files

LLVM/project 2230387flang/include/flang/Optimizer/Builder MutableBox.h, flang/include/flang/Optimizer/Dialect/CUF CUFOps.td

[flang][cuda] Split realloc for host-accessible CUDA Fortran allocatable assigns (#224195)

Allow SeparateAllocatableAssign to peel reallocation off hlfir.assign
for
managed, pinned, and unified allocatables, using cuf.alloc/cuf.free so
the
CUDA Fortran allocator is honored. Device (and other
non-host-accessible)
storage still keeps the runtime realloc path as well as derived-type.

Drop the extra op-level MemFree on cuf.free so OptimizedBufferization
can
see the freed pointer. Without that, a valueless Free effect between the
elemental and the assign forced a temporary for cases such as
`a = int(ran*100)` on a managed allocatable in the example below.

```
subroutine foo(n, a, ran)
  integer :: n

    [7 lines not shown]
DeltaFile
+51-0flang/test/HLFIR/opt-bufferization-dealloc-conflict.fir
+38-11flang/lib/Optimizer/Builder/MutableBox.cpp
+24-11flang/lib/Optimizer/HLFIR/Transforms/SeparateAllocatableAssign.cpp
+28-7flang/test/HLFIR/separate-allocatable-assign.fir
+5-2flang/include/flang/Optimizer/Builder/MutableBox.h
+1-1flang/include/flang/Optimizer/Dialect/CUF/CUFOps.td
+147-326 files

LLVM/project e35f904clang/include/clang/CIR/Dialect/IR CIRDialect.td, clang/lib/CIR/CodeGen CIRGenModule.h CIRGenFunction.h

[CIR][SYCL] Emit attributes on the SYCL kernel caller (#224667)

Emit norecurse, mustprogress, and sycl-module-id.

This matches the classic behavior that is specific to the SYCL kernel
caller.

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply at anthropic.com>
DeltaFile
+47-0clang/test/CIR/CodeGenSYCL/kernel-caller-attributes.cpp
+27-3clang/lib/CIR/CodeGen/CIRGenSYCL.cpp
+21-6clang/lib/CIR/Lowering/DirectToLLVM/LowerToLLVM.cpp
+16-0clang/lib/CIR/CodeGen/CIRGenFunction.h
+3-0clang/include/clang/CIR/Dialect/IR/CIRDialect.td
+2-0clang/lib/CIR/CodeGen/CIRGenModule.h
+116-96 files

LLVM/project 18f2436lldb/test/API/functionalities/scripted_process TestScriptedProcess.py

[lldb/test] Register the addressable_bits_scripted_process module with the script interpreter (#224803)

`test_scripted_process_addressable_bits` imported
`addressable_bits_scripted_process` into the test process, which makes
the module's constants available to the test but tells lldb's own script
interpreter nothing about it, so
`SBLaunchInfo::SetScriptedProcessClassName()` had nothing to resolve:

```
  error: failed to create ScriptedProcess: failed to create script object:
  Could not find script class:
  addressable_bits_scripted_process.AddressableBitsScriptedProcess
```

Import the file with `command script import` the way the other scripted
process tests in this file do.
DeltaFile
+5-0lldb/test/API/functionalities/scripted_process/TestScriptedProcess.py
+5-01 files

LLVM/project 5d05951flang/test/Driver bbc-openmp-target-context.f90, flang/tools/bbc bbc.cpp

Set bbc target context for semantics
DeltaFile
+17-0flang/test/Driver/bbc-openmp-target-context.f90
+2-0flang/tools/bbc/bbc.cpp
+19-02 files

LLVM/project e80cb50clang/lib/StaticAnalyzer/Checkers/WebKit RawPtrRefLocalVarsChecker.cpp

[WebKit Checkers][NFC] Extract hasGuardian from isPtrOriginSafe in RawPtrRefLocalVarsChecker (#224726)

So an upcoming borrow checker can skip it.

(A guardian variable is an independent declaration that ensures a
lifetime. Borrow checking does not accept guardian variables because
they do not convey `lifetimebound` links.)

Assisted-by: Claude
DeltaFile
+33-31clang/lib/StaticAnalyzer/Checkers/WebKit/RawPtrRefLocalVarsChecker.cpp
+33-311 files

LLVM/project c38a705lldb/source/Plugins/ObjectFile/ELF ObjectFileELF.cpp, lldb/source/Plugins/ObjectFile/Mach-O ObjectFileMachO.cpp

[lldb] Bound DataExtractor::PeekCStr to the data it reads from (#224452)

PeekCStr only validated the first byte without checking for a terminator.
Every caller then treated the result as a C string, so an unterminated
string table let strlen run past the mapping, causing a crash or
security bug.

Return an optional StringRef, produced only when the terminator lies
within the data. The length comes along with it, so callers no longer
rescan, and an absent string stays distinct from an empty one.

ValueObject was the only caller that wanted an unterminated buffer. It
scans fixed size chunks of inferior memory, so it now asks for bytes with
PeekData and bounds its own scan.

rdar://186891223
DeltaFile
+46-40lldb/source/Plugins/ObjectFile/Mach-O/ObjectFileMachO.cpp
+79-0lldb/test/Shell/ObjectFile/MachO/unterminated-symbol-name.yaml
+24-30lldb/source/Plugins/ObjectFile/ELF/ObjectFileELF.cpp
+27-24lldb/source/Utility/DataExtractor.cpp
+19-0lldb/unittests/Utility/DataExtractorTest.cpp
+14-4lldb/source/ValueObject/ValueObject.cpp
+209-985 files not shown
+242-11711 files

LLVM/project 0da0168lldb/source/Plugins/Process/FreeBSD-Kernel-Core ProcessFreeBSDKernelCore.h ProcessFreeBSDKernelCore.cpp

Revert "[lldb][FreeBSDKernel] Support cross-architecture FreeBSD kernel cores" (#224799)

Reverts llvm/llvm-project#223593

Merged by mistake.
DeltaFile
+5-33lldb/source/Plugins/Process/FreeBSD-Kernel-Core/ProcessFreeBSDKernelCore.cpp
+0-10lldb/source/Plugins/Process/FreeBSD-Kernel-Core/ProcessFreeBSDKernelCore.h
+5-432 files

LLVM/project 3088d0alldb/source/Plugins/Process/FreeBSD-Kernel-Core ProcessFreeBSDKernelCore.h ProcessFreeBSDKernelCore.cpp

[lldb][FreeBSDKernel] Support cross-architecture FreeBSD kernel cores (#223593)

FreeBSD's kvm_open(3) man page states that a valid resolver is required
to debug non-native kernel images as libkvm needs to map symbol names to
kernel virtual addresses. Since ProcessFreeBSDKernelCore aims to debug
kernel image and core dump from any architectures, pass a valid resolver
to kvm_open2().

Assisted-by: GPT
DeltaFile
+33-5lldb/source/Plugins/Process/FreeBSD-Kernel-Core/ProcessFreeBSDKernelCore.cpp
+10-0lldb/source/Plugins/Process/FreeBSD-Kernel-Core/ProcessFreeBSDKernelCore.h
+43-52 files

LLVM/project 041328elldb/source/Plugins/DynamicLoader/FreeBSD-Kernel DynamicLoaderFreeBSDKernel.cpp

[lldb][FreeBSDKernel] Avoid null dereference probing FreeBSD kernels (#223590)

CheckForKernelImageAtAddress() accepts an optional read_error pointer,
but dereferenced it unconditionally in many code paths. If read_error is
null, point read_error to a dummy local variable to prevent null
dereference.

Fixes: b3cc4804d45d6b612ac9b3cc47ebbb0da44ebc60
DeltaFile
+6-2lldb/source/Plugins/DynamicLoader/FreeBSD-Kernel/DynamicLoaderFreeBSDKernel.cpp
+6-21 files

LLVM/project 7192e07lldb/source/Plugins/Process/FreeBSD-Kernel-Core ProcessFreeBSDKernelCore.cpp

[lldb][FreeBSDKernel] Load unrelocated FreeBSD kernel sections (#223589)

Even though displacement is zero, we need to mark the module as loaded.

Fixes: 62d06083ef186abe610de4742c62c84c2f97fcd0
DeltaFile
+0-3lldb/source/Plugins/Process/FreeBSD-Kernel-Core/ProcessFreeBSDKernelCore.cpp
+0-31 files

LLVM/project d035826clang/lib/AST/ByteCode State.h State.cpp

[clang][bytecode] Flip a boolean flag meaning (#224649)

We only use this flag once, when passing it to
`setFoldFailureDiagnostic()` (where we invert it), so it doesn't make
sense to call it IsCCEDiag.
DeltaFile
+6-6clang/lib/AST/ByteCode/State.cpp
+1-1clang/lib/AST/ByteCode/State.h
+7-72 files

LLVM/project da72f2cllvm/lib/Support/Unix Path.inc

[Support] Only enable use of /proc/self/fd on Linux (#222497)

Only enable the use of /proc/self/fd on OSes that provide
the compatible /proc/self/fd interface.
DeltaFile
+3-5llvm/lib/Support/Unix/Path.inc
+3-51 files

LLVM/project 3834f58clang/lib/AST/ByteCode Interp.cpp

[clang][bytecode] Only call getSource() if we diagnose (#224663)

We can avoid this if we only return false anyway. Also rename a
variable.
DeltaFile
+7-7clang/lib/AST/ByteCode/Interp.cpp
+7-71 files