LLVM/project 611eee0 — llvm/lib/Transforms/Vectorize SLPVectorizer.cpp, llvm/lib/Transforms/Vectorize/SLPVectorizer SLPUtils.cpp SLPMemoryUtils.h

[SLP]Vectorize consecutive loads with undef lanes as a wide load

Model the lane with the absorbing constant (0 for mul/and, -1 for or) of
a copyable node as op(V, undef), so the operand column of the other
lanes gets an undef lane. Cover such undef lanes in a column of
consecutive loads with a single frozen vector load, if the whole range
is dereferenceable.

Fixes #46897

Assisted-by: Cursor

Reviewers: RKSimon

Pull Request: https://github.com/llvm/llvm-project/pull/228872
DeltaFile
+124-16llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+50-86llvm/test/Transforms/SLPVectorizer/X86/absorbing-copyable-lane.ll
+40-63llvm/test/Transforms/SLPVectorizer/X86/wide-load-absorbed-lane.ll
+54-0llvm/lib/Transforms/Vectorize/SLPVectorizer/SLPMemoryUtils.cpp
+16-0llvm/lib/Transforms/Vectorize/SLPVectorizer/SLPMemoryUtils.h
+15-0llvm/lib/Transforms/Vectorize/SLPVectorizer/SLPUtils.cpp
+299-1652 files not shown
+310-1708 files

LLVM/project 9048ff1 — llvm/lib/CodeGen/SelectionDAG TargetLowering.cpp

fix(SelectionDAG): validate identity fold demands

KnownBits returned for the LHS may only be valid for the demand
already reduced by the RHS. Check the full result demand before
folding AND/OR to the LHS.

Share the masked-bit check between both operations. Drop Disjoint
when the query rewrites an OR operand.
DeltaFile
+59-33llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
+59-331 files

LLVM/project 8ff4f38 — orc-rt/include/orc-rt/support/sps SPSSymbolLookupSet.h, orc-rt/lib/bedrock/sps NativeDylibManagerSPSCI.cpp

[orc-rt] Add SPSSymbolLookupSet typedef, clean up users. (#230876)

Existing deserializers of SymbolLookupSet were spelling out the SPS type
in full (SPSSequence<SPSTuple<SPSString, bool>>). Define an
SPSSymbolLookupSet typedef and use in instead so that deserialization
points can pick up any future changes automatically.
DeltaFile
+3-3orc-rt/lib/bedrock/sps/NativeDylibManagerSPSCI.cpp
+2-3orc-rt/test/unit/support/sps/SPSSymbolLookupSetTest.cpp
+1-2orc-rt/test/unit/bedrock/sps/NativeDylibManagerSPSCITest.cpp
+2-0orc-rt/include/orc-rt/support/sps/SPSSymbolLookupSet.h
+8-84 files

LLVM/project 8847f81 — clang/docs LanguageExtensions.md, clang/include/clang/Basic DiagnosticSemaKinds.td

Address review feedback on FP8 conversion builtins

Reuse err_builtin_invalid_arg_type for all source operand errors
instead of builtin-specific diagnostics.

Accept integer constants that fit the format width, such as 0x38,
and std::byte, so common byte values need no explicit cast.

Drop the unused OpenCL fp64 path from checkFloatingPointTypeSupport.

Document floating-point environment and fast-math behavior; trim
implementation detail from the user docs.

Change-Id: I8676a49540bdbbb8a90827c83764803eeea850a5
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
DeltaFile
+44-22clang/lib/Sema/SemaChecking.cpp
+30-9clang/test/Sema/builtins-elementwise-convert-from-arbitrary-fp.c
+10-17clang/lib/Sema/SemaType.cpp
+24-1clang/test/SemaCXX/builtins-elementwise-convert-from-arbitrary-fp.cpp
+8-7clang/docs/LanguageExtensions.md
+1-11clang/include/clang/Basic/DiagnosticSemaKinds.td
+117-674 files not shown
+131-7110 files

LLVM/project 476eb9d — clang/docs LanguageExtensions.md, clang/lib/Sema SemaType.cpp SemaChecking.cpp

[clang] Add elementwise conversions from encoded FP8 values

Add nine builtins converting Float8E5M2, Float8E4M3FN, and Float8E5M3FNU
encodings to _Float16, __bf16, or float through
llvm.convert.from.arbitrary.fp.

Accept exactly 8-bit integer scalars and generic fixed-length vectors,
preserving vector kinds and element counts without integer promotions.
Also support scalar AArch64 __mfp8 containers. Apply destination type
availability checks, including deferred offload diagnostics, and reject
constant-expression use.

Document the interface and add semantic, template, language-mode,
target-specific, and IR-generation regression coverage.
DeltaFile
+172-0clang/test/CodeGen/builtins-elementwise-convert-from-arbitrary-fp.c
+135-0clang/test/Sema/builtins-elementwise-convert-from-arbitrary-fp.c
+105-0clang/lib/Sema/SemaChecking.cpp
+97-0clang/docs/LanguageExtensions.md
+50-18clang/lib/Sema/SemaType.cpp
+44-0clang/test/Sema/builtins-elementwise-convert-from-arbitrary-fp-target.c
+603-1810 files not shown
+804-1816 files

LLVM/project 9703767 — llvm/lib/Target/AMDGPU SIInstrInfo.h GCNCreateVOPD.cpp, llvm/test/CodeGen/AMDGPU llvm.amdgcn.fdot2.f32.bf16.ll llvm.amdgcn.fdot2.ll

[AMDGPU] Form VOPD dot2 pairs with a literal in src1

A V_DOT2 with a register in src0 and a literal in src1 is matched as if
commuted, and commuted when the VOPD pair is built. This recovers the
pairs lost once MachineCSE stopped leaving those literals in src0.

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+149-0llvm/test/CodeGen/AMDGPU/vopd-dot2-commute-imm-src1.mir
+54-19llvm/lib/Target/AMDGPU/GCNVOPDUtils.cpp
+12-35llvm/test/CodeGen/AMDGPU/llvm.amdgcn.fdot2.ll
+3-6llvm/test/CodeGen/AMDGPU/llvm.amdgcn.fdot2.f32.bf16.ll
+9-0llvm/lib/Target/AMDGPU/GCNCreateVOPD.cpp
+2-2llvm/lib/Target/AMDGPU/SIInstrInfo.h
+229-621 files not shown
+232-627 files

LLVM/project c048ebc — llvm/lib/Target/AMDGPU GCNVOPDUtils.cpp, llvm/test/CodeGen/AMDGPU vopd-dot2-commute-imm-src1.mir

[AMDGPU] Address review comments

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+9-5llvm/lib/Target/AMDGPU/GCNVOPDUtils.cpp
+1-1llvm/test/CodeGen/AMDGPU/vopd-dot2-commute-imm-src1.mir
+10-62 files

LLVM/project 72e3ec8 — llvm/test/CodeGen/AMDGPU clmul.ll

[AMDGPU] Update clmul.ll for literal commute

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+4-4llvm/test/CodeGen/AMDGPU/clmul.ll
+4-41 files

LLVM/project cb403d6 — llvm/test/CodeGen/AMDGPU udiv.ll mul.ll

[AMDGPU] Update more tests from main for literal commute

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+152-152llvm/test/CodeGen/AMDGPU/frem.ll
+10-10llvm/test/CodeGen/AMDGPU/udiv.ll
+10-10llvm/test/CodeGen/AMDGPU/mul.ll
+172-1723 files

LLVM/project 2d194c6 — llvm/utils/gn/secondary/bolt/unittests/Core BUILD.gn

[gn build] Port a09df97a747bb (#230873)
DeltaFile
+1-0llvm/utils/gn/secondary/bolt/unittests/Core/BUILD.gn
+1-01 files

LLVM/project 5bc8a29 — llvm/test/CodeGen/AMDGPU commute-literal-src0-cse.mir

[AMDGPU] Drop -verify-machineinstrs from commute-literal-src0-cse.mir

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+1-1llvm/test/CodeGen/AMDGPU/commute-literal-src0-cse.mir
+1-11 files

LLVM/project 37bbd5d — llvm/test/CodeGen/AMDGPU fsub.f16.ll fmul.f16.ll, llvm/test/CodeGen/AMDGPU/GlobalISel fshr.ll fshl.ll

[AMDGPU] Update tests from main for literal commute

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+24-24llvm/test/CodeGen/AMDGPU/GlobalISel/fptosi.bf16.ll
+22-22llvm/test/CodeGen/AMDGPU/GlobalISel/fshr.ll
+22-22llvm/test/CodeGen/AMDGPU/GlobalISel/fshl.ll
+12-12llvm/test/CodeGen/AMDGPU/cvt_f32_ubyte.ll
+4-4llvm/test/CodeGen/AMDGPU/fmul.f16.ll
+2-2llvm/test/CodeGen/AMDGPU/fsub.f16.ll
+86-866 files

LLVM/project f9da2ba — llvm/lib/Target/AMDGPU SIInstrInfo.h GCNCreateVOPD.cpp, llvm/test/CodeGen/AMDGPU llvm.amdgcn.fdot2.f32.bf16.ll llvm.amdgcn.fdot2.ll

[AMDGPU] Move the VOPD dot2 changes to a separate PR

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+0-149llvm/test/CodeGen/AMDGPU/vopd-dot2-commute-imm-src1.mir
+19-54llvm/lib/Target/AMDGPU/GCNVOPDUtils.cpp
+35-12llvm/test/CodeGen/AMDGPU/llvm.amdgcn.fdot2.ll
+0-9llvm/lib/Target/AMDGPU/GCNCreateVOPD.cpp
+6-3llvm/test/CodeGen/AMDGPU/llvm.amdgcn.fdot2.f32.bf16.ll
+2-2llvm/lib/Target/AMDGPU/SIInstrInfo.h
+62-2291 files not shown
+62-2327 files

LLVM/project 73f34a0 — llvm/test/CodeGen/AMDGPU fptrunc.f16.ll

[AMDGPU] Update fptrunc.f16.ll for literal commute

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+72-72llvm/test/CodeGen/AMDGPU/fptrunc.f16.ll
+72-721 files

LLVM/project 06b3d03 — llvm/test/CodeGen/AMDGPU amdgcn.bitcast.896bit.ll amdgcn.bitcast.960bit.ll

[AMDGPU] Allow commuting immediates out of src0 when legal

isLegalToSwap refused to move any non-inline constant out of src0, so
commuting an instruction with a literal in src1 could not be undone.
AMDGPULowerVGPREncoding relies on undoing it and, on gfx1250, either hit
"Failed to restore commuted instruction" or kept the commuted
instruction with the wrong VGPR MSB mode, which made it read the wrong
VGPRs.

Allow an immediate to leave src0 when the other operand can hold it.
VOPD formation now accepts a V_DOT2 with a literal in src1 if
isLegalToSwap allows the swap, and commutes it when the pair is built,
so those pairs are still formed.

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+2,824-2,824llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+1,412-1,412llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+706-706llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.256bit.ll
+496-496llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.320bit.ll
+480-480llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.960bit.ll
+448-448llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.896bit.ll
+6,366-6,366201 files not shown
+13,734-13,261207 files

LLVM/project 16c05cd — llvm/utils/gn/secondary/clang/unittests/AST BUILD.gn

[gn build] Port dba1b67f855fb (#230874)
DeltaFile
+1-0llvm/utils/gn/secondary/clang/unittests/AST/BUILD.gn
+1-01 files

LLVM/project daf014f — llvm/utils/gn/secondary/llvm/lib/Target/AMDGPU BUILD.gn

[gn build] Port 8c4c52bb62ee (#230872)
DeltaFile
+1-0llvm/utils/gn/secondary/llvm/lib/Target/AMDGPU/BUILD.gn
+1-01 files

LLVM/project a9fd78b — llvm/utils/gn/secondary/compiler-rt/lib/builtins BUILD.gn

[gn] "port" 5a1dbdd2b16d3 (#230871)
DeltaFile
+8-0llvm/utils/gn/secondary/compiler-rt/lib/builtins/BUILD.gn
+8-01 files

LLVM/project 482b047 — mlir/lib/Dialect/Affine/IR AffineOps.cpp, mlir/lib/Dialect/ArmSVE/Transforms LegalizeVectorStorage.cpp

[mlir][NFC] Simplify using OpRewritePattern::OpRewritePattern to use Base alias (#230717)

Aligns with #158433.
DeltaFile
+10-10mlir/lib/Dialect/Math/Transforms/PolynomialApproximation.cpp
+9-9mlir/lib/Dialect/SparseTensor/Transforms/SparseTensorRewriting.cpp
+7-7mlir/lib/Dialect/Tosa/IR/TosaCanonicalizations.cpp
+7-7mlir/lib/Dialect/MemRef/Transforms/ExpandStridedMetadata.cpp
+6-6mlir/lib/Dialect/Affine/IR/AffineOps.cpp
+5-5mlir/lib/Dialect/ArmSVE/Transforms/LegalizeVectorStorage.cpp
+44-4442 files not shown
+114-11448 files

LLVM/project 2e5fa05 — llvm/test/CodeGen/AMDGPU mul.ll fcanonicalize.ll, llvm/test/CodeGen/X86 mul-constant-i64.ll

Merge branch 'main' into users/c8ef/tanpi-fix
DeltaFile
+32,088-15,710llvm/test/CodeGen/AMDGPU/frem.ll
+5,761-3,731llvm/test/CodeGen/AMDGPU/srem.ll
+6,118-2,908llvm/test/CodeGen/AMDGPU/clmul.ll
+1,502-5,317llvm/test/CodeGen/AMDGPU/fcanonicalize.ll
+1,327-1,552llvm/test/CodeGen/AMDGPU/mul.ll
+679-1,139llvm/test/CodeGen/X86/mul-constant-i64.ll
+47,475-30,3571,165 files not shown
+98,949-56,7731,171 files

LLVM/project db316ea — .github/workflows libc-freebsd-vm-tests.yml

[libc][CI] Update FreeBSD version to 15 and bump VM action (#230775)

The LLVM libc FreeBSD CI has recently been failing because upstream 
package repositories require a newer FreeBSD version than the pinned 
15.0 image.

This patch updates the FreeBSD version specification from 15.0 to 15, 
allowing the action to automatically pull the latest minor snapshot. 
It also bumps the VM action to the latest release for improved 
compatibility with Ubuntu 26.04 runners.

ref:
https://github.com/llvm/llvm-project/actions/runs/37893910030/job/113701458175?pr=230371

- before

```
  Processing entries: 
  Newer FreeBSD version for package zh-qe:

    [11 lines not shown]
DeltaFile
+2-2.github/workflows/libc-freebsd-vm-tests.yml
+2-21 files

LLVM/project dc3cff3 — llvm/lib/Target/VE VEISelLowering.cpp, llvm/test/CodeGen/VE/Scalar alloca_aligned.ll alloca.ll

VE: Remove broken nested call frame around dynamic stack allocation (#229011)

lowerDYNAMIC_STACKALLOC wrapped the __ve_grow_stack call and the
GETSTACKTOP stack-pointer read in a zero-sized CALLSEQ_START/CALLSEQ_END
pair. The call it contains emits its own CALLSEQ, so the outer bracket
only produced a nested ADJCALLSTACKDOWN 0 / ADJCALLSTACKUP 0 around the
inner ADJCALLSTACKDOWN / ADJCALLSTACKUP which is illegal.

Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
DeltaFile
+0-7llvm/lib/Target/VE/VEISelLowering.cpp
+1-1llvm/test/CodeGen/VE/Scalar/alloca_aligned.ll
+1-1llvm/test/CodeGen/VE/Scalar/alloca.ll
+2-93 files

LLVM/project a0d6f05 — clang/test/Interpreter value-print-temporaries.cpp global-dtor.cpp

[clang-repl] Mark global-dtor.cpp and value-print-temporaries.cpp unsupported under ASan (#230870)

These tests are flaky on x86_64 Linux ASan bots with `out of range of
Delta32
fixup` JITLink errors, similar to #102858, #135401, and #150242.

This likely happens because `InProcessMemoryManager` maps each
incremental
module with a separate `mmap` call, and depending on the address space
layout
some allocations appear to end up on opposite sides of ASan's large
allocator
reservation (> 2 GiB apart).

Assisted-by: Gemini
DeltaFile
+3-0clang/test/Interpreter/value-print-temporaries.cpp
+3-0clang/test/Interpreter/global-dtor.cpp
+6-02 files

LLVM/project a02a035 — clang/test/Interpreter value-print-temporaries.cpp global-dtor.cpp

[𝘀𝗽𝗿] initial version

Created using spr 1.3.7
DeltaFile
+3-0clang/test/Interpreter/value-print-temporaries.cpp
+3-0clang/test/Interpreter/global-dtor.cpp
+6-02 files

LLVM/project 7e4deb5 — llvm/test/CodeGen/AArch64 andorxor.ll, llvm/test/CodeGen/AMDGPU and.ll or.ll

fix(SelectionDAG): simplify commuted demanded bits

AND/OR demanded-bit simplification uses RHS known bits to simplify the
LHS, but does not retry the RHS using LHS known bits, making
optimizations depend on operand order.

Retry the RHS when the LHS reduces its demanded bits. Add AArch64
and AMDGPU codegen coverage.
DeltaFile
+258-286llvm/test/CodeGen/X86/float-to-arbitrary-fp.ll
+121-135llvm/test/CodeGen/X86/atomic-rm-bit-test.ll
+99-100llvm/test/CodeGen/AMDGPU/float-to-arbitrary-fp-fp8-hw.ll
+78-0llvm/test/CodeGen/AArch64/andorxor.ll
+46-0llvm/test/CodeGen/AMDGPU/or.ll
+44-0llvm/test/CodeGen/AMDGPU/and.ll
+646-52118 files not shown
+756-62524 files

LLVM/project 97f631a — llvm/lib/CodeGen/SelectionDAG TargetLowering.cpp, llvm/test/CodeGen/X86 atomic-rm-bit-test.ll

fix(SelectionDAG): isolate retry known bits

The LHS demanded-bits query already excludes bits masked by the RHS.
Its returned facts cannot safely reduce the RHS demand in turn.

Query LHS known bits independently, and preserve the original RHS facts
when a reduced-demand retry does not simplify. Defer the retry until
existing folds fail and skip constant or shared RHS operands.

Refresh the affected X86 atomic codegen checks.

Refs #230700
DeltaFile
+79-93llvm/test/CodeGen/X86/atomic-rm-bit-test.ll
+32-18llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
+111-1112 files

LLVM/project 86d42a0 — llvm/lib/CodeGen/SelectionDAG DAGCombiner.cpp, llvm/test/CodeGen/AArch64 fold-int-pow2-with-fmul-or-fdiv.ll

[DAGCombiner] Narrow the integer source of uint_to_fp (#222899)

Truncate the source of a `uint_to_fp` when it is known to fit in a
narrower
type the target can convert from directly.
For example:
```
    uitofp (and i64 %x, 255) to float
```
On AMDGPU this becomes a single `v_cvt_f32_ubyte0` instead of the
generic
i64 to f32 expansion.
DeltaFile
+208-709llvm/test/CodeGen/AMDGPU/int_to_fp_narrow_i64.ll
+314-11llvm/test/CodeGen/AMDGPU/cvt_f32_ubyte.ll
+1-27llvm/test/CodeGen/AMDGPU/fold-int-pow2-with-fmul-or-fdiv.ll
+26-0llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp
+4-4llvm/test/CodeGen/X86/fold-int-pow2-with-fmul-or-fdiv.ll
+1-1llvm/test/CodeGen/AArch64/fold-int-pow2-with-fmul-or-fdiv.ll
+554-7526 files

LLVM/project cb11298 — llvm/test/tools/llubi bitinsert_bitextract_be.ll bitinsert_bitextract_le.ll, llvm/tools/llubi/lib Context.h Context.cpp

[llubi] Add support for `bitinsert` and `bitextract`
DeltaFile
+78-0llvm/test/tools/llubi/bitinsert_bitextract_be.ll
+78-0llvm/test/tools/llubi/bitinsert_bitextract_le.ll
+29-0llvm/tools/llubi/lib/Interpreter.cpp
+19-0llvm/tools/llubi/lib/Context.cpp
+7-0llvm/tools/llubi/lib/Context.h
+211-05 files

LLVM/project c084c07 — llvm/test/tools/llubi bitcast_le.ll bitcast_be.ll, llvm/tools/llubi/lib Context.cpp

[llubi] Fix wrong assert when writing `poison` at an unaligned bit offset
DeltaFile
+5-4llvm/tools/llubi/lib/Context.cpp
+4-0llvm/test/tools/llubi/bitcast_le.ll
+4-0llvm/test/tools/llubi/bitcast_be.ll
+13-43 files

LLVM/project 82ef6e0 — llvm/tools/llubi/lib Context.cpp

Fix formatting
DeltaFile
+2-2llvm/tools/llubi/lib/Context.cpp
+2-21 files