[SLP]Vectorize consecutive loads with undef lanes as a wide load
Model the lane with the absorbing constant (0 for mul/and, -1 for or) of
a copyable node as op(V, undef), so the operand column of the other
lanes gets an undef lane. Cover such undef lanes in a column of
consecutive loads with a single frozen vector load, if the whole range
is dereferenceable.
Fixes #46897
Assisted-by: Cursor
Reviewers: RKSimon
Pull Request: https://github.com/llvm/llvm-project/pull/228872
fix(SelectionDAG): validate identity fold demands
KnownBits returned for the LHS may only be valid for the demand
already reduced by the RHS. Check the full result demand before
folding AND/OR to the LHS.
Share the masked-bit check between both operations. Drop Disjoint
when the query rewrites an OR operand.
[orc-rt] Add SPSSymbolLookupSet typedef, clean up users. (#230876)
Existing deserializers of SymbolLookupSet were spelling out the SPS type
in full (SPSSequence<SPSTuple<SPSString, bool>>). Define an
SPSSymbolLookupSet typedef and use in instead so that deserialization
points can pick up any future changes automatically.
Address review feedback on FP8 conversion builtins
Reuse err_builtin_invalid_arg_type for all source operand errors
instead of builtin-specific diagnostics.
Accept integer constants that fit the format width, such as 0x38,
and std::byte, so common byte values need no explicit cast.
Drop the unused OpenCL fp64 path from checkFloatingPointTypeSupport.
Document floating-point environment and fast-math behavior; trim
implementation detail from the user docs.
Change-Id: I8676a49540bdbbb8a90827c83764803eeea850a5
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[clang] Add elementwise conversions from encoded FP8 values
Add nine builtins converting Float8E5M2, Float8E4M3FN, and Float8E5M3FNU
encodings to _Float16, __bf16, or float through
llvm.convert.from.arbitrary.fp.
Accept exactly 8-bit integer scalars and generic fixed-length vectors,
preserving vector kinds and element counts without integer promotions.
Also support scalar AArch64 __mfp8 containers. Apply destination type
availability checks, including deferred offload diagnostics, and reject
constant-expression use.
Document the interface and add semantic, template, language-mode,
target-specific, and IR-generation regression coverage.
[AMDGPU] Form VOPD dot2 pairs with a literal in src1
A V_DOT2 with a register in src0 and a literal in src1 is matched as if
commuted, and commuted when the VOPD pair is built. This recovers the
pairs lost once MachineCSE stopped leaving those literals in src0.
Co-Authored-By: Claude <noreply at anthropic.com>
[AMDGPU] Allow commuting immediates out of src0 when legal
isLegalToSwap refused to move any non-inline constant out of src0, so
commuting an instruction with a literal in src1 could not be undone.
AMDGPULowerVGPREncoding relies on undoing it and, on gfx1250, either hit
"Failed to restore commuted instruction" or kept the commuted
instruction with the wrong VGPR MSB mode, which made it read the wrong
VGPRs.
Allow an immediate to leave src0 when the other operand can hold it.
VOPD formation now accepts a V_DOT2 with a literal in src1 if
isLegalToSwap allows the swap, and commutes it when the pair is built,
so those pairs are still formed.
Co-Authored-By: Claude <noreply at anthropic.com>
[libc][CI] Update FreeBSD version to 15 and bump VM action (#230775)
The LLVM libc FreeBSD CI has recently been failing because upstream
package repositories require a newer FreeBSD version than the pinned
15.0 image.
This patch updates the FreeBSD version specification from 15.0 to 15,
allowing the action to automatically pull the latest minor snapshot.
It also bumps the VM action to the latest release for improved
compatibility with Ubuntu 26.04 runners.
ref:
https://github.com/llvm/llvm-project/actions/runs/37893910030/job/113701458175?pr=230371
- before
```
Processing entries:
Newer FreeBSD version for package zh-qe:
[11 lines not shown]
VE: Remove broken nested call frame around dynamic stack allocation (#229011)
lowerDYNAMIC_STACKALLOC wrapped the __ve_grow_stack call and the
GETSTACKTOP stack-pointer read in a zero-sized CALLSEQ_START/CALLSEQ_END
pair. The call it contains emits its own CALLSEQ, so the outer bracket
only produced a nested ADJCALLSTACKDOWN 0 / ADJCALLSTACKUP 0 around the
inner ADJCALLSTACKDOWN / ADJCALLSTACKUP which is illegal.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
[clang-repl] Mark global-dtor.cpp and value-print-temporaries.cpp unsupported under ASan (#230870)
These tests are flaky on x86_64 Linux ASan bots with `out of range of
Delta32
fixup` JITLink errors, similar to #102858, #135401, and #150242.
This likely happens because `InProcessMemoryManager` maps each
incremental
module with a separate `mmap` call, and depending on the address space
layout
some allocations appear to end up on opposite sides of ASan's large
allocator
reservation (> 2 GiB apart).
Assisted-by: Gemini
fix(SelectionDAG): simplify commuted demanded bits
AND/OR demanded-bit simplification uses RHS known bits to simplify the
LHS, but does not retry the RHS using LHS known bits, making
optimizations depend on operand order.
Retry the RHS when the LHS reduces its demanded bits. Add AArch64
and AMDGPU codegen coverage.
fix(SelectionDAG): isolate retry known bits
The LHS demanded-bits query already excludes bits masked by the RHS.
Its returned facts cannot safely reduce the RHS demand in turn.
Query LHS known bits independently, and preserve the original RHS facts
when a reduced-demand retry does not simplify. Defer the retry until
existing folds fail and skip constant or shared RHS operands.
Refresh the affected X86 atomic codegen checks.
Refs #230700
[DAGCombiner] Narrow the integer source of uint_to_fp (#222899)
Truncate the source of a `uint_to_fp` when it is known to fit in a
narrower
type the target can convert from directly.
For example:
```
uitofp (and i64 %x, 255) to float
```
On AMDGPU this becomes a single `v_cvt_f32_ubyte0` instead of the
generic
i64 to f32 expansion.