perf(SelectionDAG): reduce KnownBits temporaries
Let SimplifyDemandedBits initialize the fold-check result. Build the
RHS demand in one APInt instead of copying both KnownBits masks.
[AMDGPU] Form VOPD dot2 pairs with a literal in src1 (#230183)
This PR restores VOPD pair formation after the legality checks
relaxation introduced by #229906. Before the legality check relaxation,
MachineCSE was commuting immediate operands from src1 to src0, and then
failing to commute them back, which inadvertently results in the
immediate operands in src0, and the VOPD pairing would succeed. After
the legality relaxation, MachineCSE is now able to successfully commute
the immediates back from src0 to src1, which breaks VOPD pairing since
the pass expected the immediates to be in src0 position. This change
adds a check in GCNVOPDUtils.cpp which checks if a commute is necessary
to allow the VOPD pairing, and then records that finding so that
GCNCreateVOPD applies the commute before creating the VOPD pair.
Co-authored by: Claude Code
---------
Co-authored-by: Claude <noreply at anthropic.com>
[AMDGPU] Allow commuting immediates out of src0 when legal (#229906)
Fixes issue introduced by #181918 on gfx10+ where an immediate can get
commuted from src0 to src1 but then fail to get commuted back to src0
due to the legality checks in `isLegalToSwap`.
This PR relaxes the checks in `isLegalToSwap`, since gfx10+ allows the
immediate to be in locations other than src0. Relaxing these checks
causes MachineCSE to also successfully commute immediate operands out of
src0, which is the reason behind all the lit tests that required
modification. The PR also adds 2 new tests.
Co-authored by: Claude Code
Fixes: LCOMPILER-2920
---------
Co-authored-by: Claude <noreply at anthropic.com>
[orc-rt] Add SPSSymbolLookupResult typedef, clean up users. (#230877)
Existing serializers of SymbolLookupResult were spelling out the SPS
type in full (SPSSequence<SPSOptional<SPSExecutorAddr>>). Define an
SPSSymbolLookupResult typedef and use in instead so that serialization
points can pick up any future changes automatically.
This is the result-side counterpart to 8ff4f386cfeb, which added a
typedef for SymbolLookupSet.
[SLP]Vectorize consecutive loads with undef lanes as a wide load
Model the lane with the absorbing constant (0 for mul/and, -1 for or) of
a copyable node as op(V, undef), so the operand column of the other
lanes gets an undef lane. Cover such undef lanes in a column of
consecutive loads with a single frozen vector load, if the whole range
is dereferenceable.
Fixes #46897
Assisted-by: Cursor
Reviewers: RKSimon
Pull Request: https://github.com/llvm/llvm-project/pull/228872
fix(SelectionDAG): validate identity fold demands
KnownBits returned for the LHS may only be valid for the demand
already reduced by the RHS. Check the full result demand before
folding AND/OR to the LHS.
Share the masked-bit check between both operations. Drop Disjoint
when the query rewrites an OR operand.
[orc-rt] Add SPSSymbolLookupSet typedef, clean up users. (#230876)
Existing deserializers of SymbolLookupSet were spelling out the SPS type
in full (SPSSequence<SPSTuple<SPSString, bool>>). Define an
SPSSymbolLookupSet typedef and use in instead so that deserialization
points can pick up any future changes automatically.
Address review feedback on FP8 conversion builtins
Reuse err_builtin_invalid_arg_type for all source operand errors
instead of builtin-specific diagnostics.
Accept integer constants that fit the format width, such as 0x38,
and std::byte, so common byte values need no explicit cast.
Drop the unused OpenCL fp64 path from checkFloatingPointTypeSupport.
Document floating-point environment and fast-math behavior; trim
implementation detail from the user docs.
Change-Id: I8676a49540bdbbb8a90827c83764803eeea850a5
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[clang] Add elementwise conversions from encoded FP8 values
Add nine builtins converting Float8E5M2, Float8E4M3FN, and Float8E5M3FNU
encodings to _Float16, __bf16, or float through
llvm.convert.from.arbitrary.fp.
Accept exactly 8-bit integer scalars and generic fixed-length vectors,
preserving vector kinds and element counts without integer promotions.
Also support scalar AArch64 __mfp8 containers. Apply destination type
availability checks, including deferred offload diagnostics, and reject
constant-expression use.
Document the interface and add semantic, template, language-mode,
target-specific, and IR-generation regression coverage.
[AMDGPU] Form VOPD dot2 pairs with a literal in src1
A V_DOT2 with a register in src0 and a literal in src1 is matched as if
commuted, and commuted when the VOPD pair is built. This recovers the
pairs lost once MachineCSE stopped leaving those literals in src0.
Co-Authored-By: Claude <noreply at anthropic.com>
[AMDGPU] Allow commuting immediates out of src0 when legal
isLegalToSwap refused to move any non-inline constant out of src0, so
commuting an instruction with a literal in src1 could not be undone.
AMDGPULowerVGPREncoding relies on undoing it and, on gfx1250, either hit
"Failed to restore commuted instruction" or kept the commuted
instruction with the wrong VGPR MSB mode, which made it read the wrong
VGPRs.
Allow an immediate to leave src0 when the other operand can hold it.
VOPD formation now accepts a V_DOT2 with a literal in src1 if
isLegalToSwap allows the swap, and commutes it when the pair is built,
so those pairs are still formed.
Co-Authored-By: Claude <noreply at anthropic.com>
[libc][CI] Update FreeBSD version to 15 and bump VM action (#230775)
The LLVM libc FreeBSD CI has recently been failing because upstream
package repositories require a newer FreeBSD version than the pinned
15.0 image.
This patch updates the FreeBSD version specification from 15.0 to 15,
allowing the action to automatically pull the latest minor snapshot.
It also bumps the VM action to the latest release for improved
compatibility with Ubuntu 26.04 runners.
ref:
https://github.com/llvm/llvm-project/actions/runs/37893910030/job/113701458175?pr=230371
- before
```
Processing entries:
Newer FreeBSD version for package zh-qe:
[11 lines not shown]
VE: Remove broken nested call frame around dynamic stack allocation (#229011)
lowerDYNAMIC_STACKALLOC wrapped the __ve_grow_stack call and the
GETSTACKTOP stack-pointer read in a zero-sized CALLSEQ_START/CALLSEQ_END
pair. The call it contains emits its own CALLSEQ, so the outer bracket
only produced a nested ADJCALLSTACKDOWN 0 / ADJCALLSTACKUP 0 around the
inner ADJCALLSTACKDOWN / ADJCALLSTACKUP which is illegal.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
[clang-repl] Mark global-dtor.cpp and value-print-temporaries.cpp unsupported under ASan (#230870)
These tests are flaky on x86_64 Linux ASan bots with `out of range of
Delta32
fixup` JITLink errors, similar to #102858, #135401, and #150242.
This likely happens because `InProcessMemoryManager` maps each
incremental
module with a separate `mmap` call, and depending on the address space
layout
some allocations appear to end up on opposite sides of ASan's large
allocator
reservation (> 2 GiB apart).
Assisted-by: Gemini
fix(SelectionDAG): simplify commuted demanded bits
AND/OR demanded-bit simplification uses RHS known bits to simplify the
LHS, but does not retry the RHS using LHS known bits, making
optimizations depend on operand order.
Retry the RHS when the LHS reduces its demanded bits. Add AArch64
and AMDGPU codegen coverage.