[RISCV] Fix FP properties for Zvbdota instructions (#228737)
Mark floating-point Zvbdota instructions as potentially raising FP
exceptions. Add an implicit FRM use to vfbdota.vv, which uses the
dynamic rounding mode.
AI Usage: Assisted by AmpCode (6.1 Sol)
Co-authored-by: Amp <amp at ampcode.com>
[clang] Replace PointerUnion::dyn_cast with llvm::dyn_cast (NFC) (#230911)
PointerUnion::dyn_cast has been soft-deprecated in favor of
llvm::dyn_cast and llvm::dyn_cast_if_present. This patch replaces the
former with llvm::dyn_cast where the operand is guaranteed to be
nonnull.
Note that PlaceholderBase::ParamOrMethod is always initialized to a
nonnull pointer in LoanManager::getOrCreatePlaceholderBase and never
modified afterward.
Assisted-by: Antigravity
[libc++] Implement LWG2991: variant copy constructor missing noexcept(see below) (#230241)
LWG2991 adds a conditional `noexcept` to the variant's copy constructor.
A small example to illustrate the issue:
```cpp
static_assert(!std::is_trivially_copy_constructible_v<std::shared_ptr<int>>);
static_assert(std::is_copy_constructible_v<std::shared_ptr<int>>);
static_assert(std::is_nothrow_copy_constructible_v<std::shared_ptr<int>>);
static_assert(!std::is_trivially_move_constructible_v<std::shared_ptr<int>>);
static_assert(std::is_move_constructible_v<std::shared_ptr<int>>);
static_assert(std::is_nothrow_move_constructible_v<std::shared_ptr<int>>);
using Variant = std::variant<int,std::shared_ptr<int>>;
static_assert(std::is_nothrow_copy_constructible_v<Variant>); // false, should be true
static_assert(std::is_nothrow_move_constructible_v<Variant>);
```
[9 lines not shown]
[llubi] Check UB before handling ret attributes (#230842)
Some code paths return AnyValue() on UB (e.g., calling function
declarations not recognized as libfunc). Check this before handling ret
attributes to avoid operating on incompatible values.
The test is generated by DeepSeek-V4.1-Flash.
[AMDGPU] Insert wmma coexec aware anti-hints rules (#226397)
This patch ports the downstream coexec anti-hint insertion to upstream
as rule-based insertion in the AMDGPU pre-ra anti-hints framework. It
adds gfx1250x anti-hint rules for wmma coexec hazards such as A/B source
war, swmmac index war, dest waw), trans source war, memory-address war
etc. Each rule builds anti-hints relationship for the register allocator
to keep hazardous registers in different physical registers so the
hazard recognizer and waitcnt inserter need less nop and waits.
Co-authored-by: Jeffrey Byrnes <jrbyrnes1989 at gmail.com>
Depends on #218075
AMDGPU: Don't use PressureDiffs with implicit allocatable physregs
canUsePressureDiffs skipped all implicit operands, so instructions
reading an allocatable physreg implicitly (e.g., vcc on v_cndmask_b32)
used the cached PressureDiff. PressureDiffs assume a single use for
physregs (see the FIXME in ScheduleDAGMILive::updatePressureDiffs), so
the pressure was added for each user. With several users in a region, the
SGPR pressure was overestimated and failed the expensive checks pressure
validation.
Fall back to the RegPressureTracker for implicit operands of
allocatable physregs, which are the ones it tracks.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[ARM] Implement CRC32 const folding (#230751)
This implements constant folding for __crc32 intrinsics in 32-bit ARM
targets when both the arguments are compile time known constants.
This is a follow up of https://github.com/llvm/llvm-project/pull/228985
which does the same for AArch64.
[Support] Fix SpecificBumpPtrAllocator move assignment leak (#230480)
Move assignment releases the destination's storage without destroying
its objects, leaking resources they own.
Call `DestroyAll()` before replacing the destination's storage.
Testing on Linux with GCC 11.4.0, Release with assertions enabled:
- `AllocatorTest.*`: 15 passed. The new regression test fails with the
original header and passes with the fix.
- Support unit tests via `llvm-lit`: 1842 passed, 13 skipped, no
failures.
- ASan/LSan reproducer: a 512-byte leak with the original header; no
sanitizer report with the fix.
- Additional ASan/LSan checks for destruction, move construction, empty
and
populated destinations, multiple slabs, and custom-sized slabs passed.
Assisted-by: Codex
[ProfileData] Move ProfCorrelatorKind out of InstrProfCorrelator (#230883)
For coverage and PGO, `clang -mllvm
-profile-correlate={debug-info,binary}` moves the per-function profile
metadata out of the memory image at runtime.
After the cl::opt to TableGen migration (#230746),
InstrumentationOptions.h includes InstrProfCorrelator.h only for this
enum, adding about 0.25s to each Instrumentation TU. Fix the compile
time regression by moving the type to its own header as a scoped enum.
LLM-aided
Simplify iterator lowering and trim duplicate tests
Use capture-only iterator callbacks and bind induction variables as block
arguments are created, eliminating the temporary induction-value vector.
Remove unused-iterator HLFIR cases covered by the per-locator and LLVM
checks. Reduce repeated stride diagnostics and remove unused declarations
from the positive AFFINITY tests.
[CIR] Implement ViewLikeOpInterface for multiple Ops
We implement ViewLikeOpInterface for dyn_cast, ptr_stride, get_member,
get_element, get_runtime_member, base_class_addr, derived_class_addr,
and memonic. Although there is no user for CIR internally as
decouplePointer returns the base and offset instead of the base
directly, it is still useful for external project to do aliasing
analysis for MemRef.
[AMDGPU] Use 8-bit barrier member count on GFX13
GFX13 widens the barrier member count in M0 to 8 bits for cluster
named barriers. s_barrier_init and s_barrier_signal_var lowering
masked it to 6 bits, truncating counts above 63.
Change-Id: I33571d2ed499ebe168571c6b4cd5d0e3340f8a0f
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[flang][cuda] share runtime type info between host and device under managed memory (#229213)
With -gpu=mem:managed, descriptors can live in managed memory, so a
descriptor built on the device may be read on the host. The type
descriptor address in its addendum then pointed at the device copy of
the type info, and host code dereferencing it crashed.
Make the host copy of the type info the single shared copy:
- Add the cuf-shared-type-info pass. It makes host type-info globals
writable, places them in the __nv_type_info section, and drops
acc.declare from type info on both host and device so that OpenACC
declare constructors no longer copy it to the device. For each type
descriptor used in the GPU module, it creates a managed pointer
global <dt>Xhostaddr<tag>. The tag is a per-unit hash, which keeps
the name unique when each unit has its own device module. The GPU
module gets a cuf.shared_type_descs dictionary mapping each type
descriptor to its pointer.
- CUFAddConstructor: add the cuda-managed-type-info option. It
[12 lines not shown]
[orc-rt] Default LockedAccess's LockT arg to std::scoped_lock (#230881)
In the common case where LockT = std::scoped_lock<std::mutex> is the
desired lock (and mutex) type, this allows us to write:
LockedAccess<T> getValue() { return { Value, Mutex }; }
without having to spell out the type for LockT.
[NVPTX] Preserve the load chain when custom-lowering i1 loads (#230497)
lowerLOADi1() rewrites an i1 load into a zext load to i16 plus a
truncate, and returns the (value, chain) pair as a MERGE_VALUES node.
LegalizeLoadOps installs that pair with
RChain = Res.getValue(1);
DAG.ReplaceAllUsesOfValueWith(SDValue(Node, 1), RChain);
so the second value of the MERGE_VALUES becomes the replacement for the
original load's chain result. Returning LD->getChain() therefore rewires
every memory operation that followed the original load to that load's
predecessor, and leaves the new zext load's chain result with no users.
The ordering edge between the new load and those memory operations is
dropped, so nothing in the DAG keeps them in order beyond whatever data
dependency happens to exist between them.
Return newLD.getValue(1) instead, so the edge is preserved.
perf(SelectionDAG): reduce KnownBits temporaries
Let SimplifyDemandedBits initialize the fold-check result. Build the
RHS demand in one APInt instead of copying both KnownBits masks.
[AMDGPU] Form VOPD dot2 pairs with a literal in src1 (#230183)
This PR restores VOPD pair formation after the legality checks
relaxation introduced by #229906. Before the legality check relaxation,
MachineCSE was commuting immediate operands from src1 to src0, and then
failing to commute them back, which inadvertently results in the
immediate operands in src0, and the VOPD pairing would succeed. After
the legality relaxation, MachineCSE is now able to successfully commute
the immediates back from src0 to src1, which breaks VOPD pairing since
the pass expected the immediates to be in src0 position. This change
adds a check in GCNVOPDUtils.cpp which checks if a commute is necessary
to allow the VOPD pairing, and then records that finding so that
GCNCreateVOPD applies the commute before creating the VOPD pair.
Co-authored by: Claude Code
---------
Co-authored-by: Claude <noreply at anthropic.com>
[AMDGPU] Allow commuting immediates out of src0 when legal (#229906)
Fixes issue introduced by #181918 on gfx10+ where an immediate can get
commuted from src0 to src1 but then fail to get commuted back to src0
due to the legality checks in `isLegalToSwap`.
This PR relaxes the checks in `isLegalToSwap`, since gfx10+ allows the
immediate to be in locations other than src0. Relaxing these checks
causes MachineCSE to also successfully commute immediate operands out of
src0, which is the reason behind all the lit tests that required
modification. The PR also adds 2 new tests.
Co-authored by: Claude Code
Fixes: LCOMPILER-2920
---------
Co-authored-by: Claude <noreply at anthropic.com>
[orc-rt] Add SPSSymbolLookupResult typedef, clean up users. (#230877)
Existing serializers of SymbolLookupResult were spelling out the SPS
type in full (SPSSequence<SPSOptional<SPSExecutorAddr>>). Define an
SPSSymbolLookupResult typedef and use in instead so that serialization
points can pick up any future changes automatically.
This is the result-side counterpart to 8ff4f386cfeb, which added a
typedef for SymbolLookupSet.
[SLP]Vectorize consecutive loads with undef lanes as a wide load
Model the lane with the absorbing constant (0 for mul/and, -1 for or) of
a copyable node as op(V, undef), so the operand column of the other
lanes gets an undef lane. Cover such undef lanes in a column of
consecutive loads with a single frozen vector load, if the whole range
is dereferenceable.
Fixes #46897
Assisted-by: Cursor
Reviewers: RKSimon
Pull Request: https://github.com/llvm/llvm-project/pull/228872
fix(SelectionDAG): validate identity fold demands
KnownBits returned for the LHS may only be valid for the demand
already reduced by the RHS. Check the full result demand before
folding AND/OR to the LHS.
Share the masked-bit check between both operations. Drop Disjoint
when the query rewrites an OR operand.
[orc-rt] Add SPSSymbolLookupSet typedef, clean up users. (#230876)
Existing deserializers of SymbolLookupSet were spelling out the SPS type
in full (SPSSequence<SPSTuple<SPSString, bool>>). Define an
SPSSymbolLookupSet typedef and use in instead so that deserialization
points can pick up any future changes automatically.
Address review feedback on FP8 conversion builtins
Reuse err_builtin_invalid_arg_type for all source operand errors
instead of builtin-specific diagnostics.
Accept integer constants that fit the format width, such as 0x38,
and std::byte, so common byte values need no explicit cast.
Drop the unused OpenCL fp64 path from checkFloatingPointTypeSupport.
Document floating-point environment and fast-math behavior; trim
implementation detail from the user docs.
Change-Id: I8676a49540bdbbb8a90827c83764803eeea850a5
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[clang] Add elementwise conversions from encoded FP8 values
Add nine builtins converting Float8E5M2, Float8E4M3FN, and Float8E5M3FNU
encodings to _Float16, __bf16, or float through
llvm.convert.from.arbitrary.fp.
Accept exactly 8-bit integer scalars and generic fixed-length vectors,
preserving vector kinds and element counts without integer promotions.
Also support scalar AArch64 __mfp8 containers. Apply destination type
availability checks, including deferred offload diagnostics, and reject
constant-expression use.
Document the interface and add semantic, template, language-mode,
target-specific, and IR-generation regression coverage.
[AMDGPU] Form VOPD dot2 pairs with a literal in src1
A V_DOT2 with a register in src0 and a literal in src1 is matched as if
commuted, and commuted when the VOPD pair is built. This recovers the
pairs lost once MachineCSE stopped leaving those literals in src0.
Co-Authored-By: Claude <noreply at anthropic.com>