LLVM/project 67e37a9 — flang/lib/Lower/OpenMP ClauseProcessor.cpp, flang/test/Lower/OpenMP affinity-part-refs.f90 depend-iterator.f90

Simplify iterator lowering and trim duplicate tests

Use capture-only iterator callbacks and bind induction variables as block
arguments are created, eliminating the temporary induction-value vector.

Remove unused-iterator HLFIR cases covered by the per-locator and LLVM
checks. Reduce repeated stride diagnostics and remove unused declarations
from the positive AFFINITY tests.
DeltaFile
+0-35flang/test/Lower/OpenMP/task-affinity.f90
+13-20flang/lib/Lower/OpenMP/ClauseProcessor.cpp
+0-32flang/test/Lower/OpenMP/depend-iterator.f90
+2-28flang/test/Semantics/OpenMP/substring-strides.f90
+6-16flang/test/Lower/OpenMP/affinity-part-refs.f90
+0-9flang/test/Semantics/OpenMP/affinity-part-refs.f90
+21-1406 files

LLVM/project e2ce52c — llvm/test/CodeGen/AMDGPU anti-hints-trans-src0.gfx1250.mir anti-hints-multi-hazard.gfx1250.mir

Addressed review, added tests detail
DeltaFile
+234-0llvm/test/CodeGen/AMDGPU/anti-hints-wmma-waw.gfx1250.mir
+209-0llvm/test/CodeGen/AMDGPU/anti-hints-wmma-war.gfx1250.mir
+69-0llvm/test/CodeGen/AMDGPU/anti-hints-addr-xcnt.gfx1250.mir
+55-5llvm/test/CodeGen/AMDGPU/anti-hints-multi-rule.gfx1250.mir
+58-0llvm/test/CodeGen/AMDGPU/anti-hints-multi-hazard.gfx1250.mir
+49-0llvm/test/CodeGen/AMDGPU/anti-hints-trans-src0.gfx1250.mir
+674-52 files not shown
+733-58 files

LLVM/project d229e35 — llvm/test/CodeGen/AMDGPU mul.ll amdgcn.bitcast.1024bit.ll

Merge remote-tracking branch 'upstream/main' into users/mssefat/anti-hints-pr5-amdgpu-pre-ra-coexec-anti-hint
DeltaFile
+32,088-15,710llvm/test/CodeGen/AMDGPU/frem.ll
+5,761-3,731llvm/test/CodeGen/AMDGPU/srem.ll
+6,118-2,908llvm/test/CodeGen/AMDGPU/clmul.ll
+1,502-5,317llvm/test/CodeGen/AMDGPU/fcanonicalize.ll
+2,824-2,824llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+1,327-1,552llvm/test/CodeGen/AMDGPU/mul.ll
+49,620-32,0421,443 files not shown
+113,961-70,5011,449 files

LLVM/project 1abc2f0 — llvm/lib/Target/AMDGPU SIISelLowering.cpp AMDGPUInstructionSelector.cpp, llvm/lib/Target/AMDGPU/Utils AMDGPUBaseInfo.h AMDGPUBaseInfo.cpp

[AMDGPU] Use 8-bit barrier member count on GFX13 (#230884)

GFX13 widens the barrier member count in M0 to 8 bits.
DeltaFile
+137-0llvm/test/CodeGen/AMDGPU/s-barrier-member-count.ll
+9-0llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.cpp
+4-2llvm/lib/Target/AMDGPU/SIISelLowering.cpp
+4-2llvm/lib/Target/AMDGPU/AMDGPUInstructionSelector.cpp
+3-0llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.h
+157-45 files

LLVM/project 1beef99 — clang/include/clang/CIR/Dialect/IR CIRDialect.h CIROps.td, clang/lib/CIR/Dialect/IR CMakeLists.txt CIRDialect.cpp

[CIR] Implement ViewLikeOpInterface for multiple Ops

We implement ViewLikeOpInterface for dyn_cast, ptr_stride, get_member,
get_element, get_runtime_member, base_class_addr, derived_class_addr,
and memonic. Although there is no user for CIR internally as
decouplePointer returns the base and offset instead of the base
directly, it is still useful for external project to do aliasing
analysis for MemRef.
DeltaFile
+22-8clang/include/clang/CIR/Dialect/IR/CIROps.td
+14-0clang/lib/CIR/Dialect/IR/CIRDialect.cpp
+1-0clang/lib/CIR/Dialect/IR/CMakeLists.txt
+1-0clang/include/clang/CIR/Dialect/IR/CIRDialect.h
+38-84 files

LLVM/project 86a2f54 — llvm/lib/Target/AMDGPU SIISelLowering.cpp AMDGPUInstructionSelector.cpp, llvm/lib/Target/AMDGPU/Utils AMDGPUBaseInfo.h AMDGPUBaseInfo.cpp

[AMDGPU] Use 8-bit barrier member count on GFX13

GFX13 widens the barrier member count in M0 to 8 bits for cluster
named barriers. s_barrier_init and s_barrier_signal_var lowering
masked it to 6 bits, truncating counts above 63.

Change-Id: I33571d2ed499ebe168571c6b4cd5d0e3340f8a0f
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
DeltaFile
+137-0llvm/test/CodeGen/AMDGPU/s-barrier-member-count.ll
+9-0llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.cpp
+4-2llvm/lib/Target/AMDGPU/SIISelLowering.cpp
+4-2llvm/lib/Target/AMDGPU/AMDGPUInstructionSelector.cpp
+3-0llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.h
+157-45 files

LLVM/project c23622f — flang/lib/Optimizer/CodeGen CodeGen.cpp, flang/lib/Optimizer/Transforms/CUDA CUFAddConstructor.cpp CUFSharedTypeInfo.cpp

[flang][cuda] share runtime type info between host and device under managed memory (#229213)

With -gpu=mem:managed, descriptors can live in managed memory, so a
descriptor built on the device may be read on the host. The type
descriptor address in its addendum then pointed at the device copy of
the type info, and host code dereferencing it crashed.

Make the host copy of the type info the single shared copy:

- Add the cuf-shared-type-info pass. It makes host type-info globals
  writable, places them in the __nv_type_info section, and drops
  acc.declare from type info on both host and device so that OpenACC
  declare constructors no longer copy it to the device. For each type
  descriptor used in the GPU module, it creates a managed pointer
  global <dt>Xhostaddr<tag>. The tag is a per-unit hash, which keeps
  the name unique when each unit has its own device module. The GPU
  module gets a cuf.shared_type_descs dictionary mapping each type
  descriptor to its pointer.
- CUFAddConstructor: add the cuda-managed-type-info option. It

    [12 lines not shown]
DeltaFile
+131-0flang/lib/Optimizer/Transforms/CUDA/CUFSharedTypeInfo.cpp
+105-0flang/test/Fir/CUDA/cuda-shared-type-info.mlir
+88-1flang/lib/Optimizer/Transforms/CUDA/CUFAddConstructor.cpp
+71-0flang/test/Fir/CUDA/cuda-shared-type-info-codegen.mlir
+61-0flang/test/Fir/CUDA/cuda-shared-type-info-rename.mlir
+58-0flang/lib/Optimizer/CodeGen/CodeGen.cpp
+514-14 files not shown
+590-110 files

LLVM/project a5de0ad — orc-rt/include/orc-rt/support LockedAccess.h

[orc-rt] Default LockedAccess's LockT arg to std::scoped_lock (#230881)

In the common case where LockT = std::scoped_lock<std::mutex> is the
desired lock (and mutex) type, this allows us to write:

  LockedAccess<T> getValue() { return { Value, Mutex }; }

without having to spell out the type for LockT.
DeltaFile
+1-1orc-rt/include/orc-rt/support/LockedAccess.h
+1-11 files

LLVM/project d283ad8 — llvm/lib/Target/NVPTX NVPTXISelLowering.cpp, llvm/test/CodeGen/NVPTX i1-load-chain.ll

[NVPTX] Preserve the load chain when custom-lowering i1 loads (#230497)

lowerLOADi1() rewrites an i1 load into a zext load to i16 plus a
truncate, and returns the (value, chain) pair as a MERGE_VALUES node.

LegalizeLoadOps installs that pair with

  RChain = Res.getValue(1);
  DAG.ReplaceAllUsesOfValueWith(SDValue(Node, 1), RChain);

so the second value of the MERGE_VALUES becomes the replacement for the
original load's chain result. Returning LD->getChain() therefore rewires
every memory operation that followed the original load to that load's
predecessor, and leaves the new zext load's chain result with no users.
The ordering edge between the new load and those memory operations is
dropped, so nothing in the DAG keeps them in order beyond whatever data
dependency happens to exist between them.

Return newLD.getValue(1) instead, so the edge is preserved.
DeltaFile
+20-0llvm/test/CodeGen/NVPTX/i1-load-chain.ll
+1-1llvm/lib/Target/NVPTX/NVPTXISelLowering.cpp
+21-12 files

LLVM/project 6d00313 — llvm/lib/CodeGen/SelectionDAG TargetLowering.cpp

perf(SelectionDAG): reduce KnownBits temporaries

Let SimplifyDemandedBits initialize the fold-check result. Build the
RHS demand in one APInt instead of copying both KnownBits masks.
DeltaFile
+10-6llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
+10-61 files

LLVM/project 92ef8fe — clang/test/Interpreter pretty-print.cpp value-print-temporaries.cpp

[𝘀𝗽𝗿] initial version

Created using spr 1.3.7
DeltaFile
+0-3clang/test/Interpreter/value-print-temporaries.cpp
+0-3clang/test/Interpreter/inline-virtual.cpp
+0-3clang/test/Interpreter/global-dtor.cpp
+0-3clang/test/Interpreter/const.cpp
+0-2clang/test/Interpreter/pretty-print.cpp
+0-145 files

LLVM/project cbdad67 — llvm/lib/Target/AMDGPU SIInstrInfo.h GCNCreateVOPD.cpp, llvm/test/CodeGen/AMDGPU llvm.amdgcn.fdot2.f32.bf16.ll llvm.amdgcn.fdot2.ll

[AMDGPU] Form VOPD dot2 pairs with a literal in src1 (#230183)

This PR restores VOPD pair formation after the legality checks
relaxation introduced by #229906. Before the legality check relaxation,
MachineCSE was commuting immediate operands from src1 to src0, and then
failing to commute them back, which inadvertently results in the
immediate operands in src0, and the VOPD pairing would succeed. After
the legality relaxation, MachineCSE is now able to successfully commute
the immediates back from src0 to src1, which breaks VOPD pairing since
the pass expected the immediates to be in src0 position. This change
adds a check in GCNVOPDUtils.cpp which checks if a commute is necessary
to allow the VOPD pairing, and then records that finding so that
GCNCreateVOPD applies the commute before creating the VOPD pair.

Co-authored by: Claude Code

---------

Co-authored-by: Claude <noreply at anthropic.com>
DeltaFile
+149-0llvm/test/CodeGen/AMDGPU/vopd-dot2-commute-imm-src1.mir
+58-19llvm/lib/Target/AMDGPU/GCNVOPDUtils.cpp
+12-35llvm/test/CodeGen/AMDGPU/llvm.amdgcn.fdot2.ll
+3-6llvm/test/CodeGen/AMDGPU/llvm.amdgcn.fdot2.f32.bf16.ll
+9-0llvm/lib/Target/AMDGPU/GCNCreateVOPD.cpp
+2-2llvm/lib/Target/AMDGPU/SIInstrInfo.h
+233-621 files not shown
+236-627 files

LLVM/project 024cc93 — llvm/test/CodeGen/AMDGPU amdgcn.bitcast.896bit.ll amdgcn.bitcast.960bit.ll

[AMDGPU] Allow commuting immediates out of src0 when legal (#229906)

Fixes issue introduced by #181918 on gfx10+ where an immediate can get
commuted from src0 to src1 but then fail to get commuted back to src0
due to the legality checks in `isLegalToSwap`.

This PR relaxes the checks in `isLegalToSwap`, since gfx10+ allows the
immediate to be in locations other than src0. Relaxing these checks
causes MachineCSE to also successfully commute immediate operands out of
src0, which is the reason behind all the lit tests that required
modification. The PR also adds 2 new tests.

Co-authored by: Claude Code

Fixes: LCOMPILER-2920

---------

Co-authored-by: Claude <noreply at anthropic.com>
DeltaFile
+2,824-2,824llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+1,412-1,412llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+706-706llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.256bit.ll
+496-496llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.320bit.ll
+480-480llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.960bit.ll
+448-448llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.896bit.ll
+6,366-6,366207 files not shown
+13,892-13,589213 files

LLVM/project 0cbb7c7 — orc-rt/include/orc-rt/support/sps SPSSymbolLookupSet.h, orc-rt/lib/bedrock/sps NativeDylibManagerSPSCI.cpp

[orc-rt] Add SPSSymbolLookupResult typedef, clean up users. (#230877)

Existing serializers of SymbolLookupResult were spelling out the SPS
type in full (SPSSequence<SPSOptional<SPSExecutorAddr>>). Define an
SPSSymbolLookupResult typedef and use in instead so that serialization
points can pick up any future changes automatically.

This is the result-side counterpart to 8ff4f386cfeb, which added a
typedef for SymbolLookupSet.
DeltaFile
+2-3orc-rt/test/unit/support/sps/SPSSymbolLookupSetTest.cpp
+2-3orc-rt/lib/bedrock/sps/NativeDylibManagerSPSCI.cpp
+2-0orc-rt/include/orc-rt/support/sps/SPSSymbolLookupSet.h
+6-63 files

LLVM/project 611eee0 — llvm/lib/Transforms/Vectorize SLPVectorizer.cpp, llvm/lib/Transforms/Vectorize/SLPVectorizer SLPUtils.cpp SLPMemoryUtils.h

[SLP]Vectorize consecutive loads with undef lanes as a wide load

Model the lane with the absorbing constant (0 for mul/and, -1 for or) of
a copyable node as op(V, undef), so the operand column of the other
lanes gets an undef lane. Cover such undef lanes in a column of
consecutive loads with a single frozen vector load, if the whole range
is dereferenceable.

Fixes #46897

Assisted-by: Cursor

Reviewers: RKSimon

Pull Request: https://github.com/llvm/llvm-project/pull/228872
DeltaFile
+124-16llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+50-86llvm/test/Transforms/SLPVectorizer/X86/absorbing-copyable-lane.ll
+40-63llvm/test/Transforms/SLPVectorizer/X86/wide-load-absorbed-lane.ll
+54-0llvm/lib/Transforms/Vectorize/SLPVectorizer/SLPMemoryUtils.cpp
+16-0llvm/lib/Transforms/Vectorize/SLPVectorizer/SLPMemoryUtils.h
+15-0llvm/lib/Transforms/Vectorize/SLPVectorizer/SLPUtils.cpp
+299-1652 files not shown
+310-1708 files

LLVM/project 9048ff1 — llvm/lib/CodeGen/SelectionDAG TargetLowering.cpp

fix(SelectionDAG): validate identity fold demands

KnownBits returned for the LHS may only be valid for the demand
already reduced by the RHS. Check the full result demand before
folding AND/OR to the LHS.

Share the masked-bit check between both operations. Drop Disjoint
when the query rewrites an OR operand.
DeltaFile
+59-33llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
+59-331 files

LLVM/project 8ff4f38 — orc-rt/include/orc-rt/support/sps SPSSymbolLookupSet.h, orc-rt/lib/bedrock/sps NativeDylibManagerSPSCI.cpp

[orc-rt] Add SPSSymbolLookupSet typedef, clean up users. (#230876)

Existing deserializers of SymbolLookupSet were spelling out the SPS type
in full (SPSSequence<SPSTuple<SPSString, bool>>). Define an
SPSSymbolLookupSet typedef and use in instead so that deserialization
points can pick up any future changes automatically.
DeltaFile
+3-3orc-rt/lib/bedrock/sps/NativeDylibManagerSPSCI.cpp
+2-3orc-rt/test/unit/support/sps/SPSSymbolLookupSetTest.cpp
+1-2orc-rt/test/unit/bedrock/sps/NativeDylibManagerSPSCITest.cpp
+2-0orc-rt/include/orc-rt/support/sps/SPSSymbolLookupSet.h
+8-84 files

LLVM/project 8847f81 — clang/docs LanguageExtensions.md, clang/include/clang/Basic DiagnosticSemaKinds.td

Address review feedback on FP8 conversion builtins

Reuse err_builtin_invalid_arg_type for all source operand errors
instead of builtin-specific diagnostics.

Accept integer constants that fit the format width, such as 0x38,
and std::byte, so common byte values need no explicit cast.

Drop the unused OpenCL fp64 path from checkFloatingPointTypeSupport.

Document floating-point environment and fast-math behavior; trim
implementation detail from the user docs.

Change-Id: I8676a49540bdbbb8a90827c83764803eeea850a5
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
DeltaFile
+44-22clang/lib/Sema/SemaChecking.cpp
+30-9clang/test/Sema/builtins-elementwise-convert-from-arbitrary-fp.c
+10-17clang/lib/Sema/SemaType.cpp
+24-1clang/test/SemaCXX/builtins-elementwise-convert-from-arbitrary-fp.cpp
+8-7clang/docs/LanguageExtensions.md
+1-11clang/include/clang/Basic/DiagnosticSemaKinds.td
+117-674 files not shown
+131-7110 files

LLVM/project 476eb9d — clang/docs LanguageExtensions.md, clang/lib/Sema SemaType.cpp SemaChecking.cpp

[clang] Add elementwise conversions from encoded FP8 values

Add nine builtins converting Float8E5M2, Float8E4M3FN, and Float8E5M3FNU
encodings to _Float16, __bf16, or float through
llvm.convert.from.arbitrary.fp.

Accept exactly 8-bit integer scalars and generic fixed-length vectors,
preserving vector kinds and element counts without integer promotions.
Also support scalar AArch64 __mfp8 containers. Apply destination type
availability checks, including deferred offload diagnostics, and reject
constant-expression use.

Document the interface and add semantic, template, language-mode,
target-specific, and IR-generation regression coverage.
DeltaFile
+172-0clang/test/CodeGen/builtins-elementwise-convert-from-arbitrary-fp.c
+135-0clang/test/Sema/builtins-elementwise-convert-from-arbitrary-fp.c
+105-0clang/lib/Sema/SemaChecking.cpp
+97-0clang/docs/LanguageExtensions.md
+50-18clang/lib/Sema/SemaType.cpp
+44-0clang/test/Sema/builtins-elementwise-convert-from-arbitrary-fp-target.c
+603-1810 files not shown
+804-1816 files

LLVM/project 9703767 — llvm/lib/Target/AMDGPU SIInstrInfo.h GCNCreateVOPD.cpp, llvm/test/CodeGen/AMDGPU llvm.amdgcn.fdot2.f32.bf16.ll llvm.amdgcn.fdot2.ll

[AMDGPU] Form VOPD dot2 pairs with a literal in src1

A V_DOT2 with a register in src0 and a literal in src1 is matched as if
commuted, and commuted when the VOPD pair is built. This recovers the
pairs lost once MachineCSE stopped leaving those literals in src0.

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+149-0llvm/test/CodeGen/AMDGPU/vopd-dot2-commute-imm-src1.mir
+54-19llvm/lib/Target/AMDGPU/GCNVOPDUtils.cpp
+12-35llvm/test/CodeGen/AMDGPU/llvm.amdgcn.fdot2.ll
+3-6llvm/test/CodeGen/AMDGPU/llvm.amdgcn.fdot2.f32.bf16.ll
+9-0llvm/lib/Target/AMDGPU/GCNCreateVOPD.cpp
+2-2llvm/lib/Target/AMDGPU/SIInstrInfo.h
+229-621 files not shown
+232-627 files

LLVM/project c048ebc — llvm/lib/Target/AMDGPU GCNVOPDUtils.cpp, llvm/test/CodeGen/AMDGPU vopd-dot2-commute-imm-src1.mir

[AMDGPU] Address review comments

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+9-5llvm/lib/Target/AMDGPU/GCNVOPDUtils.cpp
+1-1llvm/test/CodeGen/AMDGPU/vopd-dot2-commute-imm-src1.mir
+10-62 files

LLVM/project 72e3ec8 — llvm/test/CodeGen/AMDGPU clmul.ll

[AMDGPU] Update clmul.ll for literal commute

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+4-4llvm/test/CodeGen/AMDGPU/clmul.ll
+4-41 files

LLVM/project cb403d6 — llvm/test/CodeGen/AMDGPU udiv.ll mul.ll

[AMDGPU] Update more tests from main for literal commute

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+152-152llvm/test/CodeGen/AMDGPU/frem.ll
+10-10llvm/test/CodeGen/AMDGPU/udiv.ll
+10-10llvm/test/CodeGen/AMDGPU/mul.ll
+172-1723 files

LLVM/project 2d194c6 — llvm/utils/gn/secondary/bolt/unittests/Core BUILD.gn

[gn build] Port a09df97a747bb (#230873)
DeltaFile
+1-0llvm/utils/gn/secondary/bolt/unittests/Core/BUILD.gn
+1-01 files

LLVM/project 5bc8a29 — llvm/test/CodeGen/AMDGPU commute-literal-src0-cse.mir

[AMDGPU] Drop -verify-machineinstrs from commute-literal-src0-cse.mir

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+1-1llvm/test/CodeGen/AMDGPU/commute-literal-src0-cse.mir
+1-11 files

LLVM/project 37bbd5d — llvm/test/CodeGen/AMDGPU fsub.f16.ll fmul.f16.ll, llvm/test/CodeGen/AMDGPU/GlobalISel fshr.ll fshl.ll

[AMDGPU] Update tests from main for literal commute

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+24-24llvm/test/CodeGen/AMDGPU/GlobalISel/fptosi.bf16.ll
+22-22llvm/test/CodeGen/AMDGPU/GlobalISel/fshr.ll
+22-22llvm/test/CodeGen/AMDGPU/GlobalISel/fshl.ll
+12-12llvm/test/CodeGen/AMDGPU/cvt_f32_ubyte.ll
+4-4llvm/test/CodeGen/AMDGPU/fmul.f16.ll
+2-2llvm/test/CodeGen/AMDGPU/fsub.f16.ll
+86-866 files

LLVM/project f9da2ba — llvm/lib/Target/AMDGPU SIInstrInfo.h GCNCreateVOPD.cpp, llvm/test/CodeGen/AMDGPU llvm.amdgcn.fdot2.f32.bf16.ll llvm.amdgcn.fdot2.ll

[AMDGPU] Move the VOPD dot2 changes to a separate PR

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+0-149llvm/test/CodeGen/AMDGPU/vopd-dot2-commute-imm-src1.mir
+19-54llvm/lib/Target/AMDGPU/GCNVOPDUtils.cpp
+35-12llvm/test/CodeGen/AMDGPU/llvm.amdgcn.fdot2.ll
+0-9llvm/lib/Target/AMDGPU/GCNCreateVOPD.cpp
+6-3llvm/test/CodeGen/AMDGPU/llvm.amdgcn.fdot2.f32.bf16.ll
+2-2llvm/lib/Target/AMDGPU/SIInstrInfo.h
+62-2291 files not shown
+62-2327 files

LLVM/project 73f34a0 — llvm/test/CodeGen/AMDGPU fptrunc.f16.ll

[AMDGPU] Update fptrunc.f16.ll for literal commute

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+72-72llvm/test/CodeGen/AMDGPU/fptrunc.f16.ll
+72-721 files

LLVM/project 06b3d03 — llvm/test/CodeGen/AMDGPU amdgcn.bitcast.896bit.ll amdgcn.bitcast.960bit.ll

[AMDGPU] Allow commuting immediates out of src0 when legal

isLegalToSwap refused to move any non-inline constant out of src0, so
commuting an instruction with a literal in src1 could not be undone.
AMDGPULowerVGPREncoding relies on undoing it and, on gfx1250, either hit
"Failed to restore commuted instruction" or kept the commuted
instruction with the wrong VGPR MSB mode, which made it read the wrong
VGPRs.

Allow an immediate to leave src0 when the other operand can hold it.
VOPD formation now accepts a V_DOT2 with a literal in src1 if
isLegalToSwap allows the swap, and commutes it when the pair is built,
so those pairs are still formed.

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+2,824-2,824llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+1,412-1,412llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+706-706llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.256bit.ll
+496-496llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.320bit.ll
+480-480llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.960bit.ll
+448-448llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.896bit.ll
+6,366-6,366201 files not shown
+13,734-13,261207 files

LLVM/project 16c05cd — llvm/utils/gn/secondary/clang/unittests/AST BUILD.gn

[gn build] Port dba1b67f855fb (#230874)
DeltaFile
+1-0llvm/utils/gn/secondary/clang/unittests/AST/BUILD.gn
+1-01 files