LLVM/project d63124f — flang/include/flang/Evaluate tools.h, flang/lib/Evaluate tools.cpp

[flang][cuda] Look through associate names for managed and unified data (#228206)

IsCUDADeviceSymbol looks through an associate name: the name is device
data when its selector has device symbols. The managed and unified
predicates only handled object entities, so an associate name whose
selector is managed data was counted as device data that is not managed.

In host code, an element assignment such as
```
  associate(px => g%x, py => g%y)
    px%a(i,j) = r + px%s * real(py%n, 8)
  end associate
```

where `a` and `y` are managed, was then classified as a data transfer.
This
emitted a `cuf.data_transfer` from a scalar value, which the verifier
rejects.
IsCUDADataAttrSymbol now gives an associate name the attribute of the

    [2 lines not shown]
DeltaFile
+67-0flang/test/Lower/CUDA/cuda-associate-data-transfer.cuf
+56-5flang/lib/Evaluate/tools.cpp
+19-18flang/include/flang/Evaluate/tools.h
+142-233 files

LLVM/project 5263866 — clang/lib/CIR/CodeGen CIRGenFunction.cpp, clang/test/CIR/CodeGenHLSL cxx-this-lvalue.hlsl

[CIR] Support CXXThisExpr in emitLValue (#227316)

Support `CXXThisExpr` in `CIRGenFunction::emitLValue` by wrapping
`loadCXXThisAddress()` into an LValue via `makeAddrLValue()`, matching
classic Clang codegen (`CGExpr.cpp:1865`).

This enables LValue contexts for `this`, such as member access
expressions via `this.field` in languages like HLSL where `this` is
reference-like.

Fixes #227193
DeltaFile
+61-0clang/test/CIR/CodeGenHLSL/cxx-this-lvalue.hlsl
+1-2clang/lib/CIR/CodeGen/CIRGenFunction.cpp
+62-22 files

LLVM/project 79f4151 — flang/lib/Semantics check-cuda.h check-cuda.cpp, flang/test/Semantics/CUDA cuf-device-data-host-read.cuf

Revert "[flang][cuda] Diagnose host reads of device data" (#228286)

Reverts llvm/llvm-project#228271

auto-merging was on by mistake
DeltaFile
+0-133flang/lib/Semantics/check-cuda.cpp
+0-126flang/test/Semantics/CUDA/cuf-device-data-host-read.cuf
+0-11flang/lib/Semantics/check-cuda.h
+0-2703 files

LLVM/project b08cb38 — llvm/lib/CodeGen AtomicExpandPass.cpp, llvm/lib/Target/NVPTX NVPTXISelLowering.cpp

[AtomicExpand] Implement SUB → ADD(-x) and FSUB → FADD(-x) (#221425)

Implement ATOMIC_SUB → ATOMIC_ADD(-x) and ATOMIC_FSUB → ATOMIC_FADD(-x).
Many targets have atomic adds, but I don't think any have atomic subs.
This transform prevents cmpxchg expansion of atomic SUB. Enable these
transformations in NVPTX.

The integer variant exists in many different places currently. For
example, NVPTX currently expands it in DAG legalization.

begin AI generated

- SelectionDAG generic legalization:
[LegalizeDAG.cpp:3410](https://github.com/llvm/llvm-project/blob/4977a0c815c39fd90bb67c270174f2356ee9b5a7/llvm/lib/CodeGen/SelectionDAG/LegalizeDAG.cpp#L3410)
- GlobalISel generic legalization:
[LegalizerHelper.cpp:5117](https://github.com/llvm/llvm-project/blob/4977a0c815c39fd90bb67c270174f2356ee9b5a7/llvm/lib/CodeGen/GlobalISel/LegalizerHelper.cpp#L5117)
- GlobalISel outlined-atomic libcall path, mapping SUB to negated
`LDADD`:
[LegalizerHelper.cpp:919](https://github.com/llvm/llvm-project/blob/4977a0c815c39fd90bb67c270174f2356ee9b5a7/llvm/lib/CodeGen/GlobalISel/LegalizerHelper.cpp#L919)

    [42 lines not shown]
DeltaFile
+46-211llvm/test/CodeGen/NVPTX/atomicrmw-sm90.ll
+41-181llvm/test/CodeGen/NVPTX/atomicrmw-sm70.ll
+96-0llvm/test/Transforms/AtomicExpand/NVPTX/atomicrmw-sub.ll
+21-53llvm/test/CodeGen/NVPTX/atomicrmw-sm60.ll
+26-14llvm/lib/Target/NVPTX/NVPTXISelLowering.cpp
+26-0llvm/lib/CodeGen/AtomicExpandPass.cpp
+256-4591 files not shown
+260-4617 files

LLVM/project 6c1d106 — clang/lib/CIR/CodeGen CIRGenAsm.cpp CIRGenExpr.cpp, clang/lib/CIR/Dialect/IR CIRDialect.cpp

[CIR] Propagate the record address space to get_member (#226650)

Addresses: https://github.com/llvm/llvm-project/issues/226629

A member lives in its record's address space, but a few `get_member`
builders always produced a default-AS pointer. On SPIR-V that's private,
so a SYCL kernel was reading its captured pointer through a private
pointer. This patch takes the AS from the base and adds a verifier check
so we catch any stragglers.

Assisted-by: Claude / Opus 5.5
DeltaFile
+71-0clang/test/CIR/CodeGen/get-member-addrspace.cpp
+24-0clang/test/CIR/CodeGenHIP/inline-asm-multi-output-addrspace.hip
+14-0clang/test/CIR/IR/invalid-struct.cir
+6-3clang/lib/CIR/CodeGen/CIRGenExpr.cpp
+3-1clang/lib/CIR/CodeGen/CIRGenAsm.cpp
+3-0clang/lib/CIR/Dialect/IR/CIRDialect.cpp
+121-43 files not shown
+125-79 files

LLVM/project b55367c — clang/lib/CIR/CodeGen CIRGenModule.h CIRGenDecl.cpp, clang/test/CIR/CodeGen amdgpu-array-addrspace.cpp

[CIR] Cast global addresses to their declared address space (#226649)

Opened to address a portion of
https://github.com/llvm/llvm-project/issues/226629


In CUDA, `__shared__ int sh` has type `int` but lives in AS 3. Classic
codegen casts the address to the declared type's AS where it's formed,
so users just see a generic pointer. We weren't doing that, so things
like `return &sh;` bitcast the slot instead, and NVPTX never got a
`cvta.shared`. This patch does the same cast in `getAddrOfGlobalVar` and
wherever static locals are fetched.

This also drops the comment claiming lowering would emit the cast for
us. That's only true for OpenCL, where the declared type already carries
the AS. LowerToLLVM never inserts casts on its own.

Assisted-by: Claude / Opus 5.5
DeltaFile
+95-0clang/test/CIR/CodeGenCUDA/global-addrspace-cast.cu
+26-11clang/test/CIR/CodeGen/amdgpu-array-addrspace.cpp
+18-3clang/lib/CIR/CodeGen/CIRGenModule.cpp
+3-9clang/lib/CIR/CodeGen/CIRGenDecl.cpp
+3-2clang/test/CIR/CodeGenCUDA/address-spaces.cu
+4-0clang/lib/CIR/CodeGen/CIRGenModule.h
+149-251 files not shown
+151-267 files

LLVM/project 69c89de — llvm/test/CodeGen/AMDGPU slp-int-to-fp.ll

update slp-int-to-fp.ll
DeltaFile
+324-391llvm/test/CodeGen/AMDGPU/slp-int-to-fp.ll
+324-3911 files

LLVM/project 1458734 — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Analysis/CostModel/AMDGPU narrow-int-to-bfloat.ll narrow-int-to-fp.ll

[AMDGPU] Do not price the extension of a loaded i40, i48 or i56 in int to fp casts

A widened constant or invariant load still pays it.
DeltaFile
+100-100llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+24-24llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+18-6llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+142-1303 files

LLVM/project dfe42ed — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Analysis/CostModel/AMDGPU narrow-int-to-bfloat.ll narrow-int-to-fp.ll

few fixes
DeltaFile
+90-90llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+24-24llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+10-1llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+124-1153 files

LLVM/project 8f5f5ea — llvm/test/CodeGen/AMDGPU slp-int-to-fp.ll

update slp-int-to-fp.ll
DeltaFile
+128-170llvm/test/CodeGen/AMDGPU/slp-int-to-fp.ll
+128-1701 files

LLVM/project 1408c21 — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Analysis/CostModel/AMDGPU narrow-int-to-bfloat.ll narrow-int-to-fp.ll

[AMDGPU] Price scalar integer to fp casts by source width and sign

Scalar sources between a byte and 31 bits fell to the default cost of one
while the matching vector lanes were already priced, which skewed the
difference SLP weighs a bundle against. Such a source is extended before
the conversion, and what the extension takes depends on the width, on the
sign and on whether the subtarget has SDWA and 16 bit instructions.
Sources narrower than a byte are left alone, because their vector form is
not priced either.
DeltaFile
+388-388llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+72-72llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+29-9llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+489-4693 files

LLVM/project 1151a47 — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp

format
DeltaFile
+1-1llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+1-11 files

LLVM/project ea8cbfa — llvm/test/Analysis/CostModel/AMDGPU narrow-int-to-bfloat.ll narrow-int-to-fp.ll, llvm/test/CodeGen/AMDGPU itofp-odd-width-load.ll

[NFC][AMDGPU] Add more tests for int to fp casts of loaded odd width integers

Covers loads of i24, i40, i48 and i56 converted to fp, both the loads
that are split into narrower extending loads and the constant or
invariant ones that are widened to a scalar load.
DeltaFile
+466-0llvm/test/CodeGen/AMDGPU/itofp-odd-width-load.ll
+432-0llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+130-0llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+1,028-03 files

LLVM/project 63f4757 — llvm/test/CodeGen/AMDGPU slp-int-to-fp.ll

[NFC][AMDGPU] Add SLP to asm tests for int to fp casts of loaded odd width integers
DeltaFile
+855-0llvm/test/CodeGen/AMDGPU/slp-int-to-fp.ll
+855-01 files

LLVM/project 23e4666 — lld/test/ELF riscv-attributes.s

fix lld test
DeltaFile
+2-2lld/test/ELF/riscv-attributes.s
+2-21 files

LLVM/project dffdb6e — flang/lib/Semantics check-cuda.h check-cuda.cpp, flang/test/Semantics/CUDA cuf-device-data-host-read.cuf

Revert "[flang][cuda] Diagnose host reads of device data (#228271)"

This reverts commit 5c8e5ed14687b9b205ca90502976678302135b7f.
DeltaFile
+0-133flang/lib/Semantics/check-cuda.cpp
+0-126flang/test/Semantics/CUDA/cuf-device-data-host-read.cuf
+0-11flang/lib/Semantics/check-cuda.h
+0-2703 files

LLVM/project 5c8e5ed — flang/lib/Semantics check-cuda.h check-cuda.cpp, flang/test/Semantics/CUDA cuf-device-data-host-read.cuf

[flang][cuda] Diagnose host reads of device data (#228271)

Host code may only use device data in a data transfer assignment or as
an
actual argument. When device data appeared anywhere else, such as in an
IF
condition, flang accepted it silently and lowered it to a plain host
load of
device memory:

  subroutine s(h, a)
    double precision :: h, a(10)
    attributes(device) :: a
    if (a(3) > 2.5d0) h = 1.0d0
  end subroutine

The CUDA checker now emits an error when host code reads data with the
DEVICE or CONSTANT attribute in a scalar expression. These expressions
include IF and ELSE IF conditions, DO bounds, DO WHILE, SELECT CASE

    [12 lines not shown]
DeltaFile
+133-0flang/lib/Semantics/check-cuda.cpp
+126-0flang/test/Semantics/CUDA/cuf-device-data-host-read.cuf
+11-0flang/lib/Semantics/check-cuda.h
+270-03 files

LLVM/project 41ae65c — llvm/lib/Transforms/Vectorize VPlan.h

[VPlan] Add missing export to symbol used in test (#227888)
DeltaFile
+2-1llvm/lib/Transforms/Vectorize/VPlan.h
+2-11 files

LLVM/project 6a0f00d — mlir/lib/Dialect/LLVMIR/Transforms InlinerInterfaceImpl.cpp, mlir/test/Dialect/LLVMIR inlining-alias-scopes.mlir

[mlir][LLVM] Use a disjoint scope domain when inlining noalias

This matches recent changes to the LLVM inliner.

AI disclosure: Claude wrote the code, I wrote the commit message and
have done initial review.
DeltaFile
+28-25mlir/lib/Dialect/LLVMIR/Transforms/InlinerInterfaceImpl.cpp
+10-18mlir/test/Dialect/LLVMIR/inlining-alias-scopes.mlir
+38-432 files

LLVM/project 4f4537a — mlir/lib/Dialect/LLVMIR/Transforms InlinerInterfaceImpl.cpp

Update comment
DeltaFile
+2-2mlir/lib/Dialect/LLVMIR/Transforms/InlinerInterfaceImpl.cpp
+2-21 files

LLVM/project 6f00dda — mlir/test/Dialect/LLVMIR inlining-alias-scopes.mlir

Test that the inliner keeps the flag when it clones a domain

Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
DeltaFile
+35-0mlir/test/Dialect/LLVMIR/inlining-alias-scopes.mlir
+35-01 files

LLVM/project fcc4f6f — mlir/include/mlir/Dialect/LLVMIR LLVMAttrDefs.td, mlir/lib/Dialect/LLVMIR/Transforms InlinerInterfaceImpl.cpp

[mlir][LLVM] Add disjointScopes to AliasScopeDomainAttr

This also updates the MLIR-side inliner to clone disjoint domains
while cloning alias scopes, matching changes to LLVM.

AI disclosure: Claude wrote the code, I wrote the commit message and
looked at the code.
DeltaFile
+25-0mlir/test/Target/LLVMIR/Import/metadata-alias-scopes.ll
+23-0mlir/test/Target/LLVMIR/attribute-alias-scopes.mlir
+17-2mlir/include/mlir/Dialect/LLVMIR/LLVMAttrDefs.td
+4-1mlir/lib/Dialect/LLVMIR/Transforms/InlinerInterfaceImpl.cpp
+2-2mlir/lib/Target/LLVMIR/ModuleTranslation.cpp
+3-1mlir/lib/Target/LLVMIR/ModuleImport.cpp
+74-66 files

LLVM/project e2291f4 — mlir/lib/Dialect/LLVMIR/Transforms InlinerInterfaceImpl.cpp

Remove pointless comment
DeltaFile
+0-2mlir/lib/Dialect/LLVMIR/Transforms/InlinerInterfaceImpl.cpp
+0-21 files

LLVM/project 3985310 — llvm/test/CodeGen/AMDGPU lower-kernel-and-module-lds.ll lower-module-lds-via-hybrid.ll

Test fixups
DeltaFile
+28-19llvm/test/CodeGen/AMDGPU/lower-lds-struct-aa-merge.ll
+36-10llvm/test/CodeGen/AMDGPU/lower-module-lds-precise-allocate-to-module-struct.ll
+17-17llvm/test/CodeGen/AMDGPU/lower-lds-struct-aa-memcpy.ll
+15-18llvm/test/CodeGen/AMDGPU/lower-lds-struct-aa.ll
+11-13llvm/test/CodeGen/AMDGPU/lower-module-lds-via-hybrid.ll
+9-9llvm/test/CodeGen/AMDGPU/lower-kernel-and-module-lds.ll
+116-865 files not shown
+148-12011 files

LLVM/project 622513b — llvm/lib/Target/AMDGPU AMDGPULowerModuleLDSPass.cpp, llvm/test/CodeGen/AMDGPU remove-no-kernel-id-attribute.ll

[AMDGPU] Use a disjoint scope domain for merged LDS structs

When lowering LDS values, all the values are mutually disjoint, so we
can use the newly-added disjoint scopes feature to simplify the IR.

AI disclosure: Claude wrote this and I reviewed it and wrote the
 commit message
DeltaFile
+16-51llvm/lib/Target/AMDGPU/AMDGPULowerModuleLDSPass.cpp
+8-12llvm/test/CodeGen/AMDGPU/remove-no-kernel-id-attribute.ll
+24-632 files

LLVM/project 0edb0f8 — llvm/lib/Target/AMDGPU AMDGPULowerKernelArguments.cpp, llvm/test/CodeGen/AMDGPU lower-kernargs.ll si-split-load-store-alias-info.ll

[AMDGPU] Use a disjoint scope domain for noalias kernel arguments

All noalias arguments of a kernel are disjoint with each other, so we
can use a disjoint scope to save on metadata construction.

AI disclosure: Claude wrote this, I looked at it and wrote this
message.
DeltaFile
+90-105llvm/test/CodeGen/AMDGPU/lower-noalias-kernargs.ll
+19-13llvm/lib/Target/AMDGPU/AMDGPULowerKernelArguments.cpp
+11-11llvm/test/CodeGen/AMDGPU/lower-kernel-arguments-noalias-call-no-ptr-args.ll
+8-8llvm/test/CodeGen/AMDGPU/si-split-load-store-alias-info.ll
+4-4llvm/test/CodeGen/AMDGPU/lower-kernargs.ll
+132-1415 files

LLVM/project f62f1e5 — clang/test/CodeGen arm-v8.2a-neon-intrinsics-generic.c arm_neon_intrinsics.c, llvm/lib/Transforms/Utils InlineFunction.cpp

[Inliner] Use a disjoint scope domain for noalias arguments

InlineFunction creates alias.scope/noalias metadata to represent the
set of `noalias` arguments to a function. We don't need the `!noalias`
now that we have the ability to use disjoint scopes, saving us IR size
and metadata bloat.

TODO move these to a previous commit.
Also changes InstCombine to not drop the experimental.noalias.scope.decl
for disjoint scopes even if they're not mentioned in a `!noalias`, but
do still delete them if they're not used.
DeltaFile
+324-324clang/test/CodeGen/arm_neon_intrinsics.c
+72-72clang/test/CodeGen/arm-v8.2a-neon-intrinsics-generic.c
+35-30llvm/lib/Transforms/Utils/InlineFunction.cpp
+13-14llvm/test/Transforms/Inline/noalias-calls2.ll
+12-12llvm/test/Transforms/Inline/noalias2.ll
+11-11llvm/test/Transforms/PhaseOrdering/pr39282.ll
+467-4639 files not shown
+487-48315 files

LLVM/project 2430d74 — llvm/test/Transforms/Inline noalias2.ll

Test a callee that has both noalias arguments and its own scopes

Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
DeltaFile
+51-0llvm/test/Transforms/Inline/noalias2.ll
+51-01 files

LLVM/project 6ad28cd — llvm/test/CodeGen/SPIRV/extensions/SPV_INTEL_memory_access_aliasing alias-scope-shared-across-functions.ll

Fix spir-v test
DeltaFile
+1-1llvm/test/CodeGen/SPIRV/extensions/SPV_INTEL_memory_access_aliasing/alias-scope-shared-across-functions.ll
+1-11 files

LLVM/project 33d0272 — llvm/include/llvm/IR Metadata.h, llvm/include/llvm/Transforms/Utils Cloning.h

Review feedback
DeltaFile
+53-31llvm/lib/Transforms/InstCombine/InstructionCombining.cpp
+26-5llvm/test/Transforms/InstCombine/noalias-scope-decl-disjoint-domain.ll
+16-14llvm/lib/Transforms/Utils/CloneFunction.cpp
+4-3mlir/lib/Target/LLVMIR/ModuleImport.cpp
+3-1llvm/include/llvm/Transforms/Utils/Cloning.h
+3-1llvm/include/llvm/IR/Metadata.h
+105-551 files not shown
+105-567 files