LLVM/project 2fd518b — llvm/lib/Target/RISCV RISCVInstrInfoP.td

[RISCV][P-ext] Combine duplicate mhacc(su)_h_b(0/1) patterns. NFC (#230358)

Assisted-by: Claude
DeltaFile
+18-35llvm/lib/Target/RISCV/RISCVInstrInfoP.td
+18-351 files

LLVM/project a9ba4a7 — clang/test/ClangScanDeps modules-extern-unrelated.m, clang/test/Modules module-map-input-file-order-relocation-check.c

[clang][modules] Add test coverage on module map resolution, NFC (#230569)

* In ScanDeps test, capture which modulemap gets resolved as fmodule-map
dependency.
* In Modules test, capture how adding relocation checks can alter the
ordering of serialized input files.
DeltaFile
+65-0clang/test/Modules/module-map-input-file-order-relocation-check.c
+1-0clang/test/ClangScanDeps/modules-extern-unrelated.m
+66-02 files

LLVM/project 4d53a4d — clang/lib/Sema SemaHLSL.cpp, clang/test/CodeGenHLSL/semantics compute-id-16bit.hlsl

[SemaHLSL] Add missing validations of existing semantics (#224139)

This pr adds the following semantic analysis for semantics:

- Validate scalar/vector shapes, element types, and supported widths for
system semantics
 - Enforce indexing restrictions and SV_Target bounds
 - Add diag that shows previous use of overlapping semantics
 - as well as, simplifying some related sema logic.

Resolves #189765

Assisted by: GPT-6 Astra
DeltaFile
+213-112clang/lib/Sema/SemaHLSL.cpp
+86-0clang/test/SemaHLSL/Semantics/semantic-overlap.hlsl
+69-0clang/test/SemaHLSL/Semantics/target.index.hlsl
+59-0clang/test/CodeGenHLSL/semantics/compute-id-16bit.hlsl
+57-0clang/test/SemaHLSL/Semantics/semantic-64bit-types.hlsl
+37-14clang/test/SemaHLSL/Semantics/invalid_entry_parameter.hlsl
+521-12618 files not shown
+916-20324 files

LLVM/project bce11cb — mlir/lib/Dialect/LLVMIR/Transforms InlinerInterfaceImpl.cpp

Update comment
DeltaFile
+2-2mlir/lib/Dialect/LLVMIR/Transforms/InlinerInterfaceImpl.cpp
+2-21 files

LLVM/project 9d28748 — mlir/lib/Dialect/LLVMIR/Transforms InlinerInterfaceImpl.cpp, mlir/test/Dialect/LLVMIR inlining-alias-scopes.mlir

[mlir][LLVM] Use a disjoint scope domain when inlining noalias

This matches recent changes to the LLVM inliner.

AI disclosure: Claude wrote the code, I wrote the commit message and
have done initial review.
DeltaFile
+28-25mlir/lib/Dialect/LLVMIR/Transforms/InlinerInterfaceImpl.cpp
+10-18mlir/test/Dialect/LLVMIR/inlining-alias-scopes.mlir
+38-432 files

LLVM/project 9a5f86a — mlir/include/mlir/Dialect/LLVMIR LLVMAttrDefs.td, mlir/lib/Dialect/LLVMIR/Transforms InlinerInterfaceImpl.cpp

[mlir][LLVM] Add disjointScopes to AliasScopeDomainAttr

This also updates the MLIR-side inliner to clone disjoint domains
while cloning alias scopes, matching changes to LLVM.

AI disclosure: Claude wrote the code, I wrote the commit message and
looked at the code.
DeltaFile
+25-0mlir/test/Target/LLVMIR/Import/metadata-alias-scopes.ll
+23-0mlir/test/Target/LLVMIR/attribute-alias-scopes.mlir
+17-2mlir/include/mlir/Dialect/LLVMIR/LLVMAttrDefs.td
+4-1mlir/lib/Dialect/LLVMIR/Transforms/InlinerInterfaceImpl.cpp
+2-2mlir/lib/Target/LLVMIR/ModuleTranslation.cpp
+3-1mlir/lib/Target/LLVMIR/ModuleImport.cpp
+74-66 files

LLVM/project d03b1ad — mlir/lib/Dialect/LLVMIR/Transforms InlinerInterfaceImpl.cpp

Remove pointless comment
DeltaFile
+0-2mlir/lib/Dialect/LLVMIR/Transforms/InlinerInterfaceImpl.cpp
+0-21 files

LLVM/project ea86a8a — mlir/test/Dialect/LLVMIR inlining-alias-scopes.mlir, mlir/test/Target/LLVMIR attribute-alias-scopes.mlir

[AI test fix] Use property syntax for alias scope attributes

The strict-properties assembly format change means inherent attributes
such as `alignment`, `alias_scopes` and `noalias_scopes` on
`llvm.load`/`llvm.store` can no longer be parsed out of the trailing
attribute dictionary. Switch the newly added tests to the `<...>`
property syntax already used by the rest of these files.
DeltaFile
+2-2mlir/test/Target/LLVMIR/attribute-alias-scopes.mlir
+2-2mlir/test/Dialect/LLVMIR/inlining-alias-scopes.mlir
+4-42 files

LLVM/project b706d92 — mlir/test/Dialect/LLVMIR inlining-alias-scopes.mlir

Test that the inliner keeps the flag when it clones a domain

Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
DeltaFile
+35-0mlir/test/Dialect/LLVMIR/inlining-alias-scopes.mlir
+35-01 files

LLVM/project fc2d517 — llvm/test/CodeGen/AMDGPU lower-kernel-and-module-lds.ll lower-module-lds-via-hybrid.ll

Test fixups
DeltaFile
+28-19llvm/test/CodeGen/AMDGPU/lower-lds-struct-aa-merge.ll
+36-10llvm/test/CodeGen/AMDGPU/lower-module-lds-precise-allocate-to-module-struct.ll
+17-17llvm/test/CodeGen/AMDGPU/lower-lds-struct-aa-memcpy.ll
+15-18llvm/test/CodeGen/AMDGPU/lower-lds-struct-aa.ll
+11-13llvm/test/CodeGen/AMDGPU/lower-module-lds-via-hybrid.ll
+9-9llvm/test/CodeGen/AMDGPU/lower-kernel-and-module-lds.ll
+116-865 files not shown
+148-12011 files

LLVM/project 242e352 — llvm/lib/Target/AMDGPU AMDGPULowerModuleLDSPass.cpp, llvm/test/CodeGen/AMDGPU remove-no-kernel-id-attribute.ll

[AMDGPU] Use a disjoint scope domain for merged LDS structs

When lowering LDS values, all the values are mutually disjoint, so we
can use the newly-added disjoint scopes feature to simplify the IR.

AI disclosure: Claude wrote this and I reviewed it and wrote the
 commit message
DeltaFile
+16-51llvm/lib/Target/AMDGPU/AMDGPULowerModuleLDSPass.cpp
+8-12llvm/test/CodeGen/AMDGPU/remove-no-kernel-id-attribute.ll
+24-632 files

LLVM/project 6fdf983 — llvm/lib/Target/AMDGPU AMDGPULowerKernelArguments.cpp, llvm/test/CodeGen/AMDGPU lower-kernargs.ll si-split-load-store-alias-info.ll

[AMDGPU] Use a disjoint scope domain for noalias kernel arguments

All noalias arguments of a kernel are disjoint with each other, so we
can use a disjoint scope to save on metadata construction.

AI disclosure: Claude wrote this, I looked at it and wrote this
message.
DeltaFile
+90-105llvm/test/CodeGen/AMDGPU/lower-noalias-kernargs.ll
+19-13llvm/lib/Target/AMDGPU/AMDGPULowerKernelArguments.cpp
+11-11llvm/test/CodeGen/AMDGPU/lower-kernel-arguments-noalias-call-no-ptr-args.ll
+8-8llvm/test/CodeGen/AMDGPU/si-split-load-store-alias-info.ll
+4-4llvm/test/CodeGen/AMDGPU/lower-kernargs.ll
+132-1415 files

LLVM/project 07a4bea — llvm/lib/Transforms/Utils InlineFunction.cpp

Review feedback
DeltaFile
+1-4llvm/lib/Transforms/Utils/InlineFunction.cpp
+1-41 files

LLVM/project 5c89182 — llvm/test/Transforms/Inline noalias2.ll

Test a callee that has both noalias arguments and its own scopes

Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
DeltaFile
+51-0llvm/test/Transforms/Inline/noalias2.ll
+51-01 files

LLVM/project d93ab6c — clang/test/CodeGen arm-v8.2a-neon-intrinsics-generic.c arm_neon_intrinsics.c, llvm/lib/Transforms/Utils InlineFunction.cpp

[Inliner] Use a disjoint scope domain for noalias arguments

InlineFunction creates alias.scope/noalias metadata to represent the
set of `noalias` arguments to a function. We don't need the `!noalias`
now that we have the ability to use disjoint scopes, saving us IR size
and metadata bloat.

TODO move these to a previous commit.
Also changes InstCombine to not drop the experimental.noalias.scope.decl
for disjoint scopes even if they're not mentioned in a `!noalias`, but
do still delete them if they're not used.
DeltaFile
+324-324clang/test/CodeGen/arm_neon_intrinsics.c
+72-72clang/test/CodeGen/arm-v8.2a-neon-intrinsics-generic.c
+35-30llvm/lib/Transforms/Utils/InlineFunction.cpp
+13-14llvm/test/Transforms/Inline/noalias-calls2.ll
+12-12llvm/test/Transforms/Inline/noalias2.ll
+11-11llvm/test/Transforms/PhaseOrdering/pr39282.ll
+467-4639 files not shown
+487-48315 files

LLVM/project 405b8b8 — llvm/lib/Transforms/Vectorize SLPVectorizer.cpp, llvm/lib/Transforms/Vectorize/SLPVectorizer SLPUtils.h

[SLP]Do not copy disjoint to widened reduction ops

i1 leaves constrain bit 0 only, and an or of the narrowed chain may
overlap its operands. Copying the disjoint flag to the emitted wide ops
is a miscompile.

Fixes #230419

Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/230624
DeltaFile
+34-7llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+4-2llvm/lib/Transforms/Vectorize/SLPVectorizer/SLPUtils.h
+1-1llvm/test/Transforms/SLPVectorizer/X86/zext-or-nibble-reduction.ll
+1-1llvm/test/Transforms/SLPVectorizer/X86/logical-reduction-booleanized-leaves.ll
+40-114 files

LLVM/project e6ccd82 — clang/lib/CIR/CodeGen CIRGenStmtOpenMP.cpp, clang/test/CIR/CodeGenOpenMP parallel-for.c target-parallel-for.c

[CIR][OpenMP] Add support for the OpenMP 'for' directive

This patch adds support for wsloop in ClangIR: the `for` directive and its
combined forms `parallel for` and `target parallel for`. This is lowered to an
omp.wsloop + omp.loop_nest, nested utilizing the existing queue-based
decomposition.

Assisted-by: Cursor / Claude Sonnet 5 High
DeltaFile
+271-8clang/lib/CIR/CodeGen/CIRGenStmtOpenMP.cpp
+232-0clang/test/CIR/CodeGenOpenMP/for-loop-forms.c
+207-0clang/test/CIR/CodeGenOpenMP/pragma-omp-for.c
+142-0clang/test/CIR/CodeGenOpenMP/target-parallel-for.c
+62-0clang/test/CIR/CodeGenOpenMP/parallel-for.c
+914-85 files

LLVM/project ab4a90b — lldb/source/Plugins/TypeSystem/Fortran TypeSystemFortran.cpp

[lldb][Fortran][NFC] Added TODO comment for -fdefault-*-8 above GetBasicTypeFromAST
DeltaFile
+4-0lldb/source/Plugins/TypeSystem/Fortran/TypeSystemFortran.cpp
+4-01 files

LLVM/project ecce996 — llvm/lib/Target/RISCV RISCVISelLowering.cpp

[RISCV] Correct violations of SDTCisSameNumEltsAs. (#230612)

We missed a conversion to scalable vector in
lowerVectorMaskVecReduction.

lowerINSERT_SUBVECTOR used an all 1s mask with the wrong type. This made
a call to getDefaultVLOps unnecessary as both of its Mask and VL are
overridden.

Assisted-by: Claude
DeltaFile
+5-3llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+5-31 files

LLVM/project 0fbd429 — clang/test/CIR/CodeGen compound_literal.c

[CIR][NFC] Fix CIR compound literal test (#230613)

Fix CIR compound literal lit test
DeltaFile
+1-1clang/test/CIR/CodeGen/compound_literal.c
+1-11 files

LLVM/project e335041 — llvm/lib/CodeGen MachineBasicBlock.cpp

CodeGen: Compute live-outs of a split critical edge while updating LiveIntervals (#230499)

Fill in the new block's live-out set in the same loop that updates the
live intervals, instead of walking the original block's set a second
time with another liveAt query per register. A register is live out of
the new block if it is a PHI source or is live into Succ, which the
update already checks.

Instructions retired in phi-node-elimination, x86_64 -O3, on a generated
chain of N compare blocks branching to a shared PHI block:

  N     before          after           after/before
  1k       94,119,121      88,904,465   0.94
  2k      353,559,942     330,642,394   0.94
  4k    1,201,233,649   1,136,359,710   0.95
  8k    4,305,218,922   4,136,184,186   0.96
  16k  16,076,293,247  15,539,552,966   0.97

gcc-c-torture compile/20001226-1.c (liveintervals,phi-node-elimination

    [2 lines not shown]
DeltaFile
+9-11llvm/lib/CodeGen/MachineBasicBlock.cpp
+9-111 files

LLVM/project 6733e12 — llvm/test/Transforms/SLPVectorizer/X86 logical-reduction-booleanized-leaves.ll zext-or-nibble-reduction.ll

[SLP][NFC]Add tests with incorrect disjoint flag propagation, NFC



Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/230621
DeltaFile
+69-0llvm/test/Transforms/SLPVectorizer/X86/zext-or-nibble-reduction.ll
+49-0llvm/test/Transforms/SLPVectorizer/X86/logical-reduction-booleanized-leaves.ll
+118-02 files

LLVM/project 0ffc2b5 — llvm/test/TableGen HwModeBitSet.td, llvm/unittests/CodeGen TargetOptionsTest.cpp

also handle the codegen classes
DeltaFile
+39-0llvm/unittests/CodeGen/TargetOptionsTest.cpp
+7-6llvm/unittests/MC/TargetRegistry.cpp
+4-6llvm/test/TableGen/HwModeBitSet.td
+50-123 files

LLVM/project 44109cd — llvm/include/llvm/MC MCSubtargetInfo.h, llvm/lib/MC MCContext.cpp

[nspr] initial commit
DeltaFile
+24-1llvm/utils/TableGen/SubtargetEmitter.cpp
+15-0llvm/test/TableGen/HwModeBitSet.td
+2-5llvm/unittests/MC/TargetRegistry.cpp
+2-5llvm/unittests/CodeGen/TargetOptionsTest.cpp
+5-0llvm/include/llvm/MC/MCSubtargetInfo.h
+2-2llvm/lib/MC/MCContext.cpp
+50-131 files not shown
+53-147 files

LLVM/project 1881f20 — flang/include/flang/Optimizer/Dialect/CUF CUFOps.td, flang/lib/Optimizer/Dialect/CUF CUFOps.cpp

[flang][CUDA] Safely verify registered kernel symbols (#230588)

`cuf.register_kernel` resolved GPU module and kernel symbols during
nested function verification, potentially racing with concurrent changes
to the sibling GPU module. Add `SymbolUserOpInterface` and move these
lookups to `verifySymbolUses()`, where they run as part of module-level
symbol verification while the sibling IR is stable.
DeltaFile
+22-13flang/lib/Optimizer/Dialect/CUF/CUFOps.cpp
+3-1flang/include/flang/Optimizer/Dialect/CUF/CUFOps.td
+25-142 files

LLVM/project 0e411b2 — flang/lib/Lower/OpenMP OpenMP.cpp, flang/test/Lower/OpenMP variant-source-context.f90

Associate a BLOCK across ignored compiler directives

A standalone METADIRECTIVE that selects a block-associated directive
takes the BLOCK that follows it. An ignored compiler directive between
them, such as !dir$ ignored_comment, ended the search, so the selected
TARGET was empty and the BLOCK was lowered outside it.

Skip the compiler directives that semantics reports as ignored (an
unrecognized directive, a bare name/value list, LOOP COUNT, and
ASSUME_ALIGNED) and move them with the BLOCK as the DO path does. Other
compiler directives, such as PREFETCH, still end the search like any
statement.
DeltaFile
+38-0flang/test/Lower/OpenMP/variant-source-context.f90
+22-4flang/lib/Lower/OpenMP/OpenMP.cpp
+60-42 files

LLVM/project 5cbf50a — llvm/test/Transforms/PhaseOrdering/X86 hsub.ll hadd.ll

fixup! Re-generate some CHECKs
DeltaFile
+7-7llvm/test/Transforms/PhaseOrdering/X86/hsub.ll
+7-7llvm/test/Transforms/PhaseOrdering/X86/hadd.ll
+14-142 files

LLVM/project 98ba607 — flang/lib/Optimizer/Transforms/CUDA CUFAddConstructor.cpp, flang/test/Fir/CUDA cuda-managed-pointer-linkage.cuf cuda-constructor-2.f90

[flang][cuda] initialize the CUDA module after all units register managed variables (#230350)

With relocatable device code, every unit registers its managed variables
on the same CUDA module, and the runtime does not fill variables
registered after the module is initialized. Move CUFInitModule to a
second constructor with the next priority value, so that it runs after
the registration constructors of every unit in the executable or shared
library.

This will fix segmentation fault in example like: 

```
module m1
  integer, managed :: x1
contains
  attributes(global) subroutine k1()
    x1 = 1
  end subroutine
end module

    [20 lines not shown]
DeltaFile
+47-3flang/lib/Optimizer/Transforms/CUDA/CUFAddConstructor.cpp
+10-1flang/test/Fir/CUDA/cuda-constructor-2.f90
+2-0flang/test/Fir/CUDA/cuda-managed-pointer-linkage.cuf
+59-43 files

LLVM/project da913c2 — llvm/lib/Transforms/Scalar SeparateConstOffsetFromGEP.cpp, llvm/test/CodeGen/AMDGPU memintrinsic-unroll.ll memmove-var-size.ll

[SeparateConstOffsetFromGEP] Track cast state during offset extraction (#229839)

Keep the cast state used while searching for a constant offset and reuse
it when rebuilding the GEP index. This lets constants be extended or
truncated at the point where they are found, instead of carrying
separate sign/zero extension flags and then redistributing casts in a
second pass.

This also allows the RHS-of-sub zero-extension case because the constant
is zero-extended before it is negated.
DeltaFile
+112-164llvm/lib/Transforms/Scalar/SeparateConstOffsetFromGEP.cpp
+30-30llvm/test/CodeGen/AMDGPU/memmove-var-size.ll
+10-8llvm/test/Transforms/SeparateConstOffsetFromGEP/pr62379-zeroext-negative.ll
+6-6llvm/test/CodeGen/AMDGPU/memintrinsic-unroll.ll
+158-2084 files

LLVM/project 448cda5 — llvm/test/TableGen RuntimeLibcallEmitter-bad-system-library-entry-error.td RuntimeLibcallEmitter-library-dispatch.td, llvm/utils/TableGen/Basic RuntimeLibcallsEmitter.cpp

RuntimeLibcalls: Require system library members to be libraries (#230306)

Every SystemRuntimeLibrary now lists only LibcallLibrary and LibraryRef
members, so the inline path that expanded unhomed RuntimeLibcallImpl
members directly into the system's block is dead. Remove it, and error
on any member that is not a library.

This drops the SystemAvailableImpls bitset and the predicate groups
from the system setup function. Conditional and calling-convention
groups now only appear inside a LibcallLibrary, whose own setup
function emits them. The generated RuntimeLibcalls.inc is unchanged.

Co-authored-by: Claude Opus <noreply at anthropic.com>
DeltaFile
+118-108llvm/test/TableGen/RuntimeLibcallEmitter.td
+94-103llvm/test/TableGen/RuntimeLibcallEmitter-calling-conv.td
+18-117llvm/utils/TableGen/Basic/RuntimeLibcallsEmitter.cpp
+12-50llvm/test/TableGen/RuntimeLibcallEmitter-multiple-impls.td
+3-21llvm/test/TableGen/RuntimeLibcallEmitter-library-dispatch.td
+15-6llvm/test/TableGen/RuntimeLibcallEmitter-bad-system-library-entry-error.td
+260-4056 files not shown
+278-41612 files