LLVM/project cc8a48a — llvm/unittests/CAS ProgramTest.cpp

address review feedback

Created using spr 1.3.7
DeltaFile
+1-1llvm/unittests/CAS/ProgramTest.cpp
+1-11 files

LLVM/project 372d1ee — utils/bazel/llvm-project-overlay/mlir/unittests BUILD.bazel

[Bazel] Add Support dependency to amdgpu_tests (#229597)

Fixes amdgpu_tests bazel layering failure introduced in commit
45ebdb6a55db, where AMDGPUUtilsTest.cpp includes
llvm/Support/Compiler.h.
DeltaFile
+1-0utils/bazel/llvm-project-overlay/mlir/unittests/BUILD.bazel
+1-01 files

LLVM/project f970b5d — utils/bazel/llvm-project-overlay/libc BUILD.bazel

[Bazel][libc] Update platform_file to depend on dup3 (#229595)

Fixes bazel build failure introduced in commit 64d83dcba1cb, where
file.cpp switched from using dup2 to dup3.
DeltaFile
+1-1utils/bazel/llvm-project-overlay/libc/BUILD.bazel
+1-11 files

LLVM/project 93632d5 — flang/include/flang/Runtime/CUDA registration.h

[flang][cuda][NFC] Fix header for registration.h
DeltaFile
+1-1flang/include/flang/Runtime/CUDA/registration.h
+1-11 files

LLVM/project c202e0f — llvm/unittests/CAS ProgramTest.cpp

[𝘀𝗽𝗿] initial version

Created using spr 1.3.7
DeltaFile
+2-2llvm/unittests/CAS/ProgramTest.cpp
+2-21 files

LLVM/project 8aff158 — llvm/test/CodeGen/AMDGPU/GlobalISel legalize-load-local.mir regbankselect-amdgcn.s.buffer.load.ll

AMDGPU/GlobalISel: Use integer types when narrowing loads and stores

The narrowScalar mutation for loads and stores produced untyped scalars.
For FP-typed values this resulted in an untyped G_OR in the lowerLoad
expansion for unaligned private accesses, which failed to select.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+441-448llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-private.mir
+268-270llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-flat.mir
+361-0llvm/test/CodeGen/AMDGPU/GlobalISel/load-store-private-unaligned-f64.ll
+16-16llvm/test/CodeGen/AMDGPU/GlobalISel/regbankselect-amdgcn.s.buffer.load.ll
+16-16llvm/test/CodeGen/AMDGPU/GlobalISel/atomicrmw-fmin-fmax.ll
+12-12llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-local.mir
+1,114-7621 files not shown
+1,117-7657 files

LLVM/project 33f058d — llvm/lib/Target/AMDGPU SIInstrInfo.h SIInstrInfo.cpp, llvm/test/CodeGen/AMDGPU fix-sgpr-copies-f16-true16.mir fix-sgpr-copies-vgpr16-to-spgr32.ll

[AMDGPU] Fix true16 losing 16-bit subreg operands

Folding a true16 v2s copy such as %2:sreg_32 = COPY %1.lo16 rewrites
its users to read %1.lo16 and relies on legalizeOperandsVALUt16 to
legalize the narrower operand. PHI and REG_SEQUENCE operands have no
register class, so it skipped them, leaving a 16-bit input in a VGPR_32
PHI or a 32-bit REG_SEQUENCE slot. DetectDeadLanes then marked the PHI
input undef and the defining load was deleted.

Widen such operands with a REG_SEQUENCE in legalizeOperandsVALUt16,
which runs both when the copy is folded and when the user is moved to
the VALU. This miscompiled uniform i16 loads feeding PHIs on gfx1250.

Change-Id: I9ddee5b11ff8503f4b2b1d1c4c776a12b270d968
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
DeltaFile
+101-212llvm/test/CodeGen/AMDGPU/load-constant-i1.ll
+282-0llvm/test/CodeGen/AMDGPU/fix-sgpr-copies-sgpr32-to-vgpr16.ll
+125-0llvm/test/CodeGen/AMDGPU/fix-sgpr-copies-vgpr16-to-spgr32.ll
+68-13llvm/lib/Target/AMDGPU/SIInstrInfo.cpp
+75-0llvm/test/CodeGen/AMDGPU/fix-sgpr-copies-f16-true16.mir
+16-0llvm/lib/Target/AMDGPU/SIInstrInfo.h
+667-2256 files

LLVM/project 3ad4809 — flang/unittests/Evaluate uint128.cpp

Fix typo
DeltaFile
+1-1flang/unittests/Evaluate/uint128.cpp
+1-11 files

LLVM/project db34882 — llvm/lib/Target/RISCV RISCVISelLowering.cpp, llvm/test/CodeGen/RISCV fclass-select.ll

 [RISCV] Fold (sub 0, (srl (and X, (1 << ShAmt)), ShAmt)) -> (sra (shl X, ShAmt2), bits-1) (#229509)

Improves codegen of is.fpclass+select. The AND to test bits of
is.fpclass may get turned into and+srl while the select emits a
neg. Isel will turn the and+srl into shl+srl but it's too late to
fold the neg.

Assisted-by: Claude
DeltaFile
+114-0llvm/test/CodeGen/RISCV/fclass-select.ll
+23-0llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+137-02 files

LLVM/project 6c7c592 — llvm/lib/Target/AMDGPU AMDGPULegalizerInfo.cpp, llvm/test/CodeGen/AMDGPU/GlobalISel regbankselect-amdgcn.s.buffer.load.ll atomicrmw-fmin-fmax.ll

AMDGPU/GlobalISel: Use integer types when narrowing loads and stores

The narrowScalar mutation for loads and stores produced untyped scalars.
For FP-typed values this resulted in an untyped G_OR in the lowerLoad
expansion for unaligned private accesses, which failed to select.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+441-448llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-private.mir
+268-270llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-flat.mir
+361-0llvm/test/CodeGen/AMDGPU/GlobalISel/load-store-private-unaligned-f64.ll
+16-16llvm/test/CodeGen/AMDGPU/GlobalISel/regbankselect-amdgcn.s.buffer.load.ll
+16-16llvm/test/CodeGen/AMDGPU/GlobalISel/atomicrmw-fmin-fmax.ll
+3-3llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
+1,105-7536 files

LLVM/project 1dc4fa3 — llvm/lib/CodeGen/GlobalISel LegalizerHelper.cpp, llvm/test/CodeGen/AMDGPU/GlobalISel legalize-load-constant.mir legalize-load-flat.mir

GlobalISel: Use integer types when splitting loads in lowerLoad

lowerLoad built the split pieces using the destination type with the
element size changed, which preserved floating-point types. An
unaligned f64 load was decomposed into G_ZEXTLOAD, G_SHL and G_OR on
f32 and f64, which then crashed in AMDGPU RegBankLegalize. Build the
pieces as integers and bitcast to a non-integer result type, as
lowerStore already does.

The pieces are now consistently integer typed, which allows more
constants to be CSEd in the existing tests.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+3,147-2,589llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-local.mir
+2,662-2,708llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-private.mir
+2,490-2,634llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-global.mir
+1,880-1,952llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-flat.mir
+1,812-1,908llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-constant.mir
+13-12llvm/lib/CodeGen/GlobalISel/LegalizerHelper.cpp
+12,004-11,8036 files

LLVM/project b82f26b — llvm/lib/Transforms/Vectorize LoopVectorize.cpp

[LV] Use frozen start values from the main plan directly (NFC). (#229571)

Only freeze possibly-poison start values of FindIV reductions in the
main plan. When preparing the epilogue plan, set the start operand of
its FindIV reduction results directly to the value frozen for the main
plan, instead of creating Freeze recipes in the epilogue plan and
replacing them later via a map. This leaves only VPExpandSCEVRecipes to
replace in the epilogue plan's entry.
DeltaFile
+34-51llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
+34-511 files

LLVM/project cfd56f3 — llvm/test/CodeGen/AMDGPU amdgcn.bitcast.768bit.ll amdgcn.bitcast.832bit.ll

address review feedback

Created using spr 1.3.7
DeltaFile
+85,305-104,728llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+23,897-29,739llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+16,898-20,565llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.960bit.ll
+15,617-19,097llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.896bit.ll
+14,370-17,720llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.832bit.ll
+12,180-15,528llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.768bit.ll
+168,267-207,37711,737 files not shown
+776,240-588,41011,743 files

LLVM/project 9184c30 — clang/lib/CIR/Dialect/IR CIRDialect.cpp, clang/test/CIR/Transforms canonicalize.cir

[CIR] Fix use-after-free in BrOp::canonicalize (#229191)

BrOp::canonicalize took the branch's destination operands as an
OperandRange, erased the branch, and then passed the range to
mergeBlocks. The range points into the branch's operand storage, which
eraseOp frees, so mergeBlocks read freed memory whenever the merged
block had arguments. Valgrind reports it on the new test:

```
  Invalid read of size 8
     at mlir::ValueRange::dereference_iterator
     by mlir::RewriterBase::inlineBlockBefore
     by cir::BrOp::canonicalize
   Address ... is 136 bytes inside a block of size 144 free'd
     by mlir::RewriterBase::eraseOp
     by cir::BrOp::canonicalize
```

It goes unnoticed with glibc malloc, which usually leaves the freed
operands intact. Copy the operands before erasing the branch.
DeltaFile
+34-0clang/test/CIR/Transforms/canonicalize.cir
+11-1clang/lib/CIR/Dialect/IR/CIRDialect.cpp
+45-12 files

LLVM/project 44a7e32 — clang/test/CodeGen fake-use-sanitizer.cpp

[msan][test] Add MSan case to clang/test/CodeGen/fake-use-sanitizer.cpp (#229564)

Regression test for https://github.com/llvm/llvm-project/issues/225425,
showing that MSan incorrectly strictly handles fake_use.
DeltaFile
+46-0clang/test/CodeGen/fake-use-sanitizer.cpp
+46-01 files

LLVM/project 5ab70ae — clang/lib/CIR/Dialect/Transforms LoweringPrepare.cpp, clang/test/CIR/Transforms lib-opt.cir target-lowering.cir

[CIR] Register TargetLowering, LoweringPrepare and LibOpt in cir-opt

DeltaFile
+56-0clang/test/CIR/Transforms/lowering-prepare.cir
+36-0clang/test/CIR/Transforms/lowering-prepare-global-ctor.cir
+20-0clang/test/CIR/Transforms/target-lowering.cir
+14-0clang/test/CIR/Transforms/lib-opt.cir
+12-0clang/tools/cir-opt/cir-opt.cpp
+5-2clang/lib/CIR/Dialect/Transforms/LoweringPrepare.cpp
+143-21 files not shown
+148-27 files

LLVM/project 7628e26 — clang/unittests/CIR ControlFlowTest.cpp

[CIR] Update ControlFlowTest for the new global ctor syntax
DeltaFile
+1-1clang/unittests/CIR/ControlFlowTest.cpp
+1-11 files

LLVM/project 155462f — mlir/include/mlir/Conversion Passes.td, mlir/lib/Conversion/AMDGPUToROCDL AMDGPUToROCDL.cpp

[mlir] Migrate AMDGPU/ROCDL to targets, not chipset versions (#223563)

**migration tl;dr:** Replace usages of `amdgpu::Chipset` with
`ROCDL::TargetInfo`, ideally move from `chipset=` to `arch=`. If you
don't use upstream pipelines, call
'TargetInfo::migrateArchFeaturesToModuleFlags` at the appropriate
location.

Further note: if you've got a build pipeline that's getting a `gfxXXX`
name from something like `rocm_agent_enumerator`, using a full triple
name like the ones you get from `rocminfo` is preferred.

`amdgpu::Chipset` was an awkward hack that was hard to keep up to date
with changes in the compiler/new architectures, and didn't properly
support generic targets (and has been strongly disfavored by the
compiler team).

This PR replaces `amdgpu::Chipset` with `ROCDL::TargetInfo`, a
structure that uses LLVM's TargetParser and the underlying LLVM

    [47 lines not shown]
DeltaFile
+316-324mlir/lib/Conversion/AMDGPUToROCDL/AMDGPUToROCDL.cpp
+105-0mlir/test/Dialect/LLVMIR/rocdl-attach-target-arch.mlir
+87-4mlir/lib/Dialect/GPU/Transforms/ROCDLAttachTarget.cpp
+46-40mlir/lib/Dialect/AMDGPU/Transforms/EmulateAtomics.cpp
+49-24mlir/test/Dialect/AMDGPU/amdgpu-emulate-atomics.mlir
+46-15mlir/include/mlir/Conversion/Passes.td
+649-407103 files not shown
+1,205-727109 files

LLVM/project b6de87e — utils/bazel/llvm-project-overlay/clang BUILD.bazel

[Bazel] Add FlowSensitive/Models headers to clang:analysis (#229578)

Fixes bazel build failure introduced in commits e03123a63162 and
f0d6540d5307, where GtestModelHelpers.h was added to
lib/Analysis/FlowSensitive/Models and included in Models/*.cpp.
DeltaFile
+2-0utils/bazel/llvm-project-overlay/clang/BUILD.bazel
+2-01 files

LLVM/project ef9f3c1 — llvm/lib/MC GOFFObjectWriter.cpp, llvm/test/CodeGen/SystemZ zos-symbol-2.ll zos-common-global.ll

Take a different approach

HLASM derives the AMODE from the RMODE if no AMODE is explicitly
given. The same can be done in the GOFF writer, simplifying the
coding a lot.
DeltaFile
+19-0llvm/lib/MC/GOFFObjectWriter.cpp
+4-4llvm/test/CodeGen/SystemZ/zos-section-2.ll
+2-2llvm/test/CodeGen/SystemZ/zos-section-1.ll
+1-1llvm/test/CodeGen/SystemZ/zos-symbol-2.ll
+1-1llvm/test/CodeGen/SystemZ/zos-common-global.ll
+27-85 files

LLVM/project 6c0a5e1 — llvm/include/llvm/CodeGen SlotIndexes.h, llvm/lib/CodeGen RegAllocGreedy.cpp LiveIntervals.cpp

[CodeGen] Drop dead SlotIndexes before allocation

SlotIndexes keeps the index list entry of an erased instruction and only
clears its instruction pointer. Live range sizes are measured in slot
indexes and greedy ranks ranges by size, so the leftovers inflate some
ranges more than others and reorder allocation, spilling heavily on
register-starved functions.

Add SlotIndexes::compactIndexes() to erase them. Erased entries are
unlinked, so LiveIntervals first reports the indexes it holds via
appendReferencedIndexes().

Off by default behind -greedy-compact-slot-indexes, since it changes
allocation across much of the test suite.
DeltaFile
+221-0llvm/unittests/CodeGen/SlotIndexesTest.cpp
+46-0llvm/lib/CodeGen/SlotIndexes.cpp
+32-0llvm/test/CodeGen/AMDGPU/greedy-compact-slot-indexes.ll
+23-6llvm/include/llvm/CodeGen/SlotIndexes.h
+29-0llvm/lib/CodeGen/LiveIntervals.cpp
+12-1llvm/lib/CodeGen/RegAllocGreedy.cpp
+363-71 files not shown
+367-77 files

LLVM/project 64d83dc — libc/src/__support/File/linux file.cpp, libc/test/src/__support/File file_mode_test.cpp

[libc] Add 'x' and 'e' mode support to fopen() (#224207)

Add exclusive creation (x) and close-on-exec (e) support to the fopen.

Also extended the related close-on-exec handling to fdopen() and to
freopen()

Fixes #223070

Assisted-by: Codex
DeltaFile
+95-0libc/test/src/stdio/fopen_test.cpp
+47-0libc/test/src/__support/File/file_mode_test.cpp
+40-0libc/test/src/stdio/fdopen_test.cpp
+35-0libc/test/src/stdio/freopen_test.cpp
+31-2libc/src/__support/File/linux/file.cpp
+18-0libc/test/src/stdio/CMakeLists.txt
+266-23 files not shown
+282-39 files

LLVM/project 81dfeb7 — libc/src/locale newlocale.h setlocale.cpp, libc/test/src/locale CMakeLists.txt locale_test.cpp

[libc][locale] Fix nullptr argument handling in setlocale. (#229569)

`setlocale` function should accept `nullptr` as a possible value for the
second argument - in this case the
locale is not modified but queried. Make sure to return `C` locale in
this case (the only one supported by LLVM-libc).

Expand the unit test to verify `setlocale` behavior.
DeltaFile
+24-1libc/test/src/locale/locale_test.cpp
+13-3libc/src/locale/setlocale.cpp
+9-4libc/src/locale/newlocale.h
+1-0libc/test/src/locale/CMakeLists.txt
+47-84 files

LLVM/project d6b1354 — flang-rt/lib/cuda registration.cpp, flang/include/flang/Runtime/CUDA registration.h

[flang-rt][cuda] Add CUFRegisterHostMemoryRange entry point (#229572)

Extracted runtime part from #229213
DeltaFile
+25-0flang-rt/lib/cuda/registration.cpp
+5-0flang/include/flang/Runtime/CUDA/registration.h
+30-02 files

LLVM/project 4a44b71 — clang/lib/CIR/Dialect/IR CIRDialect.cpp, clang/test/CIR/CodeGen try-catch.cpp

[CIR] Classify reference-to-pointer catch parameters in CIRGen

When a handler catches a reference to a pointer, __cxa_begin_catch returns
the caught pointer by value instead of the address of the exception object,
and how the reference is bound depends on whether the pointer points to a
class.  The CIR type of the catch parameter cannot always tell whether it is
a reference to a pointer, or whether the pointee is a class.
CIRGen now records both facts from the AST in two new InitCatchKind values,
reference_to_pointer and reference_to_record_pointer.  The Itanium EH
lowering switches on the kind alone, the way classic codegen does.

Assisted-by: Claude Code / Claude Opus 5.5
DeltaFile
+258-0clang/test/CIR/IR/invalid-init-catch-param.cir
+216-0clang/test/CIR/Transforms/eh-abi-lowering-catch-param-ref.cir
+198-8clang/test/CIR/CodeGen/try-catch.cpp
+173-0clang/test/CIR/IR/construct-catch-param.cir
+107-0clang/lib/CIR/Dialect/IR/CIRDialect.cpp
+86-2clang/test/CIR/Transforms/eh-abi-lowering-construct-catch-invalid.cir
+1,038-104 files not shown
+1,173-4110 files

LLVM/project 069eb04 — llvm/lib/Target/AMDGPU GCNSubtarget.h GCNHazardRecognizer.cpp, llvm/test/CodeGen/AMDGPU postra-permlane-hazard.mir

[AMDGPU] Account for gfx950 permlane hazards during scheduling (#229517)

The [CDNA4 ISA Reference Guide, §4.5, Table 11, p.
21](https://www.amd.com/content/dam/amd/en/documents/instinct-tech-docs/instruction-set-architectures/amd-instinct-cdna4-instruction-set-architecture.pdf#page=29)
requires 2 wait states between a VALU writing a VGPR and a permlane
reading it, and 4 between a `VCMPX` writing EXEC and a permlane. As far
as I know, this is GFX950 specific.

Right now the gfx950 scheduler does not query this existing hazard
check, so final hazard processing can insert avoidable `nop`s despite
available independent instructions.

This patch expose the existing permlane hazard check in
`getHazardType()`.

Also add MIR coverage for source dependencies, unrelated writes,
available fillers, and VCMPX scheduling boundaries.

Partially addresses: #228816
DeltaFile
+127-0llvm/test/CodeGen/AMDGPU/postra-permlane-hazard.mir
+5-1llvm/lib/Target/AMDGPU/GCNHazardRecognizer.cpp
+2-0llvm/lib/Target/AMDGPU/GCNSubtarget.h
+134-13 files

LLVM/project 854d00c — llvm/lib/Target/ARM MVEGatherScatterLowering.cpp, llvm/test/CodeGen/Thumb2 mve-gather-increment.ll

[ARM] Protect against multi use increment in gather scatter combine (#227247)

Fixes #226592
DeltaFile
+93-20llvm/test/CodeGen/Thumb2/mve-gather-increment.ll
+1-9llvm/lib/Target/ARM/MVEGatherScatterLowering.cpp
+94-292 files

LLVM/project 358da84 — llvm/test/TableGen RuntimeLibcallEmitter-library-default-cc.td, llvm/utils/TableGen/Basic RuntimeLibcallsEmitter.cpp

RuntimeLibcallsEmitter: Handle calling conv in per-library functions (#223760)

Pull handling of the default calling convention into the
setAvailableLibFuncs_<name> functions, so the library logic will be
fully contained.

Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
DeltaFile
+80-5llvm/utils/TableGen/Basic/RuntimeLibcallsEmitter.cpp
+50-0llvm/test/TableGen/RuntimeLibcallEmitter-library-default-cc.td
+130-52 files

LLVM/project 8549c21 — libc/include signal.yaml, libc/include/llvm-libc-types sighandler_t.h

[libc] Use sighandler_t typedef in public headers for all platforms. (#229508)

`signal` function from ANSI C takes an argument and returns value of
type `typeof(void(int))`.

While not strictly necessary, some libc implementations provide their
own convenience typedef for the signal handler type. LLVM-libc used
`sighandler_t` previously, but only on Linux systems. This change makes
this typedef used on other systems as well, the reasons being:

* convenience for the user code, which can rely on this typedef in the
signal-handling code;
* ability to fix a gross hdrgen workaround (YAML configs and
hdrgen-based generation doesn't really support a C-style syntax for
functions-returning-function-pointers)
* ability to fix a **bug** in the generated header - `__NOEXCEPT` suffix
is translated into `noexcept` in the C++ mode, with this qualifier
applied to the _returned function pointer type_ instead of the `signal`
function itself. The "proper" way to use noexcept would be to move it

    [5 lines not shown]
DeltaFile
+3-7libc/include/signal.yaml
+0-3libc/include/llvm-libc-types/sighandler_t.h
+3-102 files

LLVM/project 4a3606e — llvm/lib/CodeGen TargetLoweringObjectFileImpl.cpp, llvm/test/CodeGen/X86 morestack-decl.ll code-model-elf-constpool-large.ll

[X86] Mark .lrodata for non-mergeable constants with SHF_X86_64_LARGE (#229279)

Under the large code model, a constant pool entry that is not mergeable
(a size other than 4, 8, 16 or 32 bytes, such as a 64-byte AVX-512
vector) is placed in .lrodata without SHF_X86_64_LARGE. Sections are
uniqued by name, so a large global emitted into .lrodata later in the
same module ends up in that unflagged section.

Assisted-By: Claude Opus 5.5
DeltaFile
+10-0llvm/test/CodeGen/X86/code-model-elf-constpool-large.ll
+3-2llvm/lib/CodeGen/TargetLoweringObjectFileImpl.cpp
+1-1llvm/test/CodeGen/X86/morestack-decl.ll
+14-33 files