LLVM/project 92dfb7a — clang/test/CodeGen fake-use-sanitizer.cpp, llvm/lib/Transforms/Instrumentation MemorySanitizer.cpp

[msan] Correctly handle llvm::fake_use (as a no-op) (#229591)

The fake_use intrinsic (used by -fextend-variable-liveness) was being
strictly handled (i.e., check that the parameter is fully initialized),
which led to false positives
(https://github.com/llvm/llvm-project/issues/225425). This patch solves
the issue by silently not instrumenting fake_use, since they are, by
definition, not real uses and therefore cannot lead to
use-of-uninitialized-memory.

Fixes: #225425
DeltaFile
+11-0llvm/lib/Transforms/Instrumentation/MemorySanitizer.cpp
+1-7clang/test/CodeGen/fake-use-sanitizer.cpp
+12-72 files

LLVM/project 38e992f — libc/src/__support/math CMakeLists.txt expf_float_eval.h, utils/bazel/llvm-project-overlay/libc BUILD.bazel

[libc][math] Improve accuracy for exp*f float-only implementations. (#227881)

Pure Estrin's scheme pushes the rounding errors a bit more than 1 ULP on
non-FMA targets for these functions.
DeltaFile
+16-9libc/src/__support/math/exp2f_float_utils.h
+3-3libc/src/__support/math/expf_float_eval.h
+1-0utils/bazel/llvm-project-overlay/libc/BUILD.bazel
+1-0libc/src/__support/math/CMakeLists.txt
+21-124 files

LLVM/project 3960c0e — llvm/lib/CodeGen/SelectionDAG SelectionDAGDumper.cpp

[SelectionDAG] Add ATOMIC_LOAD_FMAXIMUMNUM/FMINIMUMNUM to SDNode::getOperationName. (#229566)

Reorder the other FP min/max nodes to locally match the order in
ISDOpcodes.h.
DeltaFile
+4-2llvm/lib/CodeGen/SelectionDAG/SelectionDAGDumper.cpp
+4-21 files

LLVM/project 033848e — llvm/lib/CodeGen/SelectionDAG SelectionDAGDumper.cpp

[SelectionDAG] Add static_assert for ISD::BUILTIN_OP_END to SDNode::getOperationName to encourage updating when opcodes are added. (#229573)

Also add missing case for DEACTIVATION_SYMBOL.
DeltaFile
+5-0llvm/lib/CodeGen/SelectionDAG/SelectionDAGDumper.cpp
+5-01 files

LLVM/project 8b85e9d — llvm/lib/CodeGen/GlobalISel LegalizerHelper.cpp, llvm/test/CodeGen/AMDGPU/GlobalISel legalize-load-constant.mir legalize-load-flat.mir

GlobalISel: Use integer types when splitting loads in lowerLoad (#229544)

lowerLoad built the split pieces using the destination type with the
element size changed, which preserved floating-point types. An
unaligned f64 load was decomposed into G_ZEXTLOAD, G_SHL and G_OR on
f32 and f64, which then crashed in AMDGPU RegBankLegalize. Build the
pieces as integers and bitcast to a non-integer result type, as
lowerStore already does.

The pieces are now consistently integer typed, which allows more
constants to be CSEd in the existing tests.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+3,147-2,589llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-local.mir
+2,662-2,708llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-private.mir
+2,490-2,634llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-global.mir
+1,880-1,952llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-flat.mir
+1,812-1,908llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-constant.mir
+13-12llvm/lib/CodeGen/GlobalISel/LegalizerHelper.cpp
+12,004-11,8036 files

LLVM/project 65cb0aa — clang/lib/CIR/CodeGen CIRGenFunction.cpp CIRGenFunction.h, clang/test/CIR/CodeGen ctor-try-body.cpp

[CIR] Implement ctor-try-body rethrow (#229436)

This came up in a test suite, but we weren't properly re-throwing
exceptions when they were in the body of a constructor's try-body. This
patch mirrors classic-codegen's behavior reasonably well, implementing
this behavior properly.

Side note: this mirrors classic codegen's behavior of including
dtor-try-bodies too, but that isn't implemented yet, so those will just
be an NYI for now.
DeltaFile
+128-0clang/test/CIR/CodeGen/ctor-try-body.cpp
+15-5clang/lib/CIR/CodeGen/CIRGenException.cpp
+11-1clang/lib/CIR/CodeGen/CIRGenFunction.h
+2-1clang/lib/CIR/CodeGen/CIRGenFunction.cpp
+156-74 files

LLVM/project a9d5947 — flang/include/flang/Runtime/CUDA registration.h

[flang][cuda][NFC] Fix header for registration.h (#229599)
DeltaFile
+1-1flang/include/flang/Runtime/CUDA/registration.h
+1-11 files

LLVM/project a250fa8 — utils/bazel/llvm-project-overlay/mlir BUILD.bazel

[Bazel][mlir] Fixes build failures (#229605)

Fixes build failures following commit 4c520f2ae12f and commit
155462f440ff where ROCDLTargetInfo was introduced and referenced across
ROCDL/AMDGPU passes.
DeltaFile
+6-0utils/bazel/llvm-project-overlay/mlir/BUILD.bazel
+6-01 files

LLVM/project 3b5d7de — lldb/include/lldb/Target LanguageRuntime.h, lldb/source/ValueObject ValueObjectVariable.cpp

[lldb] Add LanguageRuntime::FixupVariableLocation (NFC) (#229602)

Some languages store a variable in a location whose indirection level is
only known at runtime. Swift, for example, emits resilient globals into
a fixed-size buffer; a value that doesn't fit is boxed on the heap and
the buffer holds a pointer to the box. Whether the value fits depends on
the runtime layout of a type, so the compiler can't encode it in the
DWARF location expression.

This patch adds a LanguageRuntime hook that ValueObjectVariable calls
after evaluating a variable's location, so the runtime can adjust it.
The default implementation does nothing.

Assisted-by: Claude
DeltaFile
+12-1lldb/source/ValueObject/ValueObjectVariable.cpp
+7-0lldb/include/lldb/Target/LanguageRuntime.h
+19-12 files

LLVM/project cc8a48a — llvm/unittests/CAS ProgramTest.cpp

address review feedback

Created using spr 1.3.7
DeltaFile
+1-1llvm/unittests/CAS/ProgramTest.cpp
+1-11 files

LLVM/project 372d1ee — utils/bazel/llvm-project-overlay/mlir/unittests BUILD.bazel

[Bazel] Add Support dependency to amdgpu_tests (#229597)

Fixes amdgpu_tests bazel layering failure introduced in commit
45ebdb6a55db, where AMDGPUUtilsTest.cpp includes
llvm/Support/Compiler.h.
DeltaFile
+1-0utils/bazel/llvm-project-overlay/mlir/unittests/BUILD.bazel
+1-01 files

LLVM/project f970b5d — utils/bazel/llvm-project-overlay/libc BUILD.bazel

[Bazel][libc] Update platform_file to depend on dup3 (#229595)

Fixes bazel build failure introduced in commit 64d83dcba1cb, where
file.cpp switched from using dup2 to dup3.
DeltaFile
+1-1utils/bazel/llvm-project-overlay/libc/BUILD.bazel
+1-11 files

LLVM/project 93632d5 — flang/include/flang/Runtime/CUDA registration.h

[flang][cuda][NFC] Fix header for registration.h
DeltaFile
+1-1flang/include/flang/Runtime/CUDA/registration.h
+1-11 files

LLVM/project c202e0f — llvm/unittests/CAS ProgramTest.cpp

[𝘀𝗽𝗿] initial version

Created using spr 1.3.7
DeltaFile
+2-2llvm/unittests/CAS/ProgramTest.cpp
+2-21 files

LLVM/project 8aff158 — llvm/test/CodeGen/AMDGPU/GlobalISel legalize-load-local.mir regbankselect-amdgcn.s.buffer.load.ll

AMDGPU/GlobalISel: Use integer types when narrowing loads and stores

The narrowScalar mutation for loads and stores produced untyped scalars.
For FP-typed values this resulted in an untyped G_OR in the lowerLoad
expansion for unaligned private accesses, which failed to select.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+441-448llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-private.mir
+268-270llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-flat.mir
+361-0llvm/test/CodeGen/AMDGPU/GlobalISel/load-store-private-unaligned-f64.ll
+16-16llvm/test/CodeGen/AMDGPU/GlobalISel/regbankselect-amdgcn.s.buffer.load.ll
+16-16llvm/test/CodeGen/AMDGPU/GlobalISel/atomicrmw-fmin-fmax.ll
+12-12llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-local.mir
+1,114-7621 files not shown
+1,117-7657 files

LLVM/project 33f058d — llvm/lib/Target/AMDGPU SIInstrInfo.h SIInstrInfo.cpp, llvm/test/CodeGen/AMDGPU fix-sgpr-copies-f16-true16.mir fix-sgpr-copies-vgpr16-to-spgr32.ll

[AMDGPU] Fix true16 losing 16-bit subreg operands

Folding a true16 v2s copy such as %2:sreg_32 = COPY %1.lo16 rewrites
its users to read %1.lo16 and relies on legalizeOperandsVALUt16 to
legalize the narrower operand. PHI and REG_SEQUENCE operands have no
register class, so it skipped them, leaving a 16-bit input in a VGPR_32
PHI or a 32-bit REG_SEQUENCE slot. DetectDeadLanes then marked the PHI
input undef and the defining load was deleted.

Widen such operands with a REG_SEQUENCE in legalizeOperandsVALUt16,
which runs both when the copy is folded and when the user is moved to
the VALU. This miscompiled uniform i16 loads feeding PHIs on gfx1250.

Change-Id: I9ddee5b11ff8503f4b2b1d1c4c776a12b270d968
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
DeltaFile
+101-212llvm/test/CodeGen/AMDGPU/load-constant-i1.ll
+282-0llvm/test/CodeGen/AMDGPU/fix-sgpr-copies-sgpr32-to-vgpr16.ll
+125-0llvm/test/CodeGen/AMDGPU/fix-sgpr-copies-vgpr16-to-spgr32.ll
+68-13llvm/lib/Target/AMDGPU/SIInstrInfo.cpp
+75-0llvm/test/CodeGen/AMDGPU/fix-sgpr-copies-f16-true16.mir
+16-0llvm/lib/Target/AMDGPU/SIInstrInfo.h
+667-2256 files

LLVM/project 3ad4809 — flang/unittests/Evaluate uint128.cpp

Fix typo
DeltaFile
+1-1flang/unittests/Evaluate/uint128.cpp
+1-11 files

LLVM/project db34882 — llvm/lib/Target/RISCV RISCVISelLowering.cpp, llvm/test/CodeGen/RISCV fclass-select.ll

 [RISCV] Fold (sub 0, (srl (and X, (1 << ShAmt)), ShAmt)) -> (sra (shl X, ShAmt2), bits-1) (#229509)

Improves codegen of is.fpclass+select. The AND to test bits of
is.fpclass may get turned into and+srl while the select emits a
neg. Isel will turn the and+srl into shl+srl but it's too late to
fold the neg.

Assisted-by: Claude
DeltaFile
+114-0llvm/test/CodeGen/RISCV/fclass-select.ll
+23-0llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+137-02 files

LLVM/project 6c7c592 — llvm/lib/Target/AMDGPU AMDGPULegalizerInfo.cpp, llvm/test/CodeGen/AMDGPU/GlobalISel regbankselect-amdgcn.s.buffer.load.ll atomicrmw-fmin-fmax.ll

AMDGPU/GlobalISel: Use integer types when narrowing loads and stores

The narrowScalar mutation for loads and stores produced untyped scalars.
For FP-typed values this resulted in an untyped G_OR in the lowerLoad
expansion for unaligned private accesses, which failed to select.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+441-448llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-private.mir
+268-270llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-flat.mir
+361-0llvm/test/CodeGen/AMDGPU/GlobalISel/load-store-private-unaligned-f64.ll
+16-16llvm/test/CodeGen/AMDGPU/GlobalISel/regbankselect-amdgcn.s.buffer.load.ll
+16-16llvm/test/CodeGen/AMDGPU/GlobalISel/atomicrmw-fmin-fmax.ll
+3-3llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
+1,105-7536 files

LLVM/project 1dc4fa3 — llvm/lib/CodeGen/GlobalISel LegalizerHelper.cpp, llvm/test/CodeGen/AMDGPU/GlobalISel legalize-load-constant.mir legalize-load-flat.mir

GlobalISel: Use integer types when splitting loads in lowerLoad

lowerLoad built the split pieces using the destination type with the
element size changed, which preserved floating-point types. An
unaligned f64 load was decomposed into G_ZEXTLOAD, G_SHL and G_OR on
f32 and f64, which then crashed in AMDGPU RegBankLegalize. Build the
pieces as integers and bitcast to a non-integer result type, as
lowerStore already does.

The pieces are now consistently integer typed, which allows more
constants to be CSEd in the existing tests.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+3,147-2,589llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-local.mir
+2,662-2,708llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-private.mir
+2,490-2,634llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-global.mir
+1,880-1,952llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-flat.mir
+1,812-1,908llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-constant.mir
+13-12llvm/lib/CodeGen/GlobalISel/LegalizerHelper.cpp
+12,004-11,8036 files

LLVM/project b82f26b — llvm/lib/Transforms/Vectorize LoopVectorize.cpp

[LV] Use frozen start values from the main plan directly (NFC). (#229571)

Only freeze possibly-poison start values of FindIV reductions in the
main plan. When preparing the epilogue plan, set the start operand of
its FindIV reduction results directly to the value frozen for the main
plan, instead of creating Freeze recipes in the epilogue plan and
replacing them later via a map. This leaves only VPExpandSCEVRecipes to
replace in the epilogue plan's entry.
DeltaFile
+34-51llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
+34-511 files

LLVM/project cfd56f3 — llvm/test/CodeGen/AMDGPU amdgcn.bitcast.768bit.ll amdgcn.bitcast.832bit.ll

address review feedback

Created using spr 1.3.7
DeltaFile
+85,305-104,728llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+23,897-29,739llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+16,898-20,565llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.960bit.ll
+15,617-19,097llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.896bit.ll
+14,370-17,720llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.832bit.ll
+12,180-15,528llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.768bit.ll
+168,267-207,37711,737 files not shown
+776,240-588,41011,743 files

LLVM/project 9184c30 — clang/lib/CIR/Dialect/IR CIRDialect.cpp, clang/test/CIR/Transforms canonicalize.cir

[CIR] Fix use-after-free in BrOp::canonicalize (#229191)

BrOp::canonicalize took the branch's destination operands as an
OperandRange, erased the branch, and then passed the range to
mergeBlocks. The range points into the branch's operand storage, which
eraseOp frees, so mergeBlocks read freed memory whenever the merged
block had arguments. Valgrind reports it on the new test:

```
  Invalid read of size 8
     at mlir::ValueRange::dereference_iterator
     by mlir::RewriterBase::inlineBlockBefore
     by cir::BrOp::canonicalize
   Address ... is 136 bytes inside a block of size 144 free'd
     by mlir::RewriterBase::eraseOp
     by cir::BrOp::canonicalize
```

It goes unnoticed with glibc malloc, which usually leaves the freed
operands intact. Copy the operands before erasing the branch.
DeltaFile
+34-0clang/test/CIR/Transforms/canonicalize.cir
+11-1clang/lib/CIR/Dialect/IR/CIRDialect.cpp
+45-12 files

LLVM/project 44a7e32 — clang/test/CodeGen fake-use-sanitizer.cpp

[msan][test] Add MSan case to clang/test/CodeGen/fake-use-sanitizer.cpp (#229564)

Regression test for https://github.com/llvm/llvm-project/issues/225425,
showing that MSan incorrectly strictly handles fake_use.
DeltaFile
+46-0clang/test/CodeGen/fake-use-sanitizer.cpp
+46-01 files

LLVM/project 5ab70ae — clang/lib/CIR/Dialect/Transforms LoweringPrepare.cpp, clang/test/CIR/Transforms lib-opt.cir target-lowering.cir

[CIR] Register TargetLowering, LoweringPrepare and LibOpt in cir-opt

DeltaFile
+56-0clang/test/CIR/Transforms/lowering-prepare.cir
+36-0clang/test/CIR/Transforms/lowering-prepare-global-ctor.cir
+20-0clang/test/CIR/Transforms/target-lowering.cir
+14-0clang/test/CIR/Transforms/lib-opt.cir
+12-0clang/tools/cir-opt/cir-opt.cpp
+5-2clang/lib/CIR/Dialect/Transforms/LoweringPrepare.cpp
+143-21 files not shown
+148-27 files

LLVM/project 7628e26 — clang/unittests/CIR ControlFlowTest.cpp

[CIR] Update ControlFlowTest for the new global ctor syntax
DeltaFile
+1-1clang/unittests/CIR/ControlFlowTest.cpp
+1-11 files

LLVM/project 155462f — mlir/include/mlir/Conversion Passes.td, mlir/lib/Conversion/AMDGPUToROCDL AMDGPUToROCDL.cpp

[mlir] Migrate AMDGPU/ROCDL to targets, not chipset versions (#223563)

**migration tl;dr:** Replace usages of `amdgpu::Chipset` with
`ROCDL::TargetInfo`, ideally move from `chipset=` to `arch=`. If you
don't use upstream pipelines, call
'TargetInfo::migrateArchFeaturesToModuleFlags` at the appropriate
location.

Further note: if you've got a build pipeline that's getting a `gfxXXX`
name from something like `rocm_agent_enumerator`, using a full triple
name like the ones you get from `rocminfo` is preferred.

`amdgpu::Chipset` was an awkward hack that was hard to keep up to date
with changes in the compiler/new architectures, and didn't properly
support generic targets (and has been strongly disfavored by the
compiler team).

This PR replaces `amdgpu::Chipset` with `ROCDL::TargetInfo`, a
structure that uses LLVM's TargetParser and the underlying LLVM

    [47 lines not shown]
DeltaFile
+316-324mlir/lib/Conversion/AMDGPUToROCDL/AMDGPUToROCDL.cpp
+105-0mlir/test/Dialect/LLVMIR/rocdl-attach-target-arch.mlir
+87-4mlir/lib/Dialect/GPU/Transforms/ROCDLAttachTarget.cpp
+46-40mlir/lib/Dialect/AMDGPU/Transforms/EmulateAtomics.cpp
+49-24mlir/test/Dialect/AMDGPU/amdgpu-emulate-atomics.mlir
+46-15mlir/include/mlir/Conversion/Passes.td
+649-407103 files not shown
+1,205-727109 files

LLVM/project b6de87e — utils/bazel/llvm-project-overlay/clang BUILD.bazel

[Bazel] Add FlowSensitive/Models headers to clang:analysis (#229578)

Fixes bazel build failure introduced in commits e03123a63162 and
f0d6540d5307, where GtestModelHelpers.h was added to
lib/Analysis/FlowSensitive/Models and included in Models/*.cpp.
DeltaFile
+2-0utils/bazel/llvm-project-overlay/clang/BUILD.bazel
+2-01 files

LLVM/project ef9f3c1 — llvm/lib/MC GOFFObjectWriter.cpp, llvm/test/CodeGen/SystemZ zos-symbol-2.ll zos-common-global.ll

Take a different approach

HLASM derives the AMODE from the RMODE if no AMODE is explicitly
given. The same can be done in the GOFF writer, simplifying the
coding a lot.
DeltaFile
+19-0llvm/lib/MC/GOFFObjectWriter.cpp
+4-4llvm/test/CodeGen/SystemZ/zos-section-2.ll
+2-2llvm/test/CodeGen/SystemZ/zos-section-1.ll
+1-1llvm/test/CodeGen/SystemZ/zos-symbol-2.ll
+1-1llvm/test/CodeGen/SystemZ/zos-common-global.ll
+27-85 files

LLVM/project 6c0a5e1 — llvm/include/llvm/CodeGen SlotIndexes.h, llvm/lib/CodeGen RegAllocGreedy.cpp LiveIntervals.cpp

[CodeGen] Drop dead SlotIndexes before allocation

SlotIndexes keeps the index list entry of an erased instruction and only
clears its instruction pointer. Live range sizes are measured in slot
indexes and greedy ranks ranges by size, so the leftovers inflate some
ranges more than others and reorder allocation, spilling heavily on
register-starved functions.

Add SlotIndexes::compactIndexes() to erase them. Erased entries are
unlinked, so LiveIntervals first reports the indexes it holds via
appendReferencedIndexes().

Off by default behind -greedy-compact-slot-indexes, since it changes
allocation across much of the test suite.
DeltaFile
+221-0llvm/unittests/CodeGen/SlotIndexesTest.cpp
+46-0llvm/lib/CodeGen/SlotIndexes.cpp
+32-0llvm/test/CodeGen/AMDGPU/greedy-compact-slot-indexes.ll
+23-6llvm/include/llvm/CodeGen/SlotIndexes.h
+29-0llvm/lib/CodeGen/LiveIntervals.cpp
+12-1llvm/lib/CodeGen/RegAllocGreedy.cpp
+363-71 files not shown
+367-77 files