LLVM/project a408737 — llvm/test/tools/llubi controlflow.ll, llvm/tools/llubi/lib Interpreter.cpp

[llubi] Reset retval for noop inline asm (#230849)

When a call is followed by a call to noop inline asm, `setResult` inside
`returnFromCallee` will reuse the previous return value (moved) and
trigger assertions.

The test is generated by DeepSeek-V4.1-Flash.
DeltaFile
+3-0llvm/test/tools/llubi/controlflow.ll
+1-0llvm/tools/llubi/lib/Interpreter.cpp
+4-02 files

LLVM/project d7bb842 — llvm/test/tools/llubi non_byte_size_ptr.ll, llvm/tools/llubi/lib Context.cpp

[llubi] Use correct tag bitwidth to recover provenances (#230852)

Tag always uses the pointer width rather than the padded one. Previously
the tag lookup always missed due to the width mismatch.

The test is generated by DeepSeek-V4.1-Flash.
DeltaFile
+24-0llvm/test/tools/llubi/non_byte_size_ptr.ll
+1-1llvm/tools/llubi/lib/Context.cpp
+25-12 files

LLVM/project 86ec2f8 — llvm/test/tools/llubi bitinsert_bitextract_le.ll bitinsert_bitextract_be.ll, llvm/tools/llubi/lib Context.h Context.cpp

[llubi] Add support for `bitinsert` and `bitextract`
DeltaFile
+78-0llvm/test/tools/llubi/bitinsert_bitextract_le.ll
+78-0llvm/test/tools/llubi/bitinsert_bitextract_be.ll
+29-0llvm/tools/llubi/lib/Interpreter.cpp
+19-0llvm/tools/llubi/lib/Context.cpp
+7-0llvm/tools/llubi/lib/Context.h
+211-05 files

LLVM/project 762c28a — llvm/test/tools/llubi bitcast_le.ll bitcast_be.ll, llvm/tools/llubi/lib Context.cpp

[llubi] Fix wrong assert when writing `poison` at an unaligned bit offset (#230819)

Writing `poison` that starts at a bit offset not a multiple of 8 and
spans more than one byte could trigger an assert, even though each write
stayed within a single byte.

- `Context::toBytes` marks bits as `poison` one byte at a time, from
`OffsetInBits + I`.
- The assert checked the bits from `OffsetInBits` instead, missing the
offset `I`.
DeltaFile
+5-4llvm/tools/llubi/lib/Context.cpp
+4-0llvm/test/tools/llubi/bitcast_le.ll
+4-0llvm/test/tools/llubi/bitcast_be.ll
+13-43 files

LLVM/project 964ed3b — libcxx/include/__condition_variable condition_variable.h, libcxx/test/std/thread/thread.condition/thread.condition.condvar wait_for_pred.pass.cpp wait_for.pass.cpp

[libc++][chrono][threading] Implement LWG 3504: `condition_variable::wait_for` is overspecified (#222443)

Implement the relative-to-absolute conversion mandated by LWG 3504 using
`ceil<steady_clock::duration>` to avoid precision loss with
floating-point
durations.

- Add internal `chrono::__ceil` (usable in all dialects) and make
`chrono::ceil`
  forward to it.
- Introduce `__rel_to_abs` helper and update all `wait_for` overloads on
  `condition_variable` / `condition_variable_any`.
- Move the previous nanosecond conversion into `__do_timed_wait` to 
  avoid recursion after the change.
- Add regression tests for floating-point durations.

Fixes #189807
DeltaFile
+29-25libcxx/include/__condition_variable/condition_variable.h
+24-0libcxx/test/std/thread/thread.condition/thread.condition.condvarany/wait_for.pass.cpp
+24-0libcxx/test/std/thread/thread.condition/thread.condition.condvar/wait_for.pass.cpp
+16-0libcxx/test/std/thread/thread.condition/thread.condition.condvarany/wait_for_token_pred.pass.cpp
+15-0libcxx/test/std/thread/thread.condition/thread.condition.condvarany/wait_for_pred.pass.cpp
+15-0libcxx/test/std/thread/thread.condition/thread.condition.condvar/wait_for_pred.pass.cpp
+123-253 files not shown
+136-339 files

LLVM/project 3c8fdcb — llvm/include/llvm/Analysis TargetTransformInfoImpl.h TargetTransformInfo.h, llvm/lib/Analysis TargetTransformInfo.cpp

[Analysis][ARM] Remove unused getNumBytesToPadGlobalArray (NFC) (#230912)

The last caller of TargetTransformInfo::getNumBytesToPadGlobalArray was
removed on June 30, 2025 in commit
183acdd27985afd332463e3d9fd4a2ca46d85cf1, leaving
TargetTransformInfoImplBase::getNumBytesToPadGlobalArray,
ARMTTIImpl::getNumBytesToPadGlobalArray, and the command-line option
UseWidenGlobalArrays unused as well.

Assisted-by: Antigravity
DeltaFile
+0-33llvm/lib/Target/ARM/ARMTargetTransformInfo.cpp
+0-6llvm/lib/Analysis/TargetTransformInfo.cpp
+0-5llvm/include/llvm/Analysis/TargetTransformInfoImpl.h
+0-5llvm/include/llvm/Analysis/TargetTransformInfo.h
+0-3llvm/lib/Target/ARM/ARMTargetTransformInfo.h
+0-525 files

LLVM/project ee9bb53 — llvm/lib/CodeGen TwoAddressInstructionPass.cpp, llvm/test/CodeGen/AMDGPU twoaddr-insert-subreg-undef.mir

TwoAddressInstructions: Keep undef INSERT_SUBREG lanes defined

%reg = INSERT_SUBREG undef %reg, %subreg, subidx defines all of %reg, so
reading the lanes outside subidx afterwards is valid. Rewriting it to
undef %reg.subidx = COPY %subreg narrows the definition to subidx and
leaves those reads without a live subrange, failing the "No live subrange
at use"
machine verifier check.

When a subrange outside subidx is still live past the def, insert an
IMPLICIT_DEF of the full register and drop the undef flag. In the common
case, where the lanes the INSERT_SUBREG left undefined are dead, keep the
undef flag and avoid an IMPLICIT_DEF that survives to the end of codegen
when the COPY is not coalesced.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
DeltaFile
+55-0llvm/test/CodeGen/AMDGPU/twoaddr-insert-subreg-undef.mir
+28-7llvm/lib/CodeGen/TwoAddressInstructionPass.cpp
+83-72 files

LLVM/project baf8368 — llvm/lib/CodeGen MachineBasicBlock.cpp, llvm/test/CodeGen/PowerPC common-chain.ll

CodeGen: Remove stale last-block case when splitting a critical edge

Before f264f9ad7df5, a new block at the end of the function began at the
old end index, so the isLastMBB case had to extend live-out intervals
over it. Now the new block always ends at the old end index, so they
already cover it. The extension became a no-op, and the case also
skipped trimming registers not live into the successor, leaving a dead
segment.

block-placement.ll needs a volatile store to keep its nested loop.
common-chain.ll gets shrink-wrapped.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+89-76llvm/test/CodeGen/PowerPC/common-chain.ll
+7-29llvm/lib/CodeGen/MachineBasicBlock.cpp
+36-0llvm/test/CodeGen/X86/phi-elimination-split-critical-edge-last-block.mir
+1-0llvm/test/CodeGen/X86/block-placement.ll
+133-1054 files

LLVM/project ceae108 — llvm/lib/Target/AArch64 AArch64InstrFormats.td, llvm/test/CodeGen/AArch64 shift-mod.ll

[GlobalISel] [AArch64] Support ROTR shift amount masking in SelectShiftMask  (#226579)

The RORV instructions already use the Shift multiclass, so they inherit
the shiftMask32/shiftMask64 ComplexPattern patterns added in the earlier
patches, giving them AND/ADD/SUB/NEG/MVN shift-amount optimization for
free.

This patch adds one additional pattern for the 32-bit rotate case, where
GlobalISel legalizes the shift amount with a zext to i64. The pattern
peeks through the zext to apply the 32-bit shift-amount masking:

  (i32 (rotr GPR32, (i64 (zext (i32 shiftMask32)))))

This completes the set of shift-amount optimizations (AND mask, ADD/SUB
mod, NEG, MVN, ROTR) needed to eventually replace tryShiftAmountMod, as
tracked in #224245.

Depends on https://github.com/llvm/llvm-project/pull/226273
DeltaFile
+23-0llvm/test/CodeGen/AArch64/shift-mod.ll
+5-0llvm/lib/Target/AArch64/AArch64InstrFormats.td
+28-02 files

LLVM/project bec0d06 — llvm/test/Transforms/SLPVectorizer/AMDGPU ordered-reduction-coalesced-loads.ll fmul-extract-fadd-chain.ll

update tests
DeltaFile
+0-88llvm/test/Transforms/SLPVectorizer/AMDGPU/fmul-extract-fadd-chain.ll
+16-64llvm/test/Transforms/SLPVectorizer/AMDGPU/ordered-reduction-coalesced-loads.ll
+16-1522 files

LLVM/project 2969eb4 — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Transforms/SLPVectorizer/AMDGPU alt-fmul-fadd-cost.ll fma-operand-contract-selection.ll

[AMDGPU] Limit the fmul fusion discount to a matching context type

The fmul is free when its context instruction feeds a fusable fadd or
fsub. A vector fmul priced with a scalar lane as the context inherits
that fusion only when it is emitted lane by lane. On a packed type the
scalar user does not show that the vector fmul feeds a vector fadd, and
products extracted into a scalar fadd chain keep the packed fmul while
the fma is lost.
DeltaFile
+88-0llvm/test/Transforms/SLPVectorizer/AMDGPU/fmul-extract-fadd-chain.ll
+64-16llvm/test/Transforms/SLPVectorizer/AMDGPU/ordered-reduction-coalesced-loads.ll
+38-37llvm/test/Transforms/SLPVectorizer/AMDGPU/ordered-reduction-fma-fusion.ll
+36-6llvm/test/Transforms/SLPVectorizer/AMDGPU/fma-operand-contract-selection.ll
+7-11llvm/test/Transforms/SLPVectorizer/AMDGPU/alt-fmul-fadd-cost.ll
+5-2llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+238-721 files not shown
+239-737 files

LLVM/project c510b63 — llvm/include/llvm/CodeGen TargetLoweringObjectFileImpl.h, llvm/include/llvm/Target TargetLoweringObjectFile.h

CodeGen: Merge TargetLoweringObjectFile::getModuleMetadata into initialize

getModuleMetadata had a single caller, which invoked it immediately after
Initialize. Pass the module to Initialize and fold it in. Also lowercase the
name while touching all the uses.

Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+36-26llvm/lib/CodeGen/TargetLoweringObjectFileImpl.cpp
+21-14llvm/include/llvm/CodeGen/TargetLoweringObjectFileImpl.h
+15-17llvm/lib/Target/RISCV/RISCVTargetObjectFile.cpp
+10-9llvm/include/llvm/Target/TargetLoweringObjectFile.h
+7-6llvm/lib/Target/TargetLoweringObjectFile.cpp
+3-8llvm/lib/Target/ARM/ARMTargetObjectFile.cpp
+92-8032 files not shown
+180-14238 files

LLVM/project cb23a36 — llvm/lib/CodeGen MachineModuleInfo.cpp, llvm/lib/CodeGen/AsmPrinter AsmPrinter.cpp

CodeGen: Initialize TargetLoweringObjectFile from MachineModuleInfo (#226836)

MachineModuleInfo passes TLOF to the MCContext but nothing initialized it until the 
AsmPrinter pass ran, so  every codegen pass in between saw it uninitialized. Initialize 
it from  MachineModuleInfo, and drop the calls llc and SPIRVTranslate used to work 
around this.

SPIRVTranslate's MachineModuleInfoWrapperPass was never passed on to
addPassesToEmitFile, so it was initializing a throwaway context.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+0-6llvm/lib/Target/SPIRV/SPIRVAPI.cpp
+4-0llvm/lib/CodeGen/MachineModuleInfo.cpp
+0-3llvm/tools/llc/lib/llcdriver.cpp
+0-3llvm/tools/llc/lib/NewPMDriver.cpp
+0-3llvm/lib/CodeGen/AsmPrinter/AsmPrinter.cpp
+4-155 files

LLVM/project 25a8bb5 — llvm/lib/Target/RISCV RISCVInstrInfoZvbdota.td, llvm/unittests/Target/RISCV RISCVInstrInfoTest.cpp

[RISCV] Fix FP properties for Zvbdota instructions (#228737)

Mark floating-point Zvbdota instructions as potentially raising FP
exceptions. Add an implicit FRM use to vfbdota.vv, which uses the
dynamic rounding mode.

AI Usage: Assisted by AmpCode (6.1 Sol)

Co-authored-by: Amp <amp at ampcode.com>
DeltaFile
+17-0llvm/unittests/Target/RISCV/RISCVInstrInfoTest.cpp
+4-3llvm/lib/Target/RISCV/RISCVInstrInfoZvbdota.td
+21-32 files

LLVM/project 7bb8e30 — clang/include/clang/Analysis/Analyses/LifetimeSafety Loans.h

[clang] Replace PointerUnion::dyn_cast with llvm::dyn_cast (NFC) (#230911)

PointerUnion::dyn_cast has been soft-deprecated in favor of
llvm::dyn_cast and llvm::dyn_cast_if_present.  This patch replaces the
former with llvm::dyn_cast where the operand is guaranteed to be
nonnull.

Note that PlaceholderBase::ParamOrMethod is always initialized to a
nonnull pointer in LoanManager::getOrCreatePlaceholderBase and never
modified afterward.

Assisted-by: Antigravity
DeltaFile
+2-2clang/include/clang/Analysis/Analyses/LifetimeSafety/Loans.h
+2-21 files

LLVM/project de41a35 — libcxx/docs/Status Cxx26Issues.csv, libcxx/include variant

[libc++] Implement LWG2991: variant copy constructor missing noexcept(see below) (#230241)

LWG2991 adds a conditional `noexcept` to the variant's copy constructor.
A small example to illustrate the issue:

```cpp
    static_assert(!std::is_trivially_copy_constructible_v<std::shared_ptr<int>>);
    static_assert(std::is_copy_constructible_v<std::shared_ptr<int>>);
    static_assert(std::is_nothrow_copy_constructible_v<std::shared_ptr<int>>);

    static_assert(!std::is_trivially_move_constructible_v<std::shared_ptr<int>>);
    static_assert(std::is_move_constructible_v<std::shared_ptr<int>>);
    static_assert(std::is_nothrow_move_constructible_v<std::shared_ptr<int>>);

    using Variant = std::variant<int,std::shared_ptr<int>>;
    static_assert(std::is_nothrow_copy_constructible_v<Variant>); // false, should be true
    static_assert(std::is_nothrow_move_constructible_v<Variant>);
```


    [9 lines not shown]
DeltaFile
+16-1libcxx/test/std/utilities/variant/variant.variant/variant.ctor/copy.pass.cpp
+5-4libcxx/include/variant
+1-1libcxx/docs/Status/Cxx26Issues.csv
+22-63 files

LLVM/project 118b80a — llvm/test/tools/llubi dereferenceable_unknown_func_decl.ll, llvm/tools/llubi/lib Interpreter.cpp

[llubi] Check UB before handling ret attributes (#230842)

Some code paths return AnyValue() on UB (e.g., calling function
declarations not recognized as libfunc). Check this before handling ret
attributes to avoid operating on incompatible values.

The test is generated by DeepSeek-V4.1-Flash.
DeltaFile
+11-0llvm/test/tools/llubi/dereferenceable_unknown_func_decl.ll
+2-0llvm/tools/llubi/lib/Interpreter.cpp
+13-02 files

LLVM/project 7d19366 — llvm/test/tools/llubi wide_ptr_oob.ll wide_ptr_free.ll, llvm/tools/llubi/lib Library.cpp ExecutorBase.cpp

[llubi] Correctly render large addresses (#230845)

The test is generated by DeepSeek-V4.1-Flash.
DeltaFile
+20-0llvm/test/tools/llubi/wide_ptr_free.ll
+12-7llvm/tools/llubi/lib/ExecutorBase.cpp
+15-0llvm/test/tools/llubi/wide_ptr_oob.ll
+6-6llvm/tools/llubi/lib/Library.cpp
+53-134 files

LLVM/project 0c4b521 — llvm/test/CodeGen/AMDGPU flat-saddr-load.ll anti-hints-multi-rule.gfx1250.mir

[AMDGPU] Insert wmma coexec aware anti-hints rules (#226397)

This patch ports the downstream coexec anti-hint insertion to upstream
as rule-based insertion in the AMDGPU pre-ra anti-hints framework. It
adds gfx1250x anti-hint rules for wmma coexec hazards such as A/B source
war, swmmac index war, dest waw), trans source war, memory-address war
etc. Each rule builds anti-hints relationship for the register allocator
to keep hazardous registers in different physical registers so the
hazard recognizer and waitcnt inserter need less nop and waits.

Co-authored-by:  Jeffrey Byrnes <jrbyrnes1989 at gmail.com>

Depends on #218075
DeltaFile
+2,527-0llvm/test/CodeGen/AMDGPU/anti-hints-wmma-waw.gfx1250.mir
+1,820-0llvm/test/CodeGen/AMDGPU/anti-hints-wmma-war.gfx1250.mir
+389-430llvm/test/CodeGen/AMDGPU/load-constant-i1.ll
+566-0llvm/test/CodeGen/AMDGPU/anti-hints-addr-xcnt.gfx1250.mir
+401-0llvm/test/CodeGen/AMDGPU/anti-hints-multi-rule.gfx1250.mir
+194-176llvm/test/CodeGen/AMDGPU/flat-saddr-load.ll
+5,897-60654 files not shown
+8,649-2,36060 files

LLVM/project 0f665a6 — llvm/lib/Target/AMDGPU GCNSchedStrategy.cpp, llvm/test/CodeGen/AMDGPU schedule-pressure-implicit-physreg-use.mir early-if-convert-cost.ll

AMDGPU: Don't use PressureDiffs with implicit allocatable physregs

canUsePressureDiffs skipped all implicit operands, so instructions
reading an allocatable physreg implicitly (e.g., vcc on v_cndmask_b32)
used the cached PressureDiff. PressureDiffs assume a single use for
physregs (see the FIXME in ScheduleDAGMILive::updatePressureDiffs), so
the pressure was added for each user. With several users in a region, the
SGPR pressure was overestimated and failed the expensive checks pressure
validation.

Fall back to the RegPressureTracker for implicit operands of
allocatable physregs, which are the ones it tracks.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+52-0llvm/test/CodeGen/AMDGPU/early-if-convert-cost.ll
+50-0llvm/test/CodeGen/AMDGPU/schedule-pressure-implicit-physreg-use.mir
+12-6llvm/lib/Target/AMDGPU/GCNSchedStrategy.cpp
+114-63 files

LLVM/project 5224660 — llvm/lib/Analysis ConstantFolding.cpp, llvm/test/Transforms/InstCombine/ARM crc32-const-fold.ll

[ARM] Implement CRC32 const folding (#230751)

This implements constant folding for __crc32 intrinsics in 32-bit ARM
targets when both the arguments are compile time known constants.

This is a follow up of https://github.com/llvm/llvm-project/pull/228985
which does the same for AArch64.
DeltaFile
+400-0llvm/test/Transforms/InstCombine/ARM/crc32-const-fold.ll
+12-0llvm/lib/Analysis/ConstantFolding.cpp
+412-02 files

LLVM/project a73ed0c — llvm/include/llvm/Support Allocator.h, llvm/unittests/Support AllocatorTest.cpp

[Support] Fix SpecificBumpPtrAllocator move assignment leak (#230480)

Move assignment releases the destination's storage without destroying
its objects, leaking resources they own.

Call `DestroyAll()` before replacing the destination's storage.

Testing on Linux with GCC 11.4.0, Release with assertions enabled:
- `AllocatorTest.*`: 15 passed. The new regression test fails with the
  original header and passes with the fix.
- Support unit tests via `llvm-lit`: 1842 passed, 13 skipped, no
failures.
- ASan/LSan reproducer: a 512-byte leak with the original header; no
  sanitizer report with the fix.
- Additional ASan/LSan checks for destruction, move construction, empty
and
  populated destinations, multiple slabs, and custom-sized slabs passed.

Assisted-by: Codex
DeltaFile
+22-0llvm/unittests/Support/AllocatorTest.cpp
+1-0llvm/include/llvm/Support/Allocator.h
+23-02 files

LLVM/project d05425e — llvm/include/llvm/ProfileData ProfCorrelatorKind.h InstrProfReader.h, llvm/lib/ProfileData InstrProfReader.cpp InstrProfCorrelator.cpp

[ProfileData] Move ProfCorrelatorKind out of InstrProfCorrelator (#230883)

For coverage and PGO, `clang -mllvm
-profile-correlate={debug-info,binary}` moves the per-function profile
metadata out of the memory image at runtime.

After the cl::opt to TableGen migration (#230746),
InstrumentationOptions.h includes InstrProfCorrelator.h only for this
enum, adding about 0.25s to each Instrumentation TU. Fix the compile
time regression by moving the type to its own header as a scoped enum.

LLM-aided
DeltaFile
+22-28llvm/tools/llvm-profdata/llvm-profdata.cpp
+13-12llvm/include/llvm/ProfileData/InstrProfReader.h
+8-15llvm/lib/ProfileData/InstrProfCorrelator.cpp
+20-0llvm/include/llvm/ProfileData/ProfCorrelatorKind.h
+9-10llvm/lib/Transforms/Instrumentation/InstrProfiling.cpp
+3-3llvm/lib/ProfileData/InstrProfReader.cpp
+75-684 files not shown
+80-7610 files

LLVM/project 67e37a9 — flang/lib/Lower/OpenMP ClauseProcessor.cpp, flang/test/Lower/OpenMP affinity-part-refs.f90 depend-iterator.f90

Simplify iterator lowering and trim duplicate tests

Use capture-only iterator callbacks and bind induction variables as block
arguments are created, eliminating the temporary induction-value vector.

Remove unused-iterator HLFIR cases covered by the per-locator and LLVM
checks. Reduce repeated stride diagnostics and remove unused declarations
from the positive AFFINITY tests.
DeltaFile
+0-35flang/test/Lower/OpenMP/task-affinity.f90
+13-20flang/lib/Lower/OpenMP/ClauseProcessor.cpp
+0-32flang/test/Lower/OpenMP/depend-iterator.f90
+2-28flang/test/Semantics/OpenMP/substring-strides.f90
+6-16flang/test/Lower/OpenMP/affinity-part-refs.f90
+0-9flang/test/Semantics/OpenMP/affinity-part-refs.f90
+21-1406 files

LLVM/project e2ce52c — llvm/test/CodeGen/AMDGPU anti-hints-trans-src0.gfx1250.mir anti-hints-multi-hazard.gfx1250.mir

Addressed review, added tests detail
DeltaFile
+234-0llvm/test/CodeGen/AMDGPU/anti-hints-wmma-waw.gfx1250.mir
+209-0llvm/test/CodeGen/AMDGPU/anti-hints-wmma-war.gfx1250.mir
+69-0llvm/test/CodeGen/AMDGPU/anti-hints-addr-xcnt.gfx1250.mir
+55-5llvm/test/CodeGen/AMDGPU/anti-hints-multi-rule.gfx1250.mir
+58-0llvm/test/CodeGen/AMDGPU/anti-hints-multi-hazard.gfx1250.mir
+49-0llvm/test/CodeGen/AMDGPU/anti-hints-trans-src0.gfx1250.mir
+674-52 files not shown
+733-58 files

LLVM/project d229e35 — llvm/test/CodeGen/AMDGPU mul.ll amdgcn.bitcast.1024bit.ll

Merge remote-tracking branch 'upstream/main' into users/mssefat/anti-hints-pr5-amdgpu-pre-ra-coexec-anti-hint
DeltaFile
+32,088-15,710llvm/test/CodeGen/AMDGPU/frem.ll
+5,761-3,731llvm/test/CodeGen/AMDGPU/srem.ll
+6,118-2,908llvm/test/CodeGen/AMDGPU/clmul.ll
+1,502-5,317llvm/test/CodeGen/AMDGPU/fcanonicalize.ll
+2,824-2,824llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+1,327-1,552llvm/test/CodeGen/AMDGPU/mul.ll
+49,620-32,0421,443 files not shown
+113,961-70,5011,449 files

LLVM/project 1abc2f0 — llvm/lib/Target/AMDGPU SIISelLowering.cpp AMDGPUInstructionSelector.cpp, llvm/lib/Target/AMDGPU/Utils AMDGPUBaseInfo.h AMDGPUBaseInfo.cpp

[AMDGPU] Use 8-bit barrier member count on GFX13 (#230884)

GFX13 widens the barrier member count in M0 to 8 bits.
DeltaFile
+137-0llvm/test/CodeGen/AMDGPU/s-barrier-member-count.ll
+9-0llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.cpp
+4-2llvm/lib/Target/AMDGPU/SIISelLowering.cpp
+4-2llvm/lib/Target/AMDGPU/AMDGPUInstructionSelector.cpp
+3-0llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.h
+157-45 files

LLVM/project 1beef99 — clang/include/clang/CIR/Dialect/IR CIRDialect.h CIROps.td, clang/lib/CIR/Dialect/IR CMakeLists.txt CIRDialect.cpp

[CIR] Implement ViewLikeOpInterface for multiple Ops

We implement ViewLikeOpInterface for dyn_cast, ptr_stride, get_member,
get_element, get_runtime_member, base_class_addr, derived_class_addr,
and memonic. Although there is no user for CIR internally as
decouplePointer returns the base and offset instead of the base
directly, it is still useful for external project to do aliasing
analysis for MemRef.
DeltaFile
+22-8clang/include/clang/CIR/Dialect/IR/CIROps.td
+14-0clang/lib/CIR/Dialect/IR/CIRDialect.cpp
+1-0clang/lib/CIR/Dialect/IR/CMakeLists.txt
+1-0clang/include/clang/CIR/Dialect/IR/CIRDialect.h
+38-84 files

LLVM/project 86a2f54 — llvm/lib/Target/AMDGPU SIISelLowering.cpp AMDGPUInstructionSelector.cpp, llvm/lib/Target/AMDGPU/Utils AMDGPUBaseInfo.h AMDGPUBaseInfo.cpp

[AMDGPU] Use 8-bit barrier member count on GFX13

GFX13 widens the barrier member count in M0 to 8 bits for cluster
named barriers. s_barrier_init and s_barrier_signal_var lowering
masked it to 6 bits, truncating counts above 63.

Change-Id: I33571d2ed499ebe168571c6b4cd5d0e3340f8a0f
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
DeltaFile
+137-0llvm/test/CodeGen/AMDGPU/s-barrier-member-count.ll
+9-0llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.cpp
+4-2llvm/lib/Target/AMDGPU/SIISelLowering.cpp
+4-2llvm/lib/Target/AMDGPU/AMDGPUInstructionSelector.cpp
+3-0llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.h
+157-45 files

LLVM/project c23622f — flang/lib/Optimizer/CodeGen CodeGen.cpp, flang/lib/Optimizer/Transforms/CUDA CUFAddConstructor.cpp CUFSharedTypeInfo.cpp

[flang][cuda] share runtime type info between host and device under managed memory (#229213)

With -gpu=mem:managed, descriptors can live in managed memory, so a
descriptor built on the device may be read on the host. The type
descriptor address in its addendum then pointed at the device copy of
the type info, and host code dereferencing it crashed.

Make the host copy of the type info the single shared copy:

- Add the cuf-shared-type-info pass. It makes host type-info globals
  writable, places them in the __nv_type_info section, and drops
  acc.declare from type info on both host and device so that OpenACC
  declare constructors no longer copy it to the device. For each type
  descriptor used in the GPU module, it creates a managed pointer
  global <dt>Xhostaddr<tag>. The tag is a per-unit hash, which keeps
  the name unique when each unit has its own device module. The GPU
  module gets a cuf.shared_type_descs dictionary mapping each type
  descriptor to its pointer.
- CUFAddConstructor: add the cuda-managed-type-info option. It

    [12 lines not shown]
DeltaFile
+131-0flang/lib/Optimizer/Transforms/CUDA/CUFSharedTypeInfo.cpp
+105-0flang/test/Fir/CUDA/cuda-shared-type-info.mlir
+88-1flang/lib/Optimizer/Transforms/CUDA/CUFAddConstructor.cpp
+71-0flang/test/Fir/CUDA/cuda-shared-type-info-codegen.mlir
+61-0flang/test/Fir/CUDA/cuda-shared-type-info-rename.mlir
+58-0flang/lib/Optimizer/CodeGen/CodeGen.cpp
+514-14 files not shown
+590-110 files