LLVM/project 7e75115 — clang/test/CodeGen/AArch64 neon-intrinsics-constrained.c neon-scalar-x-indexed-elem-constrained.c, clang/test/CodeGen/AArch64/neon intrinsics-constrained.c fused-multiply-constrained.c

[CIR][AArch64] Handle constrained Neon FMA and sqrt (#218307)

Propagate expression FP options through AArch64 builtin emission and
attach the active FP environment to CIR `cir.fma` and `cir.sqrt`
operations. This allows strict FP operations to lower to constrained
LLVM intrinsics while preserving unconstrained behavior.

Co-locate classic Clang, CIR, and CIR-to-LLVM coverage under
command-line strict mode. For FP16 FMA and sqrt operations, also test a
pragma overriding a command-line maytrap setting.

Place this coverage in focused `-constrained.c` files alongside the
corresponding Neon test families. Move the matching classic constrained
checks from the legacy files and remove their superseded blocks, while
retaining unrelated legacy tests. Existing unconstrained coverage
remains in the regular Neon tests.

Track FMA operands through ABI conversions and lane selection, verify
the constrained results are returned, and check that lowering does not

    [4 lines not shown]
DeltaFile
+3-277clang/test/CodeGen/AArch64/v8.2a-neon-intrinsics-constrained.c
+155-0clang/test/CodeGen/AArch64/neon/fused-multiple-fullfp16-constrained.c
+91-0clang/test/CodeGen/AArch64/neon/fused-multiply-constrained.c
+3-64clang/test/CodeGen/AArch64/neon-scalar-x-indexed-elem-constrained.c
+56-0clang/test/CodeGen/AArch64/neon/intrinsics-constrained.c
+0-40clang/test/CodeGen/AArch64/neon-intrinsics-constrained.c
+308-3814 files not shown
+359-42210 files

LLVM/project c12e2d0 — flang/test/Transforms loop-versioning-unit-slices.fir, llvm/test/CodeGen/AMDGPU/GlobalISel legalize-load-constant.mir legalize-load-flat.mir

Merge branch 'main' into users/arsenm/runtime-libcalls/schema-isolated-variant
DeltaFile
+3,103-3,156llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-private.mir
+3,159-2,601llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-local.mir
+2,490-2,634llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-global.mir
+2,148-2,222llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-flat.mir
+1,812-1,908llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-constant.mir
+1,898-0flang/test/Transforms/loop-versioning-unit-slices.fir
+14,610-12,521565 files not shown
+25,199-16,275571 files

LLVM/project 6afc47d — llvm/include/llvm/CodeGen TargetInstrInfo.h, llvm/lib/CodeGen MachineScheduler.cpp

CodeGen: Remove TRI arguments from TargetInstrInfo hooks (#228164)

Continue with cleanups enabled by #158224.. TRI can now always directly
be referenced from TargetInstrInfo

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+24-29llvm/lib/Target/X86/X86InstrInfo.cpp
+16-21llvm/include/llvm/CodeGen/TargetInstrInfo.h
+16-18llvm/lib/Target/ARM/ARMBaseInstrInfo.cpp
+12-20llvm/lib/CodeGen/MachineScheduler.cpp
+12-12llvm/lib/Target/AMDGPU/AMDGPUTargetMachine.cpp
+9-13llvm/lib/Target/X86/X86InstrInfo.h
+89-11332 files not shown
+176-23238 files

LLVM/project fde08d4 — llvm/test/CodeGen/AMDGPU store-atomic-flat.ll load-atomic-flat.ll

AMDGPU: Remove promotion of 32 and 64-bit atomic load/store types

The atomic load and store patterns now cover all 32 and 64-bit register types,
so there is no need to coerce to integer for selection. Mark the  16-bit vector
types as legal for atomic load and store rather than expanding them.

This fixes failing on atomic load and store of v2bf16 and v4bf16 which were
never promoted and ended up expanded.

The f16 and bf16 scalar cases are still promoted since the 16-bit atomic
patterns only cover i16 and i32, though that also should be fixed.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+108-0llvm/test/CodeGen/AMDGPU/load-atomic-global.ll
+90-0llvm/test/CodeGen/AMDGPU/store-atomic-global.ll
+82-0llvm/test/CodeGen/AMDGPU/load-atomic-local.ll
+74-0llvm/test/CodeGen/AMDGPU/store-atomic-local.ll
+66-0llvm/test/CodeGen/AMDGPU/load-atomic-flat.ll
+48-0llvm/test/CodeGen/AMDGPU/store-atomic-flat.ll
+468-02 files not shown
+472-388 files

LLVM/project cc67eac — clang/include/clang/Analysis/Analyses/LifetimeSafety Loans.h

[clang] Replace PointerUnion::dyn_cast with llvm::dyn_cast (NFC) (#229682)

PointerUnion::dyn_cast has been soft-deprecated in favor of
llvm::dyn_cast and llvm::dyn_cast_if_present.  This patch replaces the
former with llvm::dyn_cast where the operand is guaranteed to be
nonnull.

Note that AccessPath::Base is always initialized to a nonnull pointer.

Assisted-by: Antigravity
DeltaFile
+4-4clang/include/clang/Analysis/Analyses/LifetimeSafety/Loans.h
+4-41 files

LLVM/project c2f09f7 — llvm/lib/Target/AArch64 AArch64PredicateAsCounterLoopRewrites.cpp, llvm/test/CodeGen/AArch64 predicate-as-counter-loop-rewrites.ll

[AArch64] Use multi-vector intrinsics for masked load/store users of predicate-as-counter

If the user of the original wide mask is a masked load or store
intrinsic (matching the element size of the predicate-as-counter),
rewrite it directly to a masked multi-vector load/store.

This avoids materializing the vector mask and is easier to handle here
than later (e.g. in SelectionDAG), since we do not need to match the
concatenation of all `pext` segments of the predicate-as-counter.

Assisted-by: Codex
DeltaFile
+87-137llvm/test/CodeGen/AArch64/predicate-as-counter-loop-rewrites.ll
+110-1llvm/lib/Target/AArch64/AArch64PredicateAsCounterLoopRewrites.cpp
+197-1382 files

LLVM/project bb6f784 — llvm/lib/Target/AArch64 AArch64PredicateAsCounterLoopRewrites.cpp

Use CreateIntrinsic
DeltaFile
+11-14llvm/lib/Target/AArch64/AArch64PredicateAsCounterLoopRewrites.cpp
+11-141 files

LLVM/project d9992af — llvm/lib/Target/AArch64 AArch64PredicateAsCounterLoopRewrites.cpp, llvm/test/CodeGen/AArch64 predicate-as-counter-loop-rewrites.ll

[AArch64] Avoid materializing full masks for extractelement users of predicate-as-counter

If the user of the original wide mask is an `extractelement` and the
index is known to be within the first segment of the
predicate-as-counter, replace it with `extractelement(pext(counter, 0))`.

This avoids materializing the vector mask and produces a form that can
be folded into a conditional branch when the predicate-as-counter is
produced by a `whilelo`.

Assisted-by: Codex
DeltaFile
+89-153llvm/test/CodeGen/AArch64/predicate-as-counter-loop-rewrites.ll
+39-0llvm/lib/Target/AArch64/AArch64PredicateAsCounterLoopRewrites.cpp
+128-1532 files

LLVM/project 8f9714d — llvm/lib/Target/AArch64 AArch64PredicateAsCounterLoopRewrites.cpp

Use CreateIntrinsic
DeltaFile
+5-7llvm/lib/Target/AArch64/AArch64PredicateAsCounterLoopRewrites.cpp
+5-71 files

LLVM/project e128a79 — llvm/lib/Target/AArch64 CMakeLists.txt AArch64.h, llvm/test/CodeGen/AArch64 predicate-as-counter-loop-rewrites.ll

[AArch64] Add predicate-as-counter loop rewrite pass (#220960)

This patch adds an AArch64 IR loop pass that rewrites wide loop-carried
`llvm.get.active.lane.mask` phis to predicate-as-counter `whilelo` phis
 (for SVE2.1 or streaming SME2 targets).

Users of the original mask are preserved by materializing vector
predicates with `aarch64.sve.pext`.

The element size and vector scale (VLx2 or VLx4) of the
predicate-as-counter is inferred from the mask load/store users within
the loop. These could be optimized to multi-vector loads/stores (though
that is not included in this patch).

For example, a loop like:

```
entry:
  %step = vscale x 64

    [31 lines not shown]
DeltaFile
+675-0llvm/test/CodeGen/AArch64/predicate-as-counter-loop-rewrites.ll
+430-0llvm/lib/Target/AArch64/AArch64PredicateAsCounterLoopRewrites.cpp
+11-0llvm/lib/Target/AArch64/AArch64TargetMachine.cpp
+2-0llvm/lib/Target/AArch64/AArch64.h
+1-0llvm/lib/Target/AArch64/CMakeLists.txt
+1,119-05 files

LLVM/project 389a69b — llvm/include/llvm/Analysis CallGraphSCCPass.h, llvm/lib/Analysis CallGraphSCCPass.cpp

[Analysis] Remove unused CallGraphSCC methods (NFC) (#229683)

The last callers of CallGraphSCC::ReplaceNode and
CallGraphSCC::DeleteNode were removed on July 12, 2024 in commit
58bc98cd3abd72226cdbaa05bd92af9598d491db.

Their removal leaves the private member variable Context unused, so this
patch also removes Context and simplifies the CallGraphSCC constructor.

Assisted-by: Antigravity
DeltaFile
+1-29llvm/lib/Analysis/CallGraphSCCPass.cpp
+1-10llvm/include/llvm/Analysis/CallGraphSCCPass.h
+2-392 files

LLVM/project 246f363 — llvm/lib/CodeGen IndirectBrExpandPass.cpp, llvm/test/CodeGen/X86 opt-pipeline.ll O0-pipeline.ll

Revert "[IndirectBrExpand] Preserve profile weights (#227784)"

This reverts commit 4f0c9477072021ef6e3ce31b1240975a90bc41a0.
DeltaFile
+0-148llvm/test/Transforms/IndirectBrExpand/pgo.ll
+13-114llvm/lib/CodeGen/IndirectBrExpandPass.cpp
+1-5llvm/test/CodeGen/X86/O0-pipeline.ll
+1-4llvm/test/CodeGen/X86/opt-pipeline.ll
+1-0llvm/utils/profcheck-xfail.txt
+16-2715 files

LLVM/project 16792e0 — llvm/test/Bindings/OCaml executionengine.ml

[OCaml] Remove unused open Llvm_target in test (NFC) (#229680)

Reported at:
https://github.com/llvm/llvm-project/pull/227714#issuecomment-6026833038
DeltaFile
+0-1llvm/test/Bindings/OCaml/executionengine.ml
+0-11 files

LLVM/project 1b1dcfd — llvm/lib/Target/Mips MipsInstrInfo.cpp, llvm/test/CodeGen/Mips branch-dead-at.mir

Mips: Mark the assembler temporary def dead in inserted branches

The branch instructions reserve $at for long branch expansion, which
only uses it as a scratch. Instruction selection already marks the def
dead; branches rebuilt by later passes did not, which blocked
BranchFolding from hoisting common code past them.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+108-0llvm/test/CodeGen/Mips/branch-dead-at.mir
+11-4llvm/lib/Target/Mips/MipsInstrInfo.cpp
+3-3llvm/test/CodeGen/Mips/llvm-ir/forbidden-slot-ir.ll
+1-1llvm/test/CodeGen/Mips/Fast-ISel/branch-dead-at.ll
+123-84 files

LLVM/project ecc201f — llvm/include/llvm/IR IRBuilder.h, llvm/lib/CodeGen TypePromotion.cpp

[IRBuilder] Accept Module instead of LLVMContext in ctor (#229430)

This adds new constructor overloads, which accept `Module &` instead of
`LLVMContext &` in the ctor, and replaces nearly all usages of the old
constructors. A few more complicated cases (C API, SandboxIR context and
one unit test) are left alone for now.

The intent is to remove the old LLVMContext constructors and make
availability of the Module required at time of construction (either
directly passed to the ctor or implied by the insertion point).

The motivation for this is to ensure that IRBuilder always has an
available DataLayout. IRBuilder used to be DL-agnostic, but at this
point we've accumulated a number of key places which require a data
layout, and get it from the insertion point instead. This is not great,
because the insertion point is not actually required to be set when
producing instructions (they just won't be inserted).

Additionally, not having a required DataLayout also introduces the

    [11 lines not shown]
DeltaFile
+175-172llvm/lib/Target/Hexagon/HexagonLoopIdiomRecognition.cpp
+11-14llvm/lib/Target/AMDGPU/AMDGPULowerBufferFatPointers.cpp
+9-12llvm/lib/Target/WebAssembly/WebAssemblyLowerEmscriptenEHSjLj.cpp
+11-7llvm/lib/CodeGen/TypePromotion.cpp
+18-0llvm/include/llvm/IR/IRBuilder.h
+8-8llvm/unittests/Transforms/Utils/IntegerDivisionTest.cpp
+232-21391 files not shown
+402-38897 files

LLVM/project f4b6e8c — llvm/lib/Target/SPIRV SPIRVEmitIntrinsics.cpp, llvm/test/CodeGen/SPIRV dominator-order.ll

[SPIRV] Re-sort basic blocks after LoopSimplify for non-shader targets (#229678)

SPIRVPrepareFunctions sorts basic blocks in reverse post-order so that
dominators appear before all blocks
they dominate, as required by SPIR-V. 
However, since #187519 (4a773b9f35fc), LoopSimplifyPass runs later in
addISelPrepare for non-shader
targets as well. For shader targets, SPIRVStructurizer subsequently
re-runs sortBlocks(F),
but for non-shader targets no pass restored the block order, allowing
dedicated exit or preheader blocks
inserted by LoopSimplify to appear before their dominators.

Call sortBlocks(Func) for non-shader targets at the start of
SPIRVEmitIntrinsicsImpl::runOnFunction.
DeltaFile
+28-0llvm/test/CodeGen/SPIRV/dominator-order.ll
+5-0llvm/lib/Target/SPIRV/SPIRVEmitIntrinsics.cpp
+33-02 files

LLVM/project 7716d9e — llvm/test/CodeGen/AMDGPU global_atomics.ll store-atomic-local.ll, llvm/test/CodeGen/AMDGPU/GlobalISel inst-select-load-atomic-global.mir

AMDGPU: Use normal load/store pattern type lists for atomics

Atomic load and store of vector types are now permitted in the IR. Instead of
maintaining separate scalar-only atomic pattern lists, cover atomic load/store
in the existing per-register-type pattern loops. As a side effect
-flat-for-global is respected in more cases.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+649-0llvm/test/CodeGen/AMDGPU/load-atomic-global.ll
+398-0llvm/test/CodeGen/AMDGPU/store-atomic-global.ll
+333-0llvm/test/CodeGen/AMDGPU/load-atomic-local.ll
+250-0llvm/test/CodeGen/AMDGPU/store-atomic-local.ll
+28-34llvm/test/CodeGen/AMDGPU/global_atomics.ll
+45-15llvm/test/CodeGen/AMDGPU/GlobalISel/inst-select-load-atomic-global.mir
+1,703-498 files not shown
+1,833-17614 files

LLVM/project f2e7e3f — clang-tools-extra/clang-tidy/performance InefficientVectorOperationCheck.h InefficientVectorOperationCheck.cpp, clang-tools-extra/docs ReleaseNotes.md

[clang-tidy] Make range source classes configurable (#226193)

The check currently hardcodes the container types accepted as sources in
range-based for loops. This prevents it from recognizing
project-specific range-like containers even when they expose the
operations needed by the check.

Add a `ForRangeLoopClasses` option, analogous to `VectorLikeClasses`,
for configuring those source container classes. The existing
standard-library types remain the defaults, so the current behavior is
preserved.

The test covers a custom range-like source and verifies the generated
`reserve(range.size())` fix.

Tests:
- `ninja check-clang-extra-clang-tidy-checkers-performance` (48/48
passed)


    [7 lines not shown]
DeltaFile
+52-1clang-tools-extra/test/clang-tidy/checkers/performance/inefficient-vector-operation-vectorlike-classes.cpp
+13-8clang-tools-extra/clang-tidy/performance/InefficientVectorOperationCheck.cpp
+12-3clang-tools-extra/docs/clang-tidy/checks/performance/inefficient-vector-operation.rst
+5-0clang-tools-extra/docs/ReleaseNotes.md
+1-0clang-tools-extra/clang-tidy/performance/InefficientVectorOperationCheck.h
+83-125 files

LLVM/project 2cc7a05 — llvm/lib/CodeGen MachineCombiner.cpp, llvm/test/CodeGen/AArch64 machine-combiner-dbg-value.mir

[MachineCombiner] Don't count debug instructions in block size (#225765)

The method MachineCombinerImpl::combineInstructions made decisions based
on
 MBB->size() > inc_threshold
in two places, and since MachineBasicBlock::size() includes DBG_VALUE,
the existence of debug info can affect the resulting code.

Use sizeWithoutDebugLargerThan() instead to get the same code with and
without debug info.

I originally found this problem for my out-of-tree target. Then I used
Claude Opus 5 to convert the mir test for my out-of-tree target to one
for AArch64.

The other place, in the register pressure reduction part, was found by
code inspection. It is guarded by MustReduceRegisterPressure, which only
PowerPC uses. The PowerPC test for it was created by Claude Opus 5.
DeltaFile
+81-0llvm/test/CodeGen/PowerPC/machine-combiner-dbg-value.mir
+66-0llvm/test/CodeGen/AArch64/machine-combiner-dbg-value.mir
+2-2llvm/lib/CodeGen/MachineCombiner.cpp
+149-23 files

LLVM/project df9fdb9 — llvm/test/TableGen atomic-store-trunc-patfrags.td, llvm/utils/TableGen GlobalISelEmitter.cpp

TableGen: Allow IsTruncStore predicates on atomic PatFrags

Atomic stores can be truncating in the same way as regular stores.
Allow IsTruncStore and IsNonTruncStore to be set on atomic store
PatFrags so patterns can check the store is not truncating without
needing to specify a fixed memory size.

I still find the hierarchy of load/store PatFrags frustrating. This would be
easier if we fixed the legacy mistake of treating atomic store as an "atomic"
rather than a store.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+40-0llvm/test/TableGen/atomic-store-trunc-patfrags.td
+29-8llvm/utils/TableGen/Common/CodeGenDAGPatterns.cpp
+1-1llvm/utils/TableGen/GlobalISelEmitter.cpp
+70-93 files

LLVM/project ee22ca5 — llvm/lib/Target/AMDGPU AMDGPULegalizerInfo.cpp, llvm/test/CodeGen/AMDGPU/GlobalISel llvm.amdgcn.raw.buffer.store.format.f16.ll llvm.amdgcn.struct.buffer.store.format.f16.ll

AMDGPU/GlobalISel: Fix buffer stores of bf16 (#229673)

The 16-bit store source fixup only handled i8, i16 and f16, so bf16
values reached register bank legalization unextended and failed. Bitcast
16-bit FP sources to i16 before any-extending, rather than any-extending
the FP value directly.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+99-115llvm/test/CodeGen/AMDGPU/GlobalISel/llvm.amdgcn.raw.tbuffer.store.f16.ll
+126-5llvm/test/CodeGen/AMDGPU/GlobalISel/llvm.amdgcn.raw.buffer.store.ll
+10-12llvm/test/CodeGen/AMDGPU/GlobalISel/llvm.amdgcn.struct.buffer.store.format.f16.ll
+5-9llvm/test/CodeGen/AMDGPU/GlobalISel/llvm.amdgcn.raw.buffer.store.format.f16.ll
+7-4llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
+247-1455 files

LLVM/project b7fedff — llvm/lib/CodeGen/GlobalISel CallLowering.cpp, llvm/test/CodeGen/Generic/GlobalISel irtranslator-byte-type.ll

[GlobalISel] Ensure integer trunc and anyext in call lowering. (#229346)

This ensures that the trunc and anyext created for argument lowering
have integer types, possibly casting back to fp types. This helps the
Arm backend but also comes up for WebAssembly and AMDGPU too.
DeltaFile
+22-2llvm/lib/CodeGen/GlobalISel/CallLowering.cpp
+8-4llvm/test/CodeGen/WebAssembly/GlobalISel/irtranslator/call-basics.ll
+2-1llvm/test/CodeGen/WebAssembly/GlobalISel/irtranslator/ret-basics.ll
+2-1llvm/test/CodeGen/WebAssembly/GlobalISel/irtranslator/args.ll
+2-1llvm/test/CodeGen/Generic/GlobalISel/irtranslator-byte-type.ll
+36-95 files

LLVM/project b0ba49d — clang/include/clang/Basic DiagnosticDriverKinds.td, clang/lib/CodeGen CodeGenAction.cpp

[clang] Honor -discard-value-names for LLVM IR input (#229674)

For IR input, CodeGenAction forces setDiscardValueNames(false) to avoid
the crash in #44241, so names created by the optimizer (e.g. the
inliner's `.i`) are kept. c53d807321f6 fixed that crash in
Value::setNameImpl. Apply CodeGenOpts.DiscardValueNames after loading
the module and remove the "ignoring -fdiscard-value-names" warning.

LLM-aided. `clang -O3 -c sqlite3.bc` executes 0.7% fewer instructions.
DeltaFile
+21-0clang/test/CodeGen/discard-name-values.ll
+0-16clang/test/CodeGen/PR44896.ll
+1-8clang/lib/Driver/ToolChains/Clang.cpp
+2-3clang/lib/CodeGen/CodeGenAction.cpp
+0-3clang/include/clang/Basic/DiagnosticDriverKinds.td
+24-305 files

LLVM/project 3c336ad — clang/lib/Driver/ToolChains Darwin.cpp, clang/test/Driver darwin-static-lib-universal.c sysroot.c

clang: Do not pass an empty -syslibroot for --sysroot= on Darwin

-isysroot and DEFAULT_SYSROOT, but it checked only for the presence
of the option. An empty --sysroot= then produced -syslibroot "",
where previously it fell through to -isysroot. An empty --sysroot= is
the usual way to defeat DEFAULT_SYSROOT, so treat it as absent.

Use this to fix the darwin-static-lib tests when clang is built with
DEFAULT_SYSROOT or CLANG_USE_XCSELECT.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+11-11clang/test/Driver/darwin-static-lib.c
+12-0clang/test/Driver/sysroot.c
+2-1clang/lib/Driver/ToolChains/Darwin.cpp
+1-1clang/test/Driver/darwin-static-lib-universal.c
+26-134 files

LLVM/project 7016730 — flang/lib/Optimizer/CodeGen CodeGen.cpp, flang/test/Fir declare-codegen-debug.fir

[flang] Take the debug record's location from the declaration. (#229484)

`DeclareOpConversion` gives the `DbgDeclareOp` it creates the location
of the memref (the alloca) instead of the location of the `XDeclareOp`.
This is problematic in 2 ways.

1. The `XDeclareOp` better represents where the variable is declared.
The memref is where the variable's memory happens to be allocated, and
for a dummy argument it is a block argument, so the record ends up on
the procedure statement. For a subroutine whose dummies `n` and `a` are
declared on lines 2 and 3, their `#dbg_declare` records are both at line
1 today, and at lines 2 and 3 with this change.

2. It requires the alloca to have a valid location. If it does not, the
`DbgDeclareOp` gets none either, and a debug record without a location
is invalid, so it is dropped during the MLIR to LLVM IR translation and
the variable disappears from the debug info altogether.

This PR makes `DeclareOpConversion` use the location of the

    [4 lines not shown]
DeltaFile
+31-0flang/test/Fir/declare-codegen-debug.fir
+3-3flang/lib/Optimizer/CodeGen/CodeGen.cpp
+34-32 files

LLVM/project 4a080b6 — llvm/test/CodeGen/AMDGPU amdgcn.bitcast.768bit.ll amdgcn.bitcast.832bit.ll

Merge branch 'main' into users/arsenm/codegen/tii-remove-tri-args
DeltaFile
+56,601-75,945llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+18,219-23,687llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+13,775-17,204llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.960bit.ll
+12,841-16,119llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.896bit.ll
+11,584-14,818llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.832bit.ll
+10,483-13,715llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.768bit.ll
+123,503-161,4884,660 files not shown
+355,466-352,5334,666 files

LLVM/project 35546a4 — llvm/test/CodeGen/AMDGPU/GlobalISel legalize-load-local.mir regbankselect-amdgcn.s.buffer.load.ll

AMDGPU/GlobalISel: Use integer types when narrowing loads and stores (#229547)

The narrowScalar mutation for loads and stores produced untyped scalars.
For FP-typed values this resulted in an untyped G_OR in the lowerLoad
expansion for unaligned private accesses, which failed to select.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+441-448llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-private.mir
+268-270llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-flat.mir
+361-0llvm/test/CodeGen/AMDGPU/GlobalISel/load-store-private-unaligned-f64.ll
+16-16llvm/test/CodeGen/AMDGPU/GlobalISel/regbankselect-amdgcn.s.buffer.load.ll
+16-16llvm/test/CodeGen/AMDGPU/GlobalISel/atomicrmw-fmin-fmax.ll
+12-12llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-local.mir
+1,114-7621 files not shown
+1,117-7657 files

LLVM/project 8e28512 — llvm/lib/CodeGen MachineLICM.cpp

MachineLICM: Stop checking kill flags in register pressure estimate (#229201)

This was checking kill flags, or hasOneNonDBGUse as a kill approximation. 
The kill flag case appears to be of no practical use. InstrEmitter does not
emit the kill flag in the multiple user cases that would be required and I've
only managed to trigger a different hoisting decision with hand modified MIR.

The flag disagrees with hasOneNonDBGUse in 40 CodeGen tests (mostly
custom-inserter loops such as AMDGPU waterfalls and atomic expansions),
but removing it does not change the output of any CodeGen test.
DeltaFile
+1-5llvm/lib/CodeGen/MachineLICM.cpp
+1-51 files

LLVM/project c709385 — llvm/test/CodeGen/AArch64 windows-trap-unreachable.ll

AArch64: Test trap-after-noreturn against the exception model

Unfortunately, the AArch64TargetMachine constructor modifies
TargetOptions::TrapUnreachable if MCAsmInfo::usesWindowsCFI(). Thus, by
default, a trap after a noreturn call is emitted on COFF targets with the
default exception mode. There was no test coverage for this case's
interaction with an overridden exception model. Add the missing test in
preparation for cleaning up both the MC side exception predicates and
TrapUnreachable mutation.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+24-0llvm/test/CodeGen/AArch64/windows-trap-unreachable.ll
+24-01 files

LLVM/project 596fe13 — offload/test/offloading target-no-loop.c, offload/test/offloading/fortran target-no-loop.f90

[offload][OpenMP][NFC] Add a default-grid run to the no-loop tests

Add a run without OMP_NUM_TEAMS and OMP_TEAMS_THREAD_LIMIT to show the
kernels are promoted with the runtime's default grid as well.
DeltaFile
+9-0offload/test/offloading/target-no-loop.c
+9-0offload/test/offloading/fortran/target-no-loop.f90
+18-02 files