LLVM/project 6c978c6llvm/test/Transforms/SLPVectorizer/AArch64 loop-accumulator-reduction.ll

[SLP][NFC]Add some more tests for loop accumulated reductions, NFC



Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/221597
DeltaFile
+1,019-8llvm/test/Transforms/SLPVectorizer/AArch64/loop-accumulator-reduction.ll
+1,019-81 files

LLVM/project 3da9899flang-rt/include/flang-rt/runtime file.h connection.h, flang/include/flang/Common optional.h

[flang-rt] Initialize I/O unit storage read by short-circuit predicates (#221126)

ConnectionState and OpenFile hold common::optional members read through
predicates of the form `opt && x < *opt`, which never use an indeterminate
value in the abstract machine. Compilers do, however, routinely if-convert
the short-circuit && into a branchless compare and select, which speculates
the payload load; because these objects are placement-new'd into malloc'd
storage by UnitMap::Create(), a memory checker then reports a conditional
branch that depends on uninitialized memory. On AArch64 this fires for
every Fortran program that writes a record, giving two reports in
ExternalFileUnit::AdvanceRecord() from IsAfterEndfile() and IsAtEOF(),
while x86-64 is unaffected and libgfortran is clean on the same program and
host. The reports are false positives -- the engaged flag is 0, so both arms
of the select are 0 -- but they are unavoidable noise for anyone running
Valgrind on Fortran code.

Add common::ResetWithDefinedPayload(), which leaves an optional disengaged 
while writing its payload storage, and call it at construction for the
optionals in ConnectionAttributes, ConnectionState and OpenFile. Also

    [15 lines not shown]
DeltaFile
+25-0flang/include/flang/Common/optional.h
+14-0flang-rt/include/flang-rt/runtime/connection.h
+9-1flang-rt/include/flang-rt/runtime/file.h
+48-13 files

LLVM/project bb3fe4aflang/include/flang/Evaluate intrinsics.h, flang/lib/Evaluate intrinsics.cpp

[flang] Reclassify MVBITS, SPLIT, and TOKENIZE as SIMPLE (#205024)

F2023 makes MVBITS a simple elemental subroutine and SPLIT/TOKENIZE
simple subroutines.

This change:
- adds `simpleSubroutine` and `simpleElementalSubroutine` to
`IntrinsicClass`,
- reclassifies the MVBITS, SPLIT, and TOKENIZE intrinsic table entries,
- propagates `SIMPLE` through intrinsic resolution and procedure
characteristics.

MOVE_ALLOC is not included in this change.
DeltaFile
+49-0flang/test/Semantics/simple-intrinsics.f90
+12-7flang/lib/Evaluate/intrinsics.cpp
+8-1flang/lib/Semantics/resolve-names.cpp
+2-1flang/include/flang/Evaluate/intrinsics.h
+71-94 files

LLVM/project ec37fb6llvm/lib/Transforms/Vectorize SLPVectorizer.cpp

[SLP][NFC]Fix formatting, NFC



Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/221596
DeltaFile
+23-23llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+23-231 files

LLVM/project 4af1cd6llvm/lib/CodeGen/SelectionDAG DAGCombiner.cpp, llvm/test/CodeGen/Hexagon sffms.ll

DAGCombiner: Drop AllowFPOpFusion from visitFSUBForFMACombine

Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
DeltaFile
+3-3llvm/test/CodeGen/Hexagon/sffms.ll
+3-3llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp
+6-62 files

LLVM/project 1e4858acompiler-rt/test/asan/TestCases complete_stack_trace.c

[asan][test] Check that stack traces are not truncated (#221518)

Add a test that the access, free and malloc traces each name the whole
call
chain down to `main`, in order, and that each names its own call site in
`main`. A trace that stops early -- as every trace but the access one
does when
the fast unwinder runs on a target that chains no frames -- fails it.

`deep_stack_uaf.cpp` already covers trace depth, but it looks for three
individual frames anywhere in the trace, so it passes on a trace with
holes in
it, and it needs C++ name demangling to do that.

Split out of #220231 at reviewer request.

Assisted-by: Claude Code
DeltaFile
+58-0compiler-rt/test/asan/TestCases/complete_stack_trace.c
+58-01 files

LLVM/project 42012a9llvm/include/llvm/Option Option.h OptTable.h, llvm/lib/Option Option.cpp OptTable.cpp

[OptTable] Store Info strings in the string table (#218845)

Change HelpText, MetaVar, AliasArgs, and Values from `const char *` to
StringTable::offset, making the fields smaller, and removing dynamic
relocations in .data.rel.ro in PIC links.

Store them as StringTable::Offset, like the option names already are,
and return StringRef from getOptionHelpText() and getOptionMetaVar().
The 53 tables in the tree lose all 616 KB of .data.rel.ro, and sizeof(Info)
drops from 88 to 60; clang's table becomes 232 KB of .rodata.

An unset field and one explicitly set to the empty string, such as a
HelpText<"">, have to stay distinguishable, so the latter gets an empty
string of its own rather than offset zero.

Values declared with ValuesCode are only known to the generated code,
which supplies getOptionValuesCode() for OptTable to call; only clang has any.

Aided by Opus 5
DeltaFile
+92-80llvm/utils/TableGen/OptionParserEmitter.cpp
+44-17llvm/include/llvm/Option/OptTable.h
+21-18llvm/lib/Option/OptTable.cpp
+31-2llvm/unittests/Option/OptionParsingTest.cpp
+5-10llvm/lib/Option/Option.cpp
+10-5llvm/include/llvm/Option/Option.h
+203-1323 files not shown
+214-1359 files

LLVM/project 5735d17llvm/include/llvm/MC MCTargetOptions.h, llvm/include/llvm/Target TargetOptions.h

MC: Move DisableIntegratedAS from TargetOptions to MCTargetOptions (#221547)

The integrated assembler is only meaningful in MC, so this field belongs
in MCTargetOptions alongside the other assembler options rather than in
the codegen-level TargetOptions.

Co-authored-by: Claude (Claude-Opus-4.8)
DeltaFile
+7-11llvm/include/llvm/Target/TargetOptions.h
+0-7llvm/lib/CodeGen/CommandFlags.cpp
+7-0llvm/lib/MC/MCTargetOptionsCommandFlags.cpp
+3-0llvm/include/llvm/MC/MCTargetOptions.h
+1-1llvm/lib/LTO/LTOCodeGenerator.cpp
+1-1llvm/lib/CodeGen/CodeGenTargetMachineImpl.cpp
+19-204 files not shown
+23-2410 files

LLVM/project 2796699llvm/test/Transforms/SLPVectorizer/X86 loop-accumulator-reduction.ll

[SLP][NFC]Add more tests for loop accumulator reductions, NFC



Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/221591
DeltaFile
+1,327-0llvm/test/Transforms/SLPVectorizer/X86/loop-accumulator-reduction.ll
+1,327-01 files

LLVM/project 5840db8llvm/include/llvm/CodeGen SelectionDAG.h, llvm/lib/CodeGen/SelectionDAG SelectionDAG.cpp

[SelectionDAG] Remove dead functions (NFC) (#221541)

SelectionDAG::getBitcastedSExtOrTrunc,
SelectionDAG::getBitcastedZExtOrTrunc: Added on August 11, 2023 in
commit d26a06728da84a7302875a99ea86e887f6bc425a without any callers.

Assisted-by: Antigravity
DeltaFile
+0-30llvm/lib/CodeGen/SelectionDAG/SelectionDAG.cpp
+0-10llvm/include/llvm/CodeGen/SelectionDAG.h
+0-402 files

LLVM/project f0e701aclang/test/OpenMP interchange_codegen.cpp, llvm/test/CodeGen/AMDGPU flat-saddr-atomics.ll flat-saddr-load.ll

Merge remote-tracking branch 'upstream' into users/lukel97/loop-vectorize/simplifyRecipes-worklist
DeltaFile
+17,282-3,458llvm/test/tools/llvm-mca/AArch64/Cortex/C1Premium-sve-instructions.s
+7,983-1,591llvm/test/tools/llvm-mca/AArch64/Cortex/C1Premium-neon-instructions.s
+2,115-2,484llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+3,312-825llvm/test/CodeGen/AMDGPU/flat-saddr-load.ll
+1,704-2,400clang/test/OpenMP/interchange_codegen.cpp
+2,226-1,164llvm/test/CodeGen/AMDGPU/flat-saddr-atomics.ll
+34,622-11,9224,795 files not shown
+241,334-126,6184,801 files

LLVM/project ebb7e0clibc/src/__support/GPU allocator.cpp

[libc] Fix chunk calculation in GPU allocator (#221583)

Summary:
This would pesismistically round up 48 to 64 and the previous s0 case
was unused.
DeltaFile
+7-10libc/src/__support/GPU/allocator.cpp
+7-101 files

LLVM/project 64171f0mlir/include/mlir/Dialect/Affine/IR AffineOps.h AffineOps.td, mlir/lib/Dialect/Affine/IR CMakeLists.txt MemorySlot.cpp

[mlir][affine] Implement PromotableRegionOpInterface for AffineForOp (#221123)

The `mem2reg` pass couldn't promote memory slots accessed within an
`affine.for` because `AffineForOp` did not implement
`PromotableRegionOpInterface`, leading to the stack allocation and its
accesses to not be eliminated.

This change implements `PromotableRegionOpInterface` for `AffineForOp`,
allowing `mem2reg` to promote memory slots through affine loops.
DeltaFile
+116-0mlir/test/Dialect/Affine/mem2reg.mlir
+56-0mlir/lib/Dialect/Affine/IR/MemorySlot.cpp
+3-1mlir/include/mlir/Dialect/Affine/IR/AffineOps.td
+2-0mlir/lib/Dialect/Affine/IR/CMakeLists.txt
+2-0mlir/include/mlir/Dialect/Affine/IR/AffineOps.h
+179-15 files

LLVM/project 7f04c2bllvm/include/llvm/ADT SetOperations.h, llvm/lib/Transforms/IPO MemProfContextDisambiguation.cpp

[ADT][MemProf] Optimize set_subtract with removed-set output and use in MemProf (#221372)

Add a 3-argument set_subtract(A, B, Removed) that computes A := A - B
and records elements of B removed from A (A ^ B) in Removed. When
A.size() < B.size(), B supports contains(), and A supports remove_if(),
we iterate over A via remove_if() instead of iterating over B, improving
efficiency.

Remove the legacy 4-argument set_subtract(A, B, Removed, Remaining),
which was only used by MemProfContextDisambiguation.cpp and is now
redundant since Remaining can be updated via a separate 2-argument
set_subtract.

Update MemProfContextDisambiguation to use the 3-argument set_subtract,
and update SetOperations unit tests.
DeltaFile
+71-16llvm/unittests/ADT/SetOperationsTest.cpp
+32-6llvm/include/llvm/ADT/SetOperations.h
+8-13llvm/lib/Transforms/IPO/MemProfContextDisambiguation.cpp
+111-353 files

LLVM/project de3dd2dllvm/test/CodeGen/AMDGPU madak.ll, llvm/test/CodeGen/LoongArch float-fma.ll double-fma.ll

Reapply "DAGCombiner: Drop AllowFPOpFusion from visitFADDForFMACombine" (#221561) (#221567)

This reverts commit 502e51aa4df687807fbe51fa0b419baf65e9615f.
DeltaFile
+352-1,116llvm/test/CodeGen/LoongArch/lasx/fma-v8f32.ll
+352-1,116llvm/test/CodeGen/LoongArch/lasx/fma-v4f64.ll
+260-872llvm/test/CodeGen/LoongArch/float-fma.ll
+260-872llvm/test/CodeGen/LoongArch/double-fma.ll
+388-471llvm/test/CodeGen/AMDGPU/madak.ll
+178-564llvm/test/CodeGen/LoongArch/lsx/fma-v4f32.ll
+1,790-5,01118 files not shown
+2,495-6,19224 files

LLVM/project 75861cfllvm/lib/Transforms/Utils SimplifyLibCalls.cpp, llvm/test/Transforms/InstCombine scalbn-to-ldexp.ll

[InstCombine] Fold scalbn libcalls to llvm.ldexp (#216573)

This canonicalizes `scalbn`, `scalbnf`, and `scalbnl` libcalls to the
`llvm.ldexp` intrinsic when the call does not access memory.

LLVM floating-point types use radix 2, so `scalbn(x, n)` and `ldexp(x,
n)` produce the same numeric result.
Calls that may access memory are left unchanged because the libcall may
set `errno`, while `llvm.ldexp` does not access memory.

Fixes #216467
DeltaFile
+114-0llvm/test/Transforms/InstCombine/scalbn-to-ldexp.ll
+12-0llvm/lib/Transforms/Utils/SimplifyLibCalls.cpp
+126-02 files

LLVM/project 7333b0bllvm/test/CodeGen/AMDGPU madak.ll, llvm/test/CodeGen/LoongArch float-fma.ll double-fma.ll

Reapply "DAGCombiner: Drop AllowFPOpFusion from visitFADDForFMACombine" (#221561)

This reverts commit 502e51aa4df687807fbe51fa0b419baf65e9615f.
DeltaFile
+352-1,116llvm/test/CodeGen/LoongArch/lasx/fma-v8f32.ll
+352-1,116llvm/test/CodeGen/LoongArch/lasx/fma-v4f64.ll
+260-872llvm/test/CodeGen/LoongArch/float-fma.ll
+260-872llvm/test/CodeGen/LoongArch/double-fma.ll
+388-471llvm/test/CodeGen/AMDGPU/madak.ll
+178-564llvm/test/CodeGen/LoongArch/lsx/fma-v4f32.ll
+1,790-5,01118 files not shown
+2,495-6,19224 files

LLVM/project 502e51allvm/test/CodeGen/AMDGPU madak.ll, llvm/test/CodeGen/LoongArch float-fma.ll double-fma.ll

Revert "DAGCombiner: Drop AllowFPOpFusion from visitFADDForFMACombine" (#221561)

Reverts llvm/llvm-project#221436

Bots failing
DeltaFile
+1,116-352llvm/test/CodeGen/LoongArch/lasx/fma-v8f32.ll
+1,116-352llvm/test/CodeGen/LoongArch/lasx/fma-v4f64.ll
+872-260llvm/test/CodeGen/LoongArch/float-fma.ll
+872-260llvm/test/CodeGen/LoongArch/double-fma.ll
+471-388llvm/test/CodeGen/AMDGPU/madak.ll
+564-178llvm/test/CodeGen/LoongArch/lsx/fma-v4f32.ll
+5,011-1,79017 files not shown
+6,182-2,48523 files

LLVM/project 8d573f5orc-rt/include/orc-rt-internal/support Environment.h, orc-rt/lib/bedrock Error.cpp Logging_printf.cpp

[orc-rt] Split support sources into their own object library (#221442)

This allows support to be used by both bedrock and the upcoming SPIRE
library. Note that support is a CMake object library only, not an
archive or dylib. Clients will always target either Bedrock or SPIRE.

Moves some sources (Error, RTTI, Logging, Environment and the whole sys/
tree), and some headers (Environment.h, and the orc-rt/bedrock/sys
headers) to the support library.

The per-system source composition moves to lib/support/CMakeLists.txt.

An upcomming commit will update the unit tests to reflect this split.
DeltaFile
+0-122orc-rt/lib/bedrock/Logging_printf.cpp
+122-0orc-rt/lib/support/Logging_printf.cpp
+0-82orc-rt/lib/bedrock/Error.cpp
+82-0orc-rt/lib/support/Error.cpp
+68-0orc-rt/include/orc-rt-internal/support/Environment.h
+62-0orc-rt/lib/support/Logging.cpp
+334-20422 files not shown
+622-62928 files

LLVM/project fbf2902llvm/lib/CodeGen/SelectionDAG DAGCombiner.cpp, llvm/test/CodeGen/Hexagon sffms.ll

DAGCombiner: Drop AllowFPOpFusion from visitFSUBForFMACombine

Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
DeltaFile
+3-3llvm/test/CodeGen/Hexagon/sffms.ll
+3-3llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp
+6-62 files

LLVM/project 44888f0llvm/test/CodeGen/AMDGPU madak.ll, llvm/test/CodeGen/LoongArch float-fma.ll double-fma.ll

DAGCombiner: Drop AllowFPOpFusion from visitFADDForFMACombine (#221436)

Rewrites fp-dp3.ll to use flags on individual patterns. It weirdly
used different triples for the fp-contract on and off cases, seemingly
an artifact of the ARM64 and AArch64 merge.
    
fp-contract.cu is essentially a bugfix, the local fp contract(on) pragma
wins over the global flag now.

Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
DeltaFile
+352-1,116llvm/test/CodeGen/LoongArch/lasx/fma-v8f32.ll
+352-1,116llvm/test/CodeGen/LoongArch/lasx/fma-v4f64.ll
+260-872llvm/test/CodeGen/LoongArch/float-fma.ll
+260-872llvm/test/CodeGen/LoongArch/double-fma.ll
+388-471llvm/test/CodeGen/AMDGPU/madak.ll
+178-564llvm/test/CodeGen/LoongArch/lsx/fma-v2f64.ll
+1,790-5,01117 files not shown
+2,485-6,18223 files

LLVM/project 8d10d1fllvm/include/llvm/IR FMF.h, llvm/lib/Target/AMDGPU AMDGPUCodeGenPrepare.cpp

[AMDGPU] Add FastMathFlags::intersectValue helper for rsq ninf/nsz fix (#217724)
DeltaFile
+7-0llvm/include/llvm/IR/FMF.h
+3-2llvm/lib/Target/AMDGPU/AMDGPUCodeGenPrepare.cpp
+10-22 files

LLVM/project 36efd33llvm/test/CodeGen/VE/Scalar store_stk.ll stackframe_align.ll, llvm/test/CodeGen/VE/Vector store_stk_stvm.ll load_stk_ldvm.ll

VE: Use splitAt in expandExtendStackPseudo

Replace the manual block-splitting in expandExtendStackPseudo with
MachineBasicBlock::splitAt. Reduces boilerplate, but there's some
block renumbering churn in the output.

Co-authored-by: Claude (Claude-Opus-4.8)
DeltaFile
+57-57llvm/test/CodeGen/VE/Scalar/atomic_swap.ll
+57-57llvm/test/CodeGen/VE/Scalar/atomic_cmp_swap.ll
+48-48llvm/test/CodeGen/VE/Vector/store_stk_stvm.ll
+48-48llvm/test/CodeGen/VE/Vector/load_stk_ldvm.ll
+42-42llvm/test/CodeGen/VE/Scalar/store_stk.ll
+42-42llvm/test/CodeGen/VE/Scalar/stackframe_align.ll
+294-29460 files not shown
+753-75666 files

LLVM/project 7dbdc02llvm/lib/Target/VE VEInstrInfo.cpp, llvm/test/CodeGen/VE/Scalar builtin_sjlj.ll

VE: Compute live-ins after splitting for EXTEND_STACK expansion

expandExtendStackPseudo splits its block but left the new blocks without
live-in lists, so their uses of registers live across the split are
rejected by -verify-machineinstrs.

Co-authored-by: Claude (Claude-Opus-4.8)
DeltaFile
+5-0llvm/lib/Target/VE/VEInstrInfo.cpp
+2-2llvm/test/CodeGen/VE/Scalar/builtin_sjlj.ll
+7-22 files

LLVM/project b4cc124llvm/lib/Target/X86 X86ISelLowering.cpp

[X86] Add getUnpack wrapper to select between getUnpackl/h shuffles. NFC. (#221543)

Prep work for improving CLMULH vXi32 lowering.
DeltaFile
+10-6llvm/lib/Target/X86/X86ISelLowering.cpp
+10-61 files

LLVM/project 3afdb99llvm/include/llvm/ADT Hashing.h, llvm/include/llvm/IR Attributes.h

[IR] Unique attribute sets and lists in a UniquingSet. NFC (#221525)

Switch to UniquingSet to remove FoldingSetNodeID serialization overhead
on every AttributeSet::get and AttributeList::get.

Attribute and AttributeSet are single-pointer wrappers whose operator==
is pointer equality. Specializing `is_hashable_data` selects the fast
`hash_combine_range_impl` overload that calls `combine_bytes` directly,
skipping copying element by element.

`hash_combine_range` deduces its element type as `const T`, so
`is_hashable_data<const T>` now follows `is_hashable_data<T>` and a type
need only specialize the unqualified form.

Aided by Opus 5
DeltaFile
+6-36llvm/lib/IR/Attributes.cpp
+13-3llvm/include/llvm/IR/Attributes.h
+4-10llvm/lib/IR/AttributeImpl.h
+3-7llvm/lib/IR/LLVMContextImpl.h
+6-0llvm/unittests/ADT/HashingTest.cpp
+2-0llvm/include/llvm/ADT/Hashing.h
+34-566 files

LLVM/project f9eced2llvm/lib/Target/VE VEISelLowering.cpp

VE: Fix ill-typed setjmp result in emitEHSjLjSetJmp

Partially fixes machine verifier failures in existing tests;
they still fail due to other issues.

emitEHSjLjSetJmp materialized the 0/1 return values with LEAzii, which
defines an i64 register, into vregs with the i32 result register class.
This ill-typed MIR is rejected by -verify-machineinstrs.

Materialize the values in i64 and copy the low 32 bits (sub_i32) into the
i32 result. NFC on the emitted code.

Co-authored-by: Claude (Claude-Opus-4.8)
DeltaFile
+13-4llvm/lib/Target/VE/VEISelLowering.cpp
+13-41 files

LLVM/project 75bb9c6llvm/include/llvm/MC MCTargetOptions.h, llvm/include/llvm/Target TargetOptions.h

MC: Move DisableIntegratedAS from TargetOptions to MCTargetOptions

The integrated assembler is only meaningful in MC, so this field belongs
in MCTargetOptions alongside the other assembler options rather than in
the codegen-level TargetOptions.

Co-authored-by: Claude (Claude-Opus-4.8)
DeltaFile
+7-11llvm/include/llvm/Target/TargetOptions.h
+0-7llvm/lib/CodeGen/CommandFlags.cpp
+7-0llvm/lib/MC/MCTargetOptionsCommandFlags.cpp
+3-0llvm/include/llvm/MC/MCTargetOptions.h
+1-1llvm/lib/LTO/LTOCodeGenerator.cpp
+1-1llvm/lib/CodeGen/CodeGenTargetMachineImpl.cpp
+19-204 files not shown
+23-2410 files

LLVM/project da9625cllvm/include/llvm/Analysis ScalarEvolution.h, llvm/lib/Analysis ScalarEvolution.cpp

[SCEV] Strip unnecessary conversions around SCEVUse (NFC) (#219921)
DeltaFile
+3-14llvm/lib/Analysis/ScalarEvolution.cpp
+0-1llvm/include/llvm/Analysis/ScalarEvolution.h
+3-152 files

LLVM/project eefb335llvm/include/llvm/Support GenericLoopInfoImpl.h GenericLoopInfo.h, llvm/lib/Transforms/Scalar LoopInterchange.cpp

[LoopInfo] Merge changeTopLevelLoop and replaceChildLoopWith. NFC (#221503)

Both replace a loop among its siblings with a new one.
DeltaFile
+11-14llvm/include/llvm/Support/GenericLoopInfo.h
+0-17llvm/include/llvm/Support/GenericLoopInfoImpl.h
+1-4llvm/lib/Transforms/Utils/LoopSimplify.cpp
+1-1llvm/lib/Transforms/Scalar/LoopInterchange.cpp
+13-364 files