LLVM/project 680b975 — llvm/test/Transforms/SLPVectorizer/AArch64 reduce-add-dotprod.ll

[SLP][NFC] Add precommit test for dot-product reduction costing (#224590)

Adds a test for reduce.add(mul(ext, ext)) with and without +dotprod at a
threshold between the plain and fused costs. Baseline: both stay scalar;
a later change makes the +dotprod case vectorize..
Needed for #224066
DeltaFile
+141-0llvm/test/Transforms/SLPVectorizer/AArch64/reduce-add-dotprod.ll
+141-01 files

LLVM/project f205558 — flang/lib/Lower ConvertVariable.cpp, flang/test/Lower/CUDA cuda-program-global.cuf cuda-allocatable.cuf

[flang][cuda] Keep descriptors of pinned allocatables in host memory (#229161)

Pinned variables are host only variables. There is no need for the
descriptor of a pinned variable to be allocated in managed memory like
for device or managed variable. Keep pinned descriptor on in host
memory.
DeltaFile
+8-2flang/lib/Lower/ConvertVariable.cpp
+3-3flang/test/Lower/CUDA/cuda-gpu-managed.cuf
+3-2flang/test/Lower/OpenACC/acc-declare-cuda-pinned.f90
+3-2flang/test/Lower/CUDA/cuda-allocatable.cuf
+1-1flang/test/Lower/CUDA/cuda-program-global.cuf
+18-105 files

LLVM/project 83ec4c4 — libcxx/include/__cxx03 bitset

[libcxx] Also ignore -Wshift-count-overflow in clang (#229087)

Clang now also diagnoses this. Also suppress the warning in clang so
that the build still passes (given it is -Werror by default). This
started happening with the release version of clang 23.
DeltaFile
+1-0libcxx/include/__cxx03/bitset
+1-01 files

LLVM/project f42afb1 — compiler-rt/lib/tsan/rtl tsan_interface_ann.cpp tsan_rtl_mutex.cpp

[NFC][TSan] Move ObtainCurrentStack out of ThreadRegistryLock scope (#228794)

ObtainCurrentStack only reads the current thread's shadow stack and
allocates a VarSizeStackTrace buffer, which does not require
ThreadRegistryLock (and already runs outside ThreadRegistryLock in
ReportRace, SignalUnsafeCall, and ReportErrnoSpoiling).

Move ObtainCurrentStack (and dummy_pc in ReportDeadlock) before
ScopedReport in ReportMutexHeldWrongContext, ReportMutexMisuse,
ReportDeadlock, and ReportDestroyLocked so the stack trace buffers also
outlive ScopedReport and OutputReport (needed by #228795).

Assisted-by: Gemini
DeltaFile
+5-4compiler-rt/lib/tsan/rtl/tsan_rtl_mutex.cpp
+2-2compiler-rt/lib/tsan/rtl/tsan_interface_ann.cpp
+7-62 files

LLVM/project 89fdf0d — llvm/include/llvm/Analysis PHITransAddr.h, llvm/lib/Analysis PHITransAddr.cpp MemoryDependenceAnalysis.cpp

[GVN] Preserve !prof metadata when materializing select for load PRE

Pass the originating SelectInst through PHITransAddr, MemoryDependenceAnalysis,
and GVN so AvailableValue::MaterializeAdjustedValue can copy its profile
metadata when creating the replacement value select.
DeltaFile
+22-21llvm/lib/Transforms/Scalar/GVN.cpp
+19-11llvm/test/Transforms/GVN/PRE/pre-load-through-select.ll
+9-8llvm/include/llvm/Analysis/PHITransAddr.h
+5-4llvm/lib/Analysis/MemoryDependenceAnalysis.cpp
+0-7llvm/utils/profcheck-xfail.txt
+2-2llvm/lib/Analysis/PHITransAddr.cpp
+57-531 files not shown
+58-547 files

LLVM/project a32be17 — flang/lib/Optimizer/Builder IntrinsicCall.cpp, flang/test/Lower/Intrinsics ieee_max_min.f90 ieee_logb.f90

[flang][PPC] Implement ieee_set_flag for Linux (#224039)

This patch implements the Linux PPC specific lowering for
`ieee_set_flag`.

Calling `feraiseexcept()` may deliver `SIGFPE` when the corresponding
floating-point exception trap is enabled. To avoid this, the lowering
updates the FPSCR exception status bits directly using
`llvm.ppc.readflm` and `llvm.ppc.setflm`.

Assisted-By: IBM Bob
DeltaFile
+158-6flang/lib/Optimizer/Builder/IntrinsicCall.cpp
+90-0flang/test/Lower/Intrinsics/ieee_set_flag_ppc.f90
+1-0flang/test/Lower/Intrinsics/ieee_max_min.f90
+1-0flang/test/Lower/Intrinsics/ieee_logb.f90
+1-0flang/test/Lower/Intrinsics/ieee_flag.f90
+251-65 files

LLVM/project 523fb6e — compiler-rt/lib/tsan/rtl tsan_rtl_mutex.cpp

undo 42

Created using spr 1.3.7
DeltaFile
+1-1compiler-rt/lib/tsan/rtl/tsan_rtl_mutex.cpp
+1-11 files

LLVM/project af2d769 — compiler-rt/lib/fuzzer FuzzerDriver.cpp FuzzerLoop.cpp, compiler-rt/test/fuzzer SleepOneSecondTest.cpp stale_corpus_timeout.test

[libFuzzer] Stop fuzzing if no new corpus were found within a specified period (#176177)

This patch adds a new command-line option `-stale_corpus_timeout=N`
which allows you to specify that the fuzzer should exit if it hasn't
found a new corpus within a certain amount of time. It also adds to
stats how many seconds a new path has not been found. An analog of
`AFL_EXIT_ON_TIME`.
DeltaFile
+35-8compiler-rt/lib/fuzzer/FuzzerFork.cpp
+12-3compiler-rt/lib/fuzzer/FuzzerInternal.h
+13-0compiler-rt/test/fuzzer/stale_corpus_timeout.test
+4-2compiler-rt/lib/fuzzer/FuzzerDriver.cpp
+5-1compiler-rt/lib/fuzzer/FuzzerLoop.cpp
+4-1compiler-rt/test/fuzzer/SleepOneSecondTest.cpp
+73-153 files not shown
+81-159 files

LLVM/project c1d8247 — llvm/test/CodeGen/AMDGPU frem.ll atomic_optimizations_local_pointer.ll, llvm/test/CodeGen/AMDGPU/GlobalISel srem.i64.ll sdiv.i64.ll

rebase

Created using spr 1.3.7
DeltaFile
+733-5,435llvm/test/CodeGen/AMDGPU/atomic_optimizations_local_pointer.ll
+1,609-1,040llvm/test/CodeGen/AMDGPU/frem.ll
+986-1,144llvm/test/CodeGen/AMDGPU/GlobalISel/sdiv.i64.ll
+914-1,072llvm/test/CodeGen/AMDGPU/GlobalISel/srem.i64.ll
+728-824llvm/test/CodeGen/X86/mul-constant-i32.ll
+585-680llvm/test/CodeGen/X86/mul-constant-i64.ll
+5,555-10,1951,214 files not shown
+52,062-31,5191,220 files

LLVM/project 197fb4b — compiler-rt/lib/tsan/rtl tsan_report.cpp tsan_report.h

[NFC][TSan] Use in-class member initializers in tsan_report.h (#228780)

Use in-class member initializers for all structs and classes in
tsan_report.h and default constructors and destructor in
tsan_report.cpp.

Assisted-by: Gemini
DeltaFile
+29-28compiler-rt/lib/tsan/rtl/tsan_report.h
+5-17compiler-rt/lib/tsan/rtl/tsan_report.cpp
+34-452 files

LLVM/project 9e69e85 — mlir/lib/Dialect/Tosa/Transforms TosaProfileCompliance.cpp

[NFC][mlir][tosa] Use optnone for compliance initialization under ASan (#229292)

Similar to #223586 and #226310, compiling the ~5,000-line generated
initializer in `TosaProfileCompliance::TosaProfileCompliance()` is also
very slow under AddressSanitizer.

Mark the constructor with `optnone` under `LLVM_ADDRESS_SANITIZER_BUILD`
as well to skip expensive optimization and register allocation passes.

Assisted-by: Gemini
DeltaFile
+7-6mlir/lib/Dialect/Tosa/Transforms/TosaProfileCompliance.cpp
+7-61 files

LLVM/project 82046be — llvm/lib/Target/X86 X86TargetMachine.cpp, llvm/test/CodeGen/X86 llc-pipeline-npm.ll 2011-06-12-FastAllocSpill.ll

[X86] Let RegAllocFast lower tied operands (#228968)

... so the -O0 pipeline no longer runs TwoAddressInstructionPass.
Follow-ups will enable this for all targets that don't insert passes
with the TwoAddressInstructinPassID anchor.

https://llvm-compile-time-tracker.com/compare.php?from=8ad6c581709df648a2a1e468031a81106a7e00b4&to=33ea7b07c99b9a37f72409029867c50bd1d22a52&stat=instructions:u
-O0 -g, instructions:u geomean: -0.83% (stage1) and -0.85% (stage2).

In a llvm-test-suite build, 2088 of 2563 executables have identical
.text, while 400 have few instructions in total (-0.008%). None gained
extra instructions.

With -O2 -mllvm -regalloc=fast, .text shrinks by 0.12% but the
instruction count grows by 0.09%: RegAllocFast does not convert to
three-address form, so where a tied source stays live,
`leaq 1(%rsi), %rax` becomes `movq %rsi, %rax; incq %rax`.

Test updates:

    [13 lines not shown]
DeltaFile
+15-16llvm/test/CodeGen/X86/atomic-unordered.ll
+8-11llvm/test/CodeGen/X86/switch.ll
+0-9llvm/test/CodeGen/X86/O0-pipeline.ll
+0-2llvm/test/CodeGen/X86/llc-pipeline-npm.ll
+1-1llvm/test/CodeGen/X86/2011-06-12-FastAllocSpill.ll
+2-0llvm/lib/Target/X86/X86TargetMachine.cpp
+26-396 files

LLVM/project 7b4e138 — llvm/test/CodeGen/AMDGPU frem.ll atomic_optimizations_local_pointer.ll, llvm/test/CodeGen/AMDGPU/GlobalISel srem.i64.ll sdiv.i64.ll

Merge branch 'main' into users/vitalybuka/spr/nfctsan-use-in-class-member-initializers-for-reportdesc-and-reportmop
DeltaFile
+733-5,435llvm/test/CodeGen/AMDGPU/atomic_optimizations_local_pointer.ll
+1,609-1,040llvm/test/CodeGen/AMDGPU/frem.ll
+986-1,144llvm/test/CodeGen/AMDGPU/GlobalISel/sdiv.i64.ll
+914-1,072llvm/test/CodeGen/AMDGPU/GlobalISel/srem.i64.ll
+728-824llvm/test/CodeGen/X86/mul-constant-i32.ll
+585-680llvm/test/CodeGen/X86/mul-constant-i64.ll
+5,555-10,1951,210 files not shown
+52,037-31,4821,216 files

LLVM/project dc3e558 — clang/lib/StaticAnalyzer/Checkers/WebKit PtrTypesSemantics.cpp, clang/test/Analysis/Checkers/WebKit nodelete-annotation.cpp

[alpha.webkit.NoDeleteChecker] Handle CXXStdInitializerListExpr in trivial analysis (#224723)

TrivialFunctionAnalysisVisitor had no handler for
CXXStdInitializerListExpr, so a braced list bound to a
std::initializer_list fell through to VisitStmt and was conservatively
treated as non-trivial. This made any nodelete function containing e.g.
std::min({a, b, c}) report that it "contains code that could destruct an
object".

The backing array of a std::initializer_list is a temporary whose
lifetime ends in the enclosing function, so its elements really are
destructed there. Accept the node when the array's element type is
trivially destructible and recurse into the initializers, and keep
rejecting it otherwise.
DeltaFile
+61-0clang/test/Analysis/Checkers/WebKit/nodelete-annotation.cpp
+11-0clang/lib/StaticAnalyzer/Checkers/WebKit/PtrTypesSemantics.cpp
+72-02 files

LLVM/project 6ad06d5 — mlir/lib/Dialect/Tosa/Transforms TosaProfileCompliance.cpp

[𝘀𝗽𝗿] initial version

Created using spr 1.3.7
DeltaFile
+7-6mlir/lib/Dialect/Tosa/Transforms/TosaProfileCompliance.cpp
+7-61 files

LLVM/project 0207f04 — clang/lib/AST Type.cpp ASTContext.cpp, clang/lib/Sema SemaType.cpp

[clang] Pass only CVR array index qualifiers to ASTContext (#225582)

ArrayType stores only the CVR index qualifiers, but array declarators
may also pass `__unaligned` or `_Atomic`, so getIncompleteArrayType and
getDependentSizedArrayType insert nodes whose profile doesn't match the
lookup. Mask the qualifiers in BuildArrayType and assert in ArrayType.

Fixes #223150
Aided by Opus 5.5
DeltaFile
+10-6clang/lib/Sema/SemaType.cpp
+0-4clang/lib/AST/ASTContext.cpp
+3-0clang/test/SemaCXX/MicrosoftExtensions.cpp
+2-0clang/lib/AST/Type.cpp
+15-104 files

LLVM/project 09b7f9f — llvm/lib/Transforms/Instrumentation SanitizerCoverage.cpp, llvm/lib/Transforms/Utils Instrumentation.cpp

[SanitizerCoverage] Don't create comdat for unnamed functions (#228959)

MergeFunctions::mergeTwoFunctions may turn identical functions into a
private unnamed function, which trips `assert(F.hasName())`. Without
assertions, all unnamed functions share an empty-name IR comdat (which
yields no ELF section group per MCContext::getELFSection).

Fix #228820: Just drop comdat for unnamed functions. SanitizerCoverage
then retains the metadata sections via `llvm.used`. It's acceptable to
have redundat metadata sections with discarded associated text sections.

Improve comments at https://reviews.llvm.org/D97430 modified code.
DeltaFile
+22-7llvm/test/Instrumentation/SanitizerCoverage/interposable-symbol.ll
+8-10llvm/lib/Transforms/Instrumentation/SanitizerCoverage.cpp
+4-1llvm/lib/Transforms/Utils/Instrumentation.cpp
+34-183 files

LLVM/project c2c4780 — clang/lib/CIR/Dialect/Transforms TargetLowering.cpp, llvm/test/CodeGen/AMDGPU frem.ll

Merge branch 'main' into revert-228574-dsymutil-index-globals
DeltaFile
+1,609-1,040llvm/test/CodeGen/AMDGPU/frem.ll
+249-32llvm/test/CodeGen/NVPTX/cache-hint-transforms.ll
+279-0llvm/test/Transforms/SLPVectorizer/X86/fused-alt-fmul.ll
+215-13llvm/test/Transforms/Attributor/nofpclass-fmul.ll
+209-10llvm/test/Transforms/Attributor/nofpclass-fdiv.ll
+82-80clang/lib/CIR/Dialect/Transforms/TargetLowering.cpp
+2,643-1,17569 files not shown
+4,280-1,61775 files

LLVM/project 27c4ac8 — lld/ELF InputFiles.cpp, lld/test/ELF tls-mismatch.s

[ELF] Error on non-TLS definition with TLS reference (#228658)

postParse() reports an error when a TLS definition has a non-TLS
reference. Extend it to cover non-TLS definition with a TLS reference,
which GNU ld also errors.

Before LTO, a TLS definition in a bitcode module asm still has a type of
STT_NOTYPE. Exempt TLS references in this case.

RelocScan-time error checking is not useful - assemblers set STT_TLS
when a symbol is used by a TLS relocation.
DeltaFile
+31-1lld/test/ELF/tls-mismatch.s
+7-5lld/ELF/InputFiles.cpp
+38-62 files

LLVM/project 55334bc — llvm/lib/Transforms/Vectorize VPlanTransforms.cpp, llvm/test/Transforms/LoopVectorize/VPlan vplan-based-stride-mv.ll

Generalize stride/sext/zext replacement to live-ins rewrite
DeltaFile
+18-32llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
+3-3llvm/test/Transforms/LoopVectorize/VPlan/vplan-based-stride-mv.ll
+21-352 files

LLVM/project 75cc30c — llvm/lib/CodeGen/SelectionDAG LegalizeDAG.cpp LegalizeVectorTypes.cpp, llvm/test/CodeGen/NVPTX cache-hint-unaligned-store-aa.ll cache-hint-transforms.ll

[SelectionDAG] Preserve cache hint metadata during legalization (#225273)

We want to preserve cache hint metadata everywhere that alias analysis
metadata is preserved. Even if an operation is expanded during
legalization, it makes sense to preserve cache hint metadata on each
part. For example, a load with an eviction policy should get expanded
into multiple loads with the same eviction policy. In the future it
might make sense to provide a TTI hook for how to propagate the
metadata, but this is simplest for now, and usually correct.

To propagate metadata, I introduced `getNonRangeMMOMetadata`. There are
probably some cases here where range metadata can be preserved, but I
don't want to touch that. That can be done as a follow-up.

There are 3 cases where we were not preserving `AAInfo` where we are
now.

[Expand to a bitconvert of the value to the integer type of the same
size, then a (misaligned) int

    [7 lines not shown]
DeltaFile
+249-32llvm/test/CodeGen/NVPTX/cache-hint-transforms.ll
+136-0llvm/test/CodeGen/RISCV/mem-cache-hint-legalize-types.ll
+119-0llvm/test/CodeGen/NVPTX/cache-hint-unaligned-store-aa.ll
+33-29llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
+26-23llvm/lib/CodeGen/SelectionDAG/LegalizeVectorTypes.cpp
+17-16llvm/lib/CodeGen/SelectionDAG/LegalizeDAG.cpp
+580-1004 files not shown
+617-12810 files

LLVM/project 611f5f3 — llvm Maintainers.md

Update former maintainers (#229276)

b2ff0e8eae8236e7555afa6026e0cf251af8c4ae updated the
OpenMPOpt/Attributor maintainers but did not add Johannes to the former
maintainers list.
DeltaFile
+1-0llvm/Maintainers.md
+1-01 files

LLVM/project ad39cc1 — llvm/lib/Transforms/InstCombine InstCombineSelect.cpp, llvm/test/Transforms/InstCombine truncating-saturate.ll

[InstCombine] Fix ProfCheck for canonicalizeClampLike (#229269)

We cannot know the profile in the general case (see the added comment),
so mark the selects unknown for now. In the case that C2 is at one of
the bounds we can, but I expect this case is rare enough that handling
it explicitly doesn't make much sense.
DeltaFile
+10-4llvm/lib/Transforms/InstCombine/InstCombineSelect.cpp
+7-3llvm/test/Transforms/InstCombine/truncating-saturate.ll
+0-3llvm/utils/profcheck-xfail.txt
+17-103 files

LLVM/project 82b0594 — flang/lib/Semantics check-cuda.cpp, flang/test/Semantics/CUDA cuf-assumed-size-transfer.cuf

[flang][cuda] Reject implicit transfer of assumed-size device arrays (#229128)

An assignment in host code whose right-hand side mixes device data with
host data or constants is lowered as an implicit data transfer: each
device object is copied whole to a host temporary created from its mold,
and the expression is evaluated on the host. The size of an assumed-size
array (an assumed-size dummy or a Cray pointee declared with `(*)`) is
unknown, so the temporary cannot be created and lowering crashed with an
invalid fir.convert:

  double precision :: a(*)
  pointer (ip, a)
  attributes(device) :: a
  h = h + a(3)*2.0d0

Report a semantic error instead when such an assignment references an
assumed-size device array. Assignments that do not need an implicit
transfer are still accepted: element transfers like `h = a(3)` and
host-to-device assignments like `a(3) = h*2.0d0`.

This align flang with the reference compiler.
DeltaFile
+34-0flang/test/Semantics/CUDA/cuf-assumed-size-transfer.cuf
+15-0flang/lib/Semantics/check-cuda.cpp
+49-02 files

LLVM/project b278b73 — llvm/lib/Target/X86 X86SelectionDAGInfo.cpp X86ISelLowering.cpp

[X86] Remove an unused MVT::Other result from  VP2INTERSECT. (#228936)

Fixes verification failure in X86SelectionDAGInfo::verifyTargetNode
(https://github.com/llvm/llvm-project/issues/185649)
DeltaFile
+1-2llvm/lib/Target/X86/X86ISelLowering.cpp
+0-2llvm/lib/Target/X86/X86SelectionDAGInfo.cpp
+1-42 files

LLVM/project 06e692c — lldb/source/Core Module.cpp, lldb/test/API/functionalities/gdb_remote_client TestWasm.py

[lldb] Unload Wasm modules the stub no longer reports (#227814)

When a Wasm engine such as JavaScriptCore reloads a page, the modules it
ran go away and new instances take their place. LLDB kept two kinds of
stale module around.

ProcessGDBRemote::LoadModules never unloads the target's executable,
which no library list includes. A Wasm target has no executable, so
Target::GetExecutableModule falls back to the first module, and the
first module the stub ever reported stayed in the image list for good.
Let the dynamic loader say whether the process runs a main executable,
and have the Wasm loader say it does not.

A module reloaded under the same name matched the module read from
memory at its old address, and LLDB then moved that module to the new
one. That trips an assertion in ObjectFileWasm::SetLoadAddress, and
otherwise leaves a module whose image came from an instance that is
gone. A module read from memory now only matches a spec at the address
it was read from, like ModuleSpec::Matches already does for two specs.

rdar://175013476
DeltaFile
+94-0lldb/test/API/functionalities/gdb_remote_client/TestWasm.py
+5-0lldb/source/Core/Module.cpp
+99-02 files

LLVM/project 3c50f79 — llvm/lib/CodeGen/SelectionDAG DAGCombiner.cpp

fixup! Use plain getOpcode
DeltaFile
+1-2llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp
+1-21 files

LLVM/project 8887f3b — llvm/lib/Target/AMDGPU GCNHazardRecognizer.cpp

fixup! [AMDGPU] Measure MFMA read hazards at each producer

Turn the walk into a small class, WindowDeficitSearch, with its state as
members, and count wait states through one helper that states the
inline asm assumption.

AI-assisted.
DeltaFile
+99-71llvm/lib/Target/AMDGPU/GCNHazardRecognizer.cpp
+99-711 files

LLVM/project 3b4caff — clang/lib/Analysis/FlowSensitive/Models UncheckedStatusOrAccessModel.cpp UncheckedOptionalAccessModel.cpp, clang/unittests/Analysis/FlowSensitive MockHeaders.cpp

[FlowSensitive] [Optional] [StatusOr] handle AssertionResultExpectation

Assisted-By: Gemini
DeltaFile
+14-4clang/unittests/Analysis/FlowSensitive/MockHeaders.cpp
+9-0clang/lib/Analysis/FlowSensitive/Models/UncheckedStatusOrAccessModel.cpp
+9-0clang/lib/Analysis/FlowSensitive/Models/UncheckedOptionalAccessModel.cpp
+32-43 files

LLVM/project c48231b — clang/lib/Analysis/FlowSensitive/Models CMakeLists.txt GtestModelHelpers.h

[FlowSensitive] add handling for AssertionResultExpectation

In a follow up, this will be used by the StatusOr and Optional models.

Assisted-By: Gemini
DeltaFile
+45-0clang/lib/Analysis/FlowSensitive/Models/GtestModelHelpers.cpp
+19-0clang/lib/Analysis/FlowSensitive/Models/GtestModelHelpers.h
+2-0clang/lib/Analysis/FlowSensitive/Models/CMakeLists.txt
+66-03 files