LLVM/project 58b1026 — llvm/lib/Analysis ScalarEvolution.cpp, llvm/test/Analysis/ScalarEvolution trip-count-variable-stride-predicate.ll

[SCEV] -  Add positive-stride predicate for backedge-taken count. (#222261)

When `howManyLessThans` encounters a loop with an unknown stride that
cannot be proven finite (no `mustprogress` or side-effect-free
guarantee), SCEV currently returns `CouldNotCompute` for the
backedge-taken count. This blocks downstream consumers like the loop
vectorizer from optimizing such loops.

This patch relaxes the requirement by allowing a predicated
backedge-taken count when `AllowPredicates` is true. Instead of
requiring `loopIsFiniteByAssumption(L)` unconditionally, we add a
`Compare predicate: stride sgt 0` when finiteness cannot be proven.

A positive stride guarantees forward progress, making the BTC formula
correct. The predicate is emitted as a runtime check by consumers
(e.g., the loop vectorizer generates a guard branch before the vector
loop).



    [17 lines not shown]
DeltaFile
+102-0llvm/test/Transforms/LoopVectorize/scev-variable-stride-predicate.ll
+31-0llvm/test/Analysis/ScalarEvolution/trip-count-variable-stride-predicate.ll
+18-4llvm/lib/Analysis/ScalarEvolution.cpp
+151-43 files

LLVM/project e0bd489 — clang/docs LibASTImporter.md UsersManual.md

[docs] Pass -c when generating a PCH file (#226850)

Without options like -fsyntax-only/-E/-S/-c, the driver runs in link
mode, and a header-only command line produces a PCH only incidentally: a
linker option such as -lm or -Wl,..., including one from a config file,
adds a link job. Use -c in the PCH examples.

Also fix two broken examples: -ignore-pch is a driver option (cc1
rejects -Xclang -ignore-pch), and the relocatable PCH example is
missing -o. In LibASTImporter.md, generate the C++ AST files with
-emit-ast instead of treating .cpp files as headers.

LLM-aided
DeltaFile
+9-10clang/docs/UsersManual.md
+2-2clang/docs/LibASTImporter.md
+11-122 files

LLVM/project 090b123 — lldb/test/API/functionalities/scripted_frame_provider/pass_through_prefix frame_provider.py TestFrameProviderPassThroughPrefix.py

[lldb][test] Handle None function name in frame provider tests (#226936)

Fixes #191859.

On Arm one of the functions has no name, which is a valid state.
```
Num frames: 7
0 frame #0: 0x092a151c a.out`baz at main.c:5:3 --name:  baz
1 frame #1: 0x092a1540 a.out`bar at main.c:10:10 --name:  bar
2 frame #2: 0x092a1560 a.out`foo at main.c:15:10 --name:  foo
3 frame #3: 0x092a158c a.out`main at main.c:20:10 --name:  main
4 frame #4: 0xf7ab939a libc.so.6` --name:  None
5 frame #5: 0xf7ab943e libc.so.6`__libc_start_main + 94 --name:  __libc_start_main
6 frame #6: 0x092a1434 a.out`_start + 40 --name:  _start
```
The prefix provider needs to handle that possibility.

Though the Python type annotations on ScriptedFrame say that the name
methods return string, in reality they can return None, which the C++
sides expect. I will fix that in another PR.
DeltaFile
+0-6lldb/test/API/functionalities/scripted_frame_provider/pass_through_prefix/TestFrameProviderPassThroughPrefix.py
+3-1lldb/test/API/functionalities/scripted_frame_provider/pass_through_prefix/frame_provider.py
+3-72 files

LLVM/project c68a4ca — mlir/lib/Target/LLVMIR/Dialect/LLVMIR LLVMToLLVMIRTranslation.cpp, mlir/test/Target/LLVMIR llvmir-invalid.mlir

[LLVMIR] Directly lower global initializer GEP to constant expression (#226904)

MLIR lowers global initializers in a very unusual way, by using an
IRBuilder without insertion point, and relying on the fact that for a
well-formed initializer, all the instructions will be converted to
constant expressions and no actual instructions that require insertion
will be produced.

This causes issues for https://github.com/llvm/llvm-project/pull/226425,
which requires an insertion point for CreateGEP to determine the data
layout.

To avoid this, make convertGEPOp() directly produce a constant
expression GEP if there is no insertion point, reusing the existing code
path for inrange GEPs.
DeltaFile
+14-8mlir/lib/Target/LLVMIR/Dialect/LLVMIR/LLVMToLLVMIRTranslation.cpp
+2-2mlir/test/Target/LLVMIR/llvmir-invalid.mlir
+16-102 files

LLVM/project bb9dd27 — clang/lib/AST ExprConstant.cpp, clang/lib/AST/ByteCode InterpState.h EvalEmitter.h

[clang][bytecode] Use bytecode interpreter in toplevel `Expr::Evaluate*` functions (#226217)

To avoid all the state setup we otherwise do. This also lets us remove
the `EvalInfo::EnableNewConstInterp` flag, which was previously checked
on every `::EvaluateAsRValue` call.
DeltaFile
+123-20clang/lib/AST/ExprConstant.cpp
+0-29clang/lib/AST/ByteCode/Context.cpp
+0-13clang/lib/AST/ByteCode/InterpState.cpp
+0-6clang/lib/AST/ByteCode/EvalEmitter.cpp
+0-3clang/lib/AST/ByteCode/InterpState.h
+0-3clang/lib/AST/ByteCode/EvalEmitter.h
+123-741 files not shown
+123-767 files

LLVM/project b20896c — llvm/lib/Target/DirectX DXILFlattenArrays.cpp DXILDataScalarization.cpp, llvm/test/CodeGen/DirectX scalarize-static-array-of-float-vectors.ll bugfix_150050_data_scalarize_const_gep.ll

[DirectX] Don't use IRBuilder to create GEPs (#227000)

This switches the DirectX backend to directly call
`GetElementPtrInst::Create()` instead of creating GEPs via IRBuilder.
IRBuilder will soon start canonicalizing GEPs to ptradd representation
(with the first change being
https://github.com/llvm/llvm-project/pull/226425, affecting constant
expressions only), while DirectX needs GEPs to have specific form.

The change implemented here is a minimal one to unblock further work. In
the future, it will become impossible to construct such GEPs through any
API. The DirectX backend needs to switch to using an intrinsic to
represent any non-canonical GEPs it needs. This followup work is tracked
in https://github.com/llvm/llvm-project/issues/227005.
DeltaFile
+50-25llvm/test/CodeGen/DirectX/llc-vector-load-scalarize.ll
+6-19llvm/lib/Target/DirectX/DXILFlattenArrays.cpp
+11-14llvm/lib/Target/DirectX/DXILDataScalarization.cpp
+16-8llvm/test/CodeGen/DirectX/scalar-store.ll
+14-7llvm/test/CodeGen/DirectX/bugfix_150050_data_scalarize_const_gep.ll
+12-6llvm/test/CodeGen/DirectX/scalarize-static-array-of-float-vectors.ll
+109-797 files not shown
+145-10213 files

LLVM/project d7daf1d — llvm/tools/llvm-profgen PerfReader.cpp

[llvm][llvm-profgen] Fix printing bogus trace info on 32-bit (#226965)

And as a result, fix
llvm/test/tools/llvm-profgen/AArch64/cs-bogus-trace.test when run on Arm
32-bit:
https://lab.llvm.org/buildbot/#/builders/122/builds/59

Fixes #225569 / 28fec94991905d8dd8a08658e9590e13b7fe1af5.

`x` is for an unsigned integer but the argument given for it is a
uint64_t. This worked on 64-bit (maybe it's passed in a register), but
not on 32-bit where we got strange values printed.

Use format_hex instead, so we don't have to choose a format code.
DeltaFile
+2-2llvm/tools/llvm-profgen/PerfReader.cpp
+2-21 files

LLVM/project 6aeadac — lldb/examples/python/templates scripted_process.py

[lldb] Correct types for ScriptedFrame get_function_name get_display_function_name (#226938)

These claimed to return str but the initial value for self.name is None,
and the C++ side returns optional<string>. So I think optional string is
correct here.
DeltaFile
+2-2lldb/examples/python/templates/scripted_process.py
+2-21 files

LLVM/project c20c91f — llvm/lib/Target/AMDGPU GCNHazardRecognizer.cpp, llvm/test/CodeGen/AMDGPU wmma-coexecution-valu-hazards.mir

[AMDGPU] Fix missed WMMA C-operand co-exec hazard

The gfx1250 WMMA co-execution hazard check treats only A, B and the
SWMMAC index as registers the in-flight MMA still reads. C (src2 of a
non-SWMMAC WMMA) is missing, so a VALU scheduled into the MMA's shadow
can clobber C and the MMA consumes the new value.

This is latent while C is tied to vdst, since the existing D check then
covers it. It miscompiles where the tie does not hold: for
v_wmma_bf16f32_16x16x32_bf16, whose D is narrower than C, and for the
_threeaddr form of any WMMA.

Add src2 to the checked set for non-SWMMAC WMMAs.
DeltaFile
+171-2llvm/test/CodeGen/AMDGPU/wmma-coexecution-valu-hazards.mir
+5-0llvm/lib/Target/AMDGPU/GCNHazardRecognizer.cpp
+176-22 files

LLVM/project e8c09d1 — clang/docs InternalsManual.md, clang/include/clang/Sema Sema.h

[clang] Cache normalized constraints by expression instead of by declaration (#226620)

Sema::NormalizationCache was keyed by the constrained declaration. The
members of a class template specialization are distinct declarations for
every specialization, but they all share the uninstantiated constraint
expressions of the member of the primary template they were instantiated
from. So the constraints of e.g. the constrained constructors of
std::optional, std::span, std::pair or of the members of range adaptors
were normalized from scratch for every single specialization of those
classes that a TU uses. Normalization is expensive: it expands the whole
concept tree below the expression and substitutes the parameter mapping
at every level.

When compiling Chromium, 48%
(chrome/browser/glic/...contents_manager.cc)
to 72% (chrome/browser/ui/views/frame/browser_view.cc) of all
NormalizationCache misses were for expressions that were normalized
before
for a different declaration, and normalization is 4.5% of all compile

    [28 lines not shown]
DeltaFile
+27-12clang/lib/Sema/SemaConcept.cpp
+6-6clang/include/clang/Sema/Sema.h
+3-2clang/docs/InternalsManual.md
+36-203 files

LLVM/project 8d79b74 — polly/lib/Transform ScheduleOptimizer.cpp, polly/test/ScheduleOptimizer simplify_deps_bounded_proximity.ll

[Polly] Keep proximity dependence exact if simplifying it unbounds its distance (#225840)

Before calling the isl scheduler, `runIslScheduleOptimizer` gists the
dependences
with the iteration domains (`-polly-opt-simplify-deps=yes`, the default
since
a26db470834a, 2012). For a statement that reads a value written in one
iteration
of a loop in all later iterations, the gist also drops the upper bound
of that
loop:

{ S[k, i, k] -> S[k, i, j] : k < j <= n - 1 } becomes { S[k, i, k] ->
S[k, i, j] : j > k }

The proximity distance along j is then unbounded. isl's Pluto-like step
needs a
bound `u·n + w` on every proximity distance, finds no row that satisfies
it, and

    [67 lines not shown]
DeltaFile
+84-0polly/test/ScheduleOptimizer/simplify_deps_bounded_proximity.ll
+44-0polly/lib/Transform/ScheduleOptimizer.cpp
+128-02 files

LLVM/project a8d2c44 — llvm/lib/Target/X86 X86ISelLowering.cpp

X86: Mark eflags clobbers dead in va_arg expansion (#227197)

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+13-7llvm/lib/Target/X86/X86ISelLowering.cpp
+13-71 files

LLVM/project 992ca79 — llvm/include/llvm/IR IRBuilder.h, llvm/lib/Transforms/Vectorize VectorCombine.cpp

[VectorCombine] Pass flags during IR creation (#193271)

Since commit 777d6b5, VectorCombine has been using InstSimplifyFolder to
simplify vector instructions during IR construction. When creating a new
instruction, InstSimplifyFolder may fold the operation and return an
existing operand instead of emitting a new instruction.

In such cases, copying IR flags to the returned value is incorrect and
may unintentionally propagate flags to pre‑existing instructions,
polluting the original IR. To avoid this, flags should be passed at IR
creation rather than being set after construction. Fix #192607.
DeltaFile
+227-0llvm/test/Transforms/VectorCombine/binop-scalarize.ll
+39-12llvm/lib/Transforms/Vectorize/VectorCombine.cpp
+50-0llvm/test/Transforms/VectorCombine/cmp-scalarize.ll
+12-0llvm/test/Transforms/VectorCombine/unary-op-scalarize.ll
+7-2llvm/include/llvm/IR/IRBuilder.h
+335-145 files

LLVM/project 633223d — clang/lib/CIR/CodeGen CIRGenBuiltinAArch64.cpp, clang/test/CodeGen/AArch64 neon-intrinsics.c

[CIR][AArch64] Lower NEON vshrn* intrinsics (#226382)

## summary

part of : https://github.com/llvm/llvm-project/issues/185382

Lower all intrinsics in :
https://arm-software.github.io/acle/neon_intrinsics/advsimd.html#vector-shift-right-and-narrow

Assisted by : gpt-6-sol
DeltaFile
+173-0clang/test/CodeGen/AArch64/neon/intrinsics.c
+0-162clang/test/CodeGen/AArch64/neon-intrinsics.c
+11-5clang/lib/CIR/CodeGen/CIRGenBuiltinAArch64.cpp
+184-1673 files

LLVM/project 9b925b5 — llvm/test/CodeGen/X86 vaarg-dead-eflags.ll

Remove test
DeltaFile
+0-149llvm/test/CodeGen/X86/vaarg-dead-eflags.ll
+0-1491 files

LLVM/project 85c33ba — llvm/lib/Target/X86 X86ISelLowering.cpp, llvm/test/CodeGen/X86 vaarg-dead-eflags.ll

X86: Mark eflags clobbers dead in va_arg expansion

Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+149-0llvm/test/CodeGen/X86/vaarg-dead-eflags.ll
+13-7llvm/lib/Target/X86/X86ISelLowering.cpp
+162-72 files

LLVM/project 418d1ff — llvm/lib/CodeGen MachineBasicBlock.cpp, llvm/lib/CodeGen/MIRParser MILexer.h MILexer.cpp

MIR: Serialize MachineBasicBlock::IsEHContTarget

Co-Authored-By: Claude Sonnet 5 <noreply at anthropic.com>
DeltaFile
+16-0llvm/test/CodeGen/MIR/Generic/machine-basic-block-ehcont-target.mir
+6-0llvm/lib/CodeGen/MIRParser/MIParser.cpp
+5-0llvm/lib/CodeGen/MachineBasicBlock.cpp
+1-0llvm/lib/CodeGen/MIRParser/MILexer.h
+1-0llvm/lib/CodeGen/MIRParser/MILexer.cpp
+29-05 files

LLVM/project 37cda37 — llvm/lib/CodeGen MachineBasicBlock.cpp, llvm/lib/CodeGen/MIRParser MILexer.h MILexer.cpp

MIR: Serialize MachineBasicBlock::IsCleanupFuncletEntry

Co-authored-by: Claude Sonnet 5 <noreply at anthropic.com>
DeltaFile
+16-0llvm/test/CodeGen/MIR/Generic/machine-basic-block-cleanup-funclet-entry.mir
+6-0llvm/lib/CodeGen/MIRParser/MIParser.cpp
+5-0llvm/lib/CodeGen/MachineBasicBlock.cpp
+1-0llvm/lib/CodeGen/MIRParser/MILexer.h
+1-0llvm/lib/CodeGen/MIRParser/MILexer.cpp
+29-05 files

LLVM/project ba67350 — clang/lib/Sema SemaConcept.cpp, clang/test/CXX/temp/temp.constr/temp.constr.normal p1.cpp

[clang] Consistently cache failed constraint normalization (#227086)

Previously, if substituting the parameter mappings of a normalized
constraint failed, Sema::getNormalizedAssociatedConstraints() returned
nullptr, but it stored the partially substituted normal form in
NormalizationCache. So the first lookup for such a declaration failed,
but every later lookup returned the broken normal form, and subsumption
checking and the ambiguous-constraint diagnostics then continued with
it.

I believe this wasn't intentional:

- Before #161671 (e9972debc98c), normalization was a single step, and a
failure was cached as nullptr.

- #161671 added the parameter mapping substitution step. It inserted the
normal form into the cache before substituting, and returned nullptr if
the substitution then failed, leaving the non-null normal form in the
cache.

    [29 lines not shown]
DeltaFile
+3-8clang/test/CXX/temp/temp.constr/temp.constr.normal/p1.cpp
+3-7clang/lib/Sema/SemaConcept.cpp
+6-152 files

LLVM/project 876ff15 — llvm/lib/Target/Hexagon HexagonDepInstrInfo.td, llvm/test/MC/Hexagon v79-nontemporal.s

Add v79 scalar non-temporal forms. (#226017)

Also add support for the non-temporal forms of dcfetch and dczeroa.

Fixes https://github.com/llvm/llvm-project/issues/221485
DeltaFile
+214-0llvm/lib/Target/Hexagon/HexagonDepInstrInfo.td
+140-0llvm/test/MC/Hexagon/v79-nontemporal.s
+354-02 files

LLVM/project ef45ca7 — clang/test/CodeGen/RISCV rvp-intrinsics.c

[RISCV] Add common check-prefix to rvp-intrinsics.c. NFC (#227044)

A large number of test cases have the same IR for RV32 and RV64.
DeltaFile
+3,983-8,499clang/test/CodeGen/RISCV/rvp-intrinsics.c
+3,983-8,4991 files

LLVM/project 9fe649a — llvm/lib/CAS UnifiedOnDiskCache.cpp, llvm/test/tools/llvm-cas validate-if-needed.test

[CAS] Close the stderr temp file of the out-of-process validator (#226844)
DeltaFile
+3-4llvm/lib/CAS/UnifiedOnDiskCache.cpp
+7-0llvm/test/tools/llvm-cas/validate-if-needed.test
+10-42 files

LLVM/project e96d713 — mlir/lib/Dialect/LLVMIR/Transforms InlinerInterfaceImpl.cpp, mlir/test/Dialect/LLVMIR inlining-alias-scopes.mlir

[mlir][LLVM] Use a disjoint scope domain when inlining noalias

This matches recent changes to the LLVM inliner.

AI disclosure: Claude wrote the code, I wrote the commit message and
have done initial review.
DeltaFile
+28-25mlir/lib/Dialect/LLVMIR/Transforms/InlinerInterfaceImpl.cpp
+10-18mlir/test/Dialect/LLVMIR/inlining-alias-scopes.mlir
+38-432 files

LLVM/project 0a5eb58 — mlir/lib/Dialect/LLVMIR/Transforms InlinerInterfaceImpl.cpp

Update comment
DeltaFile
+2-2mlir/lib/Dialect/LLVMIR/Transforms/InlinerInterfaceImpl.cpp
+2-21 files

LLVM/project e99a787 — mlir/test/Dialect/LLVMIR inlining-alias-scopes.mlir

Test that the inliner keeps the flag when it clones a domain

Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
DeltaFile
+35-0mlir/test/Dialect/LLVMIR/inlining-alias-scopes.mlir
+35-01 files

LLVM/project 386fe8e — mlir/lib/Dialect/LLVMIR/Transforms InlinerInterfaceImpl.cpp

Remove pointless comment
DeltaFile
+0-2mlir/lib/Dialect/LLVMIR/Transforms/InlinerInterfaceImpl.cpp
+0-21 files

LLVM/project 7e9d23f — mlir/include/mlir/Dialect/LLVMIR LLVMAttrDefs.td, mlir/lib/Dialect/LLVMIR/Transforms InlinerInterfaceImpl.cpp

[mlir][LLVM] Add disjointScopes to AliasScopeDomainAttr

This also updates the MLIR-side inliner to clone disjoint domains
while cloning alias scopes, matching changes to LLVM.

AI disclosure: Claude wrote the code, I wrote the commit message and
looked at the code.
DeltaFile
+25-0mlir/test/Target/LLVMIR/Import/metadata-alias-scopes.ll
+23-0mlir/test/Target/LLVMIR/attribute-alias-scopes.mlir
+17-2mlir/include/mlir/Dialect/LLVMIR/LLVMAttrDefs.td
+4-1mlir/lib/Dialect/LLVMIR/Transforms/InlinerInterfaceImpl.cpp
+2-2mlir/lib/Target/LLVMIR/ModuleTranslation.cpp
+3-1mlir/lib/Target/LLVMIR/ModuleImport.cpp
+74-66 files

LLVM/project 9304616 — llvm/lib/Target/AMDGPU AMDGPULowerModuleLDSPass.cpp, llvm/test/CodeGen/AMDGPU remove-no-kernel-id-attribute.ll

[AMDGPU] Use a disjoint scope domain for merged LDS structs

When lowering LDS values, all the values are mutually disjoint, so we
can use the newly-added disjoint scopes feature to simplify the IR.

AI disclosure: Claude wrote this and I reviewed it and wrote the
 commit message
DeltaFile
+16-51llvm/lib/Target/AMDGPU/AMDGPULowerModuleLDSPass.cpp
+8-12llvm/test/CodeGen/AMDGPU/remove-no-kernel-id-attribute.ll
+24-632 files

LLVM/project e1480b4 — llvm/test/CodeGen/AMDGPU lower-kernel-and-module-lds.ll lower-module-lds-via-hybrid.ll

Test fixups
DeltaFile
+28-19llvm/test/CodeGen/AMDGPU/lower-lds-struct-aa-merge.ll
+36-10llvm/test/CodeGen/AMDGPU/lower-module-lds-precise-allocate-to-module-struct.ll
+17-17llvm/test/CodeGen/AMDGPU/lower-lds-struct-aa-memcpy.ll
+15-18llvm/test/CodeGen/AMDGPU/lower-lds-struct-aa.ll
+11-13llvm/test/CodeGen/AMDGPU/lower-module-lds-via-hybrid.ll
+9-9llvm/test/CodeGen/AMDGPU/lower-kernel-and-module-lds.ll
+116-865 files not shown
+148-12011 files

LLVM/project 8a90526 — llvm/lib/Target/AMDGPU AMDGPULowerKernelArguments.cpp, llvm/test/CodeGen/AMDGPU lower-kernargs.ll si-split-load-store-alias-info.ll

[AMDGPU] Use a disjoint scope domain for noalias kernel arguments

All noalias arguments of a kernel are disjoint with each other, so we
can use a disjoint scope to save on metadata construction.

AI disclosure: Claude wrote this, I looked at it and wrote this
message.
DeltaFile
+90-105llvm/test/CodeGen/AMDGPU/lower-noalias-kernargs.ll
+19-13llvm/lib/Target/AMDGPU/AMDGPULowerKernelArguments.cpp
+11-11llvm/test/CodeGen/AMDGPU/lower-kernel-arguments-noalias-call-no-ptr-args.ll
+8-8llvm/test/CodeGen/AMDGPU/si-split-load-store-alias-info.ll
+4-4llvm/test/CodeGen/AMDGPU/lower-kernargs.ll
+132-1415 files