LLVM/project 3f10d98 — llvm/lib/Transforms/InstCombine InstCombineCalls.cpp

[InstCombine][NFC] Use uint64_t for ExtractIdx in vector_extract fold (#225156)

## Summary
This changes `ExtractIdx` from `unsigned` to `uint64_t` in the
`Intrinsic::vector_extract` case.

    unsigned ExtractIdx = cast<ConstantInt>(Idx)->getZExtValue();

`ExtractIdx` is initialized from the extract index operand via
`getZExtValue()` , which returns `uint64_t`. In the
`get_active_lane_mask` sub-case, it is scaled and compared against the
mask's upper bound:

    if (ExtractIdx * ScaleFactor >= ALMUpperBound->getZExtValue())

Here `ALMUpperBound->getZExtValue()` returns `uint64_t`, and
`ScaleFactor` is `unsigned`.

With `ExtractIdx` declared `unsigned`, the value from `getZExtValue()`

    [9 lines not shown]
DeltaFile
+2-2llvm/lib/Transforms/InstCombine/InstCombineCalls.cpp
+2-21 files

LLVM/project 0408ea5 — llvm/lib/Target/X86 X86FastISel.cpp, llvm/test/CodeGen/X86 fast-isel-shift.ll

X86: Mark the EFLAGS def of FastISel selected shifts dead (#227060)

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+74-0llvm/test/CodeGen/X86/fast-isel-shift.ll
+2-1llvm/lib/Target/X86/X86FastISel.cpp
+76-12 files

LLVM/project d3429a7 — mlir/lib/Dialect/OpenACC/Transforms ACCCGToGPU.cpp, mlir/test/Dialect/OpenACC acc-cg-to-gpu-thread-private-storage-remark.mlir

[mlir][openacc] Update remark for privates (#227074)

Update the "thread-private" term to use "Local memory or registers" for
privates remark printing.

Co-authored-by: Yian Su <yians at nvidia.com>
DeltaFile
+3-3mlir/test/Dialect/OpenACC/acc-cg-to-gpu-thread-private-storage-remark.mlir
+1-1mlir/lib/Dialect/OpenACC/Transforms/ACCCGToGPU.cpp
+4-42 files

LLVM/project 374c5e7 — llvm/test/Transforms/SLPVectorizer/AArch64 trimmed-subtree-schedule-operands.ll

[SLP][NFC]Add a test with the regression after disabling scheduling of instructions, NFC



Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/227080
DeltaFile
+149-0llvm/test/Transforms/SLPVectorizer/AArch64/trimmed-subtree-schedule-operands.ll
+149-01 files

LLVM/project 156ba8d — llvm/lib/CodeGen/SelectionDAG LegalizeTypes.h LegalizeVectorTypes.cpp, llvm/test/CodeGen/AArch64 ldexp.ll

[SelectionDAG] Fix crash scalarizing `FLDEXP` with a legal `<1 x i1>` exponent (#224809)

Fixes #219695

When the result of an `FLDEXP` node is scalarized, we went through the
generic `ScalarizeVecRes_BinOp` handler, which asks for the scalarized
form of both operands. That only works when both operands share the
result's type action. `ldexp` is not a true binop: its exponent has its
own integer vector type, and under AVX-512 `<1 x i1>` is a legal mask
type. It was never scalarized, so it never made it into the table, and
the lookup hit `TableId should be non-zero`. It isn't specific to i1
either: on AArch64, `ldexp <1 x half>, <1 x i32>` hits the same assert
because `v1f16` is scalarized while `v1i32` is widened.

`FLDEXP` now has its own scalarization handler. The FP operand is
scalarized as before, while the exponent is legalized according to its
own type action: its scalarized value is used when one exists, otherwise
the single element is extracted from the vector. This is the same
approach `ScalarizeVecRes_UnaryOp` and `ScalarizeVecRes_SETCC` already
take for operands that don't need scalarizing, and mirrors what
`SplitVecRes_FPOp_MultiType` does for the split case.
DeltaFile
+40-0llvm/test/CodeGen/X86/ldexp.ll
+36-0llvm/test/CodeGen/AArch64/ldexp.ll
+21-1llvm/lib/CodeGen/SelectionDAG/LegalizeVectorTypes.cpp
+14-0llvm/test/CodeGen/X86/ldexp-avx512.ll
+1-0llvm/lib/CodeGen/SelectionDAG/LegalizeTypes.h
+112-15 files

LLVM/project 8959ec0 — lld/wasm InputFiles.cpp SyntheticSections.cpp, llvm/lib/Object WasmObjectFile.cpp

[WebAssembly] Add target architecture to object file format (#225979)

Store the target architecture string (`wasm32` or `wasm64`) in a new
`WASM_TARGET_ARCH` subsection in the `linking` custom section and in a
new `WASM_DYLINK_TARGET_ARCH` subsection in the `dylink.0` custom
section.

Previously, architecture detection for object files and shared libraries
relied on heuristics (such as whether memory64 was imported/defined or
presence of 64-bit relocations). For modules that do not access memory,
this resulted in wasm64 objects/shared libraries being incorrectly
detected as wasm32.

With this change:
- `WasmObjectWriter` emits the `WASM_TARGET_ARCH` subsection in
  `linking`.
- `wasm-ld` emits `WASM_DYLINK_TARGET_ARCH` when generating shared
  libraries and `WASM_TARGET_ARCH` when generating relocatable output
  (`-r`).

    [6 lines not shown]
DeltaFile
+9-9llvm/test/MC/WebAssembly/debug-info64.ll
+8-8llvm/test/MC/WebAssembly/debug-info.ll
+14-1llvm/lib/Object/WasmObjectFile.cpp
+14-0llvm/lib/ObjectYAML/WasmEmitter.cpp
+14-0lld/wasm/SyntheticSections.cpp
+2-6lld/wasm/InputFiles.cpp
+61-2448 files not shown
+131-2654 files

LLVM/project aecdab8 — llvm/lib/Target/X86 X86FastISel.cpp, llvm/test/CodeGen/X86 fast-isel-sext-dead-eflags.ll

X86: Mark EFLAGS dead on the NEG8r for sext from i1 in fast isel (#227058)

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+18-0llvm/test/CodeGen/X86/fast-isel-sext-dead-eflags.ll
+3-1llvm/lib/Target/X86/X86FastISel.cpp
+21-12 files

LLVM/project 8d1646e — llvm/test/Transforms/SLPVectorizer/AArch64 trimmed-subtree-schedule-operands.ll

[𝘀𝗽𝗿] initial version

Created using spr 1.3.7
DeltaFile
+149-0llvm/test/Transforms/SLPVectorizer/AArch64/trimmed-subtree-schedule-operands.ll
+149-01 files

LLVM/project 15d2ca3 — lldb/include/lldb lldb-enumerations.h, lldb/include/lldb/Interpreter CommandOptionArgumentTable.h

[lldb] Speed up evaluating breakpoint conditions by using DIL (#224740)

The goal of this patch is to speed up the evaluation of breakpoint
conditions. Similar to how UserExpression is used to evaluate the
condition expression, DIL lexes and parses the expression once, and then
only evaluates the AST tree on every breakpoint location hit. If DIL
fails at any step, the evaluation falls back to UserExpression, and DIL
doesn't make any new attempts on subsequent breakpoint hits. The
breakpoint default evaluation mode (DIL, UserExpression or DWIM) can be
changed by `target.breakpoints-condition-mode` setting. The mode can
also be changed separately for a specific breakpoint via command line
option `-Z (--condition-mode)` or by SB API
`SBBreakpoint::SetConditionMode`.
DeltaFile
+97-0lldb/test/API/functionalities/breakpoint/breakpoint_conditions/TestBreakpointConditions.py
+75-1lldb/source/Breakpoint/BreakpointLocation.cpp
+23-0lldb/source/API/SBBreakpoint.cpp
+16-0lldb/include/lldb/lldb-enumerations.h
+11-2lldb/source/Commands/Options.td
+10-0lldb/include/lldb/Interpreter/CommandOptionArgumentTable.h
+232-39 files not shown
+272-315 files

LLVM/project c069c0e — clang/lib/CodeGen CGHLSLBuiltins.cpp

[HLSL] Fix -Wunused-variable in #225519 (#227059)

Inline the variable names given the variable names don't add much
clarity and they aren't used anywhere outside of asserts.
DeltaFile
+3-4clang/lib/CodeGen/CGHLSLBuiltins.cpp
+3-41 files

LLVM/project 0403285 — llvm/lib/Target/X86 X86FastISel.cpp, llvm/test/CodeGen/X86 fast-isel-divrem-dead-defs.ll

X86: Mark the unused defs of FastISel selected div/rem dead (#227053)

The DIV/IDIV instructions define the quotient, the remainder and the
flags, but X86FastISel::X86SelectDivRem only reads one of the quotient
and remainder. Mark the rest dead when building the instruction instead
of relying on later recomputatios.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+79-0llvm/test/CodeGen/X86/fast-isel-divrem-dead-defs.ll
+16-7llvm/lib/Target/X86/X86FastISel.cpp
+95-72 files

LLVM/project 73014a8 — clang/docs ReleaseNotes.md, clang/lib/Sema TreeTransform.h

[Clang] Fix assertion when instantiating a matrix type with an invalid element type (#224331)

Fixes #202744

`Sema::BuildMatrixType` skips the element type check while the element
type is dependent, so `template <typename Y> using matrix_5_5 = Y
__attribute__((matrix_type(5, 5)));` gives a `ConstantMatrixType` whose
element is still unchecked. On instantiation,
`TreeTransform::RebuildConstantMatrixType` called
`ASTContext::getConstantMatrixType` directly, so nothing ever validated
the substituted type. `matrix_5_5<matrix_5_5<float>>` then hit the "need
a valid element type" assertion. The dependent-dimension case
(`matrix_type(R, C)`) was fine because `RebuildDependentSizedMatrixType`
already goes through `BuildMatrixType`.

`RebuildConstantMatrixType` now goes through `Sema::BuildMatrixType` as
well, the same way `RebuildExtVectorType` does for vectors: it takes the
attribute location and wraps the dimensions in integer literals. An
invalid instantiated element type now gets the usual `invalid matrix
element type` error instead of crashing, and the `_BitInt` element check
that this path also skipped is applied too.
DeltaFile
+31-0clang/test/SemaTemplate/matrix-type.cpp
+14-5clang/lib/Sema/TreeTransform.h
+4-0clang/docs/ReleaseNotes.md
+49-53 files

LLVM/project bbcb7e9 — llvm/include/llvm/Option ArgList.h, llvm/lib/Option ArgList.cpp

[Option] Allow subcommand names as positional arguments

ArgList::getSubCommand() treats every positional argument that matches
a subcommand name as a subcommand, and reports multiple subcommands if
more than one does. That rejects valid command lines where a later
positional argument happens to be spelled like a subcommand, e.g.

  git branch clone  # create a new branch named "clone"

Add an AllowSubCommandNamesAsPositionals parameter to getSubCommand().
When set, the first positional argument that matches a subcommand name
is the subcommand, and later ones are passed to HandleOtherPositionals
instead of being reported as multiple subcommands.
DeltaFile
+38-0llvm/unittests/Option/OptionSubCommandsTest.cpp
+7-4llvm/lib/Option/ArgList.cpp
+7-1llvm/include/llvm/Option/ArgList.h
+52-53 files

LLVM/project 8e0cefb — clang/test/CodeGen gvn-vectorization-pipeline.c, llvm/include/llvm/Transforms/Scalar GVN.h

- GVN now respects pipeline vectorization settings.
- Loop-access analysis rejects unsuitable indirect stores, retaining PRE.
DeltaFile
+294-2llvm/test/Transforms/GVN/PRE/loop-load-pre-vectorization.ll
+27-6llvm/lib/Transforms/Scalar/GVN.cpp
+14-0clang/test/CodeGen/gvn-vectorization-pipeline.c
+8-0llvm/include/llvm/Transforms/Scalar/GVN.h
+4-2llvm/lib/Passes/PassBuilderPipelines.cpp
+3-0llvm/test/Other/new-pm-print-pipeline.ll
+350-102 files not shown
+354-118 files

LLVM/project f7f538b — utils/bazel/llvm-project-overlay/lld/unittests BUILD.bazel

[bazel] Run lld unittests (#220340)
DeltaFile
+45-0utils/bazel/llvm-project-overlay/lld/unittests/BUILD.bazel
+45-01 files

LLVM/project 1d972c5 — clang/test/CIR/Transforms/abi-lowering indirect-byval.cir

[CIR] Update a test for the new fenv syntax

Assisted-by: Claude Code / Claude Opus 5.5
DeltaFile
+2-2clang/test/CIR/Transforms/abi-lowering/indirect-byval.cir
+2-21 files

LLVM/project 31ac095 — llvm/lib/IR ConstantRange.cpp, llvm/unittests/IR ConstantRangeTest.cpp

[ConstantRange] Compute exact no-wrap region w/o materializing CR. (#223969)

Inline logic into makeExactNoWrapRegion() so we do not need to construct
temporary constant ranges.

Together with using makeExecuteNoWrapRegion in SCEV, this improves
compile-time on SCEV-heavy workloads.

Analysis aided by Opus 5.

PR: https://github.com/llvm/llvm-project/pull/223969
DeltaFile
+42-3llvm/lib/IR/ConstantRange.cpp
+3-0llvm/unittests/IR/ConstantRangeTest.cpp
+45-32 files

LLVM/project c04cc6b — llvm/include/llvm/ProfileData SampleProfWriter.h SampleProf.h, llvm/lib/ProfileData SampleProfWriter.cpp

[ProfileData] Support merging MD5-based ProfileSymbolList (#226594)

This patch supports merging MD5-based ProfileSymbolList instances and
writing the merged result to an extensible binary profile.

#210235 introduced the MD5-based ProfileSymbolList section in the
Eytzinger layout, where profile merging was initially supported only
from strings to an MD5-based Eytzinger array.

This patch adds DenseSet<uint64_t> GUIDs to ProfileSymbolList to
accumulate 64-bit MD5 hashes when merging MD5-based symbol lists,
while keeping ColdGUIDTable (EytzingerTableSpan) for zero-copy lookups
during compilation.  collectGUIDs, contains, and size are updated to
query GUIDs when populated.

RFC:
https://discourse.llvm.org/t/rfc-faster-sample-profile-loading/90957/7

Assisted-by: Antigravity
DeltaFile
+45-14llvm/include/llvm/ProfileData/SampleProf.h
+38-0llvm/unittests/ProfileData/SampleProfTest.cpp
+11-4llvm/test/tools/llvm-profdata/profile-symbol-list.test
+3-1llvm/include/llvm/ProfileData/SampleProfWriter.h
+0-3llvm/lib/ProfileData/SampleProfWriter.cpp
+97-225 files

LLVM/project 7508964 — clang/docs ReleaseNotes.md

[Clang] Add missing release note entry in GH226753
DeltaFile
+5-0clang/docs/ReleaseNotes.md
+5-01 files

LLVM/project 4f70f18 — llvm/test/Analysis/CostModel/AArch64 sve-cast.ll cast.ll

[AArch64] Add extra bitcast cost coverage. NFC (#227048)
DeltaFile
+59-4llvm/test/Analysis/CostModel/AArch64/cast.ll
+16-16llvm/test/Analysis/CostModel/AArch64/sve-cast.ll
+75-202 files

LLVM/project 869989e — lldb/docs dil-expr-lang.ebnf, lldb/source/ValueObject DILParser.cpp

[lldb][NFC] Add missing composite assignments to DIL grammar (#226558)
DeltaFile
+3-0lldb/source/ValueObject/DILParser.cpp
+3-0lldb/docs/dil-expr-lang.ebnf
+6-02 files

LLVM/project 68ddb36 — llvm/test/CodeGen/ARM bitinsert-bitextract.ll, llvm/test/CodeGen/RISCV bitinsert-bitextract-fp.ll bitinsert-bitextract.ll

Merge commit 'e40e0bc36b14' into cir-callconv-byval-slot-copy
DeltaFile
+6,086-6,026llvm/test/CodeGen/RISCV/rvv/expandload.ll
+3,913-3,252llvm/test/CodeGen/RISCV/rvv/fixed-vectors-masked-gather.ll
+3,098-2,506llvm/test/CodeGen/RISCV/rvv/fixed-vectors-masked-scatter.ll
+4,294-0llvm/test/CodeGen/RISCV/bitinsert-bitextract.ll
+3,321-0llvm/test/CodeGen/ARM/bitinsert-bitextract.ll
+1,989-0llvm/test/CodeGen/RISCV/bitinsert-bitextract-fp.ll
+22,701-11,7843,614 files not shown
+139,591-57,7003,620 files

LLVM/project 111f1d4 — clang/include/clang/AST ASTConcept.h, clang/lib/Sema SemaConcept.cpp

[clang] Migrate away from PointerUnion::dyn_cast (NFC) (#226653)

Note that PointerUnion::dyn_cast has been soft deprecated in
PointerUnion.h:

  // FIXME: Replace the uses of is(), get() and dyn_cast() with
  //        isa<T>, cast<T> and the llvm::dyn_cast<T>

Literal migration would result in dyn_cast_if_present (see the
definition of PointerUnion::dyn_cast), but this patch uses dyn_cast on
UnsatisfiedConstraintRecord because it is always nonnull.
Specifically, ConstraintSatisfaction::Details only receives nonnull
pointers in the following places:

- ASTNodeImporter::ImportConstraintSatisfaction
- ConstraintSatisfactionChecker::consumeSFINAEFailure
- ConstraintSatisfactionChecker::EvaluateSlow
- ConstraintSatisfactionChecker::Evaluate
- readConstraintSatisfaction

Assisted-by: Antigravity
DeltaFile
+1-3clang/lib/Sema/SemaConcept.cpp
+1-1clang/include/clang/AST/ASTConcept.h
+2-42 files

LLVM/project d1133c1 — lldb/source/Plugins/TypeSystem CMakeLists.txt, lldb/source/Plugins/TypeSystem/Clike CMakeLists.txt TypeSystemClike.h

[lldb][Clike][NFC] Add the TypeSystemClike plugin skeleton (#222611)

Introduces an empty TypeSystem plugin infrastructure for the new
TypeSystemClike.

For the RFC with more information and background, see
https://discourse.llvm.org/t/rfc-a-faster-more-reliable-type-system-for-c-languages/91459
and #225371

assisted-by: claude
DeltaFile
+478-0lldb/source/Plugins/TypeSystem/Clike/TypeSystemClike.cpp
+200-0lldb/source/Plugins/TypeSystem/Clike/TypeSystemClike.h
+10-0lldb/source/Plugins/TypeSystem/Clike/CMakeLists.txt
+1-0lldb/source/Plugins/TypeSystem/CMakeLists.txt
+689-04 files

LLVM/project e40e0bc — llvm/test/tools/llubi intr_speculative_load.ll intr_speculative_load_ub.ll, llvm/tools/llubi/lib Interpreter.cpp

[llubi] Add support for llvm.speculative.load. (#226839)

Implement the direct form of llvm.speculative.load, where the number of
accessible bytes N is passed as an i64. Only the N accessible bytes are
read from memory and they must be in bounds of the underlying object;
all other bytes are poison. With from_end, the accessible bytes are the
last N bytes of the loaded value. It is UB if N is poison or exceeds the
size of the loaded type.

Support for the oracle form will be added as follow-up.

PR: https://github.com/llvm/llvm-project/pull/226839
DeltaFile
+86-0llvm/test/tools/llubi/intr_speculative_load_ub.ll
+56-0llvm/test/tools/llubi/intr_speculative_load.ll
+48-0llvm/tools/llubi/lib/Interpreter.cpp
+190-03 files

LLVM/project 26033a3 — clang/test/Driver riscv64-toolchain-extra.c riscv32-toolchain-extra.c

[clang][test] Pass -resource-dir in riscv*-toolchain-extra.c tests (#226952)

Fixes issue reported in
https://github.com/llvm/llvm-project/pull/216996#issuecomment-5846218483

These tests check that, with no GCC installation, the driver finds the
linker, crt0 and sysroot relative to its own bin/ directory. The
compiler-rt paths come from the resource directory, which the tests did
not specify, so they relied on the build's CLANG_RESOURCE_DIR placing it
within the fake toolchain tree.

With a relative CLANG_RESOURCE_DIR containing several "..", e.g.
../../../../lib/clang/24 as used by Gentoo, the resource directory
resolves outside of the test tree. This previously passed only because
the unnormalized path still contained "riscv64-nogcc/". Since
99988429d395 removed the dots from the path, the checks fail.

Pass an explicit -resource-dir within the test tree so the tests don't
depend on the build configuration.

Claude Code helped with this.
DeltaFile
+3-3clang/test/Driver/riscv64-toolchain-extra.c
+3-3clang/test/Driver/riscv32-toolchain-extra.c
+6-62 files

LLVM/project f32a45c — clang/test/CodeGen/AArch64/sve dup.c len.c

[clang][CIR] Add missing return stmt in tests (#226534)

This PR makes sure that tests in:
  * clang/test/CodeGen/AArch64/sve/

follow the format previously establised in:
  * clang/test/CodeGen/AArch64/sve-intrinsics/
DeltaFile
+12-0clang/test/CodeGen/AArch64/sve/len.c
+3-3clang/test/CodeGen/AArch64/sve/dup.c
+15-32 files

LLVM/project 9b9885d — llvm/lib/Transforms/InstCombine InstCombineSelect.cpp, llvm/test/Transforms/InstCombine truncating-saturate.ll

[InstCombine] Handle sext/zext of i1 "selects" in canonicalizeClampLike
DeltaFile
+17-25llvm/test/Transforms/InstCombine/truncating-saturate.ll
+3-3llvm/lib/Transforms/InstCombine/InstCombineSelect.cpp
+20-282 files

LLVM/project e8ba6bd — llvm/test/Transforms/InstCombine truncating-saturate.ll

Precommit test
DeltaFile
+34-0llvm/test/Transforms/InstCombine/truncating-saturate.ll
+34-01 files

LLVM/project fd7df11 — llvm/lib/Transforms/Scalar GVN.cpp, llvm/test/Transforms/GVN/PRE loop-load-pre-vectorization.ll

[GVN] Preserve vectorization opportunities when PREing loop loads

Loop load PRE can replace an invariant-address load with a loop-carried
PHI and conditional reload after a may-alias store. That scalar recurrence
can prevent vectorization even when runtime alias checks could disambiguate
the original accesses. Subsequent full unrolling then expands scalar code.

Conservatively preserve the header load in innermost loops whose clobber
is a conditional may-alias store through a varying pointer. Keep existing
PRE behavior for invariant-address clobbers, known aliasing, calls, ordered
memory operations, and loops with vectorization disabled or completed.
This is an opportunity heuristic, not a vectorization legality proof.
DeltaFile
+582-0llvm/test/Transforms/GVN/PRE/loop-load-pre-vectorization.ll
+55-0llvm/lib/Transforms/Scalar/GVN.cpp
+637-02 files