LLVM/project af80367 — lldb/source/Core CMakeLists.txt ModuleList.cpp, lldb/source/Plugins/TypeSystem/Clang CMakeLists.txt TypeSystemClang.cpp

[lldb] Set the default clang module cache path in TypeSystemClang
DeltaFile
+7-0lldb/source/Plugins/TypeSystem/Clang/TypeSystemClang.cpp
+0-6lldb/source/Core/ModuleList.cpp
+0-3lldb/source/Core/CMakeLists.txt
+1-0lldb/source/Plugins/TypeSystem/Clang/CMakeLists.txt
+8-94 files

LLVM/project f88e9e9 — clang-tools-extra/clang-tidy/modernize UseNullptrCheck.cpp, clang-tools-extra/docs ReleaseNotes.md

[clang-tidy] Fix modernize-use-nullptr false positive on ordering comparisons (#225585)

libstdc++ 16 renamed __cmp_cat::__unspec to __cmp_cat::__literal_zero.
Add the new name to the default IgnoredTypes.

Fixes #206245
DeltaFile
+21-0clang-tools-extra/test/clang-tidy/checkers/modernize/use-nullptr-cxx20.cpp
+6-2clang-tools-extra/clang-tidy/modernize/UseNullptrCheck.cpp
+5-0clang-tools-extra/docs/ReleaseNotes.md
+1-1clang-tools-extra/docs/clang-tidy/checks/modernize/use-nullptr.rst
+33-34 files

LLVM/project 57a4fbb — clang/lib/CIR/CodeGen CIRGenExprAggregate.cpp, clang/test/CIR/CodeGen agg-init-constexpr.cpp no-unique-address-empty-consteval.cpp

[CIR] Fix consteval store of no_unqiue_address_members (#229478)

We were unconditionally overwriting the members of a consteval type,
   even if there was padding / no_unique_address in the way.  This patch
   corrects that behavior by picking up the relevant parts of #93115
   which ensures we only write the constexpr initialized fields.
DeltaFile
+128-0clang/test/CIR/CodeGen/no-unique-address-empty-consteval.cpp
+38-3clang/lib/CIR/CodeGen/CIRGenExprAggregate.cpp
+9-5clang/test/CIR/CodeGen/agg-init-constexpr.cpp
+175-83 files

LLVM/project 3c03a42 — flang/lib/Optimizer/CodeGen LowerRepackArrays.cpp, flang/test/Integration cold_array_repacking.f90

[flang] Copy back only modified data when unpacking repacked arrays (#228543)

With `-frepack-arrays`, `fir.unpack_array` copied the contiguous
temporary back into the original array unconditionally. Since #222986, a
named constant's own storage can be the actual argument, so the
copy-back stored into read-only memory and the program crashed, even
though the callee never modified the data. For example:

```fortran
module callees
contains
  subroutine take(a)
    integer :: a(:)
    print *, sum(a)
  end subroutine
end module
program repack_const
  use callees
  integer, parameter :: c(3) = [1, 0, 2]

    [13 lines not shown]
DeltaFile
+45-30flang/test/Transforms/lower-repack-arrays.fir
+16-6flang/lib/Optimizer/CodeGen/LowerRepackArrays.cpp
+1-1flang/test/Integration/cold_array_repacking.f90
+62-373 files

LLVM/project 93f6284 — llvm/lib/Target/AArch64 AArch64ISelLowering.cpp AArch64FastISel.cpp, llvm/test/CodeGen/AArch64 ptrauth-isel.ll

AArch64: Stop setting kill flags on virtual registers before FinalizeISel (#229021)

There is no point in maintaining kill flags before register allocation anymore.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+10-10llvm/lib/Target/AArch64/AArch64FastISel.cpp
+3-3llvm/test/CodeGen/AArch64/ptrauth-isel.ll
+2-2llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
+15-153 files

LLVM/project a43c10f — clang/lib/Interpreter InterpreterValuePrinter.cpp, clang/test/Interpreter value-print-temporaries.cpp

[clang-repl] Keep the cleanups of value-printed expressions (#229675)

convertExprToValue() strips the ExprWithCleanups of the printed
expression and builds the call of __clang_Interpreter_SetValue* around
the subexpression, without the cleanups. The temporaries of the
expression then belong to no full-expression: CodeGen destroys them at
the end of the function that runs the top-level statements, which it
finishes after it has emitted the deferred declarations, so their
destructors are referenced but never emitted:

  clang-repl> struct S { ~S() {} };
  clang-repl> int f(S) { return 42; }
  clang-repl> f(S())
  JIT session error: Symbols not found: [ _ZN1SD1Ev ]

(with an assertions build, moveLazyEmissionStates() asserts instead).

Put the cleanups back around the result, so that the temporaries are
destroyed at the end of their statement. If the call cannot be built,

    [2 lines not shown]
DeltaFile
+37-0clang/test/Interpreter/value-print-temporaries.cpp
+14-4clang/lib/Interpreter/InterpreterValuePrinter.cpp
+51-42 files

LLVM/project f74b923 — llvm/lib/Target/NVPTX NVPTXISelLowering.cpp, llvm/test/CodeGen/NVPTX atomicrmw-ignore-denormal-mode.ll

[NVPTX] Honor !atomic.ignore.denormal.mode on atomicrmw fadd (#217586)

PTX atom.add has a fixed denormal behavior that the program cannot
control: atom.add.f32 flushes denormals on global memory but not on
shared, and atom.add.f16 never flushes. When that disagrees with the
function's denormal mode, the backend expands the atomic into a CAS loop
so the denormal behavior is preserved.

!atomic.ignore.denormal.mode says the denormal behavior of this
particular atomic does not matter, so use the native instruction even
when it disagrees. This is the same thing -nvptx-allow-ftz-atomics does,
except per-instruction instead of per-compilation, which lets a frontend
opt in only the operations it knows about -- notably CUDA's atomicAdd(),
which is defined in terms of atom.add.

Note that -nvptx-allow-ftz-atomics defaults to true, so the new behavior
is only observable with -nvptx-allow-ftz-atomics=false.

Co-authored-by: Artem Belevich <tra at google.com>
DeltaFile
+261-0llvm/test/CodeGen/NVPTX/atomicrmw-ignore-denormal-mode.ll
+11-3llvm/lib/Target/NVPTX/NVPTXISelLowering.cpp
+272-32 files

LLVM/project b31cb61 — llvm/test/tools/llubi lib_printf_type_mismatch.ll lib_printf_float_type.ll, llvm/tools/llubi/lib Library.cpp

[llubi] Validate argument types in `printf` (#229702)

Fix the crash when a `printf` argument has the wrong value kind for its
format specifier.
DeltaFile
+82-0llvm/test/tools/llubi/lib_printf_float_type.ll
+33-0llvm/test/tools/llubi/lib_printf_type_mismatch.ll
+14-0llvm/tools/llubi/lib/Library.cpp
+129-03 files

LLVM/project 9e43787 — clang/lib/StaticAnalyzer/Core ExprEngine.cpp

[NFC][analyzer] Allow PreStmt/PostStmt callbacks for unhandled statement kinds (#224054)

The function `shouldJustCallCheckers`, introduced recently in
e829049823905405779c6b908437b200e7ed1b2e, returns true by default and
returns false for statement kinds where it is inappropriate to call the
PreStmt and PostStmt checkers in the "normal" pattern.

The switch in `ExprEngine::Visit` starts with a large block of statement
kinds that are not evaluated in any meaningful way, i.e. they just pass
through the analyzer without modifying the analysis state.

Previously `shouldJustCallCheckers` returned false for these statement
kinds (to preserve the default behavior of the old implementation), but
this change moves them to the current default, which simplifies the code
and enables checker writers to target these statement kinds with PreCall
and PostCall callbacks. As currently there are no non-debug checkers
that target these (rare) statement kinds, this is an NFC change.
DeltaFile
+0-134clang/lib/StaticAnalyzer/Core/ExprEngine.cpp
+0-1341 files

LLVM/project c62a1a5 — mlir/include/mlir/Dialect/Linalg/Transforms Hoisting.h, mlir/lib/Dialect/Linalg/Transforms Hoisting.cpp

[mlir][linalg] Hoist masked vector transfer pairs (#221481)

This PR completes #220244.

Extend `hoistRedundantVectorTransfers` to hoist `vector.mask`-wrapped
transfer pairs (and lone reads), moving the whole mask op to the loop
boundary and pairing a masked read only with a masked write carrying the
same mask.

Assisted-by: Claude Code (Anthropic)
DeltaFile
+304-0mlir/test/Dialect/Linalg/hoist-redundant-vector-transfers.mlir
+98-12mlir/lib/Dialect/Linalg/Transforms/Hoisting.cpp
+6-0mlir/include/mlir/Dialect/Linalg/Transforms/Hoisting.h
+408-123 files

LLVM/project 1d3c9f0 — llvm/lib/Target/AArch64 AArch64ISelDAGToDAG.cpp, llvm/lib/Target/AArch64/GISel AArch64InstructionSelector.cpp

[GlobalISel] [AArch64] Generate MVN for shift by (N - X) in SelectShiftMask (#226273)

AArch64 shift instructions only use the low log2(ShiftWidth) bits of the
shift amount. When shifting by N-X where N == -1 mod ShiftWidth, this is
equivalent to shifting by ~X, so we can generate a NOT (MVN) instead of
a SUB from a constant.

For example, shl x, (sub 63, y) becomes shl x, mvn(y).

This implements both SelectionDAG and GlobalISel. The SelectionDAG side
uses ORNWrr/ORNXrr with the zero register; the GlobalISel side builds
the same in the ComplexPattern renderer and constrains its operands with
constrainSelectedInstRegOperands.

Part of the incremental work tracked in #224245 to replace
tryShiftAmountMod with ComplexPattern-based isel.
DeltaFile
+24-0llvm/lib/Target/AArch64/GISel/AArch64InstructionSelector.cpp
+23-0llvm/test/CodeGen/AArch64/shift-mod.ll
+20-0llvm/lib/Target/AArch64/AArch64ISelDAGToDAG.cpp
+67-03 files

LLVM/project 270bcfa — llvm/test/CodeGen/AMDGPU gds-load-store.ll, llvm/test/CodeGen/AMDGPU/GlobalISel legalize-load-constant.mir legalize-load-flat.mir

Merge branch 'users/nicebert/clang-openmp-no-loop-core' into users/nicebert/offload-openmp-no-loop-default-grid
DeltaFile
+3,103-3,156llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-private.mir
+3,245-2,647llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-local.mir
+2,490-2,634llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-global.mir
+4,568-0llvm/test/CodeGen/AMDGPU/gds-load-store.ll
+2,148-2,222llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-flat.mir
+1,812-1,908llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-constant.mir
+17,366-12,5671,094 files not shown
+53,367-23,6911,100 files

LLVM/project cd6eae7 — llvm/lib/Target/AMDGPU AMDGPULegalizerInfo.cpp, llvm/test/CodeGen/AMDGPU/GlobalISel legalize-fmul.mir legalize-fadd.mir

[AMDGPU][GlobalISel] Legalize bf16 for fadd, fmul, fma, fsub, fmin, fmax, and fcanonicalize (#227901)

Legalize bf16 operations for fadd, fmul, fma, fsub, fmin, fmax, and
fcanonicalize including strict variants.

Promote bf16 to v2bf16 for GFX1250+.  Otherwise expand to f32.

Enable testing for GFX9, GFX12 for fmin, fmax, and fsub which was
blocked by a lack of bf16 support.

---------

Signed-off-by: John Lu <John.Lu at amd.com>
DeltaFile
+970-0llvm/test/CodeGen/AMDGPU/GlobalISel/fminnum.bf16.ll
+590-2llvm/test/CodeGen/AMDGPU/GlobalISel/fsub.bf16.ll
+7-30llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-fma.mir
+6-26llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-fmul.mir
+6-26llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-fadd.mir
+23-9llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
+1,602-939 files not shown
+1,732-19515 files

LLVM/project 115061f — llvm/test/CodeGen/AMDGPU gds-load-store.ll, llvm/test/CodeGen/AMDGPU/GlobalISel legalize-load-constant.mir legalize-load-flat.mir

Merge branch 'users/nicebert/clang-openmp-no-loop-core' into users/nicebert/offload-openmp-no-loop-default-grid
DeltaFile
+3,103-3,156llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-private.mir
+3,245-2,647llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-local.mir
+2,490-2,634llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-global.mir
+4,568-0llvm/test/CodeGen/AMDGPU/gds-load-store.ll
+2,148-2,222llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-flat.mir
+1,812-1,908llvm/test/CodeGen/AMDGPU/GlobalISel/legalize-load-constant.mir
+17,366-12,5671,093 files not shown
+53,367-23,5531,099 files

LLVM/project d66ab4b — llvm/lib/Target/SPIRV SPIRVNonSemanticDebugHandler.cpp, llvm/test/CodeGen/SPIRV/debug-info debug-scope-lexical-block-file.ll

[SPIR-V] Fix DebugScope resolving DILexicalBlockFile to DebugCompilationUnit (#228259)

resolveScope() did not handle DILexicalBlockFile, which carries
file/discriminator info. Its SPIR-V counterpart
(DebugLexicalBlockDiscriminator) cannot be used as a scope, so
resolveScope() must unwrap it. When
emitDebugScopeForInstruction() encountered a DILocation whose scope
was a DILexicalBlockFile (wrapping a DISubprogram), resolveScope() fell
through to the compile-unit fallback. This emitted a DebugScope
referencing DebugCompilationUnit instead of DebugFunction, producing
incorrect debug scope resolution.

Fix by unwrapping DILexicalBlockFile to its underlying scope at the top
of resolveScope().

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply at anthropic.com>
DeltaFile
+67-0llvm/test/CodeGen/SPIRV/debug-info/debug-scope-lexical-block-file.ll
+4-0llvm/lib/Target/SPIRV/SPIRVNonSemanticDebugHandler.cpp
+71-02 files

LLVM/project f560269 — clang/test/OpenMP target_loop_promotion.cpp target_teams_generic_loop_codegen_as_parallel_for.cpp

[clang][OpenMP] Add strided-loop SPMD kernel promotion

A target teams distribute parallel for is only promoted when every
iteration is guaranteed a thread. Without the oversubscription
assumptions, or with a num_teams clause, it is emitted as nested
worksharing loops, while Flang runs the same region as a grid-stride
loop.

Promote these kernels too, so the loop body strides across the grid,
and tag them SPMD_STRIDED_LOOP so the launch can be told apart from
plain SPMD. This allows teams loop and target teams loop to always get
optimized. Since those can carry lastprivate(conditional:), the
promoted path has to track it.
DeltaFile
+3,247-5,702clang/test/OpenMP/nvptx_SPMD_codegen.cpp
+664-1,243clang/test/OpenMP/nvptx_target_teams_generic_loop_codegen.cpp
+595-1,174clang/test/OpenMP/nvptx_target_teams_distribute_parallel_for_codegen.cpp
+348-658clang/test/OpenMP/nvptx_target_teams_distribute_parallel_for_simd_codegen.cpp
+209-343clang/test/OpenMP/target_teams_generic_loop_codegen_as_parallel_for.cpp
+395-0clang/test/OpenMP/target_loop_promotion.cpp
+5,458-9,12019 files not shown
+5,839-9,90525 files

LLVM/project 8a34b53 — llvm/test/CodeGen/PowerPC aix-ifunc-toc-restore-query-neg.ll

PPC: Fix unnecessary host requirement in test (#229749)
DeltaFile
+6-7llvm/test/CodeGen/PowerPC/aix-ifunc-toc-restore-query-neg.ll
+6-71 files

LLVM/project 0fee4dd — clang/test/CIR/CodeGenCoroutines coro-agg.cpp, libcxx/test/std/ranges/range.adaptors/range.enumerate/sentinel equal.pass.cpp

Merge branch 'main' into users/arsenm/runtime-libcalls/attribute-defaultcc-reference-library
DeltaFile
+452-561llvm/test/CodeGen/AMDGPU/slp-int-to-fp.ll
+518-446llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+360-0clang/test/CIR/CodeGenCoroutines/coro-agg.cpp
+89-153llvm/test/CodeGen/AArch64/predicate-as-counter-loop-rewrites.ll
+122-90llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+178-33libcxx/test/std/ranges/range.adaptors/range.enumerate/sentinel/equal.pass.cpp
+1,719-1,283122 files not shown
+3,786-1,944128 files

LLVM/project 330d833 — clang-tools-extra/clang-tidy/bugprone EasilySwappableParametersCheck.cpp, clang-tools-extra/docs ReleaseNotes.md

[clang-tidy] Fix empty notes in bugprone-easily-swappable-parameters (#229753)

With ModelImplicitConversions enabled, a pair of parameters involving a
type alias whose underlying types are only related through an implicit
conversion has no common type. The check still emitted the "after
resolving type aliases" note for such pairs, but with an empty message.
This commit fixes the problem.
DeltaFile
+14-0clang-tools-extra/test/clang-tidy/checkers/bugprone/easily-swappable-parameters-implicits.cpp
+7-5clang-tools-extra/clang-tidy/bugprone/EasilySwappableParametersCheck.cpp
+5-0clang-tools-extra/docs/ReleaseNotes.md
+26-53 files

LLVM/project 5aacef3 — llvm/lib/Target/PowerPC PPCInstrInfo.cpp, llvm/test/CodeGen/PowerPC sext_elimination.mir

PPC: Fix EXTSW elimination promoting a subregister operand (#208051)

promoteInstr32To64ForElimEXTSW copies operands from the 32-bit
instruction verbatim into its promoted 64-bit form. When an operand
reads the sub_32 subregister of a 64-bit register, the promoted
instruction (which takes a full register) ended up with an illegal
subregister use and failed the machine verifier.

Drop the sub_32 subregister and use the original full register, which
provides the low 32 bits the promoted instruction operates on. This
avoids  verifier error regressions in a future change.

Co-authored-by: Claude Opus 4.8 <noreply at anthropic.com>
DeltaFile
+36-10llvm/test/CodeGen/PowerPC/sext_elimination.mir
+12-0llvm/lib/Target/PowerPC/PPCInstrInfo.cpp
+48-102 files

LLVM/project 3da3149 — llvm/lib/Target/ARM ARMBaseRegisterInfo.cpp, llvm/test/CodeGen/ARM inline-asm-clobber.ll

[ARM] don't warn when clobbering dregs (#229157)

These registers only exist with the `+d32` feature, and are reserved
without it. But a function without that feature can be inlined into a
function with the feature, so it is reasonable that front ends mark
these registers as clobbered regardless of whether the feature is
available.

This change is in line with e.g. x86_64 just ignoring a clobber of
`zmm0` even when `avx512` is not enabled. There are some other overrides
of `isAsmClobberable` that handle adjacent cases but none (that I could
find) that specifically handle target features. Still this looks like a
natural place to put this logic.

This came up in 

- https://github.com/rust-lang/rust/issues/163490
DeltaFile
+15-0llvm/test/CodeGen/ARM/inline-asm-clobber.ll
+11-0llvm/lib/Target/ARM/ARMBaseRegisterInfo.cpp
+26-02 files

LLVM/project 0f8f16d — llvm/lib/Target/AArch64 AArch64ISelLowering.cpp AArch64FastISel.cpp, llvm/test/CodeGen/AArch64 ptrauth-isel.ll

AArch64: Stop setting kill flags on virtual registers before FinalizeISel

There is no point in maintaining kill flags before register allocation
anymore.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+10-10llvm/lib/Target/AArch64/AArch64FastISel.cpp
+3-3llvm/test/CodeGen/AArch64/ptrauth-isel.ll
+2-2llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
+15-153 files

LLVM/project a1dd1ea — clang/lib/CIR/CodeGen CIRGenModule.cpp, clang/test/CIR/CodeGen global-init-order.cpp

[CIR] Correct globals-with-ctor/dtor init ordering (#229608)

LoweringPrepare emits our globals in the order of CIR-declaration.
However, we emit them in order of first-reference, not definition, so
the order of the inits was not correct.

This patch ensures that any global where it matters (that is, has a
ctor/dtor) is re-ordered to be in the order of definition, by moving it
to the 'end' of the list when we try to emit its definition.

The result is that the order of reference (such as if a declaration is
    referenced from another global!) doesn't cause us to get the init
order/dependencies wrong.
DeltaFile
+59-0clang/test/CIR/CodeGen/global-init-order.cpp
+18-1clang/lib/CIR/CodeGen/CIRGenModule.cpp
+77-12 files

LLVM/project 192e33e — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Analysis/CostModel/AMDGPU narrow-int-to-bfloat.ll narrow-int-to-fp.ll

[AMDGPU] Price scalar integer to fp casts by source width and sign (#225343)

Scalar sources between a byte and 31 bits fell to the default cost of one
while the matching vector lanes were already priced, which skewed the
difference SLP weighs a bundle against. Such a source is extended before
the conversion, and what the extension takes depends on the width, on
the sign and on whether the subtarget has SDWA and 16 bit instructions.
Sources narrower than a byte are left alone, because their vector form
is not priced either.


Assisted-by: Claude Code Opus 5
DeltaFile
+452-561llvm/test/CodeGen/AMDGPU/slp-int-to-fp.ll
+518-446llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+122-90llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+56-10llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+1,148-1,1074 files

LLVM/project 8db65ef — llvm/lib/Target/Mips MipsInstrInfo.cpp, llvm/test/CodeGen/Mips branch-dead-at.mir

Mips: Mark the assembler temporary def dead in inserted branches (#229690)

The branch instructions reserve $at for long branch expansion, which only uses it 
as a scratch. Instruction selection already marks the def dead; branches rebuilt by 
later passes did not, which blocked BranchFolding from hoisting common code 
past them.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+108-0llvm/test/CodeGen/Mips/branch-dead-at.mir
+11-4llvm/lib/Target/Mips/MipsInstrInfo.cpp
+3-3llvm/test/CodeGen/Mips/llvm-ir/forbidden-slot-ir.ll
+1-1llvm/test/CodeGen/Mips/Fast-ISel/branch-dead-at.ll
+123-84 files

LLVM/project 62c1f75 — offload/plugins-nextgen/level_zero/include L0Kernel.h L0Plugin.h, offload/plugins-nextgen/level_zero/src L0Kernel.cpp

[Offload][L0] Include context allocations in kernel indirect access flags (#229384)

#224930 moved L0-backed liboffload allocations into the per-context
allocators of `LevelZeroPluginContextTy`. However,
`L0KernelTy::setIndirectFlags` was never updated to count flags from
these new allocators, and as a result `olMemAlloc` allocations no longer
count towards kernel indirect access flags. This is significant for
libsycl - for instance, a kernel lambda is passed there as a single
byval struct, and therefore all captured USM pointers are necessarily
indirect access.

This bug was revealed by a buildbot failure when
https://github.com/llvm/llvm-project/pull/227589 was merged and later
reverted. The libsycl tests on the buildbot (Arc B580) only passed
because the counters in `L0ContextTy::HostMemAllocator` are
uninitialized and happened to contain leftover heap data, which set
SHARED by accident. #227589 changed the size of L0ContextTy, the
counters came out as zero, and 15 libsycl tests failed. That PR was
reverted in #228485. Its change was correct and can reland once this one

    [9 lines not shown]
DeltaFile
+27-0offload/unittests/OffloadAPI/kernel/olLaunchKernel.cpp
+12-0offload/unittests/OffloadAPI/device_code/byval_ptr.cpp
+10-0offload/plugins-nextgen/level_zero/include/L0Plugin.h
+6-1offload/plugins-nextgen/level_zero/src/L0Kernel.cpp
+4-1offload/plugins-nextgen/level_zero/include/L0Kernel.h
+2-0offload/unittests/OffloadAPI/device_code/CMakeLists.txt
+61-21 files not shown
+62-37 files

LLVM/project 7ca0478 — llvm/lib/IR Verifier.cpp, llvm/test/Verifier invalid-vp-intrinsics.ll insert-extract-intrinsics-invalid.ll

[Verifier] Check intrinsic signature validity before handwritten code (#229716)

Some handwritten code in `visitIntrinsicCall` assumes the parameter is
always a vector. However, the signature check is only performed on the
intrinsic declaration, while `visitIntrinsicCall` runs on the call site.
This causes the verifier to crash first without any diagnostic message.

This patch calls `isSignatureValid` before handwritten code, so the
logic below can safely assume some basic properties guarded by
constraints in .td files. When the check fails, it exits silently, as
the error message is already reported at the declaration.

Some dyn_cast/type checking code below is also simplified/eliminated.

This was discovered while I asked the LLM to hack llubi. It created some
inputs with an invalid intrinsic like `llvm.vscale.v2i64`, bypassing the
verifier... A follow-up patch will strengthen type constraints in
Intrinsics.td.


    [4 lines not shown]
DeltaFile
+11-42llvm/lib/IR/Verifier.cpp
+23-0llvm/test/Verifier/insert-extract-intrinsics-invalid.ll
+10-0llvm/test/Verifier/invalid-vp-intrinsics.ll
+44-423 files

LLVM/project a413043 — llvm/lib/Target/AArch64 AArch64PredicateAsCounterLoopRewrites.cpp

Use CreateIntrinsic
DeltaFile
+11-14llvm/lib/Target/AArch64/AArch64PredicateAsCounterLoopRewrites.cpp
+11-141 files

LLVM/project a44cd9b — llvm/lib/Target/AArch64 AArch64PredicateAsCounterLoopRewrites.cpp, llvm/test/CodeGen/AArch64 predicate-as-counter-loop-rewrites.ll

[AArch64] Use multi-vector intrinsics for masked load/store users of predicate-as-counter

If the user of the original wide mask is a masked load or store
intrinsic (matching the element size of the predicate-as-counter),
rewrite it directly to a masked multi-vector load/store.

This avoids materializing the vector mask and is easier to handle here
than later (e.g. in SelectionDAG), since we do not need to match the
concatenation of all `pext` segments of the predicate-as-counter.

Assisted-by: Codex
DeltaFile
+87-137llvm/test/CodeGen/AArch64/predicate-as-counter-loop-rewrites.ll
+110-1llvm/lib/Target/AArch64/AArch64PredicateAsCounterLoopRewrites.cpp
+197-1382 files

LLVM/project 9ee0c90 — llvm/lib/Target/AArch64 AArch64PredicateAsCounterLoopRewrites.cpp, llvm/test/CodeGen/AArch64 predicate-as-counter-loop-rewrites.ll

[AArch64] Avoid materializing full masks for extractelement users of predicate-as-counter (#220961)

If the user of the original wide mask is an `extractelement` and the
index is known to be within the first segment of the
predicate-as-counter, replace it with `extractelement(pext(counter,
0))`.

This avoids materializing the vector mask and produces a form that can
be folded into a conditional branch when the predicate-as-counter is
produced by a `whilelo`.

Assisted-by: Codex
DeltaFile
+89-153llvm/test/CodeGen/AArch64/predicate-as-counter-loop-rewrites.ll
+37-0llvm/lib/Target/AArch64/AArch64PredicateAsCounterLoopRewrites.cpp
+126-1532 files