LLVM/project 8219f5c — llvm/test/Instrumentation/MemorySanitizer masked-divrem.ll

[msan] add tests for llvm.masked.{udiv,sdiv,urem,srem} (#227462)

In preparation for #225363, which changes behavior.
DeltaFile
+307-0llvm/test/Instrumentation/MemorySanitizer/masked-divrem.ll
+307-01 files

LLVM/project a903200 — llvm/include/llvm/IR Constants.h

Fix typo in comment for ConstantInt get method (#227493)

Originally found this typo from docs generated by Doxygen.
DeltaFile
+1-1llvm/include/llvm/IR/Constants.h
+1-11 files

LLVM/project 9b63886 — clang/include/clang/CIR/Dialect/IR CIROps.td, clang/lib/CIR/CodeGen CIRGenBuiltin.cpp

fixup! [CIR] Support type deduction builder for BinaryFPToFPBuiltinOp
DeltaFile
+2-5clang/lib/CIR/CodeGen/CIRGenBuiltin.cpp
+0-4clang/include/clang/CIR/Dialect/IR/CIROps.td
+2-92 files

LLVM/project 69385fc — orc-rt/cmake OrcRTTesting.cmake, orc-rt/test/regression README.md lit.cfg.py

[orc-rt] Make split-file a required regression test tool. (#227510)

split-file is an LLVM utility, like FileCheck and not, so it's available
wherever they are. Require it in the same way, rather than gating tests
on a split-file lit feature.
DeltaFile
+14-1orc-rt/cmake/OrcRTTesting.cmake
+6-8orc-rt/test/regression/lit.cfg.py
+0-2orc-rt/test/regression/README.md
+0-1orc-rt/test/regression/languages/c/addressing/cross-object/hidden-data-load.test
+0-1orc-rt/test/regression/languages/c/addressing/cross-object/global-data-load.test
+0-1orc-rt/test/regression/languages/c/addressing/cross-object/global-array-element-pointer-in-data.test
+20-146 files

LLVM/project 9af3f9f — llvm/test/TableGen HwModeBitSet.td, llvm/unittests/MC TargetRegistry.cpp

[MC] Add baseline test for MCContext::getSubtargetCopy HwMode loss

No change intended here, just adding test coverage showing that
MCContext::getSubtargetCopy currently slices away the target's
<Target>GenMCSubtargetInfo subclass and causes getHwMode() and
getHwModeSet() to return 0.

This commit was created with the help of AI tools
DeltaFile
+34-0llvm/unittests/MC/TargetRegistry.cpp
+9-0llvm/test/TableGen/HwModeBitSet.td
+43-02 files

LLVM/project 5978944 — llvm/test/tools/llvm-objdump/ELF/RISCV mapping-sym-isa-disassembly.s mapping-sym-isa-mattr-conflict.s, llvm/tools/llvm-objdump llvm-objdump.cpp

[llvm-objdump][RISC-V] Do not leak file-level features into $x<ISA> regions

Previously, RISCVISATargetCache::get initialized per-region features from
Base.SubtargetInfo->getFeatureString(), which already included the
file-level Tag_RISCV_arch and --mattr flags. Because RISCVISAInfo::toFeatures()
only emits "+ext" for extensions present in the "$x<ISA>" mapping symbol
(and never "-ext"), any extension enabled in Tag_RISCV_arch or via a
conflicting --mattr remained enabled even in regions whose "$x<ISA>"
mapping symbol disabled it.

Start from an empty SubtargetFeatures instead and populate it from the
parsed "$x<ISA>" mapping symbol (plus non-conflicting --mattr flags).

This commit was created with the help of AI tools
DeltaFile
+12-15llvm/tools/llvm-objdump/llvm-objdump.cpp
+10-0llvm/test/tools/llvm-objdump/ELF/RISCV/mapping-sym-isa-mattr-conflict.s
+2-2llvm/test/tools/llvm-objdump/ELF/RISCV/mapping-sym-isa-disassembly.s
+24-173 files

LLVM/project f0e6410 — llvm/lib/Target/RISCV RISCVInstrFormats.td, llvm/test/MC/RISCV rvi-pseudos-invalid.s

[RISC-V][MC] Reject x0 as the temporary register for store pseudos

The zero register as the temporary for the store pseudo is illegal since
it is used to synthesize the store address and using zero would mean
that the auipc result is ignored and we store to an invalid location.
DeltaFile
+31-14llvm/test/MC/RISCV/rvi-pseudos-invalid.s
+2-2llvm/lib/Target/RISCV/RISCVInstrFormats.td
+33-162 files

LLVM/project a2e1ea9 — bolt/runtime CMakeLists.txt

[BOLT] Ensure runtime is built with uncompressed debug symbols (#218126)

Currently, if the runtime is built with compressed debug symbols it
causes an out-of-bounds read and segfault when instrumenting a binary.

This is due to JITLink currently not supporting SHF_COMPRESSED symbols
as it would need to decompress the symbols first to update the reloc
symbols

As some distributions are starting to enable compressed debug sections
in their toolchains by default, ensure the runtime is built without them
until JITLink supports updating SHF_COMPRESSED relocs.
DeltaFile
+3-0bolt/runtime/CMakeLists.txt
+3-01 files

LLVM/project 53504f5 — llvm/test/TableGen RegClassByHwModeCompressPat4Modes.td RegClassByHwModeCompressPatError.td, llvm/utils/TableGen CompressInstEmitter.cpp

[TableGen] Support HwMode registers in CompressInstEmitter

Previously, CompressInstEmitter only handled plain Register and
RegisterClass records when matching and validating operands in
CompressPat definitions. Attempting to use a RegisterByHwMode (such as a
mode-dependent stack pointer) or a RegClassByHwMode with mode-dependent
registers failed because CompressInstEmitter could not resolve the
underlying register or register class for a given subtarget configuration.

Introduce HwModePredicates in CodeGenHwModes to resolve HwModeSelect
records using the subtarget features required by a CompressPat (with
predicates implying the absence of all non-default modes resolving to
DefaultMode). This allows using a single instruction definition for
compressed instructions like RISC-V C_ADDI4SPN/C_ADDI16SP across HwModes
instead of duplicating the instruction definitions for each mode.

This commit was created with the help of AI tools
DeltaFile
+194-0llvm/utils/TableGen/Common/CodeGenHwModes.cpp
+99-83llvm/utils/TableGen/CompressInstEmitter.cpp
+154-1llvm/test/TableGen/RegClassByHwModeCompressPat.td
+129-0llvm/test/TableGen/RegClassByHwModeCompressPatError.td
+116-0llvm/test/TableGen/RegClassByHwModeCompressPat4Modes.td
+50-0llvm/utils/TableGen/Common/SubtargetFeatureInfo.cpp
+742-842 files not shown
+785-848 files

LLVM/project 48b6484 — clang/lib/CIR/Dialect/Transforms CallConvLoweringPass.cpp, clang/lib/CIR/Dialect/Transforms/TargetLowering CIRABIRewriteContext.h CIRABIRewriteContext.cpp

[CIR] Assert that getIndirectCalleeType is given an indirect call

Assisted-by: Claude Code / claude-opus-5-5
DeltaFile
+1-2clang/lib/CIR/Dialect/Transforms/TargetLowering/CIRABIRewriteContext.h
+1-2clang/lib/CIR/Dialect/Transforms/TargetLowering/CIRABIRewriteContext.cpp
+1-1clang/lib/CIR/Dialect/Transforms/CallConvLoweringPass.cpp
+3-53 files

LLVM/project d883453 — orc-rt/test/regression README.md

[orc-rt] Remove duplicate section from regression test README. NFC. (#227503)

The "Tests with more than one source file" section appeared twice.
Remove the older copy.
DeltaFile
+0-24orc-rt/test/regression/README.md
+0-241 files

LLVM/project c529383 — llvm/test/CodeGen/X86 cmpxchg-dead-eflags.ll

Remove test
DeltaFile
+0-138llvm/test/CodeGen/X86/cmpxchg-dead-eflags.ll
+0-1381 files

LLVM/project 88f2fef — llvm/lib/Target/X86 X86ISelLowering.cpp, llvm/test/CodeGen/X86 cmpxchg-dead-eflags.ll

X86: Avoid keeping EFLAGS live for an unused cmpxchg success value

Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+138-0llvm/test/CodeGen/X86/cmpxchg-dead-eflags.ll
+13-1llvm/lib/Target/X86/X86ISelLowering.cpp
+151-12 files

LLVM/project 531fe44 — mlir/include/mlir/Dialect/XeGPU/uArch uArchBase.h, mlir/lib/Conversion/VectorToXeGPU VectorToXeGPU.cpp

[mlir][xegpu] Check uArch block shapes in VectorToXeGPU transfer lowering (#217179)

vector.transfer_read/transfer_write were lowered to xegpu.load_nd /
xegpu.store_nd whenever the target chip was pvc, bmg or cri and the
transfer had rank >= 2, without asking whether the target's subgroup 2D
block instructions can access the requested tile at all.
Changing it to query the uArch 2D block load/store description up front
and keep the block path only for tiles whose two innermost dims are each
a multiple of a supported block size- the property the later layout
propagation and blocking passes rely on when they split a tile into
hardware-sized blocks. Everything else takes the scattered
load_gather/store_scatter fallback that both patterns already have.
Deriving 2D block support from the uArch instruction registry also
replaces the hardcoded chip list, resolving the TODO left there; the
three chips that registry covers are the same three that were listed.


assisted by claude
DeltaFile
+70-58mlir/test/Conversion/VectorToXeGPU/transfer-read-to-xegpu.mlir
+93-33mlir/lib/Conversion/VectorToXeGPU/VectorToXeGPU.cpp
+44-21mlir/test/Conversion/VectorToXeGPU/transfer-write-to-xegpu.mlir
+17-13mlir/include/mlir/Dialect/XeGPU/uArch/uArchBase.h
+224-1254 files

LLVM/project c2ccc57 — utils/bazel/llvm-project-overlay/mlir BUILD.bazel

[bazel][MLIR] Fix 98331a0a49e6d4a86d2e39ecc71025ce3b85acd9 (#227500)
DeltaFile
+2-0utils/bazel/llvm-project-overlay/mlir/BUILD.bazel
+2-01 files

LLVM/project b8c4e9f — lldb/test/API/tools/lldb-dap/cancel TestDAP_cancel.py, lldb/test/API/tools/lldb-dap/disconnect TestDAP_disconnect.py

[lldb-dap] End the session before sending the disconnect response (#227106)

Previously the `disconnect` request asks the debug adapter to disconnect
from the debuggee, ending the debug session, and then to shut down.
lldb-dap responded right after killing or detaching the process. So the
`exited` event, terminatedCommands and exitedCommands can be sent after
`disconnect` could reaches the client.

`DisconnectRequestHandler` now stores its response in
`DAP::on_session_end`. `DAP::Loop()` sends it once the session ended.
The teardown/cleanup is now:
- The event threads stop, after they report the exit of a killed
process.
- `terminated` Event is sent, if the event thread didn't send it.
- The transport thread stops.
- If there is a pending `configuration_done` we reply.
- The requests queued behind `disconnect` request are cancelled.
- The `disconnect` response is sent last.
- The debugger is destroyed.

    [17 lines not shown]
DeltaFile
+107-42lldb/tools/lldb-dap/DAP.cpp
+83-9lldb/unittests/DAP/Handler/DisconnectTest.cpp
+35-4lldb/test/API/tools/lldb-dap/disconnect/TestDAP_disconnect.py
+36-0lldb/unittests/DAP/TestBase.h
+33-1lldb/test/API/tools/lldb-dap/cancel/TestDAP_cancel.py
+22-5lldb/tools/lldb-dap/Handler/RequestHandler.h
+316-615 files not shown
+367-7411 files

LLVM/project 40edcbb — llvm/test/Transforms/IndirectBrExpand basic.ll

[IndirectBrExpand] Use UTC for basic.ll

To make it easier to modify in the future.

Reviewers: aeubanks, mtrofin

Pull Request: https://github.com/llvm/llvm-project/pull/227467
DeltaFile
+42-21llvm/test/Transforms/IndirectBrExpand/basic.ll
+42-211 files

LLVM/project 8630009 — clang/docs ReleaseNotes.md, clang/lib/AST ItaniumMangle.cpp

[Clang] Do not assume existing substition for template specialization

Mangled symbol is not always available when having template
specialization. For example, a template alias mangled the type when
actually doing substitution. In previous code, B in A<B>::foo() is
always mangled when reaching A<B>. When doing substition insides A<T>,
B is always available. Alias is not the case here, B is mangled when
reaching A<T> inside.

As now, TemplateName::SubstTemplateTemplateParm might needs to be
resolved further as B is not always defined, we remove it from the
switch case and find recursively like the origianl path.

Assisted-by: Claude # Tests, ReleaseNotes
DeltaFile
+57-52clang/lib/AST/ItaniumMangle.cpp
+15-0clang/test/CodeGenCXX/mangle-template.cpp
+4-0clang/docs/ReleaseNotes.md
+76-523 files

LLVM/project 0cf5430 — bolt/lib/Rewrite DWARFRewriter.cpp, bolt/test/AArch64 go_dwarf.test dwarf2-highpc-addr-form.s

[BOLT] Emit DW_AT_high_pc matching the class of its form (#227066)

DW_AT_high_pc holding an offset from DW_AT_low_pc is a DWARF 4 addition;
in DWARF 2 and 3 the attribute is always class address.

updateLowPCHighPC() always wrote "HighPC - LowPC" while reusing whatever
form the input DIE had, so a DWARF 2 CU using DW_FORM_addr ended up
storing a length where an end address is expected.

BOLT should write the end address when the form is DW_FORM_addr, and
the length otherwise. The default form for a new attribute is now
DW_FORM_addr below DWARF 4, and size needs to be widened to 64 bits so
DW_FORM_data8 won't be truncated. The FunctionRanges.empty() code path
no longer treats a raw DW_FORM_addr high_pc as a size.

Assited-by: opus
DeltaFile
+140-0bolt/test/AArch64/dwarf2-highpc-addr-form.s
+25-8bolt/lib/Rewrite/DWARFRewriter.cpp
+6-1bolt/test/AArch64/go_dwarf.test
+171-93 files

LLVM/project f35bd33 — libc/src/string/memory_utils/aarch64 inline_memcpy.h

[libc][aarch64] Use inline_memcpy_aligned_access_64bit under -mstrict-align (#227428)

Under `-mstrict-align`, clang cannot determine both `src` and `dst` are
aligned. The end of `inline_memcpy_aarch64` aligns `src` but it cannot
confirm `dst` is aligned. As a result,
`builtin::Memcpy<64>::loop_and_tail(dst, src, count)` lowers to
`__builtin_memcpy_inline(..., 64)` but clang emits many many single-byte
load/store instructions to ensure correctness. This can be very slow, so
if unaligned accesses are not supported, opt for adjusting both `src`
and `dst` such that we can use 8-byte loads/stores.
`inline_memcpy_aligned_access_64bit` already does this so we can just
reuse that.
DeltaFile
+17-0libc/src/string/memory_utils/aarch64/inline_memcpy.h
+17-01 files

LLVM/project b41ef0f — llvm/test/CodeGen/Hexagon intrinsics-v67.ll

[Hexagon,test] Enable +audio for intrinsics-v67.ll (#227362)

The test uses llvm.hexagon.M7.vdmpy{,.acc}, which lower to
M7_dcmpyrwc{,_acc}. Those instructions require the audio feature, so add
-mattr=+audio to the RUN line to avoid a predicate-check crash in the
asm printer.
DeltaFile
+1-1llvm/test/CodeGen/Hexagon/intrinsics-v67.ll
+1-11 files

LLVM/project b8952a1 — clang/include/clang/Basic AttrDocs.td

[clang][docs] Use doc links to HLSL/ResourceTypes in AttrDocs.td (#227407)

#222504 made the docs build warn about absolute links to pages of the
same Sphinx project, and the docs build turns warnings into errors. The
HLSL resource attribute docs added in #213346 linked to
https://clang.llvm.org/docs/HLSL/ResourceTypes.html, which after #226729
broke the clang docs build (docs-clang-html, docs-clang-man). Use the
{doc} role instead, like the rest of AttrDocs.td.
DeltaFile
+8-8clang/include/clang/Basic/AttrDocs.td
+8-81 files

LLVM/project ba19a43 — lldb/include/lldb/Core Debugger.h, lldb/source/Commands CommandObjectScripting.cpp

[lldb] Add label, title and divider ANSI settings (#226606)

Commands that print structured listings hard-coded their colors, but
terminal themes render ANSI colors differently, so users couldn't adjust
them the way they can with the prompt, progress or autosuggestion
colors.

Add `label-ansi-prefix`, `title-ansi-prefix` and `divider-ansi-prefix`
settings, each with a matching suffix, for the field labels, entry
titles and dividers such listings are made of. Their defaults are the
colors `scripting extension list` uses today, and it now reads them from
these settings.

Signed-off-by: Med Ismail Bennani <ismail at bennani.ma>
DeltaFile
+48-0lldb/test/API/commands/scripting/extension/TestScriptingExtensionListColors.py
+36-0lldb/source/Core/Debugger.cpp
+21-12lldb/source/Commands/CommandObjectScripting.cpp
+24-0lldb/source/Core/CoreProperties.td
+12-0lldb/include/lldb/Core/Debugger.h
+141-125 files

LLVM/project 83826c2 — flang/include/flang/Optimizer/OpenACC Passes.td, flang/lib/Optimizer/OpenACC/Transforms CMakeLists.txt ACCEraseUnusedKernelAllocations.cpp

[flang][openacc] Add ACCEraseUnusedKernelAllocations pass (#227481)

fir.declare carries a debug-memory effect and fir.freemem is a real
free, so an unused fir.allocmem inside acc.compute_region stays live
through ordinary dead-code elimination and is lowered to a checked
device malloc.

Delete that allocation when it has no uses, or when every use is
fir.freemem, a ViewLikeOpInterface view such as fir.convert, or
fir.declare of that storage. A load, store, or other memory use keeps
it. An unused private-recipe allocation is removed the same way as an
unused source array.
DeltaFile
+122-0flang/lib/Optimizer/OpenACC/Transforms/ACCEraseUnusedKernelAllocations.cpp
+51-0flang/test/Fir/OpenACC/acc-erase-unused-kernel-allocations.mlir
+15-0flang/include/flang/Optimizer/OpenACC/Passes.td
+2-0flang/lib/Optimizer/OpenACC/Transforms/CMakeLists.txt
+190-04 files

LLVM/project 280a20b — mlir/include/mlir/Dialect/GPU/Pipelines Passes.h, mlir/lib/Dialect/GPU/Pipelines GPUToXeVMPipeline.cpp

[MLIR][XeGPU][XeVM] Lower to XeVM: support skipping backend. (#227125)

And skip backend invocation for slow tests.
DeltaFile
+4-1mlir/include/mlir/Dialect/GPU/Pipelines/Passes.h
+2-2mlir/test/Integration/Dialect/XeGPU/WG/simple_mxfp_gemm_quantizeA_F4.mlir
+3-1mlir/lib/Dialect/GPU/Pipelines/GPUToXeVMPipeline.cpp
+1-1mlir/test/Integration/Dialect/XeGPU/WG/simple_mxfp_gemm_quantizeA_F8.mlir
+10-54 files

LLVM/project 1179810 — clang/test/CodeGen/RISCV rvp-intrinsics.c, llvm/test/CodeGen/AMDGPU load-constant-i1.ll amdgcn.bitcast.512bit.ll

Merge branch 'main' into users/adams381/cir-callconv-byval-slot-copy
DeltaFile
+42,441-42,426llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+7,527-7,623llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+4,115-8,359clang/test/CodeGen/RISCV/rvp-intrinsics.c
+4,036-4,380llvm/test/CodeGen/RISCV/rvv/clmul-sdnode.ll
+2,544-2,872llvm/test/CodeGen/RISCV/rvv/clmulh-sdnode.ll
+2,603-2,582llvm/test/CodeGen/AMDGPU/load-constant-i1.ll
+63,266-68,2421,637 files not shown
+154,927-141,4821,643 files

LLVM/project 4fec5ba — clang/test/CodeGen/RISCV rvp-intrinsics.c, llvm/test/CodeGen/AMDGPU load-constant-i1.ll amdgcn.bitcast.512bit.ll

Merge branch 'main' into users/adams381/cir-callconv-coerce-return-slot
DeltaFile
+42,441-42,426llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+7,527-7,623llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+4,115-8,359clang/test/CodeGen/RISCV/rvp-intrinsics.c
+4,036-4,380llvm/test/CodeGen/RISCV/rvv/clmul-sdnode.ll
+2,544-2,872llvm/test/CodeGen/RISCV/rvv/clmulh-sdnode.ll
+2,603-2,582llvm/test/CodeGen/AMDGPU/load-constant-i1.ll
+63,266-68,2421,467 files not shown
+148,098-139,5181,473 files

LLVM/project 8a7dedc — clang/test/CodeGen/RISCV rvp-intrinsics.c, llvm/test/CodeGen/AMDGPU load-constant-i1.ll amdgcn.bitcast.512bit.ll

Merge branch 'main' into users/adams381/cir-callconv-indirect-variadic
DeltaFile
+42,441-42,426llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+7,527-7,623llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+4,115-8,359clang/test/CodeGen/RISCV/rvp-intrinsics.c
+4,036-4,380llvm/test/CodeGen/RISCV/rvv/clmul-sdnode.ll
+2,544-2,872llvm/test/CodeGen/RISCV/rvv/clmulh-sdnode.ll
+2,603-2,582llvm/test/CodeGen/AMDGPU/load-constant-i1.ll
+63,266-68,2421,474 files not shown
+148,847-140,3841,480 files

LLVM/project 7e0802b — clang/test/CodeGen/AArch64/sve len.c

[clang][CIR] Match CIR return value with OGCG using mem2reg (#227241)

This fixes the test failure introduced by
[226534](https://github.com/llvm/llvm-project/pull/226534)

Currently, the `sve/len.c` test fails because of this difference between
OGCG and CIR:

```
; classic codegen                  ; CIR (-fclangir)
%2 = mul nuw i64 %1, 16            %6 = mul nuw i64 %5, 16
ret i64 %2                         store i64 %6, ptr %3      ; retval slot
                                   %7 = load i64, ptr %3
                                   ret i64 %7
```
So, adding the return value checks as they are breaks the test.

This adds `mem2reg`, so that the CIR output matches OGCG and the same
return value check works for both.
DeltaFile
+8-6clang/test/CodeGen/AArch64/sve/len.c
+8-61 files

LLVM/project 8081af8 — clang/lib/CodeGen/TargetBuiltins RISCV.cpp, clang/lib/Headers riscv_packed_simd.h

[RISCV][P-ext] Add scalar multiply high intrinsics (#225561)

Add intrinsics, Clang builtins and SelectionDAG support for the scalar
multiply high operations.
Support
[multiply-high](https://github.com/riscv/riscv-p-spec/blob/master/P-ext-intrinsics.adoc#multiply-high).

On RV32 the non-rounding forms map to the M-extension mulh/mulhu/mulhsu
instructions and the rounding forms to the P-extension
mulhr/mulhru/mulhrsu instructions.
On RV64 the scalar spellings reuse the packed pmulh.w family on the low
words via the existing PMULH*_W patterns.
DeltaFile
+96-0clang/test/CodeGen/RISCV/rvp-intrinsics.c
+85-0llvm/test/CodeGen/RISCV/rvp-simd-32.ll
+50-0llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+43-0cross-project-tests/intrinsic-header-tests/riscv_packed_simd.c
+32-0clang/lib/CodeGen/TargetBuiltins/RISCV.cpp
+14-0clang/lib/Headers/riscv_packed_simd.h
+320-02 files not shown
+340-08 files