LLVM/project a72b3da — llvm/cmake/modules MLGOLower.cmake, llvm/lib/Analysis CMakeLists.txt

[mlgo] Allow passing pre-emitc-ed models
DeltaFile
+79-58llvm/cmake/modules/MLGOLower.cmake
+98-0llvm/lib/Analysis/models/inline-oz-test-model.inc
+49-0llvm/lib/Analysis/models/regalloc-eviction-test-model.inc
+13-12llvm/lib/CodeGen/CMakeLists.txt
+13-12llvm/lib/Analysis/CMakeLists.txt
+3-21llvm/lib/CodeGen/MLRegAllocEvictAdvisor.cpp
+255-10320 files not shown
+323-19826 files

LLVM/project 77bd783 — llvm/include/llvm/CodeGen LiveVariables.h, llvm/lib/CodeGen LiveVariables.cpp

LiveVariables: Remove dead live-in handling and Defs plumbing

No physical register is tracked at the start of a block, so handling
the block live-ins was a no-op. The Defs list was only appended for
instruction defs, which runOnInstr already collects.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+9-24llvm/lib/CodeGen/LiveVariables.cpp
+3-5llvm/include/llvm/CodeGen/LiveVariables.h
+12-292 files

LLVM/project 8a7c6e0 — llvm/include/llvm/CodeGen LiveVariables.h, llvm/lib/CodeGen LiveVariables.cpp

LiveVariables: Only visit tracked physical registers

Keep a bitvector of physical registers with a recorded def or use in
the current block. Register mask handling, the end of block scan, and
the per-block reset now only visit those registers instead of every
register. This is significant for targets with many registers, such as
AMDGPU.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+19-8llvm/lib/CodeGen/LiveVariables.cpp
+3-0llvm/include/llvm/CodeGen/LiveVariables.h
+22-82 files

LLVM/project c76cf30 — llvm/include/llvm/CodeGen LiveVariables.h, llvm/lib/CodeGen LiveVariables.cpp

CodeGen: Strip LiveVariables down to dead flag computation

This analysis is dead and there are no more explicit uses. There are still
passes implicitly relying on adjustments of dead flags. Missing dead
flags are added, and implicit-def operands are added for partially dead
physical registers.

The whole pass should be deleted, but it's taking a while to get all the dead
flag changes through the rest of the compiler. As a stop-gap to try to recover
some compile time regression, and avoiding new users appearing, strip the pass
down to only commputing the dead flags.

The main side effect of this is kill flags are no longer made accurate, which
is the source of the test churn.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+53-519llvm/lib/CodeGen/LiveVariables.cpp
+7-231llvm/include/llvm/CodeGen/LiveVariables.h
+34-34llvm/test/CodeGen/PowerPC/aix-cc-abi-mir.ll
+0-64llvm/unittests/MI/LiveIntervalTest.cpp
+26-26llvm/test/CodeGen/AMDGPU/waterfall-loop-exec-update-terminator.mir
+16-16llvm/test/CodeGen/AArch64/GlobalISel/arm64-pcsections.ll
+136-89034 files not shown
+216-1,00040 files

LLVM/project 4a428b1 — clang/include/clang/CIR/Dialect/IR CIROps.td, clang/lib/CIR/Dialect/Transforms LoweringPrepare.cpp

[CIR] Register a static's extended temporaries where they are built (#229896)

Temporaries whose lifetime a static extends were all destroyed together,
in one batch registered after the static's whole initializer ran. Now
each one is registered for destruction right after it is built, the way
classic codegen does it. That fixes a verifier failure when a static
extends more than one temporary, a variable's own destructor going
missing, objects destroyed in the wrong order, and a temporary destroyed
even when it was never built.

Assisted-by: Claude Code / claude-opus-5-5
DeltaFile
+225-127clang/test/CIR/CodeGen/global-temp-dtor.cpp
+226-34clang/test/CIR/CodeGen/local-static-temp-dtor.cpp
+218-0clang/test/CIR/IR/invalid-register-exit-dtor.cir
+136-64clang/lib/CIR/Dialect/Transforms/LoweringPrepare.cpp
+75-0clang/test/CIR/IR/register-exit-dtor.cir
+47-2clang/include/clang/CIR/Dialect/IR/CIROps.td
+927-2274 files not shown
+979-26210 files

LLVM/project d3d55f2 — clang/lib/Format TokenAnnotator.cpp ContinuationIndenter.cpp, clang/unittests/Format FormatTest.cpp

[clang-format] Fix BeforeLambdaBody causing incorrect argument wrapping (#182319)

Fixes #182274
DeltaFile
+52-0clang/unittests/Format/FormatTest.cpp
+13-5clang/lib/Format/ContinuationIndenter.cpp
+11-0clang/lib/Format/TokenAnnotator.cpp
+76-53 files

LLVM/project 36613f8 — mlir/include/mlir/Dialect/OpenACC OpenACCOps.td

[nfc][mlir][acc] Improve acc.on_device documentation (#230143)

This updates the operation documentation because as per OpenACC spec, it
is not just a runtime call, but can be evaluated at compile time. Also
add the spec section.
DeltaFile
+11-4mlir/include/mlir/Dialect/OpenACC/OpenACCOps.td
+11-41 files

LLVM/project 0c14766 — llvm/test/CodeGen/AMDGPU commute-literal-src0-cse.mir

Merge branch 'users/adelejjeh/amdgpu-vgpr-msb-commute-literal' into users/adelejjeh/amdgpu-vopd-dot2-commute-literal
DeltaFile
+1-1llvm/test/CodeGen/AMDGPU/commute-literal-src0-cse.mir
+1-11 files

LLVM/project bbfafe0 — llvm/lib/Target/AMDGPU GCNVOPDUtils.cpp, llvm/test/CodeGen/AMDGPU vopd-dot2-commute-imm-src1.mir

[AMDGPU] Address review comments

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+9-5llvm/lib/Target/AMDGPU/GCNVOPDUtils.cpp
+1-1llvm/test/CodeGen/AMDGPU/vopd-dot2-commute-imm-src1.mir
+10-62 files

LLVM/project d4abf56 — llvm/include/llvm/Target TargetOptions.h, llvm/lib/CodeGen BasicBlockSections.cpp

 [CodeGen] Honor -function-splitting=none in BasicBlockSections (#227880)

With -function-splitting=none, functions with a basic block sections
profile are still laid out according to the profile, but all their basic
blocks are now emitted in a single section, instead of the profile's
clusters and a cold section.
DeltaFile
+13-5llvm/lib/CodeGen/BasicBlockSections.cpp
+15-0llvm/test/CodeGen/X86/basic-block-sections-clusters.ll
+15-0llvm/test/CodeGen/Generic/machine-function-splitter.ll
+0-1llvm/include/llvm/Target/TargetOptions.h
+43-64 files

LLVM/project cb75841 — mlir/include/mlir/Dialect/LLVMIR LLVMOps.td, mlir/lib/Dialect/LLVMIR/IR LLVMDialect.cpp

[mlir][LLVM] Verify llvm.invoke callees like llvm.call (#229954)

Generalize CallOp's symbol-use verification into a template shared with
InvokeOp, which now implements SymbolUserOpInterface. Direct invokes get
the same checking calls have: callee resolution, operand and result
types, vararg attributes, and the debug-location rule.

This newly rejects malformed direct invokes that used to be accepted and
translate. `mlir/test/Dialect/LLVMIR/roundtrip.mlir` had such an invoke
and now uses a matching callee.

Assisted-by: Muse Code (Muse).

---------

Co-authored-by: Bruno Cardoso Lopes <bruno.cardosolopes at gmail.com>
DeltaFile
+134-7mlir/test/Dialect/LLVMIR/invalid.mlir
+43-31mlir/lib/Dialect/LLVMIR/IR/LLVMDialect.cpp
+39-0mlir/test/Dialect/LLVMIR/invalid-call-location.mlir
+4-2mlir/test/Dialect/LLVMIR/roundtrip.mlir
+1-0mlir/include/mlir/Dialect/LLVMIR/LLVMOps.td
+221-405 files

LLVM/project 09f8e3d — mlir/include/mlir/Target/LLVMIR ModuleImport.h, mlir/lib/Target/LLVMIR ModuleImport.cpp

[mlir][LLVMIR] Bound import diagnostics and share one slot tracker (#229956)

Import diagnostics render LLVM entities unbounded: a global prints its
entire initializer (a merged vtable reaches megabytes), and every
rendering builds its own ModuleSlotTracker, whose module walk makes a
per-instruction diagnostic quadratic in module size. To the scale we
link and use the LLVM IR dialect, this has bitten us many times
downstream, usually leading to OOM's and other silly behaviors from
something that isn't supported anyways. This solution has helped us
finding and fixing a lot of missing attributes, which will be
contributed soon.

Cap each rendering at 256 characters and print globals by name only,
diagnostics only need to identify the entity. Share one lazily created
tracker across the import instead: printing through it numbers exactly
like the per-call tracker, the module is never mutated during import so
slots stay valid, and each print establishes its own function context.
Thread the tracker through the module-flag converters so every
diagnostic uses the same numbering.

    [5 lines not shown]
DeltaFile
+154-97mlir/lib/Target/LLVMIR/ModuleImport.cpp
+34-0mlir/test/Target/LLVMIR/Import/import-failure.ll
+16-0mlir/include/mlir/Target/LLVMIR/ModuleImport.h
+7-0mlir/test/Target/LLVMIR/Import/global-variables.ll
+2-2mlir/test/Target/LLVMIR/Import/function-metadata.ll
+213-995 files

LLVM/project 986f495 — clang/include/clang/CIR/FrontendAction CIRGenAction.h, clang/lib/CIR/FrontendAction CIRGenAction.cpp

[CIR] Enable bytecode emission through the driver and cc1 (#229607)

Teach the driver to serialize CIR directly to bytecode: new EmitCIRBC
frontend action kind, `TY_CIRBC` (.cirbc) driver type, and
EmitCIRBCAction writing the in-memory module with
`mlir::writeBytecodeToFile` instead of printing text. `.cirbc` inputs
map to the CIR language like .cir.

Tests cover cc1 emission with magic-byte and read-back pins, .cir-input
and .cirbc-input round trips, and driver forwarding, phase selection,
default output name, and flag-implied pipeline selection.

I started proper bytecode encoding support in #229586, so this is a step
into using that through the driver.
DeltaFile
+28-0clang/test/CIR/Driver/emit-cir-bc.c
+19-0clang/lib/CIR/FrontendAction/CIRGenAction.cpp
+15-0clang/test/CIR/emit-actions.cpp
+10-2clang/lib/FrontendTool/ExecuteCompilerInvocation.cpp
+8-0clang/include/clang/CIR/FrontendAction/CIRGenAction.h
+6-1clang/lib/Driver/Driver.cpp
+86-39 files not shown
+111-515 files

LLVM/project ae2fee0 — llvm/lib/Target/AMDGPU AMDGPULegalizerInfo.cpp, llvm/test/CodeGen/AMDGPU/GlobalISel fcmp.bf16.ll

AMDGPU/GlobalISel: Legalize bf16 g_fcmp (#227909)

Legalize bf16 fcmp by widening to f32.
DeltaFile
+410-0llvm/test/CodeGen/AMDGPU/GlobalISel/fcmp.bf16.ll
+4-1llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
+414-12 files

LLVM/project 229beb8 — compiler-rt/lib/csan csan.cpp, compiler-rt/test/csan access-sizes.cpp

Use unaligned loads when snapshotting watched values
DeltaFile
+22-0compiler-rt/test/csan/access-sizes.cpp
+3-3compiler-rt/lib/csan/csan.cpp
+25-32 files

LLVM/project 20e22f9 — compiler-rt/lib/csan csan_gpu.cpp csan.cpp, compiler-rt/test/csan volatile.cpp

Remove separate volatile entry points
DeltaFile
+49-0compiler-rt/test/csan/volatile.cpp
+22-0compiler-rt/test/csan/AMDGPU/volatile.hip
+0-4compiler-rt/lib/csan/csan_gpu.cpp
+0-4compiler-rt/lib/csan/csan.cpp
+71-84 files

LLVM/project a49fa31 — compiler-rt/lib/csan CMakeLists.txt, compiler-rt/lib/csan/offload csan_offload_preinit.cpp CMakeLists.txt

[compiler-rt] Remove dlsym interceptor and support `-shared-libsan` for CSan

Summary:
Follow the UBSan offload runtime. Offload now resolves HSA through the
global scope, so the `dlsym` interceptor is no longer needed. The real
HSA entry points are still taken from the loaded HSA library rather than
`RTLD_NEXT`, since every DSO with a static runtime exports the same
wrappers and they would otherwise chain back into each other.

Build `libclang_rt.csan.so` with the offload objects folded in. The
exported HSA wrappers report failure when HSA is absent and warn when HSA
was loaded ahead of the runtime. The preinit hook moves to a separate
`csan_offload-preinit` archive for executables.
DeltaFile
+34-98compiler-rt/lib/csan/offload/csan_offload_hsa_interceptors.cpp
+84-9compiler-rt/lib/csan/CMakeLists.txt
+35-0compiler-rt/test/csan/hsa-load-order.cpp
+22-4compiler-rt/lib/csan/offload/CMakeLists.txt
+24-0compiler-rt/test/csan/hsa-missing.cpp
+22-0compiler-rt/lib/csan/offload/csan_offload_preinit.cpp
+221-1115 files not shown
+241-11111 files

LLVM/project b4a8e89 — compiler-rt/cmake config-ix.cmake, compiler-rt/cmake/Modules AllSupportedArchDefs.cmake

comments
DeltaFile
+13-9compiler-rt/lib/csan/csan_gpu.cpp
+3-3compiler-rt/cmake/config-ix.cmake
+2-2compiler-rt/lib/csan/CMakeLists.txt
+1-1compiler-rt/test/csan/CMakeLists.txt
+1-1compiler-rt/lib/csan/offload/CMakeLists.txt
+1-0compiler-rt/cmake/Modules/AllSupportedArchDefs.cmake
+21-166 files

LLVM/project 79e313e — compiler-rt/lib/csan csan_watch.h csan_report.cpp, compiler-rt/lib/csan/offload csan_offload_report.cpp csan_offload_hsa_interceptors.cpp

[compiler-rt] Add 'csan' library for the concurrency sanitizer

Summary:
Adds the runtime for the concurrency sanitizer, both CPU and GPU.
Fundamentally, this works using the following pseudocode:

```c
static u64 watchpoints[N]; // Hash-indexed, zero is empty.

// Emitted before the access, so we never trip on our own write.
void check_access(volatile void *addr, u32 size, u32 type) {
    // Every access probes. A read conflicts only with a watched write, a
    // write conflicts with either.
    if (u64 *wp = find_watchpoint(addr, size, type))
        consume(wp, this_pc()); // Hand our location to the owner.

    if (!should_sample()) // Wave-uniform, 1-in-N chance.
        return;


    [17 lines not shown]
DeltaFile
+456-0compiler-rt/lib/csan/csan_gpu.cpp
+370-0compiler-rt/lib/csan/offload/csan_offload_hsa_interceptors.cpp
+334-0compiler-rt/lib/csan/csan.cpp
+279-0compiler-rt/lib/csan/csan_report.cpp
+184-0compiler-rt/lib/csan/offload/csan_offload_report.cpp
+163-0compiler-rt/lib/csan/csan_watch.h
+1,786-044 files not shown
+2,940-250 files

LLVM/project 98a1002 — clang/lib/Driver/ToolChains CommonArgs.cpp, clang/test/Driver fsanitize.c fsanitize-concurrency-offload.c

[Clang] Support `-shared-libsan` for offload CSan

Summary:
Follow the UBSan handling. The shared runtime embeds the HSA
interceptors, so `csan_offload` is only linked with static runtimes and
executables pull in `csan_offload-preinit` to initialize early.
DeltaFile
+38-0clang/test/Driver/fsanitize-concurrency-offload.c
+9-4clang/lib/Driver/ToolChains/CommonArgs.cpp
+3-1clang/test/Driver/fsanitize.c
+50-53 files

LLVM/project 7a8299c — clang/docs ConcurrencySanitizer.md, clang/lib/CodeGen CodeGenFunction.cpp

[Clang] Add support for the `-fsanitize=concurrency` runtime

Summary:
Add the frontend sanitizer kind, function attributes, pass pipeline
integration, predefined macro, driver handling, and documentation for
ConcurrencySanitizer.
DeltaFile
+207-0clang/docs/ConcurrencySanitizer.md
+73-0clang/test/CodeGen/sanitize-concurrency.c
+32-8clang/lib/Driver/ToolChains/CommonArgs.cpp
+15-0clang/test/Driver/fsanitize.c
+10-4clang/lib/Driver/SanitizerArgs.cpp
+10-3clang/lib/CodeGen/CodeGenFunction.cpp
+347-1510 files not shown
+375-1616 files

LLVM/project 3fc6756 — clang/test/CodeGen sanitize-concurrency.c, clang/test/CodeGen/Inputs sanitize-concurrency-ignorelist.txt

ignore
DeltaFile
+29-0clang/test/CodeGen/sanitize-concurrency.c
+2-0clang/test/CodeGen/Inputs/sanitize-concurrency-ignorelist.txt
+31-02 files

LLVM/project 6106770 — clang/lib/Driver SanitizerArgs.cpp, clang/lib/Driver/ToolChains AMDGPU.cpp CommonArgs.cpp

Comments
DeltaFile
+48-0clang/test/Driver/fsanitize-concurrency-offload.c
+27-0clang/test/Driver/fsanitize.c
+16-7clang/lib/Driver/ToolChains/Clang.cpp
+11-4clang/lib/Driver/SanitizerArgs.cpp
+5-6clang/lib/Driver/ToolChains/CommonArgs.cpp
+3-2clang/lib/Driver/ToolChains/AMDGPU.cpp
+110-195 files not shown
+112-2111 files

LLVM/project 8f18381 — offload/plugins-nextgen/common/include PluginInterface.h, offload/plugins-nextgen/common/src PluginInterface.cpp

[Offload] Add OFFLOAD_FORCE_BLOCKING force-synchronization escape hatch (#222635)

- Add OF_ForceSyncOps BoolEnvar and forceSyncOps() on GenericDeviceTy
with actual spelling: OFFLOAD_FORCE_BOCKING
- Add shouldForceSync predicate and drain external async info objects in
AsyncInfoWrapperTy::finalize() when the flag is set

Assisted-by: Claude Code
DeltaFile
+16-0offload/plugins-nextgen/common/include/PluginInterface.h
+11-0offload/unittests/OffloadAPI/common/Fixtures.hpp
+8-0offload/test/unit/lit.cfg.py
+7-0offload/plugins-nextgen/common/src/PluginInterface.cpp
+4-0offload/test/lit.cfg
+3-0offload/unittests/OffloadAPI/memory/olMemFill.cpp
+49-02 files not shown
+51-08 files

LLVM/project ed85b09 — llvm/test/CodeGen/AMDGPU commute-literal-src0-cse.mir

[AMDGPU] Drop -verify-machineinstrs from commute-literal-src0-cse.mir

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+1-1llvm/test/CodeGen/AMDGPU/commute-literal-src0-cse.mir
+1-11 files

LLVM/project d753324 — llvm/lib/Transforms/Instrumentation ConcurrencySanitizer.cpp, llvm/test/Instrumentation/ConcurrencySanitizer selection.ll

Yaxunl comments
DeltaFile
+35-7llvm/lib/Transforms/Instrumentation/ConcurrencySanitizer.cpp
+22-0llvm/test/Instrumentation/ConcurrencySanitizer/selection.ll
+57-72 files

LLVM/project 89942af — llvm/test/Transforms/SLPVectorizer/AMDGPU ordered-reduction-coalesced-loads.ll fmul-extract-fadd-chain.ll

update tests
DeltaFile
+0-88llvm/test/Transforms/SLPVectorizer/AMDGPU/fmul-extract-fadd-chain.ll
+16-64llvm/test/Transforms/SLPVectorizer/AMDGPU/ordered-reduction-coalesced-loads.ll
+16-1522 files

LLVM/project 7aac451 — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Transforms/SLPVectorizer/AMDGPU alt-fmul-fadd-cost.ll fma-operand-contract-selection.ll

[AMDGPU] Limit the fmul fusion discount to a matching context type

The fmul is free when its context instruction feeds a fusable fadd or
fsub. A vector fmul priced with a scalar lane as the context inherits
that fusion only when it is emitted lane by lane. On a packed type the
scalar user does not show that the vector fmul feeds a vector fadd, and
products extracted into a scalar fadd chain keep the packed fmul while
the fma is lost.
DeltaFile
+88-0llvm/test/Transforms/SLPVectorizer/AMDGPU/fmul-extract-fadd-chain.ll
+64-16llvm/test/Transforms/SLPVectorizer/AMDGPU/ordered-reduction-coalesced-loads.ll
+38-37llvm/test/Transforms/SLPVectorizer/AMDGPU/ordered-reduction-fma-fusion.ll
+36-6llvm/test/Transforms/SLPVectorizer/AMDGPU/fma-operand-contract-selection.ll
+7-11llvm/test/Transforms/SLPVectorizer/AMDGPU/alt-fmul-fadd-cost.ll
+5-2llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+238-721 files not shown
+239-737 files

LLVM/project 660a0e6 — llvm/test/Transforms/SLPVectorizer/AMDGPU fdiv-afn-reduction.ll reassoc-reduction-coalesced-loads.ll, llvm/test/Transforms/SLPVectorizer/X86 udiv-strictfp.ll

[NFC][SLP] Precommit tests for the phantom load saving (#228901)

The idea of what the tests show. On AMDGPU, consecutive scalar loads
coalesce in the backend, so the saving SLP counts for vectorizing them
is phantom.

Assisted-By: Claude Code Opus 5
DeltaFile
+582-0llvm/test/Transforms/SLPVectorizer/AMDGPU/ordered-reduction-coalesced-loads.ll
+512-0llvm/test/Transforms/SLPVectorizer/AMDGPU/int-reduction-coalesced-loads.ll
+448-0llvm/test/Transforms/SLPVectorizer/AMDGPU/elementwise-coalesced-loads.ll
+154-0llvm/test/Transforms/SLPVectorizer/AMDGPU/reassoc-reduction-coalesced-loads.ll
+112-0llvm/test/Transforms/SLPVectorizer/X86/udiv-strictfp.ll
+70-0llvm/test/Transforms/SLPVectorizer/AMDGPU/fdiv-afn-reduction.ll
+1,878-06 files

LLVM/project 7eded9f — utils/bazel/llvm-project-overlay/clang BUILD.bazel

[Bazel] Fixes a789eb1 (#230272)

This fixes a789eb1d665bb015820ce32db87105749ea537f6 (#223451).

Buildkite error link:
https://buildkite.com/llvm-project/upstream-bazel/builds?commit=a789eb1d665bb015820ce32db87105749ea537f6

Co-authored-by: Google Bazel Bot <google-bazel-bot at google.com>
DeltaFile
+1-0utils/bazel/llvm-project-overlay/clang/BUILD.bazel
+1-01 files