LLVM/project 5b3ae00 —

[NVPTX] Consolidate AsmPrinter type and global variable printing (#222240)
DeltaFile
+0-00 files

LLVM/project bc1a551 — llvm/lib/Target/NVPTX NVPTXUtilities.h NVPTXAsmPrinter.cpp, llvm/test/CodeGen/NVPTX param-align.ll ptr-global.ll

[NVPTX] Consolidate AsmPrinter type and global variable printing (#222240)
DeltaFile
+219-461llvm/lib/Target/NVPTX/NVPTXAsmPrinter.cpp
+75-0llvm/test/CodeGen/NVPTX/fp-global.ll
+37-0llvm/test/CodeGen/NVPTX/odd-integer-width.ll
+22-0llvm/test/CodeGen/NVPTX/ptr-global.ll
+8-8llvm/lib/Target/NVPTX/NVPTXUtilities.h
+5-5llvm/test/CodeGen/NVPTX/param-align.ll
+366-4748 files not shown
+381-48214 files

LLVM/project 3865d82 — llvm/lib/CAS PluginCAS.cpp

Include the plugin's error message with an unknown validation result
DeltaFile
+9-5llvm/lib/CAS/PluginCAS.cpp
+9-51 files

LLVM/project 9f00966 — mlir/include/mlir/Dialect/OpenACC OpenACCCGAttributes.td OpenACCUtilsCG.h, mlir/lib/Dialect/OpenACC/Transforms ACCComputeLowering.cpp

[mlir][OpenACC] Record gang(static:) chunk size on lowered loops (#227037)

Lower OpenACC `gang(static:)` through compute lowering as
`acc.chunk_size`. A constant size is stored directly; `static:*` and a
non-constant size are recorded as -1. `gang(static:)` also maps to gang
dimension 1.
DeltaFile
+55-0mlir/test/Dialect/OpenACC/acc-compute-lowering-loop.mlir
+23-2mlir/lib/Dialect/OpenACC/Transforms/ACCComputeLowering.cpp
+20-0mlir/lib/Dialect/OpenACC/Utils/OpenACCUtilsCG.cpp
+16-0mlir/unittests/Dialect/OpenACC/OpenACCUtilsCGTest.cpp
+15-0mlir/include/mlir/Dialect/OpenACC/OpenACCUtilsCG.h
+11-2mlir/include/mlir/Dialect/OpenACC/OpenACCCGAttributes.td
+140-46 files

LLVM/project f95e5ea — llvm/utils/gn/secondary/clang/include/clang/Config BUILD.gn

[gn] Fix Interpreter/emulated-tls.cpp on mac (#227132)

In, #225475 (5dfe8605911e), CLANG_HAVE_EMUTLS_GET_ADDRESS got hardcoded
to 0 in the GN build, while the CMake build has a config-time check for
__emutls_get_address.

On mac and linux, compiler-rt provides that symbol, so set
CLANG_HAVE_EMUTLS_GET_ADDRESS to 1 there. That way, __emutls_get_address
is force-linked as absolute symbol and things are happy.

(With it set to 0. clang-repl dlsym()s for the symbol, which works on
Linux where it's found in libgcc_s.so.1, but it doesn't work on mac.)
DeltaFile
+2-1llvm/utils/gn/secondary/clang/include/clang/Config/BUILD.gn
+2-11 files

LLVM/project 34ae391 — llvm/test/CodeGen/AMDGPU bitcast-vector-extract.ll

[AMDGPU] Replace GCN-NOT with autogen checks in test

Change-Id: I21350cf041ea4f2fa22d152ba372da6095e8b762
DeltaFile
+235-30llvm/test/CodeGen/AMDGPU/bitcast-vector-extract.ll
+235-301 files

LLVM/project f6f868b — llvm/include/llvm-c/CAS PluginAPI_functions.h, llvm/include/llvm/CAS ObjectStore.h

Address review feedback
DeltaFile
+7-1llvm/lib/CAS/PluginCAS.cpp
+3-3llvm/include/llvm-c/CAS/PluginAPI_functions.h
+2-2llvm/include/llvm/CAS/ObjectStore.h
+12-63 files

LLVM/project f26ba92 — llvm/tools/llvm-cas llvm-cas.cpp

Rebase on updated base
DeltaFile
+8-6llvm/tools/llvm-cas/llvm-cas.cpp
+8-61 files

LLVM/project 14cd556 — llvm/tools/llvm-cas llvm-cas.cpp

Simplify redirect selection

Created using spr 1.3.7
DeltaFile
+8-6llvm/tools/llvm-cas/llvm-cas.cpp
+8-61 files

LLVM/project 54343ee — llvm/include/llvm/CAS UnifiedOnDiskCache.h, llvm/lib/CAS UnifiedOnDiskCache.cpp

Rebase on updated base
DeltaFile
+13-17llvm/lib/CAS/UnifiedOnDiskCache.cpp
+28-0llvm/test/tools/llvm-cas/validation.test
+20-7llvm/tools/llvm-cas/llvm-cas.cpp
+10-0llvm/unittests/CAS/UnifiedOnDiskCacheTest.cpp
+2-2llvm/include/llvm/CAS/UnifiedOnDiskCache.h
+3-0llvm/tools/llvm-cas/Options.td
+76-266 files

LLVM/project 0426fd3 — llvm/include/llvm/CAS UnifiedOnDiskCache.h, llvm/lib/CAS UnifiedOnDiskCache.cpp

Address review feedback

Created using spr 1.3.7
DeltaFile
+13-17llvm/lib/CAS/UnifiedOnDiskCache.cpp
+28-0llvm/test/tools/llvm-cas/validation.test
+20-7llvm/tools/llvm-cas/llvm-cas.cpp
+10-0llvm/unittests/CAS/UnifiedOnDiskCacheTest.cpp
+2-2llvm/include/llvm/CAS/UnifiedOnDiskCache.h
+3-0llvm/tools/llvm-cas/Options.td
+76-266 files

LLVM/project 6fee55d — flang/include/flang/Support Fortran-features.h, flang/lib/Semantics check-declarations.cpp

[flang][Semantics] Warn on BIND(C) interfaces with assumed-shape/rank dummies (#225965)

## Motivation

Fortran 2018 §18.3.6 requires assumed-shape, deferred-shape, and
assumed-rank `BIND(C)` dummy arguments to be passed using a CFI
descriptor (`CFI_cdesc_t`). Some existing C/C++ interfaces predate this
requirement and expect such an argument to be passed by bare address
instead. Declaring a `BIND(C)` Fortran interface against such code
currently produces a silent calling-convention mismatch — no diagnostic
warns the user that the ABI they've declared doesn't match what
pre-2018 C/C++ code expects.

## What this PR does

Adds a portability warning in `CheckSubprogram`
(`flang/lib/Semantics/check-declarations.cpp`) that fires when a
`BIND(C)` interface declares an assumed-shape or assumed-rank dummy
argument, under a new `UsageWarning` category, `BindCArrayDescriptor`.

    [40 lines not shown]
DeltaFile
+72-0flang/test/Semantics/bind-c20.f90
+25-0flang/test/Semantics/bind-c21.f90
+23-0flang/lib/Semantics/check-declarations.cpp
+1-1flang/include/flang/Support/Fortran-features.h
+121-14 files

LLVM/project e5cc9fb — llvm/lib/CodeGen MachineBasicBlock.cpp, llvm/lib/CodeGen/MIRParser MILexer.h MILexer.cpp

MIR: Serialize MachineBasicBlock::MaxBytesForAlignment

Fix missing serialization of another field. The alignment was
already handled. The name is a bit verbose. Some places call it
"MaxSkip" which matches the name of the 2nd operand to the .p2align
directive this corresponds to.

Co-Authored-By: Claude Sonnet 5 <noreply at anthropic.com>
DeltaFile
+55-0llvm/test/CodeGen/MIR/Generic/machine-basic-block-max-bytes-for-alignment-errors.mir
+31-1llvm/test/CodeGen/MIR/Generic/basic-blocks.mir
+24-0llvm/lib/CodeGen/MIRParser/MIParser.cpp
+2-0llvm/lib/CodeGen/MachineBasicBlock.cpp
+1-0llvm/lib/CodeGen/MIRParser/MILexer.h
+1-0llvm/lib/CodeGen/MIRParser/MILexer.cpp
+114-16 files

LLVM/project 60f7179 — .github/workflows libcxx-pr-benchmark.yml libcxx-pr-test-tools.yml

[libc++] Move test tools CI job back to k8s runner sets (#226564)

While the issue with k8s runners in #226230 is still not resolved, it is
better to target the k8s runners and get transient failures than to
target the old runners and have the jobs hang forever (there seems to be
no runners registered in llvm-premerge-libcxx-runners).

Co-authored-by: Aiden Grossman <aidengrossman at google.com>
DeltaFile
+5-2.github/workflows/libcxx-pr-test-tools.yml
+0-1.github/workflows/libcxx-pr-benchmark.yml
+5-32 files

LLVM/project c6b15d3 — lld/test/wasm check-arch-32-in-64.test check-arch-64-in-32.test, lld/wasm InputFiles.cpp

[lld][WebAssembly] Fix error message when linking wasm64 file with -mwasm32 (#227091)

When `-mwasm32` is explicitly passed and a wasm64 object file is linked,
the error message previously stated:
"wasm32 object file can't be linked in wasm64 mode". Fix this to report
that the wasm64 object file cannot be linked in wasm32 mode.
DeltaFile
+4-2lld/wasm/InputFiles.cpp
+4-1lld/test/wasm/check-arch-64-in-32.test
+1-1lld/test/wasm/check-arch-32-in-64.test
+9-43 files

LLVM/project 64907b6 — libcxxabi/src cxa_personality.cpp, libunwind/include unwind_wasm.h

fix Wasm exceptions + coop threading + shared libraries (#222747)

Prior to this commit, the combination of Wasm exception handling,
cooperative multithreading, and shared libraries was broken.
Specifically, the code generation in `WasmEHPrepare.cpp` involved
direct, cross-library access to `libunwind.so`'s thread-local
`__wasm_lpad_context` variable. However, the ABI used for cooperative
multithreading does not support cross-library access to thread-local
variables.

The solution used here is to add a new `_Unwind_GetWasmLPadContext`
function to `libunwind.so` and use that to get address of the
`__wasm_lpad_context` for the current thread, both in the code generated
by `WasmEHPrepare.cpp` and in the `__gxx_wasm_personality_v0` function
defined in `cxa_personality.cpp`. I've used this strategy
unconditionally for all targets, regardless of whether cooperative
multithreading and/or position-independent are enabled. If desired (e.g.
for performance or code complexity reasons), I could make it conditional
on both of those features being enabled and fall back to using

    [2 lines not shown]
DeltaFile
+29-27llvm/lib/CodeGen/WasmEHPrepare.cpp
+8-7llvm/test/CodeGen/WebAssembly/eh-lsda.ll
+8-6llvm/test/CodeGen/WebAssembly/wasm-eh-prepare.ll
+5-3libcxxabi/src/cxa_personality.cpp
+6-2libunwind/src/Unwind-wasm.c
+4-3libunwind/include/unwind_wasm.h
+60-483 files not shown
+67-519 files

LLVM/project c7bed2c — llvm/lib/Target/AMDGPU VOP3PInstructions.td, llvm/test/CodeGen/AMDGPU frem.ll mad-mix-lo-bf16.ll

[AMDGPU] Fold fpround of fadd and fsub into v_mad/fma_mixlo and mixhi

MadFmaMixFP32Pats turns (fadd x, y) into (fma x, 1.0, y) and (fsub x, y)
into (fma (-y), 1.0, x) so the mix instructions absorb the operation along
with the f16 or bf16 source modifiers. MadFmaMixFP16Pats and
MadFmaMixFP16Pats_t16 only did this for fmul, so a rounded result still
needed a separate convert for a rounding the mix instructions perform
themselves.

Unlike the f32 patterns these do not require an operand to be an fpextend
of an f16, since an fpround on the result always removes the convert. The
rewrite is exact because the mix instructions round the f32 result again
when they write the 16-bit destination, so it stays f32_to_f16(fma(x, 1.0,
y)).

Assisted-by: Claude Code Opus 5
DeltaFile
+192-285llvm/test/CodeGen/AMDGPU/GlobalISel/fdiv.f16.ll
+61-311llvm/test/CodeGen/AMDGPU/mad-mix-lo.ll
+62-145llvm/test/CodeGen/AMDGPU/mad-mix-hi.ll
+55-36llvm/test/CodeGen/AMDGPU/mad-mix-lo-bf16.ll
+26-52llvm/test/CodeGen/AMDGPU/frem.ll
+70-0llvm/lib/Target/AMDGPU/VOP3PInstructions.td
+466-8294 files not shown
+499-88510 files

LLVM/project 5c881d4 — llvm/include/llvm/ProfileData SampleProf.h, llvm/lib/ProfileData SampleProf.cpp SampleProfReader.cpp

[ProfileData] Only keep module functions when reading ProfileSymbolList

When a sample profile is loaded for a module (SampleProfileLoader), the
profile symbol list is only ever queried for functions of that module:
`PSL->contains(F.getName())` in SampleProfileLoader and
`PSL->contains(CanonFName)` in SampleProfileMatcher. Yet the string-based
reader inserts every symbol of the profiled binary into a DenseSet, in every
compile and every ThinLTO backend that loads the profile.

When the reader has a module, build a small set of that module's function
names (raw and canonical) and only add matching list entries. The list is
still scanned, but nothing outside the module is inserted, so there is no
large hash table to build. Readers without a module (llvm-profdata) still
load the full list. The MD5 symbol list is unaffected.

In a fleet-wide CPU profile of a production clang,
`ProfileSymbolList::read` accounted for 1.7% of all clang cycles and 10% of
ThinLTO backend cycles.


    [11 lines not shown]
DeltaFile
+71-0llvm/test/Transforms/SampleProfile/pseudo-probe-stale-profile-symbol-list.ll
+10-6llvm/lib/ProfileData/SampleProf.cpp
+15-1llvm/lib/ProfileData/SampleProfReader.cpp
+3-1llvm/include/llvm/ProfileData/SampleProf.h
+3-0llvm/test/Transforms/SampleProfile/Inputs/pseudo-probe-stale-profile-symbol-list.text
+102-85 files

LLVM/project c87ad07 — llvm/lib/Transforms/Vectorize/SLPVectorizer SLPMemoryUtils.cpp

[SLP][NFC]Fix build for MSVC compiler, NFC

Reported in https://lab.llvm.org/buildbot/#/builders/46/builds/42017

Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/227129
DeltaFile
+3-3llvm/lib/Transforms/Vectorize/SLPVectorizer/SLPMemoryUtils.cpp
+3-31 files

LLVM/project 9e09945 — llvm/lib/Transforms/Vectorize/SLPVectorizer SLPMemoryUtils.cpp

[𝘀𝗽𝗿] initial version

Created using spr 1.3.7
DeltaFile
+3-3llvm/lib/Transforms/Vectorize/SLPVectorizer/SLPMemoryUtils.cpp
+3-31 files

LLVM/project c021bb4 —

[gn] bump deployment target to macOS 13 (#227124)

macOS 13 is four years old by now.

This has the effect that lld starts defaulting to chained fixups with
this. Chained fixups reduces `clang --version` from 4 ms to 3.2 ms. (Not
that it matters.)

(Without #227120, chained fixups reduce `clang --version` from 12.5 ms
to 10.9 ms.)

No behavior change.
DeltaFile
+0-00 files

LLVM/project 4ca8a56 — clang/lib/CodeGen CGObjCRuntime.cpp, clang/test/CodeGenObjC exceptions-seh.m

[ObjC][SEH] Fix clang crash when using finally statements (#176779)

When targeting a platform that does not have funclet-based EH, we push
the finally cleanup (normal edge) and catchall (unwind edge) onto the
EHStack _before_ pushing all catch handlers. The try statement is then
emitted and catch handlers popped from EHStack. Last, the finally
cleanup is popped from EHStack.

For funclet-based EH, we outline and push the finally funclet (of type
`NormalAndEHCleanup`) onto the EHStack _after_ pushing the catch
handlers and never pop it. This results in a crash during codegen when
we try to emit the catch handlers. Not popping the finally cleanup from
the EHStack results in incorrect calls to cleanup handlers in nested
try/catch/finally statements.

I fixed the two issues by:
1. Pushing the finally cleanup first, and
2. Popping it at the end of `CGObjCRuntime::EmitTryCatchStmt`.

Fixes #51899
DeltaFile
+53-0clang/test/CodeGenObjC/exceptions-seh.m
+27-2clang/lib/CodeGen/CGObjCRuntime.cpp
+80-22 files

LLVM/project 6ab3be2 — llvm/utils/gn/build mac_sdk.gni

[gn] bump deployment target to macOS 13 (#227124)

macOS 13 is four years old by now.

This has the effect that lld starts defaulting to chained fixups with
this. Chained fixups reduces `clang --version` from 4 ms to 3.2 ms. (Not
that it matters.)

(Without #227120, chained fixups reduce `clang --version` from 12.5 ms
to 10.9 ms.)

No behavior change.
DeltaFile
+1-1llvm/utils/gn/build/mac_sdk.gni
+1-11 files

LLVM/project a8debc5 — mlir/include/mlir/Target/Cpp CppEmitter.h

Add registerToCppTranslation to CppEmmitter.h (#226337)

Currently if one wants to register `mlir-to-cpp` out of tree, they must
include `mlir/InitAllTranslations.h` and depend transitively on all
translation targets (in bazel, `@llvm-project//mlir:AllTranslations`).

This change adds the registration declaration to `CppEmitter.h` so that
one can depend just on the `MLIRTargetCpp` target (or
`@llvm-project//mlir:TargetCpp` in bazel).

This matches the organization of the SMTLib codegen registration in
`mlir/include/mlir/Target/SMTLIB/ExportSMTLIB.h` (though some other
targets like `IRDLToCpp` do it differently).
DeltaFile
+4-0mlir/include/mlir/Target/Cpp/CppEmitter.h
+4-01 files

LLVM/project 285a865 — llvm/utils/gn/build BUILDCONFIG.gn BUILD.gn, llvm/utils/gn/secondary/clang/tools/clang-repl BUILD.gn

[gn] Build executables without exported symbols on macOS (#227120)

Speeds up `clang --version` from 12 ms to 4 ms on my system. 8 ms faster
startup isn't a lot, but there's also no reason not to do it.

No intended behavior change.
DeltaFile
+7-2llvm/utils/gn/secondary/clang/tools/clang-repl/BUILD.gn
+7-0llvm/utils/gn/build/BUILD.gn
+3-1llvm/utils/gn/build/BUILDCONFIG.gn
+4-0llvm/utils/gn/secondary/llvm/unittests/Passes/Plugins/BUILD.gn
+4-0llvm/utils/gn/secondary/llvm/tools/llvm-jitlink/BUILD.gn
+4-0llvm/utils/gn/secondary/llvm/tools/lli/BUILD.gn
+29-38 files not shown
+41-414 files

LLVM/project a69848d — llvm/utils profcheck-xfail.txt

[profcheck] Exclude find-first-byte-nested.ll (#227121)

We only fixed x86 for LoopIdiom. PR #225576 added a test that looks like
it's just exposing existing propagation issues.
DeltaFile
+1-0llvm/utils/profcheck-xfail.txt
+1-01 files

LLVM/project f9d9b42 — clang/lib/CodeGen BackendConsumer.h CodeGenAction.cpp, clang/lib/Interpreter Interpreter.cpp DeviceOffload.h

Revert "releand "[clang-repl] Implement IncrementalHIPDeviceParser for HIP de…"

This reverts commit 5c20fe98552af8fdbcd7ce714c1fad1f1e731a65.
DeltaFile
+8-209clang/lib/Interpreter/DeviceOffload.cpp
+9-55clang/lib/Interpreter/DeviceOffload.h
+0-50clang/unittests/Basic/TargetIDTest.cpp
+0-9clang/lib/CodeGen/CodeGenAction.cpp
+0-8clang/lib/CodeGen/BackendConsumer.h
+6-1clang/lib/Interpreter/Interpreter.cpp
+23-3324 files not shown
+25-34210 files

LLVM/project b28ced4 — llvm/lib/Transforms/Vectorize SLPVectorizer.cpp, llvm/test/CodeGen/AMDGPU bitinsert-bitextract.ll

Rebase

Created using spr 1.3.7
DeltaFile
+4,294-0llvm/test/CodeGen/RISCV/bitinsert-bitextract.ll
+3,321-0llvm/test/CodeGen/ARM/bitinsert-bitextract.ll
+1,989-0llvm/test/CodeGen/RISCV/bitinsert-bitextract-fp.ll
+1,976-0llvm/test/CodeGen/AMDGPU/bitinsert-bitextract.ll
+1,011-908llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+1,552-0llvm/test/CodeGen/ARM/bitinsert-bitextract-fp.ll
+14,143-9081,794 files not shown
+66,080-18,8631,800 files

LLVM/project c5f99a1 — clang/test/CodeGen/X86 sse41-builtins-constrained.c sse41-builtins.c

[clang][X86] Fix round builtins tests (#226708)

Fixed minor issues in round builtins tests. Noticed them while working on #215787.
DeltaFile
+9-9clang/test/CodeGen/X86/sse41-builtins.c
+6-6clang/test/CodeGen/X86/sse41-builtins-constrained.c
+15-152 files

LLVM/project 650a674 — clang/test/CIR/CodeGen pragma-fenv_access.c, llvm/include/llvm/ADT DenseMap.h

Rebase

Created using spr 1.3.7
DeltaFile
+977-870llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+1,308-0llvm/test/CodeGen/AMDGPU/llvm.amdgcn.cvt.scale.pk32.gfx13.ll
+339-431llvm/include/llvm/ADT/DenseMap.h
+367-367clang/test/CIR/CodeGen/pragma-fenv_access.c
+696-0llvm/test/Transforms/SLPVectorizer/X86/strength-reducible-address.ll
+694-0llvm/test/CodeGen/AMDGPU/rewrite-vgpr-mfma-to-agpr-spill-multi-store-codegen.ll
+4,381-1,668752 files not shown
+19,313-6,537758 files