LLVM/project dbedf98 — flang/lib/Evaluate tools.cpp, flang/lib/Lower ConvertConstant.cpp ConvertCall.cpp

[flang] Lower enumeration types as named records with a type descriptor

Represent an F2023 enumeration type as !fir.type<...{__ordinal:i32}>
with a real .dt descriptor, instead of a bare i32. Enumeration values
can now be boxed, passed as polymorphic (SELECT TYPE, ALLOCATE, I/O),
and are distinct from INTEGER in descriptors. Ordinary scalar uses
access __ordinal directly through hlfir.designate, with no extra
boxing.

This new approach should address all the review findings thus far.
DeltaFile
+172-336flang/lib/Lower/ConvertExprToHLFIR.cpp
+189-263flang/test/Lower/enumeration-type.f90
+400-0flang/test/Lower/enumeration-type-next-previous.f90
+188-0flang/lib/Lower/ConvertCall.cpp
+27-29flang/lib/Evaluate/tools.cpp
+0-38flang/lib/Lower/ConvertConstant.cpp
+976-6668 files not shown
+1,046-70714 files

LLVM/project c16c939 — bolt/lib/Core Exceptions.cpp, bolt/lib/Rewrite RewriteInstance.cpp

[BOLT] Support DW_EH_PE_sdata8 encoding in .eh_frame_hdr (#227847)

BOLT always wrote .eh_frame_hdr with 4-byte offsets, truncating those
beyond
2GB. Use 8-byte encoding when needed, as lld does since #179089.

Assisted-By: Opus 5.5
DeltaFile
+97-0bolt/test/X86/eh-frame-hdr-sdata8.s
+34-19bolt/lib/Core/Exceptions.cpp
+5-0bolt/lib/Rewrite/RewriteInstance.cpp
+136-193 files

LLVM/project 285f5ab — lld/ELF/Arch AArch64.cpp, lld/test/ELF aarch64-tls-le.s aarch64-tls-le-ldst.s

[lld][AArch64] Support R_AARCH64_TLSLE_LDST*_TPREL_LO12 (#227629)

Follow-up to #227173. Support the non-NC local-exec TLS load/store
relocations, emitted for `ldr/str xN, [xN, :tprel_lo12:sym]` (previously
"unknown relocation"). They resolve like the `_NC` variants — imm12
holds bits 11:scale of the TP offset — plus an unsigned 12-bit range
check on the full offset. Boundaries and encodings verified identical to
GNU ld 2.46.

Testing: new `aarch64-tls-le-ldst.s` (encodings, out-of-range boundary,
alignment); extended `aarch64-tls-le.s`; full `lld/test/ELF` passes.
DeltaFile
+73-0lld/test/ELF/aarch64-tls-le-ldst.s
+20-0lld/ELF/Arch/AArch64.cpp
+16-1lld/test/ELF/aarch64-tls-le.s
+109-13 files

LLVM/project a90d9d2 — mlir/include/mlir/Dialect/Vector/Transforms Passes.td, mlir/lib/Dialect/Vector/Transforms VectorMaskElimination.cpp

[mlir][vector] Add an eliminate-vector-masks pass (#226517)
DeltaFile
+41-0mlir/lib/Dialect/Vector/Transforms/VectorMaskElimination.cpp
+0-35mlir/test/lib/Dialect/Vector/TestVectorTransforms.cpp
+24-0mlir/include/mlir/Dialect/Vector/Transforms/Passes.td
+21-0mlir/test/Dialect/Vector/eliminate-masks-invalid-options.mlir
+2-2mlir/test/Dialect/Vector/eliminate-masks.mlir
+88-375 files

LLVM/project 92d88bb — clang/docs ReleaseNotes.md, clang/lib/AST ExprConstant.cpp

Revert "[clang] Avoid stack exhaustion in recursive `constexpr` calls" (#228110)

Reverts llvm/llvm-project#201706

See
https://github.com/llvm/llvm-project/pull/201706#issuecomment-5934225006
DeltaFile
+6-31clang/lib/AST/ExprConstant.cpp
+0-16clang/test/SemaCXX/constexpr-float-call-stack.cpp
+0-11clang/test/SemaCXX/constexpr-call-stack.cpp
+0-3clang/docs/ReleaseNotes.md
+6-614 files

LLVM/project 2cd234b — clang/lib/CodeGen CGStmtOpenMP.cpp, clang/test/OpenMP teams_generic_loop_reduction_distribute_codegen.cpp

[clang][OpenMP] Don't use fused dist schedule for teams loop emitted as distribute

Fix teams loop reductions lowered as 'distribute' lose their loop.

Claude assisted with this patch.
DeltaFile
+63-0clang/test/OpenMP/teams_generic_loop_reduction_distribute_codegen.cpp
+9-0clang/lib/CodeGen/CGStmtOpenMP.cpp
+72-02 files

LLVM/project c6c841a — lldb/include/lldb/Host/common NativeThreadProtocol.h, lldb/source/Plugins/Process/Windows/Common NativeThreadWindows.h NativeThreadWindows.cpp

[lldb-server] Handle jThreadExtendedInfo
DeltaFile
+49-0lldb/source/Plugins/Process/gdb-remote/GDBRemoteCommunicationServerLLGS.cpp
+4-0lldb/source/Plugins/Process/Windows/Common/NativeThreadWindows.cpp
+3-0lldb/include/lldb/Host/common/NativeThreadProtocol.h
+2-0lldb/source/Utility/StringExtractorGDBRemote.cpp
+2-0lldb/source/Plugins/Process/gdb-remote/GDBRemoteCommunicationServerLLGS.h
+2-0lldb/source/Plugins/Process/Windows/Common/NativeThreadWindows.h
+62-01 files not shown
+63-07 files

LLVM/project 9e15ea8 — lldb/source/Host/windows HostThreadWindows.cpp

[lldb][Windows] Fix import of NtQueryInformationThread (#227471)

`NtQueryInformationThread` is from `ntdll.dll`, not `Kernel32.dll`. We
never tested this. I'll add a test in the next PR when testing the
gdb-remote handler.
DeltaFile
+1-1lldb/source/Host/windows/HostThreadWindows.cpp
+1-11 files

LLVM/project 4bd8908 — llvm/lib/Target/RISCV RISCVISelLowering.cpp

[RISCV][P-ext] Fix confusing variable names. NFC (#227973)
DeltaFile
+3-3llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+3-31 files

LLVM/project 73ed9a9 — mlir/lib/Dialect/Vector/Transforms VectorDistribute.cpp, mlir/test/Dialect/Vector vector-warp-distribute.mlir

[mlir][Vector] Reject scalable reduction in warp distribution (#225267)

`WarpOpReduction` distributes a `vector.reduction` across warp lanes by
computing `numElements = vectorType.getShape()[0] / warpSize` and
building the per-lane type as a plain `VectorType::get({numElements},
...)`, with no check for a scalable operand.

For a scalable reduction vector, e.g. `vector<[32]xf32>` with warp size
32, this silently drops the scalable marker: the per-lane type becomes
the fixed `vector<1xf32>` instead of `vector<[1]xf32>`, so only 32 total
elements get reduced across lanes instead of `32*vscale`.

Reject a scalable reduction operand before any IR is created, and add a
negative test.
DeltaFile
+14-0mlir/test/Dialect/Vector/vector-warp-distribute.mlir
+3-0mlir/lib/Dialect/Vector/Transforms/VectorDistribute.cpp
+17-02 files

LLVM/project 85c141a — llvm/lib/Target/NVPTX NVPTXTargetMachine.cpp NVPTXCodeGenPassBuilder.cpp, llvm/test/CodeGen/NVPTX llc-pipeline-npm.ll

NVPTX: Drop LiveVariables from the register allocation pipeline (#225183)

The optimized RegAlloc pipeline ran LiveVariables only to satisfy 
PHIElimination and TwoAddressInstruction, both of which no longer need it. 
Remove the LiveVariables run (and, in the new pass manager, the 
UnreachableMachineBlockElim that was there only as a LiveVariables 
prerequisite).

Co-authored-by: Claude (Claude-Opus-4.8)
DeltaFile
+0-7llvm/lib/Target/NVPTX/NVPTXCodeGenPassBuilder.cpp
+0-2llvm/test/CodeGen/NVPTX/llc-pipeline-npm.ll
+0-1llvm/lib/Target/NVPTX/NVPTXTargetMachine.cpp
+0-103 files

LLVM/project a36eb0c — lldb/test/API/api/multithreaded TestMultithreaded.py

[lldb][test] Disable test_python_stop_hook on AArch64 Linux (#228114)

It has beeen flakey on our buildbot. See
https://github.com/llvm/llvm-project/issues/225860 for details.
DeltaFile
+2-0lldb/test/API/api/multithreaded/TestMultithreaded.py
+2-01 files

LLVM/project cce31d7 — llvm/test/Analysis/LoopAccessAnalysis runtime-checks-may-not-return-call.ll

[LAA] Add tests with loops with may-not-return calls (NFC) (#228102)

Add tests for runtime-check bounds of loops containing a call that may
not return.
DeltaFile
+229-0llvm/test/Analysis/LoopAccessAnalysis/runtime-checks-may-not-return-call.ll
+229-01 files

LLVM/project 8cb1d15 — clang/lib/Interpreter Interpreter.cpp, clang/test/Interpreter/CUDA device-module-verifier.cu

[clang-repl] Flush CUDA device bootstrap module before the first PTU (#226975)

The host path sets the bootstrap module aside with CacheCodeGenModule()
once the initial action has run, so the first PTU starts from a fresh
module. The CUDA device path skipped this step. The device module that
HandleTranslationUnit had already finalized during bootstrap stayed
current, and the first device PTU finalized it a second time, resulting
in CodeGen adding every module flag twice. The IR verifier then rejects
the module.

To reproduce, on a build with assertions, or with
`-fverify-intermediate-code` (hidden on release builds since the driver
disables the verifier there), the first input to `clang-repl --cuda`
fails:

```
  module flag identifiers must be unique (or of 'require' type)
  !"nvvm-reflect-ftz"
  module flag identifiers must be unique (or of 'require' type)

    [12 lines not shown]
DeltaFile
+76-0clang/unittests/Interpreter/DeviceOffloadTest.cpp
+14-0clang/test/Interpreter/CUDA/device-module-verifier.cu
+3-0clang/lib/Interpreter/Interpreter.cpp
+1-0clang/unittests/Interpreter/CMakeLists.txt
+94-04 files

LLVM/project a7f8ead — llvm/lib/Target/AMDGPU SIFoldOperands.cpp SIInstrInfo.cpp, llvm/lib/Target/RISCV RISCVInstrInfo.cpp

CodeGen: Drop the LiveVariables parameter from convertToThreeAddress (#225182)

This was used for analysis updates, but now the analysis is being removed.

Co-authored-by: Claude (Claude-Opus-4.8)
DeltaFile
+12-66llvm/lib/Target/X86/X86InstrInfo.cpp
+0-17llvm/lib/Target/AMDGPU/SIInstrInfo.cpp
+0-11llvm/lib/Target/RISCV/RISCVInstrInfo.cpp
+1-10llvm/lib/Target/SystemZ/SystemZInstrInfo.cpp
+2-4llvm/lib/Target/X86/X86InstrInfo.h
+2-2llvm/lib/Target/AMDGPU/SIFoldOperands.cpp
+17-1106 files not shown
+22-11712 files

LLVM/project 2641947 — lldb/include/lldb/Host/common NativeThreadProtocol.h, lldb/source/Plugins/Process/Windows/Common NativeThreadWindows.h NativeThreadWindows.cpp

[lldb-server] Handle jThreadExtendedInfo
DeltaFile
+49-0lldb/source/Plugins/Process/gdb-remote/GDBRemoteCommunicationServerLLGS.cpp
+4-0lldb/source/Plugins/Process/Windows/Common/NativeThreadWindows.cpp
+3-0lldb/include/lldb/Host/common/NativeThreadProtocol.h
+2-0lldb/source/Utility/StringExtractorGDBRemote.cpp
+2-0lldb/source/Plugins/Process/gdb-remote/GDBRemoteCommunicationServerLLGS.h
+2-0lldb/source/Plugins/Process/Windows/Common/NativeThreadWindows.h
+62-01 files not shown
+63-07 files

LLVM/project 243c774 — lldb/source/Host/windows HostThreadWindows.cpp

[lldb][Windows] Fix import of NtQueryInformationThread
DeltaFile
+1-1lldb/source/Host/windows/HostThreadWindows.cpp
+1-11 files

LLVM/project bd036a9 — lldb/include/lldb/Host HostNativeThreadBase.h, lldb/include/lldb/Host/windows HostThreadWindows.h

[lldb][Windows] Move extended info to host thread (#227470)

I want to add support for reading thread locals on Windows. To do this,
we need to know the address of the TEB. This was implemented for
`TargetThreadWindows`, but when using lldb-server, we didn't have this
info. As a first step, move this to the `HostThreadWindows`, so both the
target thread and native thread can call it.
DeltaFile
+41-0lldb/source/Host/windows/HostThreadWindows.cpp
+1-37lldb/source/Plugins/Process/Windows/Common/TargetThreadWindows.cpp
+4-0lldb/source/Host/common/HostNativeThreadBase.cpp
+2-0lldb/include/lldb/Host/windows/HostThreadWindows.h
+2-0lldb/include/lldb/Host/HostNativeThreadBase.h
+50-375 files

LLVM/project a0cbc28 — llvm/test/CodeGen/AMDGPU amdgcn.bitcast.832bit.ll amdgcn.bitcast.768bit.ll

Merge branch 'main' into users/adams381/llvmabi-x86-vector-abi-size
DeltaFile
+13,422-13,644llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+4,545-4,824llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+3,581-3,841llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.256bit.ll
+2,765-3,518llvm/test/CodeGen/AMDGPU/flat_atomics_i64.ll
+2,962-3,123llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.768bit.ll
+2,852-2,999llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.832bit.ll
+30,127-31,949853 files not shown
+100,746-83,768859 files

LLVM/project 556b112 — llvm/lib/Transforms/Instrumentation DataFlowSanitizer.cpp, llvm/test/Instrumentation/DataFlowSanitizer basic.ll abilist_aggregate.ll

[DataFlowSanitizer] Properly add ext attributes on arguments as needed. (#225443)

TargetLibraryInfo is used to compute the extension attributes, in part by a new
getExtAttrForI8Param() method. It currently always returns ZExt (or SExt) but is
used so that a target can easily override this if needed.
DeltaFile
+104-60llvm/lib/Transforms/Instrumentation/DataFlowSanitizer.cpp
+36-0llvm/test/Instrumentation/DataFlowSanitizer/instrumented-args-exts.ll
+16-16llvm/test/Instrumentation/DataFlowSanitizer/origin_abilist.ll
+6-6llvm/test/Instrumentation/DataFlowSanitizer/shadow-args-zext.ll
+4-4llvm/test/Instrumentation/DataFlowSanitizer/basic.ll
+4-4llvm/test/Instrumentation/DataFlowSanitizer/abilist_aggregate.ll
+170-903 files not shown
+186-949 files

LLVM/project fe0f2bb — clang/include/clang/Basic DiagnosticLexKinds.td, clang/test/C/C23 n2549.c

[clang] Suppress binary literal warnings in system macros (#228095)

Follow up #228035

Close #192490
DeltaFile
+42-3clang/test/C/C23/n2549.c
+5-4clang/include/clang/Basic/DiagnosticLexKinds.td
+47-72 files

LLVM/project a28ddaa — llvm/include/llvm/CodeGen SpillPlacement.h

[CodeGen] Remove unused SpillPlacement::Linked (NFC) (#227989)

The last use was removed on May 19, 2016 in commit
b926bdac4c18e0f31d827dec482f207856e88e1e.

Assisted-by: Antigravity
DeltaFile
+0-3llvm/include/llvm/CodeGen/SpillPlacement.h
+0-31 files

LLVM/project 4eb6e48 — llvm/include/llvm/Frontend/OpenMP OMPIRBuilder.h, llvm/lib/Frontend/OpenMP OMPIRBuilder.cpp

[OpenMP] Fix missing implicit barriers in device worksharing loops (#227735)

Device worksharing lowering discarded `NeedsBarrier`, omitting implicit
barriers after loops without `nowait`. Forward the requirement and emit
the barrier outside the outlined loop body, ensuring all participating
threads synchronise.

Co-authored-by: Codex <codex at openai.com>
DeltaFile
+57-2mlir/test/Target/LLVMIR/omptarget-wsloop.mlir
+24-2llvm/unittests/Frontend/OpenMPIRBuilderTest.cpp
+18-3llvm/lib/Frontend/OpenMP/OMPIRBuilder.cpp
+5-4llvm/include/llvm/Frontend/OpenMP/OMPIRBuilder.h
+104-114 files

LLVM/project 8da79c7 — llvm/lib/Target/AArch64 MachineSMEABIPass.cpp, llvm/test/CodeGen/AArch64 sme-lazy-sve-nzcv-live.mir sme-abi-eh-liveins.mir

AArch64: Mark status flag clobbers dead in SME ABI pass (#228042)

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+20-0llvm/test/CodeGen/AArch64/machine-sme-abi-dead-nzcv.ll
+5-5llvm/test/CodeGen/AArch64/machine-sme-abi-find-insert-pt.mir
+6-3llvm/lib/Target/AArch64/MachineSMEABIPass.cpp
+2-2llvm/test/CodeGen/AArch64/sme-lazy-sve-nzcv-live.mir
+2-2llvm/test/CodeGen/AArch64/sme-abi-eh-liveins.mir
+2-2llvm/test/CodeGen/AArch64/aarch64-sme-za-call-lowering.ll
+37-141 files not shown
+38-157 files

LLVM/project 70a0078 — llvm/include/llvm/Analysis LoopAccessAnalysis.h, llvm/lib/Analysis LoopAccessAnalysis.cpp

[LAA] Add stencil group merging to reduce runtime pointer checks (#187252)

Take this loop, where S is only known at runtime:

  for (i = 0; i < N; i++)
    Out[i] = In[i - S] + In[i - 1] + In[i] + In[i + 1] + In[i + S];

groupChecks puts In[i - 1], In[i] and In[i + 1] in one group, because
their bounds differ by a constant. In[i - S] and In[i + S] differ from
the rest by a multiple of S, so each stays in its own group. That is
3 groups for In and 3 checks against Out. With two or three strides,
as in 3D stencils, the count grows fast. The Einstein Toolkit / Cactus
CCZ4 code has loops with thousands of checks, and the vectorizer gives
up on them.

This patch adds mergeStencilGroups, which runs after groupChecks. For
the loop above it makes one group for In, from In[i - S] to In[i + S].
One check against Out is enough, plus a check that 1 <= S <= Max.


    [27 lines not shown]
DeltaFile
+4,039-0llvm/test/Analysis/LoopAccessAnalysis/stencil-group-merging.ll
+791-3llvm/lib/Analysis/LoopAccessAnalysis.cpp
+158-0llvm/test/Analysis/LoopAccessAnalysis/stencil-group-merging-limits.ll
+157-0llvm/test/Analysis/LoopAccessAnalysis/stencil-group-merging-i128.ll
+17-3llvm/include/llvm/Analysis/LoopAccessAnalysis.h
+3-7llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
+5,165-136 files

LLVM/project c97e17a — offload/plugins-nextgen/level_zero/include L0Program.h L0Plugin.h, offload/plugins-nextgen/level_zero/src L0Plugin.cpp L0Program.cpp

[Offload][L0][NFC] Remove old ELF format support (#228011)

After we switched to using Offload Binary (for OpenMP) or direct SPIR-V
images (for SYCL) this support is not used anymore.

Assisted by Claude.
DeltaFile
+1-303offload/plugins-nextgen/level_zero/src/L0Program.cpp
+0-7offload/plugins-nextgen/level_zero/src/L0Plugin.cpp
+4-1offload/plugins-nextgen/level_zero/include/L0Plugin.h
+0-2offload/plugins-nextgen/level_zero/include/L0Program.h
+5-3134 files

LLVM/project 16bd62a — flang/include/flang/Common uint128.h

Remove unrelated changes
DeltaFile
+16-13flang/include/flang/Common/uint128.h
+16-131 files

LLVM/project 0ce2cdc — llvm/utils/TableGen AsmWriterEmitter.cpp

[TableGen] Allow AsmWriter to generate uint64_t tables. NFC. (#227641)

Each OpInfo entry in the generated AMDGPUInstPrinter::getMnemonic
carries 8 bytes of data. Previously GenAsmWriter would split that into
two uint32_t tables for no good reason. Generating a single uint64_t
table makes for shorter output and slightly better generated code.
DeltaFile
+5-4llvm/utils/TableGen/AsmWriterEmitter.cpp
+5-41 files

LLVM/project 9576254 — llvm/lib/IR IRBuilder.cpp, llvm/test/CodeGen/AMDGPU promote-alloca-fp-ptr-cast.ll

[AMDGPU][SROA] Expand cast chain handling to floating point types

CreateBitPreservingCastChain does not produce inttoptr or ptrtoint for floating-point types. When promoting structs like { float, float } to <2 x float>, this can lead to pointers being bitcast directory to <2 x float> which is invalid. This change expands the use of the intermediate to these cases
DeltaFile
+131-0llvm/test/CodeGen/AMDGPU/promote-alloca-fp-ptr-cast.ll
+30-0llvm/unittests/IR/IRBuilderTest.cpp
+7-8llvm/lib/IR/IRBuilder.cpp
+168-83 files

LLVM/project 895bd20 — clang/docs ReleaseNotes.md, clang/lib/AST ExprConstant.cpp

Revert "[clang] Avoid stack exhaustion in recursive `constexpr` calls (#201706)"

This reverts commit 0a9f172d35d472b2b421a6aae19f6aae9e40b805.
DeltaFile
+6-31clang/lib/AST/ExprConstant.cpp
+0-16clang/test/SemaCXX/constexpr-float-call-stack.cpp
+0-11clang/test/SemaCXX/constexpr-call-stack.cpp
+0-3clang/docs/ReleaseNotes.md
+6-614 files