LLVM/project ac7853d — llvm/lib/Target/NVPTX NVPTXInstrInfo.td NVPTXIntrinsics.td, llvm/test/CodeGen/MIR/NVPTX floating-point-immediate-operands.mir

[NVPTX] Generalize flexible operands to floating-point instructions (#230258)

Extend the register-or-immediate operand approach to floating-point
instructions and intrinsics.
DeltaFile
+445-525llvm/lib/Target/NVPTX/NVPTXIntrinsics.td
+221-405llvm/lib/Target/NVPTX/NVPTXInstrInfo.td
+12-14llvm/test/CodeGen/NVPTX/frem.ll
+10-10llvm/test/CodeGen/MIR/NVPTX/floating-point-immediate-operands.mir
+8-9llvm/test/CodeGen/NVPTX/f16x2-instructions.ll
+5-7llvm/test/CodeGen/NVPTX/div.ll
+701-9705 files not shown
+713-98711 files

LLVM/project bd02a13 — clang/lib/CIR/CodeGen CIRGenStmtOpenMP.cpp, clang/test/CIR/CodeGenOpenMP parallel-for.c for-loop-review-fixes.c

[CIR][OpenMP] Add support for the OpenMP 'for' directive

This patch adds support for wsloop in ClangIR: the `for` directive and its
combined forms `parallel for` and `target parallel for`. This is lowered to an
omp.wsloop + omp.loop_nest, nested utilizing the existing queue-based
decomposition.

Assisted-by: Cursor / Claude Sonnet 5 High
DeltaFile
+271-8clang/lib/CIR/CodeGen/CIRGenStmtOpenMP.cpp
+207-0clang/test/CIR/CodeGenOpenMP/pragma-omp-for.c
+161-0clang/test/CIR/CodeGenOpenMP/for-loop-forms.c
+146-0clang/test/CIR/CodeGenOpenMP/target-parallel-for.c
+123-0clang/test/CIR/CodeGenOpenMP/for-loop-review-fixes.c
+62-0clang/test/CIR/CodeGenOpenMP/parallel-for.c
+970-86 files

LLVM/project 8f53f0b — llvm CMakeLists.txt

[mlgo] flush bot cmake caches after #227941
DeltaFile
+6-0llvm/CMakeLists.txt
+6-01 files

LLVM/project 68f9424 — llvm/lib/Transforms/Vectorize SLPVectorizer.cpp, llvm/test/Transforms/SLPVectorizer int_sideeffect.ll

[SLP]Fix deps for stores ahead of may-throw calls

Make stores-after-may-throw insts control dependent on the next
may-throw instruction.

Fixes #230402

Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/230567
DeltaFile
+40-2llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+11-4llvm/test/Transforms/SLPVectorizer/X86/store-across-may-throw-call.ll
+11-2llvm/test/Transforms/SLPVectorizer/int_sideeffect.ll
+62-83 files

LLVM/project f541760 — clang/lib/CIR/Dialect/IR CIRTypes.cpp, clang/lib/CIR/Dialect/Transforms/TargetLowering CIRABIRewriteContext.cpp

[CIR] Fix layout and attributes for non-power-of-two vectors

Struct members, array elements and union storage are now laid out by
alloc size, as in LLVM, so a three-element vector takes its full 16
bytes instead of 12.  Pointer differences and the x86_64 va_arg stride
use the alloc size too.

The calling-convention pass now places noundef and nofpclass the way
classic does.  A coercion that widens the value drops noundef.  Each
half of a flattened argument keeps the argument's attributes.  An
indirect argument's pointer no longer carries nofpclass.

Assisted-by: Cursor / Claude Opus 5.5
DeltaFile
+300-0clang/test/CIR/CodeGen/call-conv-lowering-x86_64-npot-vector.c
+96-74clang/lib/CIR/Dialect/Transforms/TargetLowering/CIRABIRewriteContext.cpp
+85-74clang/test/CIR/CodeGen/attr-noundef.cpp
+31-0clang/test/CIR/Transforms/abi-lowering/x86_64-vector.cir
+14-14clang/test/CIR/CodeGen/call-conv-lowering-x86_64-vec3.c
+13-8clang/lib/CIR/Dialect/IR/CIRTypes.cpp
+539-17010 files not shown
+592-20816 files

LLVM/project e6cdca0 — llvm/utils/UnicodeData CMakeLists.txt

[llvm][UnicodeData] Check for LIBXML_READER_ENABLED before building UnicodeCharSetsGenerator (#230254)

Since LIBXML_READER_ENABLED is a required feature for building
UnicodeCharSetsGenerator, I added a check in the CMakeLists to ensure
that if there's a static library of libxml2 that library has
LIBXML_READER_ENABLED.

I double checked and this looks like the only required libxml2 feature
for UnicodeCharSetsGenerator.
DeltaFile
+20-5llvm/utils/UnicodeData/CMakeLists.txt
+20-51 files

LLVM/project c4dc9d1 — flang/lib/Lower ConvertExprToHLFIR.cpp ConvertCall.cpp, flang/test/Lower enumeration-type-next-previous.f90 enumeration-type.f90

Vector subscript gather now accepts parameters.
Cleaned up code.
DeltaFile
+66-0flang/test/Semantics/enumeration-type-intrinsics-valid.f90
+25-0flang/test/Lower/enumeration-type.f90
+17-0flang/test/Lower/enumeration-type-next-previous.f90
+2-2flang/lib/Lower/ConvertCall.cpp
+1-1flang/lib/Lower/ConvertExprToHLFIR.cpp
+111-35 files

LLVM/project 3b5a9f1 — llvm/test/Transforms/SLPVectorizer/X86 store-across-may-throw-call.ll

[SLP][NFC]Add a test with incorrect store vectorization between throwing insts, NFC



Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/230559
DeltaFile
+100-0llvm/test/Transforms/SLPVectorizer/X86/store-across-may-throw-call.ll
+100-01 files

LLVM/project 1c5ce01 — clang/docs/_static custom.css

[docs] Fix good/bad table rendering when the theme is auto-selected to be dark (#230545)

When the system prefers dark mode and Furo is set to follow system-wide
preferences (i.e. `auto`), switch to dark-styled table rendering.

Before:
<img width="945" height="370" alt="image"
src="https://github.com/user-attachments/assets/e0e33120-52a2-49da-84c3-59a70f3a212b"
/>

After:
<img width="943" height="377" alt="image"
src="https://github.com/user-attachments/assets/60f84107-cc67-49bb-b050-7559c2f81e76"
/>

Fixes #228952
DeltaFile
+9-0clang/docs/_static/custom.css
+9-01 files

LLVM/project a24f8ed — llvm/lib/CodeGen/SelectionDAG LegalizeTypes.h LegalizeVectorTypes.cpp, llvm/test/CodeGen/AArch64 mask-beforefirst.ll

[DAG] Implement scalarization for mask_beforefirst

It's just NOT of the input lane.
DeltaFile
+9-0llvm/test/CodeGen/AArch64/mask-beforefirst.ll
+9-0llvm/lib/CodeGen/SelectionDAG/LegalizeVectorTypes.cpp
+1-0llvm/lib/CodeGen/SelectionDAG/LegalizeTypes.h
+19-03 files

LLVM/project 643e9dc — libc/src/__support/math exp2f_double_eval.h exp10f_double_eval.h, libc/test/src/math/performance_testing CMakeLists.txt exp2f_perf.cpp

[libc][math] Optimize exp2f and exp10f hot paths (#224178)

Route common finite inputs directly to the existing degree-5 kernels so
exp2f avoids its exceptional-path frame and exp10f keeps cold handling
out of line. Preserve the proven polynomial arithmetic and special
cases.

Add targeted MPFR regression inputs and official range-based performance
coverage for both functions.

PerfTest medians from alternating baseline/patch runs were:
```

|                      |baseline  | patch     | speedup|
+----------------------+----------+-----------+--------+
|exp2f hot normal      |2.33805   | 1.92944   | 17.02% |
|exp2f close to one    |2.28649   | 1.76252   | 22.78% |
|exp10f hot normal     |2.83441   | 2.36588   | 15.80% |
|exp10f close to one   |2.89851   | 2.25680   | 20.67% |

    [6 lines not shown]
DeltaFile
+79-68libc/src/__support/math/exp10f_double_eval.h
+50-46libc/src/__support/math/exp2f_double_eval.h
+52-0libc/test/src/math/performance_testing/exp10f_perf.cpp
+37-3libc/test/src/math/performance_testing/exp2f_perf.cpp
+11-0libc/test/src/math/performance_testing/CMakeLists.txt
+229-1175 files

LLVM/project b9f89c4 — clang/lib/CIR/CodeGen CIRGenRecordLayoutBuilder.cpp, clang/lib/CIR/Dialect/Transforms/TargetLowering CIRABIRewriteContext.cpp

[CIR] Support bool vectors in x86_64 calling-convention lowering (#230280)

Bool vectors are now sized one bit per element, as clang does, and are
passed and returned the way classic CodeGen passes them, including in
records.

Assisted-by: Cursor / Claude Opus 5.5
DeltaFile
+246-0clang/test/CIR/CodeGen/call-conv-lowering-x86_64-bool-vector.c
+74-52clang/lib/CIR/Dialect/Transforms/TargetLowering/CIRABIRewriteContext.cpp
+40-0clang/test/CIR/CodeGen/bool-vector-record-storage-nyi.c
+38-0clang/test/CIR/CodeGen/call-conv-lowering-x86_64-bool-vector-return.cpp
+28-0clang/test/CIR/Transforms/abi-lowering/x86_64-vector.cir
+24-0clang/lib/CIR/CodeGen/CIRGenRecordLayoutBuilder.cpp
+450-527 files not shown
+504-7413 files

LLVM/project 9aa425a — llvm/lib/CodeGen/SelectionDAG LegalizeVectorOps.cpp

[DAG] Remove unused include in LegalizeVectorOps. NFC (#230552)

Followup to #223935
DeltaFile
+0-1llvm/lib/CodeGen/SelectionDAG/LegalizeVectorOps.cpp
+0-11 files

LLVM/project 4207032 — llvm/lib/CodeGen MLRegAllocEvictAdvisor.cpp, llvm/test/CodeGen/MLRegAlloc dev-mode-logging.ll

[MLGO] Do not cap evictions when logging default advisor decisions (#230502)

Otherwise trace collection hits the Regs[CandidatePos].second assertion
on large functions (ones with many evictions)
DeltaFile
+7-0llvm/test/CodeGen/MLRegAlloc/dev-mode-logging.ll
+4-2llvm/lib/CodeGen/MLRegAllocEvictAdvisor.cpp
+11-22 files

LLVM/project 6339dfe — flang/lib/Semantics check-io.cpp

Drop check-io.cpp from this PR.
DeltaFile
+2-2flang/lib/Semantics/check-io.cpp
+2-21 files

LLVM/project abcdf64 — llvm/docs LangRef.md

[LangRef] Clarify poison input elements give poison in llvm.mask.beforefirst

This matches #223935 and reflects the generic expansion.
DeltaFile
+2-0llvm/docs/LangRef.md
+2-01 files

LLVM/project 51941c4 — llvm/test/Verifier sat-intrinsics.ll scatter_gather.ll

[Verifier][NFC] Merge negative intrinsic signature tests into a single file (#230160)

See the comment
https://github.com/llvm/llvm-project/pull/229880#pullrequestreview-5447604943.

Aided by DeepSeek-V4.1-Flash.
DeltaFile
+740-0llvm/test/Verifier/invalid_intrinsic_signatures.ll
+0-172llvm/test/Verifier/intrinsic-bad-arg-type1.ll
+0-101llvm/test/Verifier/invalid-vp-intrinsics.ll
+0-79llvm/test/Verifier/intrinsic-arg-overloading-struct-ret.ll
+0-67llvm/test/Verifier/scatter_gather.ll
+0-44llvm/test/Verifier/sat-intrinsics.ll
+740-46328 files not shown
+741-91034 files

LLVM/project c90b641 — llvm/cmake/modules MLGOLower.cmake, llvm/lib/Analysis CMakeLists.txt

[mlgo] Allow passing pre-emitc-ed models (#227941)

Support pre-lowering models and then passing them via the exact same mechanism - i.e. `LLVM_MLGO_MODELS`. The extension for the pre-generated ones needs to be `.inc`. High level, this just skips trying to run the mlir toolchain over those. Mixing `.inc` and `.mlir` is supported. The mlir toolchain isn't required unless `.mlir` are passed in the list.

  
As a result we can test the AOT case in regular builds. We just always append to the `LLVM_MLGO_MODELS`list the test models, with an "ugly" command line flag (a `_test` prefix). Each pass just lists the mlir test model and its corresponding .inc as part of the call to `MLGOLower`.

The bulk of the change is changing tests accordingly, and the addition of the same models we use in the mlir case, but EmitC-ed.

A subsequent change will remove listing the mlir models in llvm-zorg, since they now get auto-appended to the list when the mlir tools are specified. In the interim (after this change lands but before we change zorg) the ml-rel bot won't get red because we register the models under a dfferent name on zorg.

Issue #199007
DeltaFile
+79-58llvm/cmake/modules/MLGOLower.cmake
+98-0llvm/lib/Analysis/models/inline-oz-test-model.inc
+49-0llvm/lib/Analysis/models/regalloc-eviction-test-model.inc
+13-12llvm/lib/CodeGen/CMakeLists.txt
+13-12llvm/lib/Analysis/CMakeLists.txt
+3-21llvm/lib/CodeGen/MLRegAllocEvictAdvisor.cpp
+255-10320 files not shown
+324-19526 files

LLVM/project d5804b8 — mlir/lib/Dialect/X86/Utils X86Utils.cpp, mlir/test/Dialect/X86/AMX vector-contract-to-tiled-dp.mlir

[mlir][x86] Fix result tracing through loops (#230229)

Fixes the search for the write of a contraction result in the AMX
lowering when the result is passed through loop args.

The value was followed to the wrong loop result, which either found the
wrong write or crashed when the loop has few results.

Assisted-by: Claude
DeltaFile
+192-0mlir/test/Dialect/X86/AMX/vector-contract-to-tiled-dp.mlir
+4-2mlir/lib/Dialect/X86/Utils/X86Utils.cpp
+196-22 files

LLVM/project f5bd08d — llvm/lib/Target/AArch64 AArch64InstrAtomics.td

[AArch64] Reduce AddedComplexity value on atomic hint patterns (NFC) (#230478)

Also removes the GISelPredicateCode and replaces it with GISelShouldIgnore
until support is added for the atomic hint instructions in GlobalISel.

Improves CTMark geomean by -0.12% on aarch64-O0-g:

https://llvm-compile-time-tracker.com/compare.php?from=5c31fc696c2dca47458b0da7a05d0b487386242b&to=671900e9ad62a3221d5de0ff05fa9099aa5f7a0b&stat=instructions:u
DeltaFile
+1-4llvm/lib/Target/AArch64/AArch64InstrAtomics.td
+1-41 files

LLVM/project a5d2f05 — mlir/lib/Conversion/ArithToLLVM ArithToLLVM.cpp, mlir/test/Conversion/ArithToLLVM constant-index-bitwidth.mlir index-bitwidth.mlir

[mlir][ArithToLLVM] Fix index lowering for addui_extended (#223265)

Lowering `arith.addui_extended` with `index` operands fails because the
LLVM
result struct uses the unconverted sum type, even though the operands
have
already been converted.

Convert the sum type before constructing the LLVM result struct. Extend
the
existing index-bitwidth tests to cover `arith.addui_extended` at 32-,
64-, and
128-bit widths, reusing their RUN lines. Rename
`constant-index-bitwidth.mlir`
to `index-bitwidth.mlir` to reflect the broader coverage.

This preserves the operation's existing `index` support and uses the
converted
index width for the overflow intrinsic. Related discussion of `index`

    [12 lines not shown]
DeltaFile
+81-0mlir/test/Conversion/ArithToLLVM/index-bitwidth.mlir
+0-67mlir/test/Conversion/ArithToLLVM/constant-index-bitwidth.mlir
+2-1mlir/lib/Conversion/ArithToLLVM/ArithToLLVM.cpp
+83-683 files

LLVM/project b676f4b — lldb/source/Plugins/SymbolFile/NativePDB SymbolFileNativePDB.cpp, lldb/test/API/functionalities/target_var TestTargetVar.py

[lldb][NativePDB] Only list a compile unit's own global variables (#230126)

`SymbolFileNativePDB::ParseVariablesForCompileUnit` adds every global
data symbol of the globals stream to whichever compile unit it's asked
about. The globals stream covers the whole PDB, so every compile unit
reports all globals of the program, including the CRT's, and target
variable list hundreds of them.

This patch only keeps the variables whose owning compile unit is the one
being parsed.

Requires:
- https://github.com/llvm/llvm-project/pull/230132

Fixes `TestTargetVar` with PDB debug info.

rdar://189620344
DeltaFile
+2-1lldb/source/Plugins/SymbolFile/NativePDB/SymbolFileNativePDB.cpp
+2-0lldb/test/API/functionalities/target_var/TestTargetVar.py
+4-12 files

LLVM/project d962cab — lldb/source/Plugins/SymbolFile/NativePDB SymbolFileNativePDB.cpp, lldb/test/API/lang/cpp/class_static TestStaticVariables.py

[lldb][NativePDB] Look up global variables by qualified name (#230134)

`SymbolFileNativePDB::FindGlobalVariables` looks the name up in an index
keyed by basename, so a qualified name such as `A::g_points` never
matches (`target variable A::g_points` fails with `"can't find global
variable"`).

This patch splits the name with
`CPlusPlusLanguage::ExtractContextAndIdentifier`, looks up the basename,
and when the name was qualified only keeps variables whose qualified
name contains it, as `SymbolFileDWARF::FindGlobalVariables` does.

Fixes `TestStaticVariables` with PDB debug info.

rdar://189618458
DeltaFile
+11-1lldb/source/Plugins/SymbolFile/NativePDB/SymbolFileNativePDB.cpp
+2-0lldb/test/API/lang/cpp/class_static/TestStaticVariables.py
+13-12 files

LLVM/project 362016a — lldb/source/Plugins/SymbolFile/NativePDB SymbolFileNativePDB.cpp, lldb/test/API/lang/cpp/function_refs TestFunctionRefs.py main.cpp

[lldb][NativePDB] Use the public symbol as the mangled name of globals (#230132)

`CreateGlobalVariable` passes `"::" + name` as the mangled name of every
global, and `Variable::GetName` prefers the mangled name, so every
global shows a leading `:: ((int) ::C::abc = 123, (&::ref = ...)`,
SBValue::GetName() returning "::i")`.

This patch uses the `S_PUB32` at the variable's address as the mangled
name when there is one, and no mangled name otherwise, which matches
what lldb shows for MSVC ABI globals with DWARF. NativePDB shell tests
are updated, and `TestFunctionRefs` now accepts both forms (the DWARF
output depends on the C++ ABI), which also removes its Windows XFAIL.

rdar://189620487
DeltaFile
+60-60lldb/test/Shell/SymbolFile/NativePDB/globals-fundamental.cpp
+12-12lldb/test/API/lang/cpp/function_refs/main.cpp
+9-5lldb/source/Plugins/SymbolFile/NativePDB/SymbolFileNativePDB.cpp
+4-4lldb/test/Shell/SymbolFile/NativePDB/function-types-builtins.cpp
+1-3lldb/test/API/lang/cpp/function_refs/TestFunctionRefs.py
+1-1lldb/test/Shell/SymbolFile/NativePDB/udt-layout.test
+87-856 files

LLVM/project f77d267 — bolt/test/AArch64 relax-calls.s relax-calls-large-binary.s

[BOLT][AArch64] Shrink large binary call relaxation test (#230526)

The test intermittently times out on the bolt-aarch64-ubuntu-nfc
builder, so I am making the input binary a tad smaller.
DeltaFile
+0-81bolt/test/AArch64/relax-calls.s
+81-0bolt/test/AArch64/relax-calls-large-binary.s
+81-812 files

LLVM/project 477b9ad — offload/include/OpenMP OffloadRTL.h, offload/libomptarget OffloadRTL.cpp

add fullstops
DeltaFile
+1-1offload/libomptarget/OffloadRTL.cpp
+1-1offload/include/OpenMP/OffloadRTL.h
+2-22 files

LLVM/project 1d38b0b — offload/include PluginManager.h, offload/include/OpenMP OffloadRTL.h

[offload][nfc] Pull OpenMP's InteropTbl out of PluginManager

PluginManager is shared with OpenACC, so the OpenMP interop table moves
to OmpPluginManager in libomptarget. The OpenMP PM is defined there, and
interop cleanup is registered from initRuntime.
DeltaFile
+12-4offload/libomptarget/OffloadRTL.cpp
+10-1offload/include/OpenMP/OffloadRTL.h
+0-7offload/libompaccsupport/PluginManager.cpp
+0-5offload/include/PluginManager.h
+22-174 files

LLVM/project 97f3dc8 — offload/include PluginManager.h, offload/include/OpenMP OffloadRTL.h

make loadImagesOntoDevice a member

wip
DeltaFile
+7-7offload/libompaccsupport/PluginManager.cpp
+3-6offload/include/PluginManager.h
+6-0offload/include/OpenMP/OffloadRTL.h
+16-133 files

LLVM/project d32e4b7 — llvm/lib/Target/AMDGPU GCNVOPDUtils.cpp, llvm/test/CodeGen/AMDGPU vopd-dot2-commute-imm-src1.mir

[AMDGPU] Address review comments

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+9-5llvm/lib/Target/AMDGPU/GCNVOPDUtils.cpp
+1-1llvm/test/CodeGen/AMDGPU/vopd-dot2-commute-imm-src1.mir
+10-62 files

LLVM/project 9a547a8 — llvm/lib/Target/AMDGPU SIInstrInfo.h GCNCreateVOPD.cpp, llvm/test/CodeGen/AMDGPU llvm.amdgcn.fdot2.f32.bf16.ll llvm.amdgcn.fdot2.ll

[AMDGPU] Form VOPD dot2 pairs with a literal in src1

A V_DOT2 with a register in src0 and a literal in src1 is matched as if
commuted, and commuted when the VOPD pair is built. This recovers the
pairs lost once MachineCSE stopped leaving those literals in src0.

Co-Authored-By: Claude <noreply at anthropic.com>
DeltaFile
+149-0llvm/test/CodeGen/AMDGPU/vopd-dot2-commute-imm-src1.mir
+54-19llvm/lib/Target/AMDGPU/GCNVOPDUtils.cpp
+12-35llvm/test/CodeGen/AMDGPU/llvm.amdgcn.fdot2.ll
+3-6llvm/test/CodeGen/AMDGPU/llvm.amdgcn.fdot2.f32.bf16.ll
+9-0llvm/lib/Target/AMDGPU/GCNCreateVOPD.cpp
+2-2llvm/lib/Target/AMDGPU/SIInstrInfo.h
+229-621 files not shown
+232-627 files