LLVM/project c5f99a1 — clang/test/CodeGen/X86 sse41-builtins-constrained.c sse41-builtins.c

[clang][X86] Fix round builtins tests (#226708)

Fixed minor issues in round builtins tests. Noticed them while working on #215787.
DeltaFile
+9-9clang/test/CodeGen/X86/sse41-builtins.c
+6-6clang/test/CodeGen/X86/sse41-builtins-constrained.c
+15-152 files

LLVM/project 650a674 — clang/test/CIR/CodeGen pragma-fenv_access.c, llvm/include/llvm/ADT DenseMap.h

Rebase

Created using spr 1.3.7
DeltaFile
+977-870llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+1,308-0llvm/test/CodeGen/AMDGPU/llvm.amdgcn.cvt.scale.pk32.gfx13.ll
+339-431llvm/include/llvm/ADT/DenseMap.h
+367-367clang/test/CIR/CodeGen/pragma-fenv_access.c
+696-0llvm/test/Transforms/SLPVectorizer/X86/strength-reducible-address.ll
+694-0llvm/test/CodeGen/AMDGPU/rewrite-vgpr-mfma-to-agpr-spill-multi-store-codegen.ll
+4,381-1,668752 files not shown
+19,313-6,537758 files

LLVM/project 9146ac4 — llvm/lib/Target/AMDGPU GCNHazardRecognizer.h GCNHazardRecognizer.cpp, llvm/test/CodeGen/AMDGPU vperm-pk16-postmisched-hazard.mir vperm-pk16-sched-softcost.mir

[AMDGPU] Prefer a safe V_PERM_PK16 follower in the scheduler (gfx1250/gfx1251)

Stacked on the post-RA V_PERM_PK16 hazard fixup. V_PERM_PK16 must be
immediately followed by a "safe" instruction (see
SIInstrInfo::isVPermPk16SafeInstr) or the post-RA fixup has to insert a
forced-EXEC V_NOP. Teach GCNHazardRecognizer to bias a safe follower into
the slot right after a V_PERM_PK16 so that V_NOP can be avoided.

Assisted-by: Opus 4.8 Medium
DeltaFile
+56-0llvm/test/CodeGen/AMDGPU/vperm-pk16-sched-softcost.mir
+46-3llvm/lib/Target/AMDGPU/GCNHazardRecognizer.cpp
+33-0llvm/test/CodeGen/AMDGPU/vperm-pk16-postmisched-hazard.mir
+21-0llvm/lib/Target/AMDGPU/GCNHazardRecognizer.h
+156-34 files

LLVM/project dc306a2 — clang/lib/CodeGen CGExprScalar.cpp, clang/test/CodeGenHLSL/BasicFeatures VectorElementwiseCast.hlsl MatrixElementTypeCast.hlsl

[HLSL] Build elementwise cast results from poison (#225591)

I noticed this unnecessary alloca while doing this pr:
https://github.com/llvm/llvm-project/pull/225519

The change is to initialize vector and matrix elementwise cast results
with poison instead of loading uninitialized temporary storage.

We do this because every result element is overwritten before use,
making the temporary allocation and load unnecessary.
DeltaFile
+77-93clang/test/CodeGenHLSL/BasicFeatures/MatrixElementTypeCast.hlsl
+7-21clang/test/CodeGenHLSL/BasicFeatures/VectorElementwiseCast.hlsl
+2-4clang/lib/CodeGen/CGExprScalar.cpp
+86-1183 files

LLVM/project cdbca60 — llvm/lib/Target/NVPTX NVPTXTargetTransformInfo.h, llvm/lib/Transforms/Scalar InferAddressSpaces.cpp

[InferAddressSpaces] Check volatile support in the destination AS (#224702)

Perviously we nonsensically checked the old address space...

Avoid introducing a regression by adding ADDRESS_SPACE_SHARED_CLUSTER to
NVPTX hasVolatileVariant.
DeltaFile
+167-0llvm/test/Transforms/InferAddressSpaces/NVPTX/volatile.ll
+58-0llvm/test/CodeGen/NVPTX/infer-volatile-address-spaces.ll
+9-10llvm/lib/Transforms/Scalar/InferAddressSpaces.cpp
+8-8llvm/lib/Target/NVPTX/NVPTXTargetTransformInfo.h
+9-6llvm/test/CodeGen/NVPTX/lower-byval-args.ll
+8-3llvm/test/CodeGen/NVPTX/local-stack-frame.ll
+259-271 files not shown
+263-297 files

LLVM/project 5c0be58 — clang/cmake/modules CMakeLists.txt

fixup! [CMake] Link MLIR if CLANG_ENABLE_CIR
DeltaFile
+1-1clang/cmake/modules/CMakeLists.txt
+1-11 files

LLVM/project 385d074 — llvm/utils/gn/secondary/clang/unittests/Basic BUILD.gn

[gn build] Port 5c20fe98552a (#227114)
DeltaFile
+1-0llvm/utils/gn/secondary/clang/unittests/Basic/BUILD.gn
+1-01 files

LLVM/project 8d752fb — llvm/lib/Transforms/Vectorize SLPVectorizer.cpp, llvm/lib/Transforms/Vectorize/SLPVectorizer SLPMemoryUtils.cpp SLPCompatibilityAnalysis.cpp

[SLP][NFC]Broader use of filter/isa ranges

Most of loops of the form "for (...) { if (cond) continue; ... }"
become make_filter_range, and dyn_cast-and-skip loops become
make_isa_range.

Assisted-by: Cursor

Reviewers: bababuck

Pull Request: https://github.com/llvm/llvm-project/pull/226535
DeltaFile
+689-820llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+19-24llvm/lib/Transforms/Vectorize/SLPVectorizer/SLPUtils.cpp
+16-16llvm/lib/Transforms/Vectorize/SLPVectorizer/SLPShuffleAnalysis.h
+6-11llvm/lib/Transforms/Vectorize/SLPVectorizer/SLPCompatibilityAnalysis.cpp
+3-3llvm/lib/Transforms/Vectorize/SLPVectorizer/SLPMemoryUtils.cpp
+733-8745 files

LLVM/project 348d23c — libcxx/include/__chrono convert_to_tm.h

[libc++][chrono] Initialize optional tm_zone member in __convert_to_tm (#220068)

Defensively initialize `struct tm`'s optional `tm_zone` member in
`__convert_to_tm` before passing it to `strftime` or `time_put`.
    
`tm_zone` is a BSD extension standardized in POSIX.1-2024, but is not
part
of standard C or C++ (and is absent on Windows). Instead of guarding
with
`#ifdef __GLIBC__`, use `if constexpr (requires ...)` so it is
initialized
on all platforms that provide it (such as Bionic, musl, macOS, and BSDs,
whether typed as `const char*` or `char*`).
    
Assisted-by: Gemini
DeltaFile
+4-6libcxx/include/__chrono/convert_to_tm.h
+4-61 files

LLVM/project 4f92dab — mlir/docs/Tools mlir-reduce.md, mlir/lib/Reducer ReductionNode.cpp

[MLIR] Fix mlir-reduce splitting smallest range instead of the largest one (#214738)

The function `max_element` expects comparison lambda function to return
true if first argument is **less** than the second one, which is the
opposite to the current code. 
DeltaFile
+17-0mlir/test/mlir-reduce/reduction-tree/except-last.mlir
+6-5mlir/docs/Tools/mlir-reduce.md
+5-4mlir/test/mlir-reduce/reduction-tree/doc-example.mlir
+7-0mlir/test/mlir-reduce/script/except-last.sh
+1-1mlir/lib/Reducer/ReductionNode.cpp
+36-105 files

LLVM/project dd2841c — clang/cmake/modules CMakeLists.txt

fixup! [CMake] Link MLIR if CLANG_ENABLE_CIR
DeltaFile
+15-2clang/cmake/modules/CMakeLists.txt
+15-21 files

LLVM/project 6674876 — clang/test/CIR/CodeGen pragma-fenv_access.c, libcxx/docs/DesignDocs AtomicDesign.md AtomicDesign.rst

Rebase

Created using spr 1.3.7
DeltaFile
+0-797libcxx/docs/DesignDocs/AtomicDesign.rst
+790-0libcxx/docs/DesignDocs/AtomicDesign.md
+464-298llvm/include/llvm/ADT/DenseMap.h
+367-367clang/test/CIR/CodeGen/pragma-fenv_access.c
+696-0llvm/test/Transforms/SLPVectorizer/X86/strength-reducible-address.ll
+694-0llvm/test/CodeGen/AMDGPU/rewrite-vgpr-mfma-to-agpr-spill-multi-store-codegen.ll
+3,011-1,4621,380 files not shown
+32,735-12,6291,386 files

LLVM/project 16036fd — libc/src/grp CMakeLists.txt getgrouplist.h, libc/test/src/grp CMakeLists.txt getgrouplist_test.cpp

[libc] Add getgrouplist entrypoint (#226958)

Add the getgrouplist entrypoint from <grp.h> (BSD extension), which
scans the group database to obtain the list of groups to which a user
belongs.

The base group passed by the caller is included unconditionally, and
supplementary groups for the user are gathered with duplicate group IDs
suppressed. A small-buffer-optimised container avoids heap allocations
for users belonging to up to 32 groups, falling back to dynamic
allocation when more groups are present. Note that this fallback does
not use AllocChecker because it relies on realloc to grow the buffer.
The lookup uses a scoped database stream to avoid disturbing concurrent
iteration.

Assisted-by: Automated tooling, human reviewed.
DeltaFile
+216-0libc/test/src/grp/getgrouplist_test.cpp
+121-0libc/src/grp/grp_utils.cpp
+55-0libc/src/grp/getgrouplist.cpp
+29-0libc/test/src/grp/CMakeLists.txt
+26-0libc/src/grp/getgrouplist.h
+24-0libc/src/grp/CMakeLists.txt
+471-07 files not shown
+491-013 files

LLVM/project 3c5f2ac — llvm/test/CodeGen/AMDGPU llvm.amdgcn.intersect_ray.ll calling-conventions.ll

improve optimization, added missing testlines, and address comment
DeltaFile
+4,860-4,866llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+1,684-1,713llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+2,413-368llvm/test/CodeGen/AMDGPU/immv216.ll
+1,222-1,071llvm/test/CodeGen/AMDGPU/atomic_optimizations_global_pointer.ll
+384-779llvm/test/CodeGen/AMDGPU/calling-conventions.ll
+603-6llvm/test/CodeGen/AMDGPU/llvm.amdgcn.intersect_ray.ll
+11,166-8,80327 files not shown
+12,696-10,16233 files

LLVM/project bf1009b — llvm/lib/CodeGen/SelectionDAG FastISel.cpp, llvm/lib/Target/X86 X86FastISel.cpp

FastISel: Assert the emitted instruction defines the result (#226502)

The fallback path copied the result out of implicit_defs()[0], assuming
the first implicit physical register def is the result. That is an X86
assumption about MUL/IMUL, and it is unreachable for all but
fastEmitInst_r: FastISelEmitter skips any instruction whose first
operand is not an output register, so every opcode reaching these
helpers from generated code has an explicit def.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+30-95llvm/lib/CodeGen/SelectionDAG/FastISel.cpp
+57-28llvm/lib/Target/X86/X86FastISel.cpp
+4-2llvm/utils/TableGen/FastISelEmitter.cpp
+91-1253 files

LLVM/project 8085363 — offload/test/offloading shared_lib_global_var.c

[offload][omp] Mark shlib_global_var test unsupported for NVIDIA (#227098)

The shared library test introduced in #226980 seems to fail with NVIDIA
backend. Marking it unsupported.
DeltaFile
+1-0offload/test/offloading/shared_lib_global_var.c
+1-01 files

LLVM/project 2aa363f — llvm/lib/Transforms/Vectorize VPlanUtils.cpp, llvm/test/Transforms/LoopVectorize/VPlan execution-frequencies-match-bfi.ll

[VPlan] Compute edge probabilities via getEdgeProbabilitiesFromWeights. (#226981)

Use getEdgeProbabilitiesFromWeights (added in
https://github.com/llvm/llvm-project/pull/226560) to compute edge
probabilities like BFI.

This makes sure the VPlan-based logic combines weights of parallel edges
like BFI.

PR: https://github.com/llvm/llvm-project/pull/226981
DeltaFile
+16-35llvm/lib/Transforms/Vectorize/VPlanUtils.cpp
+8-11llvm/test/Transforms/LoopVectorize/VPlan/execution-frequencies-match-bfi.ll
+24-462 files

LLVM/project d8a2498 — llvm/include/llvm/IR Verifier.h, llvm/lib/Analysis TypeBasedAliasAnalysis.cpp

[Verifier] Validate !tbaa.struct metadata (#225910)

Check that !tbaa.struct operands come in (offset, size, tag) triples
with constant offset and size.
DeltaFile
+39-0llvm/lib/IR/Verifier.cpp
+17-5llvm/test/Verifier/tbaa-struct.ll
+10-9llvm/test/Transforms/InstCombine/struct-assign-tbaa.ll
+3-6llvm/lib/Analysis/TypeBasedAliasAnalysis.cpp
+1-1llvm/test/Transforms/AtomicExpand/AMDGPU/expand-atomic-i16.ll
+1-0llvm/include/llvm/IR/Verifier.h
+71-216 files

LLVM/project 0a7dc52 — llvm/include/llvm/ADT DenseMap.h

[ADT] Remove CRTP from DenseMapBase (NFC) (#227063)

This patch removes CRTP from DenseMapBase by replacing DerivedT with
StorageT (DenseMapStorage or SmallDenseMapStorage).

DenseMapBase now owns the Storage member by composition and provides
the entire user-facing map interface -- from constructors, the
destructor, and operator= to find, try_emplace, and erase.  DenseMap
and SmallDenseMap simply specialize DenseMapBase with their respective
storage types.

This completes the effort to replace CRTP in DenseMapBase with
composition (see #168255, #226664, and #226882).

Assisted-by: Antigravity
DeltaFile
+95-230llvm/include/llvm/ADT/DenseMap.h
+95-2301 files

LLVM/project 5ab44c4 — clang/include/clang/Basic BuiltinsAMDGPU.td BuiltinsAMDGPUDocs.td, clang/test/CodeGenOpenCL builtins-amdgcn-gfx13-w32-err.cl builtins-amdgcn-gfx13-err.cl

[AMDGPU] Add intrinsics and builtins for v_cvt_scale_pk32_* instructions (#222617)
DeltaFile
+1,308-0llvm/test/CodeGen/AMDGPU/llvm.amdgcn.cvt.scale.pk32.gfx13.ll
+108-0clang/include/clang/Basic/BuiltinsAMDGPUDocs.td
+58-0clang/test/CodeGenOpenCL/builtins-amdgcn-gfx13.cl
+26-0clang/test/CodeGenOpenCL/builtins-amdgcn-gfx13-err.cl
+24-0clang/include/clang/Basic/BuiltinsAMDGPU.td
+20-0clang/test/CodeGenOpenCL/builtins-amdgcn-gfx13-w32-err.cl
+1,544-07 files not shown
+1,596-813 files

LLVM/project 620ad7e — llvm/lib/Transforms/InstCombine InstCombineCasts.cpp, llvm/test/Transforms/InstCombine truncating-saturate.ll

[InstCombine] Match swapped form of truncating saturation clamp (#226614)

Extend the fold added in #189703 to also handle the inverted select:
trunc (select (icmp ugt A, DestTy_umax), sext(icmp sgt A, 0), A) -->
trunc (smin (smax (0, A), DestTy_umax))

InstCombine canonicalizes (A & NegPow2) != 0 into the ult form with
swapped select operands, but if SCCP first rewrites the compare as
icmp uge A, C, InstCombine only turns it into icmp ugt and never swaps
the select, so the original fold is missed.

While here, match the compare constant with m_APInt instead of
m_Constant + getUniqueInteger, and build TruncatedMax directly with
APInt::getLowBitsSet. Comparing the constant against TruncatedMax + 1
or TruncatedMax makes the separate zero check unnecessary.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
DeltaFile
+130-0llvm/test/Transforms/InstCombine/truncating-saturate.ll
+20-12llvm/lib/Transforms/InstCombine/InstCombineCasts.cpp
+150-122 files

LLVM/project d66a193 — clang/lib/CIR/CodeGen CIRGenAtomic.cpp, clang/test/CIR/CodeGen atomic.c

[CIR] Correct 2 lowering bugs of atomic cmp-xchng builtins (#227054)

This patch fixese two bugs that showed up in a benchmark.

First; convertToAtomicIntPointer was zero-filling the source object
directly, rather than the temporary. The result was that anything that
would not be overwritten thanks to the power-of-2 write, would be
incorrect, and corrupted.

Second; emitAtomicCmpXchg didn't set the 'old' value back into the real
object. This ends up doing an additional argument on this function that
better matches classic-codegen.
DeltaFile
+77-54clang/lib/CIR/CodeGen/CIRGenAtomic.cpp
+66-6clang/test/CIR/CodeGen/atomic.c
+143-602 files

LLVM/project 844f933 — libc/src/__support/pwd flat_file_db.h, libc/src/pwd CMakeLists.txt getpwent_r.h

[libc] Add getpwent_r entrypoint (#226966)

Add the reentrant password database iteration entrypoint getpwent_r.

getpwent_r reads the next password database record from the stream into
the caller-supplied struct passwd and buffer, returning 0 on success,
ENOENT at end-of-file, or an error number (such as ERANGE) on failure.
When a record exceeds the caller buffer size, FlatFileDatabase::getnext
rewinds the stream to the beginning of that record so that a subsequent
retry with a larger buffer reads the same entry.

* Add getpwent_r entrypoint and pwd::read_next fixed-buffer overload
* Rewind stream on ERANGE in FlatFileDatabase::getnext(EntryType *,
span<char>)
* Define getpwent_r in include/pwd.yaml and Linux entrypoints.txt
* Add unit tests for getpwent_r

Assisted-by: Automated tooling, human reviewed.
DeltaFile
+188-0libc/test/src/pwd/getpwent_r_test.cpp
+46-0libc/src/pwd/getpwent_r.cpp
+28-0libc/src/pwd/getpwent_r.h
+26-0libc/test/src/pwd/CMakeLists.txt
+14-3libc/src/__support/pwd/flat_file_db.h
+17-0libc/src/pwd/CMakeLists.txt
+319-38 files not shown
+341-314 files

LLVM/project 57012c6 — clang/lib/CIR/Dialect/Transforms/TargetLowering CIRABIRewriteContext.cpp

[CIR] Fix order of creation so that lit test will not fail (#226706)

This patch fixes the order of creation, otherwise the compiler may
evaluate one before the other and the lit test fail.
DeltaFile
+3-2clang/lib/CIR/Dialect/Transforms/TargetLowering/CIRABIRewriteContext.cpp
+3-21 files

LLVM/project 638ffb6 — llvm/lib/CodeGen/SelectionDAG FastISel.cpp, llvm/utils/TableGen FastISelEmitter.cpp

Stop routing def-less instructions through fastEmitInst_*
DeltaFile
+2-10llvm/lib/CodeGen/SelectionDAG/FastISel.cpp
+4-2llvm/utils/TableGen/FastISelEmitter.cpp
+6-122 files

LLVM/project 86c1241 — llvm/lib/CodeGen/SelectionDAG FastISel.cpp

FastISel: Assert the emitted instruction defines the result

The fallback path copied the result out of implicit_defs()[0], assuming
the first implicit physical register def is the result. That is an X86
assumption about MUL/IMUL, and it is unreachable for all but
fastEmitInst_r: FastISelEmitter skips any instruction whose first
operand is not an output register, so every opcode reaching these
helpers from generated code has an explicit def.

Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+28-85llvm/lib/CodeGen/SelectionDAG/FastISel.cpp
+28-851 files

LLVM/project 53d24b5 — llvm/lib/Target/X86 X86FastISel.cpp

Assert fastEmitInst_rrrr's instruction defines the result
DeltaFile
+6-16llvm/lib/Target/X86/X86FastISel.cpp
+6-161 files

LLVM/project ebdd073 — llvm/lib/Target/X86 X86FastISel.cpp

Select 8-bit multiply in X86FastISel
DeltaFile
+22-0llvm/lib/Target/X86/X86FastISel.cpp
+22-01 files

LLVM/project 24b12df — llvm/lib/Target/X86 X86FastISel.cpp

Add a FastISel helper for with-overflow multiply emission
DeltaFile
+29-12llvm/lib/Target/X86/X86FastISel.cpp
+29-121 files

LLVM/project e5b60e5 — llvm/include/llvm/CodeGen SDPatternMatch.h, llvm/lib/CodeGen/SelectionDAG DAGCombiner.cpp

[SDPatternMatch] Make m_SetCC work like m_ICmp from IR PatternMatch. (#226623)

The condition code is stored an operand, but we don't need to expose
that to the interface.

This adds 2 signatures of m_Setcc, one that takes 2 operands and matches
any condition code and one that takes the matched condition code by
reference. For m_SpecificCondCode cases, I've added m_SpecificSetCC.

Similar changes have been applied to m_SelectCC and m_SelectCCLike.

Out of tree targets will need to update to the new interface.

Assisted-by: Claude
DeltaFile
+119-49llvm/include/llvm/CodeGen/SDPatternMatch.h
+71-24llvm/unittests/CodeGen/SelectionDAGPatternMatchTest.cpp
+22-24llvm/lib/CodeGen/SelectionDAG/DAGCombiner.cpp
+20-21llvm/lib/Target/X86/X86ISelLowering.cpp
+8-9llvm/lib/Target/NVPTX/NVPTXISelLowering.cpp
+7-8llvm/lib/Target/WebAssembly/WebAssemblyISelLowering.cpp
+247-1352 files not shown
+258-1458 files