LLVM/project fadd7e6orc-rt/include/orc-rt/support Compiler.h

[orc-rt] Remove unused ORC_RT_HAS_CPP_ATTRIBUTE (#220747)
DeltaFile
+0-10orc-rt/include/orc-rt/support/Compiler.h
+0-101 files

LLVM/project f33ef3allvm/lib/Target/AArch64 AArch64SchedC1Ultra.td AArch64SchedC1Premium.td, llvm/test/CodeGen/AArch64 misched-c1sme.ll

[AArch64] Use NoSchedPred for SME instructions in C1 scheduling models. (#220553)

We can have the scheduling model enabled without SME using -mtune, which
means that no scheduling information was present for any instructions
that execute in either SME or the core. AFAICT the predicate should be
NoSchedPred, as any instructions should be using the non-streaming
scheduling info when not in a SME function.

Fixes #220070
Fixes #220067
DeltaFile
+24-0llvm/test/CodeGen/AArch64/misched-c1sme.ll
+1-1llvm/lib/Target/AArch64/AArch64SchedC1Ultra.td
+1-1llvm/lib/Target/AArch64/AArch64SchedC1Premium.td
+26-23 files

LLVM/project 70519bcorc-rt/include/orc-rt-c/support Logging.h Compiler.h

[orc-rt] Drop the _C_ prefix from ORC_RT_C_FORMAT_PRINTF (#220592)

The macro is not C-specific. Rename it to ORC_RT_FORMAT_PRINTF and
update its two uses in Logging.h.
DeltaFile
+3-3orc-rt/include/orc-rt-c/support/Compiler.h
+2-2orc-rt/include/orc-rt-c/support/Logging.h
+5-52 files

LLVM/project aba5ec2llvm/lib/Target/NVPTX NVPTXTargetTransformInfo.cpp, llvm/test/Transforms/InstCombine/NVPTX nvvm-intrins.ll

[NVPTX] Fold abs into redux intrinsics

Fold llvm.fabs into the absolute-value variants of floating-point
redux min/max intrinsics during InstCombine.
DeltaFile
+37-1llvm/lib/Target/NVPTX/NVPTXTargetTransformInfo.cpp
+32-0llvm/test/Transforms/InstCombine/NVPTX/nvvm-intrins.ll
+69-12 files

LLVM/project 53b30bfllvm/lib/Target/AMDGPU SOPInstructions.td

[AMDGPU] Use named operands in SOP1_Real. NFC (#220705)
DeltaFile
+14-13llvm/lib/Target/AMDGPU/SOPInstructions.td
+14-131 files

LLVM/project 81bf35dflang/lib/Optimizer/Transforms/CUDA CUFAllocDelay.cpp, flang/test/Transforms/CUF cuf-alloc-delay.fir

[flang][cuda] Delay descriptor alloc when addressed reused on host/device (#220534)

CSE can share one fir.coordinate_of between the host-association capture
store and a later fir.load. Treating that coordinate_of as a real use
made cuf-alloc-delay think the movable group depended on an operand at
the sink point, so the device descriptor stayed at function entry and
cudaMallocManaged ran before cudaSetDevice.

Count only users of the slot address that actually read it. Stores that
populate the tuple still sink with the allocation group.
DeltaFile
+38-1flang/test/Transforms/CUF/cuf-alloc-delay.fir
+11-11flang/lib/Optimizer/Transforms/CUDA/CUFAllocDelay.cpp
+49-122 files

LLVM/project 3b28702llvm/lib/Transforms/Vectorize VPlanTransforms.cpp

[VPlan] Allow non-live-in IV offsets when simplifying latch cond (NFC). (#220734)

simplifyBranchConditionForVFAndUF matches the canonical IV increment
plus an offset, which epilogue vectorization adds to resume the
canonical IV at the vector trip count of the main vector loop. Require
the offset to be defined outside the vector loop region instead of
requiring it to be a live-in; that is what makes it available in the
preheader..

This is NFC today, but prepares for modeling the full epilogue skeleton
in VPlan, which requires adding phi nodes in the preheader before
execute.
DeltaFile
+8-3llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
+8-31 files

LLVM/project 3dcc5b3llvm/test/Transforms/LoopVersioning preserved-analyses.ll

[LoopVersioning] Add missing verify-analysis-invalidation=false to test. (#220735)

Add -verify-analysis-invalidation=false to test added in
https://github.com/llvm/llvm-project/pull/220537 to fix expensive check
failures due to extra verification passes.
DeltaFile
+1-1llvm/test/Transforms/LoopVersioning/preserved-analyses.ll
+1-11 files

LLVM/project 32e83ealldb/source/Plugins/Platform/WebAssembly PlatformWasmProperties.td PlatformWasm.h, lldb/unittests/Platform CMakeLists.txt PlatformWasmTest.cpp

[lldb] Pass Wasm runtime-args before the port argument (#220700)

A runtime that dispatches on a leading subcommand, such as WasmKit's
`wasmkit run`, could not be driven directly: runtime-args landed after
the port argument, so the subcommand did too and the runtime rejected
it. Naming the subcommand required a wrapper script. Move runtime-args
ahead of the port argument so the setting can carry it.

Extract the command line assembly into PlatformWasm::MakeRuntimeCommand
so the ordering is covered by unit tests, and clarify that port-arg has
to carry its value in the same argument.
DeltaFile
+98-0lldb/unittests/Platform/PlatformWasmTest.cpp
+37-23lldb/source/Plugins/Platform/WebAssembly/PlatformWasm.cpp
+10-0lldb/source/Plugins/Platform/WebAssembly/PlatformWasm.h
+5-2lldb/source/Plugins/Platform/WebAssembly/PlatformWasmProperties.td
+5-0llvm/docs/ReleaseNotes.md
+2-0lldb/unittests/Platform/CMakeLists.txt
+157-252 files not shown
+158-288 files

LLVM/project 80c5684llvm/include/llvm/IR Module.h, llvm/unittests/IR ModuleTest.cpp

[llvm] Use ValueMap for ValueToGUIDMap (#220682)

ValueToGUIDMap currently uses a DenseMap which does not properly track
the deletion of Values, leaving dangling pointers in the map.

This change fixes this by using ValueMap. FollowRAUW is set to false to
match the current behavior of DenseMap.

This bug was discovered by a sanity test for deterministic compilation
where, depending on the allocator state, a newly allocated Value could
re-use the address of a previously deleted one, incorrectly inheriting
the GUID.
DeltaFile
+24-0llvm/unittests/IR/ModuleTest.cpp
+8-2llvm/include/llvm/IR/Module.h
+32-22 files

LLVM/project 48111b6llvm/test/Transforms/LoopVectorize div-exact.ll if-pred-stores.ll, llvm/test/Transforms/LoopVectorize/AArch64 conditional-branches-cost.ll

[VPlan] Skip branch term in masks for some preserved uniform edges
DeltaFile
+28-2,414llvm/test/Transforms/LoopVectorize/predicator.ll
+6-425llvm/test/Transforms/LoopVectorize/X86/cost-conditional-branches.ll
+27-120llvm/test/Transforms/LoopVectorize/if-pred-stores.ll
+16-130llvm/test/Transforms/LoopVectorize/div-exact.ll
+32-93llvm/test/Transforms/LoopVectorize/X86/predicated-replicate-feeding-cast.ll
+102-9llvm/test/Transforms/LoopVectorize/AArch64/conditional-branches-cost.ll
+211-3,19133 files not shown
+667-3,78039 files

LLVM/project b9a2f44llvm/lib/Transforms/Vectorize VPlanPredicator.cpp, llvm/test/Transforms/LoopVectorize hoist-predicated-loads.ll if-pred-stores.ll

[VPlan][Predicator] Preserve some uniform control flow

Implements "Partial Control-Flow Linearization" by Simon Moll and
Sebastian Hack.

That should allow implementation of an alternative to
https://github.com/llvm/llvm-project/pull/141900 based on this
functionality (see BOSCC in the paper).
DeltaFile
+1,514-204llvm/test/Transforms/LoopVectorize/VPlan/predicator.ll
+1,492-137llvm/test/Transforms/LoopVectorize/predicator.ll
+184-128llvm/test/Transforms/LoopVectorize/X86/cost-conditional-branches.ll
+236-18llvm/lib/Transforms/Vectorize/VPlanPredicator.cpp
+81-45llvm/test/Transforms/LoopVectorize/if-pred-stores.ll
+102-10llvm/test/Transforms/LoopVectorize/hoist-predicated-loads.ll
+3,609-54238 files not shown
+4,284-79344 files

LLVM/project 71a85eflibclc/clc/lib/generic/math clc_remquo_stret.inc

[libclc] Fix remainder calculation in clc_remquo for subnormals (#217925)

The remainder t was previously computed using:

    __CLC_GENTYPE t = __clc_mad(y, -__CLC_CONVERT_GENTYPE(qsgn), x);

Multiplying y by -qsgn (+-1.0) introduces an unnecessary intermediate
multiplication step. On platforms or execution modes where subnormals
are flushed to zero computing `y * -qsgn` can prematurely flush a
subnormal `y` to zero, resulting in `0.0 + x = x` instead of computing
the subtraction `x - y` (or `x + y`).

Replace `__clc_mad` with a direct addition/subtraction based on the sign
of the quotient:

    __CLC_GENTYPE t = qsgn > 0 ? (x - y) : (x + y);

This avoids multiplication by +-1.0, eliminates unwanted subnormal
flushing on intermediate products in FTZ modes, and computes the exact
remainder.
DeltaFile
+1-1libclc/clc/lib/generic/math/clc_remquo_stret.inc
+1-11 files

LLVM/project 7ec504ellvm/lib/Transforms/Vectorize VPlanPredicator.cpp

Implement non-uniform part of partial linearization algorithm
DeltaFile
+57-12llvm/lib/Transforms/Vectorize/VPlanPredicator.cpp
+57-121 files

LLVM/project 6346f20llvm/lib/Transforms/Vectorize VPlanRecipes.cpp

Don't crash dumping malformed phis
DeltaFile
+20-0llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
+20-01 files

LLVM/project cdf854bllvm/lib/Transforms/Vectorize VPlanPredicator.cpp

[AI] Move convertPhisToBlends to post-linearization
DeltaFile
+28-9llvm/lib/Transforms/Vectorize/VPlanPredicator.cpp
+28-91 files

LLVM/project ac5fb45llvm/lib/Target/RISCV RISCVISelLowering.cpp, llvm/test/CodeGen/RISCV/rvv convert-from-arbitrary-fp.ll fixed-vector-convert-from-arbitrary-fp.ll

[RISCV] Lower CONVERT_FROM_ARBITRARY_FP of Float8E5M2 with Zvfofp8min
DeltaFile
+486-0llvm/test/CodeGen/RISCV/rvv/fixed-vector-convert-from-arbitrary-fp.ll
+54-0llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+24-0llvm/test/CodeGen/RISCV/rvv/convert-from-arbitrary-fp.ll
+564-03 files

LLVM/project 1ab6947llvm/lib/CodeGen/SelectionDAG TargetLowering.cpp, llvm/test/CodeGen/RISCV convert-from-arbitrary-fp.ll

[SelectionDAG] Do not use illegal type when expanding `CONVERT_FROM_ARBITRARY_FP` (#219597)

During `CONVERT_FROM_ARBITRARY_FP`'s expansion, it'll try to create
intermediate integer values with the same width as the floating point
result. However, that integer type might not be legal, and would cause
problem when dealing with scalar version of `CONVERT_FROM_ARBITRARY_FP`.

For example, in the attached LIT tests, it'll generate something like
```
t58: f32 = convert_from_arbitrary_fp t57, TargetConstant:i32<7>
```
after type legalization. While f32 is a legal type, its integer
counterpart with the same width, i32, is not a legal type in RV64.

This patch fixes such problem by using the legal type for those
intermediate values, if the legal type is wider.
DeltaFile
+137-0llvm/test/CodeGen/RISCV/rvv/fixed-vector-convert-from-arbitrary-fp.ll
+69-0llvm/test/CodeGen/RISCV/convert-from-arbitrary-fp.ll
+42-9llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
+248-93 files

LLVM/project 022758bllvm/lib/Transforms/Vectorize VPlanUtils.h VPlanUtils.cpp, llvm/unittests/Transforms/Vectorize VPlanTest.cpp

Luke's reconstructSSA (#212209)
DeltaFile
+268-0llvm/unittests/Transforms/Vectorize/VPlanTest.cpp
+37-0llvm/lib/Transforms/Vectorize/VPlanUtils.cpp
+9-0llvm/lib/Transforms/Vectorize/VPlanUtils.h
+314-03 files

LLVM/project 567a9f0llvm/test/Transforms/LoopVectorize/VPlan predicator.ll

Create actual test functions (AI-assisted)
DeltaFile
+478-25llvm/test/Transforms/LoopVectorize/VPlan/predicator.ll
+478-251 files

LLVM/project 41437a5llvm/test/Transforms/LoopVectorize predicator.ll, llvm/test/Transforms/LoopVectorize/VPlan predicator.ll

Copy to LoopVectorize/predicator.ll and generate CHECKs in both
DeltaFile
+1,077-0llvm/test/Transforms/LoopVectorize/predicator.ll
+547-1llvm/test/Transforms/LoopVectorize/VPlan/predicator.ll
+1,624-12 files

LLVM/project 72a805allvm/test/Transforms/LoopVectorize/VPlan predicator.ll

Predicator tests for uniform control flow preservation
DeltaFile
+168-0llvm/test/Transforms/LoopVectorize/VPlan/predicator.ll
+168-01 files

LLVM/project 5594056llvm/lib/Transforms/Vectorize VPlanPredicator.cpp, llvm/test/Transforms/LoopVectorize predicator.ll

[VPlan] Use compact RPOT instead of just RPOT

This is necessary for the future partial linearization change, but I
wanted to commit this bit independently because it changes some tests on
itself and could potentially provide more blend optimization
opportunites (at least I hoped) but that didn't seem to happen.
DeltaFile
+58-15llvm/lib/Transforms/Vectorize/VPlanPredicator.cpp
+32-32llvm/test/Transforms/LoopVectorize/VPlan/predicator.ll
+4-4llvm/test/Transforms/LoopVectorize/predicator.ll
+4-4llvm/test/Transforms/LoopVectorize/X86/predicate-switch.ll
+98-554 files

LLVM/project e947d45llvm/lib/Transforms/Vectorize VPlanTransforms.cpp

Drop `removeCommonBlendMask` from `simplifyBlends` - noop now
DeltaFile
+0-19llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
+0-191 files

LLVM/project 3677f04llvm/lib/Transforms/Vectorize VPlanPredicator.cpp, llvm/test/Transforms/LoopVectorize pr43166-fold-tail-by-masking.ll

[VPlan] Optimize away common blend mask in predicator
DeltaFile
+12-105llvm/test/Transforms/LoopVectorize/AArch64/conditional-branches-cost.ll
+5-16llvm/test/Transforms/LoopVectorize/pr43166-fold-tail-by-masking.ll
+16-2llvm/lib/Transforms/Vectorize/VPlanPredicator.cpp
+8-7llvm/test/Transforms/LoopVectorize/VPlan/conditional-scalar-assignment-vplan.ll
+6-8llvm/test/Transforms/LoopVectorize/AArch64/masked-call-scalarize.ll
+8-6llvm/test/Transforms/LoopVectorize/RISCV/divrem.ll
+55-14422 files not shown
+90-18828 files

LLVM/project 54c6e0ellvm/include/llvm/TargetParser AMDGPUTargetParser.h, llvm/lib/Target/AMDGPU GCNSubtarget.cpp

[AMDGPU] Add getLocalMemorySize to TargetParser

Add getLocalMemorySize and getAddressableLocalMemorySize, both taking a
GPUKind or a Triple::SubArchType, so the LDS a work-group gets can be
queried from a GPU name alone without an MCSubtargetInfo. The first
returns the physical block available in the current mode, the second
caps it at what one work-group can address, mirroring the IsaInfo pair.

The number of SIMDs a work-group runs on is a per-kernel mode rather
than a property of the GPU, so it stays a parameter. GCNSubtarget
initializes its cached sizes from the new entry points. There is no
functional change.

Change-Id: Ib71428b66032a231ed491d6294122b359abfe7e2
Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
DeltaFile
+56-0llvm/unittests/TargetParser/TargetParserTest.cpp
+30-0llvm/lib/TargetParser/AMDGPUTargetParser.cpp
+13-0llvm/include/llvm/TargetParser/AMDGPUTargetParser.h
+4-3llvm/lib/Target/AMDGPU/GCNSubtarget.cpp
+103-34 files

LLVM/project 4fbc7d8mlir/lib/Dialect/OpenACC/IR OpenACC.cpp, mlir/test/Dialect/OpenACC invalid.mlir

[mlir][OpenACC] Reject unsupported routine bind attributes (#220710)

Require each routine bind item to be a symbol reference or string
attribute. Other successfully parsed attributes left the kind
discriminator uninitialized and were silently dropped.

Found by Coverity.

Assisted-by: Codex
DeltaFile
+7-3mlir/lib/Dialect/OpenACC/IR/OpenACC.cpp
+8-0mlir/test/Dialect/OpenACC/invalid.mlir
+15-32 files

LLVM/project 54f8aa7offload/liboffload/include OffloadImpl.hpp, offload/liboffload/src OffloadImpl.cpp

[Offload] Address static analyzer hits (#220725)

Address uninitialized variable and unreachable code.
DeltaFile
+0-10offload/liboffload/src/OffloadImpl.cpp
+1-1offload/liboffload/include/OffloadImpl.hpp
+1-112 files

LLVM/project fc1ab1fclang-tools-extra/include-cleaner/unittests AnalysisTest.cpp, clang/include/clang/Tooling/Inclusions HeaderIncludes.h

[clang][include cleaner] Fix MainHeader insertion issue (#212852)

When doing multiple header insertions (like from include cleaner) there
was an issue with MainHeaders inserted in the wrong location.
HeaderIncludes now supports a bulk insertion where it sorts the
insertions appropriately before inserting them to make sure that the
MainHeader ends up in the correct location.

Format.cpp has been updated to use the bulk insertion method correctly.
Tests added to verify.
DeltaFile
+121-0clang/unittests/Tooling/HeaderIncludesTest.cpp
+80-4clang/lib/Tooling/Inclusions/HeaderIncludes.cpp
+79-2clang-tools-extra/include-cleaner/unittests/AnalysisTest.cpp
+39-4clang/include/clang/Tooling/Inclusions/HeaderIncludes.h
+33-0clang/unittests/Format/CleanupTest.cpp
+12-13clang/lib/Format/Format.cpp
+364-236 files

LLVM/project 96cb9c4mlir/include/mlir/Conversion/ArithCommon AttrToLLVMConverter.h, mlir/lib/Conversion/XeVMToLLVM XeVMToLLVM.cpp

[mlir][LLVM][GPU] Migrate to explicit split inherent/discardable attribute APIs access (#218921)

Use discardable attribute APIs and typed operation accessors throughout
the LLVM and GPU dialect families, their conversions, translations, and
tests.

Assisted-by: Codex
DeltaFile
+55-0mlir/unittests/Dialect/LLVMIR/LLVMAttrsTest.cpp
+17-36mlir/include/mlir/Conversion/ArithCommon/AttrToLLVMConverter.h
+26-18mlir/lib/Dialect/GPU/IR/GPUDialect.cpp
+40-0mlir/test/Conversion/ArithToLLVM/attribute-storage.mlir
+25-9mlir/lib/Conversion/XeVMToLLVM/XeVMToLLVM.cpp
+17-12mlir/lib/Dialect/LLVMIR/IR/LLVMDialect.cpp
+180-7558 files not shown
+488-25264 files