LLVM/project 98ae570llvm/lib/Transforms/Vectorize VPlan.h VPlanRecipes.cpp

[VPlan] Remove the unused CalculateTripCountMinusVF opcode (NFC) (#216675)

bcc272b3220f ("[LV] Remove DataAndControlFlowWithoutRuntimeCheck. NFC",
#183762) removed the only createNaryOp building this opcode, leaving
behind the enum entry and its cases for type inference, operand count,
scalar generation, lowering and printing. Nothing constructs it, so no
plan can contain it and no test prints it.
DeltaFile
+0-17llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
+0-1llvm/lib/Transforms/Vectorize/VPlan.h
+0-182 files

LLVM/project 31a1886llvm/lib/Target/PowerPC PPCInstrInfo.cpp, llvm/test/CodeGen/PowerPC mi-peephole-forwarding-undef.mir

PowerPC: Fix MI peephole crash on an undef forwarding operand

Found by AI while working on something else.

Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
DeltaFile
+25-0llvm/test/CodeGen/PowerPC/mi-peephole-forwarding-undef.mir
+2-0llvm/lib/Target/PowerPC/PPCInstrInfo.cpp
+27-02 files

LLVM/project a14f953clang/lib/Analysis/LifetimeSafety FactsGenerator.cpp Origins.cpp, clang/test/Sema/LifetimeSafety safety.cpp capture-by.cpp

Support [[clang::lifetime_capture_by(X)]] in Plain Containers  (#204361)

This PR implements support for `[[clang::lifetime_capture_by(X)]]` to
enable tracking lifetimes for plain structs and containers like
`std::vector` without requiring manual `[[gsl::Pointer]]` or
`[[gsl::Owner]]` annotations. The implementation extends the
`LifetimeAnnotatedOriginTypeCollector` to register types in capture_by
contracts for origin tracking.

This PR also enables the intra-procedural analysis in existing tests
using -Wlifetime-safety and updated expectations to handle the more
detailed flow-sensitive diagnostics.

```cpp
struct MyContainer {
  const char* stored_ptr;
};

void captureInto(std::string_view v [[clang::lifetime_capture_by(c)]], MyContainer& c);

    [22 lines not shown]
DeltaFile
+295-83clang/test/Sema/LifetimeSafety/capture-by.cpp
+32-1clang/lib/Analysis/LifetimeSafety/Origins.cpp
+15-7clang/lib/Analysis/LifetimeSafety/FactsGenerator.cpp
+3-4clang/test/Sema/LifetimeSafety/safety.cpp
+345-954 files

LLVM/project 35aea1dllvm/lib/CodeGen/GlobalISel GIMatchTableExecutor.cpp

GlobalISel: Use MIPatternMatch in GIMatchTableExecutor (#216601)

Replace the getVRegDef + opcode-check idiom in isBaseWithConstantOffset
with mi_match using m_GPtrAdd and m_GConstant.

Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
DeltaFile
+4-10llvm/lib/CodeGen/GlobalISel/GIMatchTableExecutor.cpp
+4-101 files

LLVM/project 4487de6llvm/include/llvm/IR RuntimeLibcalls.td, llvm/test/CodeGen/X86 fp128-libcalls-longdouble.ll

RuntimeLibcalls: Provide fp128 long double libcalls on X86 (#216622)

16a8d8d038a3  removed the l-suffixed long double math functions
from the default set and re-added them per-target gated on
isLongDoubleF128, but X86 was not given the re-add. On targets
whose long double is fp128 (e.g. x86_64 Android/OHOS) this dropped
the fp128 l-suffixed libcalls.

Fixes the regression reported on #214944.

Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
DeltaFile
+98-0llvm/test/CodeGen/X86/fp128-libcalls-longdouble.ll
+8-0llvm/include/llvm/IR/RuntimeLibcalls.td
+106-02 files

LLVM/project f7c03ecllvm/lib/Target/AMDGPU AMDGPUISelLowering.cpp, llvm/test/CodeGen/AMDGPU fmed3.ll

[AMDGPU] Fix isKnownNeverNaN for FMIN_LEGACY/FMAX_LEGACY (#216338)

These compare-selects return one of the operands bit-for-bit, so a
signaling NaN operand passes through unquieted

Recurse into both operands instead of assuming never-sNaN
DeltaFile
+208-0llvm/test/CodeGen/AMDGPU/fmed3.ll
+3-8llvm/lib/Target/AMDGPU/AMDGPUISelLowering.cpp
+211-82 files

LLVM/project 378f2d8bolt/lib/Core Relocation.cpp, bolt/test/AArch64 tls.c

[AArch64][BOLT] Fold local-exec TLS relocations into loads and stores (#215531)

Emit the low part of the 12- and 24-bit local-exec sequences as an
ADDlow so that it folds into the addressing mode of a following load or
store, saving one instruction.

This re-lands the local-exec part of r327316 (7bc64bd889ad), whose ELF
changes r327503 (bde677289acc) reverted because neither LLD nor GNU
binutils implemented the relevant
R_AARCH64_TLSLE_LDST*_TPREL_LO12 relocations at the time.

LLD and GNU bfd now handle the non-checking 8- to 64-bit variants used
by the default 24-bit sequence. GNU bfd also handles the checked
variants used by the 12-bit sequence; LLD does not support those, but it
did not support the predecessor checked ADD relocation either. GNU bfd
still has no LDST128 support, so 128-bit accesses stay unfolded.

Teach BOLT to classify the folded TLS load/store relocations so it can
process binaries linked with --emit-relocs.
DeltaFile
+47-6llvm/test/CodeGen/AArch64/arm64-tls-local-exec.ll
+18-1llvm/lib/Target/AArch64/AArch64ISelDAGToDAG.cpp
+5-8llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
+12-0bolt/test/AArch64/tls.c
+10-0llvm/test/CodeGen/AArch64/win-tls.ll
+8-0bolt/lib/Core/Relocation.cpp
+100-156 files

LLVM/project b834b9ellvm/lib/Target/ARM ARMTargetMachine.cpp, llvm/test/CodeGen/ARM arm-eabi.ll eabihf-no-fpregs.ll

[ARM] Emit an error when the hard-float ABI is enabled but can't be used (#111334)

Prior to this, compiling for an eabihf target with a CPU lacking
floating-point registers would silently use the soft-float ABI instead,
even though the Arm attributes section would still have
"Tag_ABI_VFP_args: VFP registers", which leads to silent ABI mismatches
at link time.

Update various tests that were using inconsistent ABI/PCS and features.

Change ARMTargetLowering::getEffectiveCallingConv from private to public
and modify it to pass through unrecognized calling conventions. Now that
ARMBaseTargetMachine::createMachineFunctionInfo calls it, leaving a
fatal error would change the behavior of IR passes that do not need to
lower calling conventions. Unrecognized calling conventions are still
treated as errors in lowering.

Fixes #110383.
DeltaFile
+61-57llvm/test/CodeGen/ARM/byval_struct_copy_tailcall.ll
+21-21llvm/test/CodeGen/ARM/constantfp.ll
+38-2llvm/lib/Target/ARM/ARMTargetMachine.cpp
+39-0llvm/test/CodeGen/ARM/eabihf-no-fpregs.ll
+12-12llvm/test/CodeGen/Thumb2/cde-gpr.ll
+12-12llvm/test/CodeGen/ARM/arm-eabi.ll
+183-10432 files not shown
+271-16438 files

LLVM/project 746a343libc/src/time asctime_r.cpp asctime.cpp, libc/test/src/time CMakeLists.txt

[libc] Fix {asc,c,gm,mk}time(_r)? tests and spurios snprintf call (#216423)

time_test_utils was depending on a non-existent library, which caused
these tests to be auto-skipped. Fixing that exposed the fact that some
of the tests don't build (in hermetic mode) due to a snprintf
dependency concealed behind a __builtin_snprintf in asctime.

This patch addresses the existing TODO by moving
asctime to asctime_utils.h (avoiding a dependency loop)
and implementing it via strftime_main.

Assisted by Gemini.
DeltaFile
+59-0libc/src/time/asctime_utils.h
+0-30libc/src/time/time_utils.h
+22-4libc/src/time/CMakeLists.txt
+7-1libc/test/src/time/CMakeLists.txt
+1-1libc/src/time/asctime_r.cpp
+1-1libc/src/time/asctime.cpp
+90-372 files not shown
+92-378 files

LLVM/project f769478clang/lib/CodeGen CodeGenModule.cpp, clang/test/CodeGen amdgpu-builtin-processor-is.c amdgpu-builtin-is-invocable.c

clang/AMDGPU: Don't emit target-features on AMDGCN-flavored SPIR-V

The spirv64-amd-amdhsa target unions every GPU's features in its feature
map so it can report builtins as available. The CodeGen doesn't have
any use of the target-features. Putting it into the IR just results
in an annoying to update test every time a new feature is added. The
ultimate SPIRV codegen doesn't do anything with it, and if it did
survive to AMDGPU codegen, it would be actively harmful.

This isn't an ideal solution. The target-features spam is also
noisy and useless in the AMDGPU case, but solving that is more
intricate because we do currently rely on this for some features,
most notably the wavesize.

Co-authored-by: Claude (Claude-Opus-4.8)
DeltaFile
+5-0clang/lib/CodeGen/CodeGenModule.cpp
+2-2clang/test/CodeGenCXX/dynamic-cast-address-space.cpp
+1-1clang/test/CodeGen/amdgpu-builtin-processor-is.c
+1-1clang/test/CodeGen/amdgpu-builtin-is-invocable.c
+9-44 files

LLVM/project 9d81ccamlir/include/mlir/Dialect/X86 X86.td, mlir/test/Dialect/X86/AMX legalize-for-llvm.mlir

[mlir][x86] Fix - Instrincs selection for hf8 and bf8. (#216647)

This patch fixes the error instrincs selection for `bf8` and `hf8`
types.

- Intel AMX naming: `bf8 == E5M2`, `hf8 == E4M3FN`.
DeltaFile
+4-4mlir/test/Dialect/X86/AMX/legalize-for-llvm.mlir
+3-3mlir/test/Target/LLVMIR/amx.mlir
+2-2mlir/include/mlir/Dialect/X86/X86.td
+9-93 files

LLVM/project 8861fc6clang/lib/Basic/Targets AMDGPU.h AMDGPU.cpp, clang/lib/Driver/ToolChains CommonArgs.cpp AMDGPU.cpp

clang/AMDGPU: Use feature bitset instead of ArchAttr

Convert from the legacy getArchAttrAMDGCN manual bitmask checks to using
the new generated bitset. These are the easy cases. sramecc and xnack
require more supporting work so will be done later.

Co-authored-by: Claude (Claude-Opus-4.8)
DeltaFile
+11-9clang/lib/Driver/ToolChains/AMDGPU.cpp
+6-2clang/lib/Basic/Targets/AMDGPU.cpp
+4-2clang/lib/Basic/Targets/AMDGPU.h
+2-2clang/lib/Driver/ToolChains/CommonArgs.cpp
+2-1llvm/lib/Target/AMDGPU/AMDGPU.td
+25-165 files

LLVM/project cfd277eclang/lib/CIR/CodeGen CIRGenBuiltinAArch64.cpp, clang/test/CodeGen/AArch64 neon-fcvt-intrinsics.c

[clang][CIR][AArch64] Add lowering for conversion intrinsics (#211609)

This PR adds lowering for intrinsic from the following groups:
* https://arm-software.github.io/acle/neon_intrinsics/advsimd.html#conversions

It continues the work started in #190961, #193273, #199990 and #209252.
This PR implements the remaining conversions truncating to zero:
  * vcvts_s32_f32
  * vcvts_s64_f32
  * vcvts_u32_f32
  * vcvts_u64_f32

The corresponding tests are moved from:
  * clang/test/CodeGen/AArch64/

to:
  * clang/test/CodeGen/AArch64/neon/

The lowering follows the existing implementation in
CodeGen/TargetBuiltins/ARM.cpp
DeltaFile
+44-0clang/test/CodeGen/AArch64/neon/intrinsics.c
+0-41clang/test/CodeGen/AArch64/neon-fcvt-intrinsics.c
+4-0clang/lib/CIR/CodeGen/CIRGenBuiltinAArch64.cpp
+48-413 files

LLVM/project 2cee0cbllvm/include/llvm/CodeGen/GlobalISel MIPatternMatch.h, llvm/lib/CodeGen/GlobalISel CombinerHelper.cpp

GlobalISel: Match loads by pointer operand in CombinerHelper

Add a load matcher that binds the pointer operand (like IR's m_Load), with
optional outputs for the load instruction and its MachineMemOperand via m_MMO.
Use it to replace the getVRegDef + dyn_cast idiom in the load combines.

Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
DeltaFile
+59-0llvm/include/llvm/CodeGen/GlobalISel/MIPatternMatch.h
+18-20llvm/lib/CodeGen/GlobalISel/CombinerHelper.cpp
+77-202 files

LLVM/project 5696ae1llvm/include/llvm/CodeGen/GlobalISel MIPatternMatch.h, llvm/lib/CodeGen/GlobalISel CombinerHelperCasts.cpp CombinerHelperVectorOps.cpp

GlobalISel: Migrate misc. CombinerHelper def checks to MIPatternMatch

Replace getVRegDef + cast/opcode-check idioms across CombinerHelper
with mi_match, adding named instruction binders and operand-form matchers
as needed.

Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
DeltaFile
+208-157llvm/lib/CodeGen/GlobalISel/CombinerHelper.cpp
+47-7llvm/include/llvm/CodeGen/GlobalISel/MIPatternMatch.h
+25-19llvm/lib/CodeGen/GlobalISel/CombinerHelperVectorOps.cpp
+6-4llvm/lib/CodeGen/GlobalISel/CombinerHelperCasts.cpp
+286-1874 files

LLVM/project 184f413mlir/lib/Dialect/Affine/Transforms SuperVectorize.cpp, mlir/test/Dialect/Affine/SuperVectorize invalid_missing_vector_size.mlir invalid-zero-size.mlir

[mlir][affine] Require a vector size in affine-super-vectorize (#171110)

Diagnose invocations that omit the `virtual-vector-size` option. A
vector rank is required to construct a vectorization pattern, so
accepting an empty size list would silently leave the input unchanged.

Check that vector sizes are present and positive before applying
rank-dependent constraints.

Fixes #114528
DeltaFile
+11-5mlir/lib/Dialect/Affine/Transforms/SuperVectorize.cpp
+10-0mlir/test/Dialect/Affine/SuperVectorize/invalid_zero_size.mlir
+0-9mlir/test/Dialect/Affine/SuperVectorize/invalid-zero-size.mlir
+6-0mlir/test/Dialect/Affine/SuperVectorize/invalid_missing_vector_size.mlir
+27-144 files

LLVM/project 2fc8f5ellvm/lib/CodeGen/GlobalISel GIMatchTableExecutor.cpp

GlobalISel: Use MIPatternMatch in GIMatchTableExecutor

Replace the getVRegDef + opcode-check idiom in isBaseWithConstantOffset with
mi_match using m_GPtrAdd and m_GConstant.

Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
DeltaFile
+4-10llvm/lib/CodeGen/GlobalISel/GIMatchTableExecutor.cpp
+4-101 files

LLVM/project 24045belldb/include/lldb/ValueObject ValueObjectRegister.h, lldb/source/ValueObject ValueObjectRegister.cpp

[lldb] Fix GetIndexOfChildWithName and GetChildMemberWithName on register sets (#212727)

And GetChildMemberWithName which had the same issue.

Fixes #211787.

Both of these methods were doing a lookup on the register info array as
a whole, rather than the subset of indexes into that array. That subset
of indexes is the "register set".

This lead to problems like this where index and name getters disagreed:
```
>>> lldb.frame.GetRegisters()[1].GetChildAtIndex(0)
(unsigned char __attribute__((ext_vector_type(16)))) v0 = (0x2f, 0x2f, 0x2f, 0x2f, 0x2f, 0x2f,
 0x2f, 0x2f, 0x2f, 0x2f, 0x2f, 0x2f, 0x2f, 0x2f, 0x2f, 0x2f)
>>> lldb.frame.GetRegisters()[1].GetIndexOfChildWithName("v0")
63
```
GetChildAtIndex told us that v0 was at index 0, but looking up v0 by

    [18 lines not shown]
DeltaFile
+85-0lldb/test/API/python_api/value/TestValueAPI.py
+29-15lldb/source/ValueObject/ValueObjectRegister.cpp
+3-16lldb/test/API/linux/aarch64/aarch32_compat/TestAArch64LinuxAArch32Compat.py
+11-0llvm/docs/ReleaseNotes.md
+4-4lldb/test/API/functionalities/postmortem/minidump-new/TestMiniDumpNew.py
+3-0lldb/include/lldb/ValueObject/ValueObjectRegister.h
+135-356 files

LLVM/project 58a2155llvm/include/llvm/CodeGen TargetLowering.h, llvm/lib/CodeGen/SelectionDAG SelectionDAGBuilder.cpp

Update for comments
DeltaFile
+5-5llvm/include/llvm/CodeGen/TargetLowering.h
+2-2llvm/lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp
+7-72 files

LLVM/project 07e3a4ellvm/include/llvm/CodeGen/GlobalISel MIPatternMatch.h, llvm/lib/CodeGen/GlobalISel CombinerHelper.cpp

GlobalISel: Introduce m_GPtrAdd flags matcher in CombinerHelper (#216600)

Add an optional MIFlags output operand to the binary-op matcher and a
m_GPtrAdd(L, R, m_MIFlags(F)) overload, and use it to replace getVRegDef
+ opcode checks.

Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
DeltaFile
+24-1llvm/include/llvm/CodeGen/GlobalISel/MIPatternMatch.h
+4-5llvm/lib/CodeGen/GlobalISel/CombinerHelper.cpp
+28-62 files

LLVM/project 4593ff7lldb/test/API/tools/lldb-dap/databreakpoint TestDAP_setDataBreakpoints.py, lldb/tools/lldb-dap DAP.h Watchpoint.h

[lldb-dap] Preserve watchpoints from console (#215228)

We should not delete watchpoints created via LLDB console when
processing DAP `setDataBreakpoints` request.
DeltaFile
+158-2lldb/test/API/tools/lldb-dap/databreakpoint/TestDAP_setDataBreakpoints.py
+48-11lldb/tools/lldb-dap/Handler/SetDataBreakpointsRequestHandler.cpp
+11-0lldb/tools/lldb-dap/Watchpoint.cpp
+4-0lldb/tools/lldb-dap/Watchpoint.h
+3-0lldb/tools/lldb-dap/DAP.h
+224-135 files

LLVM/project 0ee6dedflang/lib/Lower/OpenMP OpenMP.cpp, flang/test/Lower/OpenMP omp-declarative-allocate-module.f90

[Flang][OpenMP] PoC module support for allocate directives

This patch implements partial support for `allocate` on Fortran
module variables, based on adding global constructor functions for each
impacted variable.

Shared as a proof of concept, because I have a few concerns about it:
  1. It appears that Clang ignores `allocate` directives on global
     variables instead. Is that the expected behavior?
  2. The existing implementation for `allocate` in Flang doesn't
     actually impact where the memory used for a variable resides. It
     allocates/deallocates extra memory for it using OpenMP internal
     compiler calls but then that storage is never used. The original
     alloca is still used. This addition suffers from the same issue:
     global constructors allocate extra memory that is never used to
     update in any way the associated global variable or its users.
  3. No `omp.allocate_free` (should be `omp.allocate.free`) can be added
     by this approach.
  4. The representation of `omp.allocate_dir` (should be `omp.allocate`)

    [10 lines not shown]
DeltaFile
+110-30flang/lib/Lower/OpenMP/OpenMP.cpp
+42-0flang/test/Lower/OpenMP/omp-declarative-allocate-module.f90
+152-302 files

LLVM/project 952515aflang/lib/Lower/OpenMP OpenMP.cpp, flang/test/Lower/OpenMP/Todo allocate-module.f90

[Flang][OpenMP] Prevent allocate directive ICE on module variables (#216021)

The current lowering implementation for `allocate` directives assumes
the MLIR function in which it is creating operations will still be there
by finalization time, so that it can add a deallocation call.

When lowering Fortran modules, this is not the case (lowering happens in
a temporary dummy function) and it results in a compiler crash while
running cleanup callbacks. This patch adds a TODO for this case.
DeltaFile
+9-0flang/lib/Lower/OpenMP/OpenMP.cpp
+9-0flang/test/Lower/OpenMP/Todo/allocate-module.f90
+18-02 files

LLVM/project 6628701llvm/test/CodeGen/X86 fp128-libcalls-longdouble.ll fp128-libcalls-gnu.ll

Move the test to a separate file, this isn't gnu
DeltaFile
+4-264llvm/test/CodeGen/X86/fp128-libcalls-gnu.ll
+98-0llvm/test/CodeGen/X86/fp128-libcalls-longdouble.ll
+102-2642 files

LLVM/project c8a0460mlir/test lit.cfg.py, mlir/test/Conversion/SCFToAffine scf-to-affine.mlir

[mlir] Filter out failing tests when expensive checks are ON (#216323)

Adds logic to conditionally disable tests that fail when expensive
checks are enabled,
*  -DMLIR_ENABLE_EXPENSIVE_PATTERN_API_CHECKS=ON. 

When the expensive API checks are disabled, the newly marked tests
are run as usual.

This is a temporary measure to enable the introduction of a buildbot
that will run with expensive API checks enabled. No new tests disabled with
expensive checks should be added, i.e. tests with 
 * `XFAIL: mlir-expensive-checks`.
 
The existing failures marked in this PR should be fixed.

GitHub issue that reported these failures prior to this PR:
  * https://github.com/llvm/llvm-project/issues/163599
DeltaFile
+3-1mlir/test/Conversion/SCFToOpenMP/vector-reduction.mlir
+3-1mlir/test/Conversion/SCFToAffine/scf-to-affine.mlir
+2-1mlir/test/Integration/Dialect/Linalg/CPU/ArmSVE/pack-scalable-inner-tile.mlir
+3-0mlir/test/lit.cfg.py
+2-0mlir/test/Conversion/SCFToOpenMP/scf-to-openmp.mlir
+2-0mlir/test/Conversion/SCFToOpenMP/reductions.mlir
+15-341 files not shown
+95-347 files

LLVM/project 988e509libc/src/__support/FPUtil FPBits.h, libc/src/__support/macros/properties types.h

[libc] Fix FreeBSD build for 53-bit-rounded fp80s (#216332)

Fix build issues on FreeBSD, which reports `LDBL_MANT_DIG == 53` for
`long double` on some targets despite using fp80 as the underlying type.
This is because it stores the value as an fp80, but configures the FPU
to round the mantissa to 53-bits. However, this causes some parts of
libc to misidentify the fp80 as an fp64 since both use a 53-bit
mantissa. This led to build issues on FreeBSD when trying to bitcast the
12-byte `FPBits<long double>` to an 8-byte fp64 value.

This is fixed by checking both the mantissa size and the exponent range
when determining the correct format for `long double`. Also, move the
check for this to a single place in `types.h`, rather than re-checking
the `LDBL_MANT_DIG` and `LDBL_MAX_EXP` values in `FPBits.h`.
DeltaFile
+8-7libc/src/__support/FPUtil/FPBits.h
+10-3libc/src/__support/macros/properties/types.h
+18-102 files

LLVM/project 898b018llvm/lib/Support APFloat.cpp, llvm/unittests/ADT APFloatTest.cpp

[APFloat] Report the sign and the zero a conversion cannot represent (#216056)

`APFloat::convert` reports through `losesInfo` what rounding lost, but
not what
the target format has no encoding for at all. Two properties of a format
are not
rounding:

| property | formats today | what happens |
|---|---|---|
| `hasSignedRepr == false` | `f8E8M0FNU`, `f8E5M3FNU` | the sign bit is
carried into a format with no room for it |
| `hasZero == false` | `f8E8M0FNU` | zero is replaced by the smallest
normalized value, 2^-127 |

Both were reported as `opOK` with `losesInfo == false`. Callers gate on
`losesInfo` -- that is how `arith.truncf`'s folder decides whether a
constant
fold is legal -- so they kept a value the format cannot hold.

    [84 lines not shown]
DeltaFile
+71-2llvm/unittests/ADT/APFloatTest.cpp
+40-28llvm/lib/Support/APFloat.cpp
+24-0mlir/test/Dialect/Arith/canonicalize.mlir
+18-0mlir/test/IR/invalid-builtin-attributes.mlir
+9-0mlir/lib/AsmParser/AttributeParser.cpp
+8-0mlir/lib/AsmParser/Parser.cpp
+170-306 files

LLVM/project 32efbc2mlir/include/mlir/Dialect/LLVMIR NVVMOps.td, mlir/lib/Dialect/LLVMIR/IR NVVMDialect.cpp

[MLIR][NVVM] Spell strict assembly properties directly

Bind every NVVM inherent property in its operation assembly format and
re-enable strict property parsing for the dialect. Use direct named clauses
for declarative formats and custom MMA parsers while retaining dictionaries
for discardable attributes.

Assisted-by: Codex
DeltaFile
+619-5mlir/lib/Dialect/LLVMIR/IR/NVVMDialect.cpp
+350-157mlir/include/mlir/Dialect/LLVMIR/NVVMOps.td
+156-156mlir/test/Target/LLVMIR/nvvm/tcgen05-mma-tensor.mlir
+156-156mlir/test/Target/LLVMIR/nvvm/tcgen05-mma-sp-tensor.mlir
+128-128mlir/test/Target/LLVMIR/nvvm/tma_store_reduce.mlir
+29-203mlir/test/Dialect/LLVMIR/nvvm-mma-sparse-blockscale.mlir
+1,438-805102 files not shown
+3,725-4,037108 files

LLVM/project 79679bemlir/include/mlir/Dialect/LLVMIR NVVMOps.td, mlir/lib/Dialect/LLVMIR/IR NVVMDialect.cpp

[MLIR][NVVM] Add asynchronous store Ops (#210931)

This change adds the `store.async.global` and `store.async.shared`
ops to the NVVM dialect to perform asynchronous stores to global
or shared-cluster address spaces.

PTX Spec References:
1.
[`st.async`](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#data-movement-and-conversion-instructions-st-async)
2.
[`multimem.st.async`](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#data-movement-and-conversion-instructions-multimem-st-async)
DeltaFile
+51-0mlir/include/mlir/Dialect/LLVMIR/NVVMOps.td
+42-0mlir/test/Target/LLVMIR/nvvm/store_async_global.mlir
+40-0mlir/lib/Dialect/LLVMIR/IR/NVVMDialect.cpp
+23-0mlir/test/Target/LLVMIR/nvvm/store_async_global_invalid.mlir
+17-0mlir/test/Target/LLVMIR/nvvm/store_async_shared.mlir
+173-05 files

LLVM/project 3944056mlir/lib/ExecutionEngine LevelZeroRuntimeWrappers.cpp

[mlir][gpu] Fix L0_SAFE_CALL in LevelZero runtime (#215308)

Using NULL as `RTContext` can lead to crashes when calling
`zeDriverGetLastErrorDescription`.
Now using `getRtContext()` instead.
DeltaFile
+17-2mlir/lib/ExecutionEngine/LevelZeroRuntimeWrappers.cpp
+17-21 files