LLVM/project de2e0aa — clang/docs ReleaseNotes.md, clang/lib/Sema SemaDeclAttr.cpp

[Clang][Smea] Accept weak reference after declaration

GCC accepts a weak reference even after there is a declaration. It is
fine for us to just append a weak attribute in the previous
declaration directly.
DeltaFile
+37-0clang/test/CodeGen/pragma-weak.c
+25-0clang/lib/Sema/SemaDeclAttr.cpp
+17-0clang/test/Sema/pragma-weak.c
+3-0clang/docs/ReleaseNotes.md
+82-04 files

LLVM/project 47d237b — mlir/include/mlir/Dialect/XeGPU/Utils XeGPUUtils.h, mlir/lib/Dialect/XeGPU/Transforms XeGPUPropagateLayout.cpp XeGPUSgToLaneDistribute.cpp

[mlir][XeGPU] Add an SLM round-trip fallback for sg-to-lane convert_layout (#227884)

This PR adds a general fallback lowering for xegpu.convert_layout in the
sg-to-lane distribution pass, which round-trips the value through shared
local memory.

The existing lane-level lowerings are all special cases: the layouts
fold into each other, or they differ in a way a xegpu.lane_shuffle or a
gpu.shuffle can express. Anything else failed to legalize and stopped
the pipeline — for instance a conversion that only moves the lane_data
of a dimension distributed over part of the subgroup:

```mlir
xegpu.convert_layout %src
  <{input_layout  = #xegpu.layout<lane_layout = [1, 2, 8], lane_data = [1, 1, 4]>,
    target_layout = #xegpu.layout<lane_layout = [1, 2, 8], lane_data = [1, 1, 1]>}>
  : vector<1x2x32xbf16>
```


    [17 lines not shown]
DeltaFile
+202-0mlir/lib/Dialect/XeGPU/Transforms/XeGPUSgToLaneDistribute.cpp
+51-0mlir/test/Dialect/XeGPU/sg-to-lane-distribute-unit.mlir
+18-0mlir/lib/Dialect/XeGPU/Utils/XeGPUUtils.cpp
+4-12mlir/lib/Dialect/XeGPU/Transforms/XeGPUPropagateLayout.cpp
+7-0mlir/include/mlir/Dialect/XeGPU/Utils/XeGPUUtils.h
+282-125 files

LLVM/project b5ab963 — mlir/include/mlir/Dialect/XeGPU/Transforms XeGPULayoutImpl.h, mlir/lib/Dialect/XeGPU/Transforms XeGPUPropagateLayout.cpp XeGPULayoutImpl.cpp

[mlir][xegpu] Fill split-group inner dims on shape_cast result layout (#227560)

This PR set the result layout for shape_cast when it can't infer the
layout from incoming consumer layout. For example,
```mlir
vector.shape_cast %0 : vector<16x1024xbf16> to vector<16x32x32xbf16>
// consumer layout: inst_data = [1, 2, 8], lane_layout = [1, 2, 8], lane_data = [1, 1, 1]
```
Taking the consumer layout as is doesn't work for layout propagation.
The lanes cover 2 along dim 1 and only 8 of dim 2's 32 — a 2x8 box.
Collapsing dims 1 and 2 gives inst_data = [1, 16] on the 16x1024 source,
which is 16 contiguous elements. Those are not the same elements: the
2x8 box is two runs of 8, 32 apart.

This PR fixes the issue by stretching the lane_data to cover the
innermost dims of source shape, to set up result layout so that each
lane gets the same data before and after the shapecast, which is the
condition that shapecast op can be successfully distributed.
```

    [10 lines not shown]
DeltaFile
+95-0mlir/lib/Dialect/XeGPU/Transforms/XeGPULayoutImpl.cpp
+64-0mlir/test/Dialect/XeGPU/propagate-layout-inst-data-invalid.mlir
+40-0mlir/include/mlir/Dialect/XeGPU/Transforms/XeGPULayoutImpl.h
+36-0mlir/test/Dialect/XeGPU/propagate-layout-inst-data.mlir
+17-2mlir/lib/Dialect/XeGPU/Transforms/XeGPUPropagateLayout.cpp
+252-25 files

LLVM/project 1372974 — compiler-rt/lib/tsan/rtl tsan_rtl_mutex.cpp

[NFC][TSan] Move RestoreStack before ScopedReport in ReportDestroyLocked (#228643)

Run RestoreStack in its own lock scope before constructing ScopedReport
in ReportDestroyLocked (matching ReportRace). This avoids acquiring
ScopedErrorReportLock or symbolizing the current stack if RestoreStack
fails, and avoids holding slot_lock and slot_mtx while populating the
ScopedReport.

Assisted-by: Gemini
DeltaFile
+17-14compiler-rt/lib/tsan/rtl/tsan_rtl_mutex.cpp
+17-141 files

LLVM/project e835954 — llvm/lib/CodeGen CodeGenOptions.cpp CodeGenOptions.h, utils/bazel/llvm-project-overlay/llvm BUILD.bazel

[CodeGen] Declare command line options in TableGen

Move the cl::opts of TargetPassConfig.cpp and CodeGenPrepare.cpp into
CodeGenOptions.td, private to lib/CodeGen; other files will follow.
-enable-machine-outliner (cl::ValueOptional) and -regalloc
(RegisterPassParser) stay cl::opt.

CodeGenPrepare and its addressing-mode helpers hold
`const CodeGenOptions &Opts`; TargetPassConfig functions read
CodeGenOptions::Global. getCGPassBuilderOption() converts
std::optional<bool> members to the cl::boolOrDefault fields of the
public CGPassBuilderOption. -basic-block-section-match-infer, which was
not cl::Hidden, is now listed by -help-hidden only.

Aided by Opus 5.5
DeltaFile
+142-318llvm/lib/CodeGen/TargetPassConfig.cpp
+96-221llvm/lib/CodeGen/CodeGenPrepare.cpp
+200-0llvm/lib/CodeGen/CodeGenOptions.td
+19-0llvm/lib/CodeGen/CodeGenOptions.h
+15-0llvm/lib/CodeGen/CodeGenOptions.cpp
+11-0utils/bazel/llvm-project-overlay/llvm/BUILD.bazel
+483-5393 files not shown
+503-5449 files

LLVM/project 922e22c — llvm/lib/CodeGen RegAllocFast.cpp

[CodeGen] Remove unused -rafast-ignore-missing-defs (#228808)

Unused since the 2020 RegAllocFast rewrite (c8757ff3aa7d) introduced it.
DeltaFile
+1-5llvm/lib/CodeGen/RegAllocFast.cpp
+1-51 files

LLVM/project 552d320 — llvm/lib/Target/X86 X86InstrAVX512.td X86ISelLowering.cpp

[X86] Correct the SDTypeProfile for VNNI instructions. (#228745)

Operand 1 and 2 have i8 or i16 elements while the result and operand 0
have i32 elements. The type profile previously said all operands were
the same type.

I've split the type profile to capture the i8 and i16 element size
accurately.

Assisted-by: Claude
DeltaFile
+53-53llvm/lib/Target/X86/X86InstrSSE.td
+27-21llvm/lib/Target/X86/X86InstrFragmentsSIMD.td
+9-8llvm/lib/Target/X86/X86ISelLowering.cpp
+4-3llvm/lib/Target/X86/X86InstrAVX512.td
+93-854 files

LLVM/project ed17053 — mlir/lib/Conversion/NVGPUToNVVM NVGPUToNVVM.cpp, mlir/lib/Dialect/NVGPU/IR NVGPUDialect.cpp

[mlir][nvgpu] Align integer WGMMA type checks with i8 operands (#212215)

Integer WGMMA uses i8 operands with an i32 accumulator, but the NVGPU
verifier and K-shape selection still checked for i16.

Changes both checks to i8 and adds tests for rejected i16 inputs and the
existing i8 "not supported yet" limitation.
DeltaFile
+22-0mlir/test/Dialect/NVGPU/invalid.mlir
+1-1mlir/lib/Dialect/NVGPU/IR/NVGPUDialect.cpp
+1-1mlir/lib/Conversion/NVGPUToNVVM/NVGPUToNVVM.cpp
+24-23 files

LLVM/project 5729303 — compiler-rt/lib/tsan/rtl tsan_interface_ann.cpp tsan_rtl_mutex.cpp

[NFC][TSan] Move ObtainCurrentStack out of ThreadRegistryLock scope

ObtainCurrentStack only reads the current thread's shadow stack and
allocates a VarSizeStackTrace buffer, which does not require
ThreadRegistryLock (and already runs outside ThreadRegistryLock in
ReportRace, SignalUnsafeCall, and ReportErrnoSpoiling).

Move ObtainCurrentStack (and dummy_pc in ReportDeadlock) before
ScopedReport in ReportMutexHeldWrongContext, ReportMutexMisuse,
ReportDeadlock, and ReportDestroyLocked so the stack trace buffers also
outlive ScopedReport and OutputReport.

Assisted-by: Gemini

Pull Request: https://github.com/llvm/llvm-project/pull/228794
DeltaFile
+6-5compiler-rt/lib/tsan/rtl/tsan_rtl_mutex.cpp
+2-2compiler-rt/lib/tsan/rtl/tsan_interface_ann.cpp
+8-72 files

LLVM/project b4525ca — compiler-rt/lib/tsan/rtl tsan_rtl_mutex.cpp

[NFC][TSan] Move RestoreStack before ScopedReport in ReportDestroyLocked

Run RestoreStack in its own lock scope before constructing ScopedReport
in ReportDestroyLocked (matching ReportRace). This avoids acquiring
ScopedErrorReportLock or symbolizing the current stack if RestoreStack
fails, and avoids holding slot_lock and slot_mtx while populating the
ScopedReport.

Assisted-by: Gemini

Pull Request: https://github.com/llvm/llvm-project/pull/228643
DeltaFile
+17-14compiler-rt/lib/tsan/rtl/tsan_rtl_mutex.cpp
+17-141 files

LLVM/project d8f3268 — compiler-rt/lib/tsan/rtl tsan_report.h tsan_rtl_report.cpp

[TSan] Defer symbolization in AddStack and AddSleep to SymbolizeStackElems

PR #151495 delayed symbolization for memory accesses, locations,
threads, and mutexes until OutputReport (after ThreadRegistryLock is
released), but missed ScopedReport::AddStack and ScopedReport::AddSleep.
As a result, ReportRace (via AddSleep) and ReportMutexMisuse,
ReportDeadlock, and ReportDestroyLocked (via AddStack) still invoked the
symbolizer while holding ThreadRegistryLock.

Store the unsymbolized stack traces and sleep stack ID in ReportDesc and
symbolize them in ScopedReport::SymbolizeStackElems().

Assisted-by: Gemini

Pull Request: https://github.com/llvm/llvm-project/pull/228795
DeltaFile
+15-4compiler-rt/lib/tsan/rtl/tsan_rtl_report.cpp
+7-0compiler-rt/lib/tsan/rtl/tsan_report.h
+22-42 files

LLVM/project cc6d1dc — compiler-rt/lib/tsan/rtl tsan_mman.cpp tsan_rtl_thread.cpp

[NFC][TSan] Allocate ScopedReport as a stack variable

Now that ScopedReport is constructed before acquiring ThreadRegistryLock
or slot locks across all reporting functions, it no longer needs to be
constructed inside the lock scope via placement new on __builtin_alloca
storage.

Declare ScopedReport as a normal stack variable before the lock scope
and remove the manual destructor calls.

Assisted-by: Gemini

Pull Request: https://github.com/llvm/llvm-project/pull/228637
DeltaFile
+17-33compiler-rt/lib/tsan/rtl/tsan_rtl_mutex.cpp
+8-13compiler-rt/lib/tsan/rtl/tsan_rtl_report.cpp
+4-9compiler-rt/lib/tsan/rtl/tsan_rtl_thread.cpp
+4-9compiler-rt/lib/tsan/rtl/tsan_interface_ann.cpp
+4-9compiler-rt/lib/tsan/rtl/tsan_interceptors_posix.cpp
+3-8compiler-rt/lib/tsan/rtl/tsan_mman.cpp
+40-816 files

LLVM/project 7ccb220 — compiler-rt/lib/tsan/rtl tsan_interface_ann.cpp tsan_rtl_mutex.cpp

[NFC][TSan] Move ObtainCurrentStack out of ThreadRegistryLock scope

ObtainCurrentStack only reads the current thread's shadow stack and
allocates a VarSizeStackTrace buffer, which does not require
ThreadRegistryLock (and already runs outside ThreadRegistryLock in
ReportRace, SignalUnsafeCall, and ReportErrnoSpoiling).

Move ObtainCurrentStack (and dummy_pc in ReportDeadlock) before
ScopedReport in ReportMutexHeldWrongContext, ReportMutexMisuse,
ReportDeadlock, and ReportDestroyLocked so the stack trace buffers also
outlive ScopedReport and OutputReport.

Assisted-by: Gemini

Pull Request: https://github.com/llvm/llvm-project/pull/228794
DeltaFile
+6-5compiler-rt/lib/tsan/rtl/tsan_rtl_mutex.cpp
+2-2compiler-rt/lib/tsan/rtl/tsan_interface_ann.cpp
+8-72 files

LLVM/project 120e630 — compiler-rt/lib/fuzzer FuzzerTracePC.cpp

[libFuzzer] Remove unused helper `GetPreviousInstructionPc` (#228453)

This function hasn't been used since 2891b257c24c ("[libFuzzer] remove
stale code"), and it contains a typo in the RISCV #ifdef arm. Removing
this altogether.

This was found when surveying RISCV-specific code in [a downstream
project](https://searchfox.org/firefox-main/rev/084057e952e7dbf376f6c3765ad242aec3785dc6/tools/fuzzing/libfuzzer/FuzzerTracePC.cpp#141)
that vendored libFuzzer.
DeltaFile
+0-19compiler-rt/lib/fuzzer/FuzzerTracePC.cpp
+0-191 files

LLVM/project 6fae12c — compiler-rt/lib/tsan/rtl tsan_rtl_thread.cpp tsan_mman.cpp

[TSan] Lock ScopedErrorReportLock before slot and thread_registry locks

OutputReport runs while ScopedErrorReportLock is held after slot_mtx and
thread_registry have been unlocked. Because code executed during
OutputReport (symbolizer, callbacks, or signal handlers) can acquire
slot_mtx or thread_registry, ScopedErrorReportLock must precede slot and
thread_registry locks in the lock hierarchy to avoid AB-BA deadlocks
between concurrent reports or fork().

- Move ScopedErrorReportLock::Lock() before slot.mtx, thread_registry,
  and slot_mtx in ForkBefore (and unlock in reverse order in ForkAfter).
- Replace ctx->thread_registry.CheckLocked() in ScopedReportBase's
  constructor with CheckedMutex::CheckNoLocks(), and add CheckLocked() to
  AddThread(const ThreadContext *) and CheckNoLocks() to OutputReport.
- Construct ScopedReport before acquiring ThreadRegistryLock across all
  reporting functions, and close the RestoreStack lock scope before
  constructing ScopedReport in ReportRace.

Assisted-by: Gemini

    [2 lines not shown]
DeltaFile
+22-17compiler-rt/lib/tsan/rtl/tsan_rtl_report.cpp
+3-3compiler-rt/lib/tsan/rtl/tsan_rtl_mutex.cpp
+2-2compiler-rt/lib/tsan/rtl/tsan_rtl.cpp
+1-1compiler-rt/lib/tsan/rtl/tsan_rtl_thread.cpp
+1-1compiler-rt/lib/tsan/rtl/tsan_mman.cpp
+1-1compiler-rt/lib/tsan/rtl/tsan_interface_ann.cpp
+30-251 files not shown
+31-267 files

LLVM/project f19701f — compiler-rt/lib/tsan/rtl tsan_rtl_mutex.cpp

[NFC][TSan] Move RestoreStack before ScopedReport in ReportDestroyLocked

Run RestoreStack in its own lock scope before constructing ScopedReport
in ReportDestroyLocked (matching ReportRace). This avoids acquiring
ScopedErrorReportLock or symbolizing the current stack if RestoreStack
fails, and avoids holding slot_lock and slot_mtx while populating the
ScopedReport.

Assisted-by: Gemini

Pull Request: https://github.com/llvm/llvm-project/pull/228643
DeltaFile
+17-14compiler-rt/lib/tsan/rtl/tsan_rtl_mutex.cpp
+17-141 files

LLVM/project 9cd841a — compiler-rt/lib/tsan/rtl tsan_report.cpp tsan_report.h

[NFC][TSan] Use in-class member initializers in tsan_report.h

Use in-class member initializers for all structs and classes in
tsan_report.h and default constructors and destructor in
tsan_report.cpp.

Assisted-by: Gemini

Pull Request: https://github.com/llvm/llvm-project/pull/228780
DeltaFile
+29-28compiler-rt/lib/tsan/rtl/tsan_report.h
+5-17compiler-rt/lib/tsan/rtl/tsan_report.cpp
+34-452 files

LLVM/project 4bbc53c — orc-rt/test/unit CommonTestUtils.h

[orc-rt] Make noErrors report errors as test failures (#228807)

noErrors used cantFail, so an unexpected error aborted the whole test
binary without naming the test that hit it, and with assertions disabled
was dropped silently. Use EXPECT_THAT_ERROR instead, so the error is
reported as a failure of the current test and the remaining tests still
run.
DeltaFile
+8-4orc-rt/test/unit/CommonTestUtils.h
+8-41 files

LLVM/project e94b748 — compiler-rt/lib/tsan/rtl tsan_rtl.h tsan_rtl_report.cpp

[NFC][TSan] Merge ScopedReportBase into ScopedReport (#228640)

ScopedReport is the only subclass of ScopedReportBase and nothing uses
ScopedReportBase directly. Merge ScopedReportBase into ScopedReport.

Assisted-by: Gemini
DeltaFile
+16-21compiler-rt/lib/tsan/rtl/tsan_rtl_report.cpp
+7-16compiler-rt/lib/tsan/rtl/tsan_rtl.h
+23-372 files

LLVM/project f3a1a8c — llvm/lib/Transforms/Vectorize SLPVectorizer.cpp, llvm/test/Transforms/SLPVectorizer/X86 reduction-erased-copyable-cast-main-op.ll

[SLP]Fix crash on deleted main op of copyable state in reductions

Vectorizing one group of reduced values can erase the main operation of
a later group's copyable state, leaving it with dropped operands.

Fixes #228768

Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/228806
DeltaFile
+56-0llvm/test/Transforms/SLPVectorizer/X86/reduction-erased-copyable-cast-main-op.ll
+5-0llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+61-02 files

LLVM/project 0579ef5 — orc-rt/test/unit/support/sps SPSWrapperFunctionTest.cpp

[orc-rt] Use the Error matchers in SPSWrapperFunctionTest (#228804)

Use the Error matchers introduced in 4c8a437d0487 to clean up error
tests in SPSWrapperFunctionTest.
DeltaFile
+45-16orc-rt/test/unit/support/sps/SPSWrapperFunctionTest.cpp
+45-161 files

LLVM/project ee2abab — compiler-rt/lib/tsan/rtl tsan_interface_ann.cpp tsan_rtl_mutex.cpp

[NFC][TSan] Move ObtainCurrentStack out of ThreadRegistryLock scope

ObtainCurrentStack only reads the current thread's shadow stack and
allocates a VarSizeStackTrace buffer, which does not require
ThreadRegistryLock (and already runs outside ThreadRegistryLock in
ReportRace, SignalUnsafeCall, ReportErrnoSpoiling, and
ReportDestroyLocked).

Move ObtainCurrentStack (and dummy_pc in ReportDeadlock) before
ScopedReport in ReportMutexHeldWrongContext, ReportMutexMisuse, and
ReportDeadlock so the stack trace buffers also outlive ScopedReport and
OutputReport.

Assisted-by: Gemini

Pull Request: https://github.com/llvm/llvm-project/pull/228794
DeltaFile
+3-3compiler-rt/lib/tsan/rtl/tsan_rtl_mutex.cpp
+2-2compiler-rt/lib/tsan/rtl/tsan_interface_ann.cpp
+5-52 files

LLVM/project ca508d9 — compiler-rt/lib/tsan/rtl tsan_rtl_mutex.cpp

[NFC][TSan] Move RestoreStack before ScopedReport in ReportDestroyLocked

Run RestoreStack in its own lock scope before constructing ScopedReport
in ReportDestroyLocked (matching ReportRace). This avoids acquiring
ScopedErrorReportLock or symbolizing the current stack if RestoreStack
fails, and avoids holding slot_lock and slot_mtx while populating the
ScopedReport.

Assisted-by: Gemini

Pull Request: https://github.com/llvm/llvm-project/pull/228643
DeltaFile
+17-13compiler-rt/lib/tsan/rtl/tsan_rtl_mutex.cpp
+17-131 files

LLVM/project d707fb3 — compiler-rt/lib/tsan/rtl tsan_rtl.h tsan_rtl_report.cpp

[NFC][TSan] Merge ScopedReportBase into ScopedReport

ScopedReport is the only subclass of ScopedReportBase and nothing uses
ScopedReportBase directly. Merge ScopedReportBase into ScopedReport.

Assisted-by: Gemini

Pull Request: https://github.com/llvm/llvm-project/pull/228640
DeltaFile
+16-21compiler-rt/lib/tsan/rtl/tsan_rtl_report.cpp
+7-16compiler-rt/lib/tsan/rtl/tsan_rtl.h
+23-372 files

LLVM/project 265a53e — orc-rt/test/unit/support ProxyTest.cpp

[orc-rt] Use the Error matchers in ProxyTest (#228799)

Use the Error matchers introduced in 4c8a437d0487 to clean up error
tests in ProxyTest.
DeltaFile
+16-3orc-rt/test/unit/support/ProxyTest.cpp
+16-31 files

LLVM/project aa50c97 — llvm/lib/CodeGen CodeGenOptions.cpp CodeGenOptions.h, utils/bazel/llvm-project-overlay/llvm BUILD.bazel

[CodeGen] Declare command line options in TableGen

Move the cl::opts of TargetPassConfig.cpp and CodeGenPrepare.cpp into
CodeGenOptions.td, private to lib/CodeGen; other files will follow.
-enable-machine-outliner (cl::ValueOptional) and -regalloc
(RegisterPassParser) stay cl::opt.

CodeGenPrepare and its addressing-mode helpers hold
`const CodeGenOptions &Opts`; TargetPassConfig functions read
CodeGenOptions::Global. getCGPassBuilderOption() converts
std::optional<bool> members to the cl::boolOrDefault fields of the
public CGPassBuilderOption. -basic-block-section-match-infer, which was
not cl::Hidden, is now listed by -help-hidden only.

Aided by Opus 5.5
DeltaFile
+142-318llvm/lib/CodeGen/TargetPassConfig.cpp
+96-221llvm/lib/CodeGen/CodeGenPrepare.cpp
+200-0llvm/lib/CodeGen/CodeGenOptions.td
+19-0llvm/lib/CodeGen/CodeGenOptions.h
+15-0llvm/lib/CodeGen/CodeGenOptions.cpp
+11-0utils/bazel/llvm-project-overlay/llvm/BUILD.bazel
+483-5393 files not shown
+503-5449 files

LLVM/project 5bdcaa7 — mlir/docs Canonicalization.md, mlir/docs/DefiningDialects _index.md

[mlir][tblgen] Warn about the deprecated multi-result fold form

The legacy form `LogicalResult fold(FoldAdaptor,
SmallVectorImpl<OpFoldResult> &)` will be removed. This patch makes
`mlir-tblgen -gen-op-decls` warn when a dialect keeps `useOpFoldResults`
at 0 and has an op with `hasFolder` that does not have exactly one fixed
result. The warning points at the dialect definition. A note names the
first op that uses the legacy form.

The warning comes once for each dialect in one `-gen-op-decls` run. A
dialect with ops in more than one `.td` file can warn once for each file
that contains such an op. `-gen-op-defs` does not warn.

No in-tree dialect warns, because each in-tree dialect with such an op
already sets the bit.

The diagnostic follows the `-on-deprecated` option: `none` silences it,
`warn` (the default) warns, and `error` reports an error and fails the
run. `MlirTblgenMain.h` exposes the option value through

    [3 lines not shown]
DeltaFile
+88-0mlir/test/mlir-tblgen/op-fold-results-deprecation.td
+45-3mlir/tools/mlir-tblgen/OpDefinitionsGen.cpp
+6-3mlir/lib/Tools/mlir-tblgen/MlirTblgenMain.cpp
+6-0mlir/include/mlir/Tools/mlir-tblgen/MlirTblgenMain.h
+3-1mlir/docs/Canonicalization.md
+4-0mlir/docs/DefiningDialects/_index.md
+152-76 files

LLVM/project 204c379 — clang/test/CIR/Transforms canonicalize.cir, mlir/include/mlir/IR ExtensibleDialect.h

[mlir] Deprecate the legacy fold APIs with a results vector

The legacy `Operation::fold` overloads with a
`SmallVectorImpl<OpFoldResult> &` parameter drop partial folds. The
overloads that return `OpFoldResults` keep them.

This patch moves every in-tree caller of the legacy overloads to the
overloads that return `OpFoldResults`, except the unit tests of the
legacy APIs. `OpFoldResultsTest.cpp` suppresses the deprecation warnings
for these tests. Then the patch marks these APIs as deprecated: the two
legacy `Operation::fold` overloads, the legacy general `foldTrait` form,
the legacy fold hook overloads of `DynamicOpDefinition`, and the
`DynamicOpDefinition::LegacyFoldHookFn` alias. ODS cannot put an
attribute on an interface method, so only the documentation marks the
legacy `DialectFoldInterface::fold` method as deprecated.

The new code in `cir::CastOp::fold` also fixes two bugs. The fold
crashed when the source of an integral cast was a block argument, and it
read the fold result of result 0 when the source was a different result

    [17 lines not shown]
DeltaFile
+14-11mlir/lib/IR/ExtensibleDialect.cpp
+17-0clang/test/CIR/Transforms/canonicalize.cir
+11-4mlir/lib/IR/Operation.cpp
+6-8mlir/lib/Dialect/Affine/IR/AffineOps.cpp
+11-2mlir/include/mlir/IR/ExtensibleDialect.h
+8-3mlir/unittests/IR/OpFoldResultsTest.cpp
+67-289 files not shown
+103-5015 files

LLVM/project 17969ad — clang/include/clang/CIR/Dialect/IR CIRDialect.td, clang/lib/CIR/Dialect/IR CIRDialect.cpp

[mlir][CIR] Use OpFoldResults for the cir.scope fold

The CIR dialect now sets the `useOpFoldResults` bit. ODS then declares
`OpFoldResults fold(FoldAdaptor)` for each CIR op that does not have
exactly one fixed result. `cir.scope` is the only such op with a fold.
Its fold now returns the yielded value directly. The behavior does not
change. The existing test `clang/test/CIR/Transforms/canonicalize.cir`
covers this fold.

Signed-off-by: Víctor Pérez Carrasco <victor.pc.upm at gmail.com>
DeltaFile
+2-4clang/lib/CIR/Dialect/IR/CIRDialect.cpp
+2-0clang/include/clang/CIR/Dialect/IR/CIRDialect.td
+4-42 files

LLVM/project 5653c18 — mlir/lib/Dialect/Math/IR MathOps.cpp, mlir/lib/Dialect/Shape/IR Shape.cpp

[mlir] Use OpFoldResults in more upstream dialects

The scf, shape, sparse_tensor, gpu, math, and builtin dialects now set
the `useOpFoldResults` bit. ODS then declares `OpFoldResults
fold(FoldAdaptor)` for each of their ops that does not have exactly one
fixed result. This patch moves the seven folds of such ops to the new
form: `scf.if`, `shape.split_at`, `sparse_tensor.crd_translate`,
`gpu.memcpy`, `gpu.memset`, `math.sincos`, and
`builtin.unrealized_conversion_cast`.

The behavior of the folds does not change, except in the graph-region
case below. Each fold builds the same normalized `OpFoldResults` object
that the legacy adapter built from the old return. The existing tests of
each dialect cover these folds.

In a graph region, an operand that `unrealized_conversion_cast` or
`sparse_tensor.crd_translate` forwards can be another result of the
same op that the same fold replaces. The new form drops such a fold.
For a forward chain, the legacy form gave a correct result by the order

    [22 lines not shown]
DeltaFile
+37-0mlir/test/Dialect/Builtin/canonicalize.mlir
+16-0mlir/test/Dialect/SparseTensor/fold.mlir
+6-9mlir/lib/Dialect/SparseTensor/IR/SparseTensorDialect.cpp
+4-9mlir/lib/IR/BuiltinDialect.cpp
+3-7mlir/lib/Dialect/Math/IR/MathOps.cpp
+3-5mlir/lib/Dialect/Shape/IR/Shape.cpp
+69-308 files not shown
+79-3614 files