LLVM/project 60ac629 — llvm/lib/Transforms/Scalar TailRecursionElimination.cpp, llvm/test/Transforms/TailCallElim return-value-select-pgo.ll

[TailCallElim] Add profile annotations to return value selects

Tail call elimination in some cases can create selects on possible
return values conditioned on whether or not execution is currently in
what was a recursive call. That is equal to the probability with which
we recurse, which in turn can be computed from the block frequencies of
blocks that recurse and blocks that directly return.

Reviewers: mtrofin

Reviewed By: mtrofin

Pull Request: https://github.com/llvm/llvm-project/pull/202518
DeltaFile
+121-0llvm/test/Transforms/TailCallElim/return-value-select-pgo.ll
+29-0llvm/lib/Transforms/Scalar/TailRecursionElimination.cpp
+0-7llvm/utils/profcheck-xfail.txt
+150-73 files

LLVM/project 62e8956 — libc/test/src/math/exhaustive cos.wc sin.wc, llvm/lib/Support UnicodeNameToCodepointGenerated.cpp

feedback

Created using spr 1.3.7
DeltaFile
+1,091,085-0libc/test/src/math/exhaustive/sin.wc
+1,090,178-0libc/test/src/math/exhaustive/cos.wc
+82,648-81,221llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+31,001-87,165llvm/test/CodeGen/RISCV/rvv/clmulh-sdnode.ll
+40,941-24,498llvm/test/CodeGen/RISCV/clmul.ll
+24,053-23,916llvm/lib/Support/UnicodeNameToCodepointGenerated.cpp
+2,359,906-216,80049,344 files not shown
+6,170,671-2,241,57949,350 files

LLVM/project 5e07950 — libc/test/src/math/exhaustive cos.wc sin.wc, llvm/lib/Support UnicodeNameToCodepointGenerated.cpp

[𝘀𝗽𝗿] changes introduced through rebase

Created using spr 1.3.7

[skip ci]
DeltaFile
+1,091,085-0libc/test/src/math/exhaustive/sin.wc
+1,090,178-0libc/test/src/math/exhaustive/cos.wc
+82,648-81,221llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+31,001-87,165llvm/test/CodeGen/RISCV/rvv/clmulh-sdnode.ll
+40,941-24,498llvm/test/CodeGen/RISCV/clmul.ll
+24,053-23,916llvm/lib/Support/UnicodeNameToCodepointGenerated.cpp
+2,359,906-216,80049,343 files not shown
+6,170,666-2,241,56949,349 files

LLVM/project efea234 — llvm/lib/CodeGen/SelectionDAG TargetLowering.cpp, llvm/test/CodeGen/AMDGPU llvm.amdgcn.mul.i24.ll rotl.ll

[SelectionDAG] Handle constants in SimplifyMultipleUseDemandedBits

Replace a non-zero constant with zero when none of its set bits are
demanded.

This allows users of `SimplifyMultipleUseDemandedBits` to eliminate
irrelevant constant bits while preserving the convention that a null
SDValue indicates no simplification.
DeltaFile
+38-45llvm/test/CodeGen/AMDGPU/rotl.ll
+3-6llvm/test/CodeGen/AMDGPU/llvm.amdgcn.mul.i24.ll
+3-4llvm/test/CodeGen/PowerPC/ppc-rotate-clear.ll
+2-4llvm/test/CodeGen/SystemZ/shift-08.ll
+2-4llvm/test/CodeGen/SystemZ/shift-04.ll
+6-0llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
+54-634 files not shown
+58-7110 files

LLVM/project 17cfb00 — llvm/lib/Target/AMDGPU AMDGPUISelLowering.cpp, llvm/test/CodeGen/AMDGPU llvm.amdgcn.mul.i24.ll

Check one user
DeltaFile
+44-0llvm/test/CodeGen/AMDGPU/llvm.amdgcn.mul.i24.ll
+2-2llvm/lib/Target/AMDGPU/AMDGPUISelLowering.cpp
+46-22 files

LLVM/project 31d7b51 — clang/docs ReleaseNotes.md, clang/test/AST/ByteCode new-delete.cpp

[Clang] Add missing release note entry in #226753 (#227078)

As per the feedback from #226753, we add release note for GH-212211.
Also move the test to new-delete.cpp.

Assisted-by: Claude
DeltaFile
+0-29clang/test/SemaCXX/new-nothrow-by-value.cpp
+17-0clang/test/AST/ByteCode/new-delete.cpp
+5-0clang/docs/ReleaseNotes.md
+22-293 files

LLVM/project 2fc9681 — llvm/lib/CodeGen SpillPlacement.cpp, llvm/test/CodeGen/AMDGPU wave-profile-spill-zero-cost.mir wave-profile-spill.mir

[CodeGen] Enable validated AMDGPU wave spill costs by default

Use available validated wave counts for AMDGPU spill placement without an
explicit opt-in. Keep the existing target, mapping and normalization checks
and the ordinary block-frequency fallback for unavailable or rejected data.

Retain -enable-wave-profiled-spill=false for debugging and matched performance
comparisons. Test default behavior, explicit disabling, equivalence to explicit
enabling, fallback cases and the positive frequency floor.
DeltaFile
+14-11llvm/test/CodeGen/AMDGPU/wave-profile-spill.mir
+1-1llvm/test/CodeGen/AMDGPU/wave-profile-spill-zero-cost.mir
+1-1llvm/lib/CodeGen/SpillPlacement.cpp
+16-133 files

LLVM/project f896873 — clang/lib/Headers riscv_packed_simd.h, clang/test/CodeGen/RISCV rvp-intrinsics.c

[RISCV][P-ext] Support the missing packed pair (#226992)

This patch supports the packed pair for v2i16.
Some are inconsistent with the spec because existing shuffle-lowering
transforms them into equivalent narrowing shifts or zips.
DeltaFile
+176-0clang/test/CodeGen/RISCV/rvp-intrinsics.c
+52-0cross-project-tests/intrinsic-header-tests/riscv_packed_simd.c
+46-0llvm/test/CodeGen/RISCV/rvp-simd-32.ll
+29-1clang/lib/Headers/riscv_packed_simd.h
+6-0llvm/lib/Target/RISCV/RISCVInstrInfoP.td
+2-1llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+311-26 files

LLVM/project 449af54 — llvm/lib/Target/AMDGPU AMDGPUInstCombineIntrinsic.cpp, llvm/test/Transforms/InstCombine/AMDGPU llvm.amdgcn.sudot.ll

[AMDGPU] Fold sudot intrinsics with a zero multiplicand

Replace sudot4 and sudot8 with their accumulator when either
multiplicand is zero.

For example:
```
  sudot4(sign0, x, sign1, 0, acc, clamp)
  ->
  acc
```
DeltaFile
+2-4llvm/test/Transforms/InstCombine/AMDGPU/llvm.amdgcn.sudot.ll
+3-0llvm/lib/Target/AMDGPU/AMDGPUInstCombineIntrinsic.cpp
+5-42 files

LLVM/project eabb900 — llvm/lib/Target/AMDGPU AMDGPUInstCombineIntrinsic.cpp, llvm/test/Transforms/InstCombine/AMDGPU llvm.amdgcn.sudot.ll

[AMDGPU] Canonicalize constant operands of sudot intrinsics (#226478)

Move a constant multiplicand and its sign flag to the second operand.

For example:
```
  sudot4(true, 1, false, x, acc, clamp)
  ->
  sudot4(false, x, true, 1, acc, clamp)
```
DeltaFile
+15-0llvm/lib/Target/AMDGPU/AMDGPUInstCombineIntrinsic.cpp
+4-4llvm/test/Transforms/InstCombine/AMDGPU/llvm.amdgcn.sudot.ll
+19-42 files

LLVM/project 14a00a9 — clang/docs ReleaseNotes.md, clang/test/AST/ByteCode new-delete.cpp

fixup! [Clang] Add missing release note entry in GH226753
DeltaFile
+0-29clang/test/SemaCXX/new-nothrow-by-value.cpp
+17-0clang/test/AST/ByteCode/new-delete.cpp
+1-1clang/docs/ReleaseNotes.md
+18-303 files

LLVM/project 94b07b7 — llvm/lib/CodeGen SpillPlacement.cpp, llvm/test/CodeGen/AMDGPU wave-profile-spill-zero-cost.mir

[CodeGen] Keep wave-profiled spill frequencies positive

SpillPlacement expects positive block weights, but a valid wave profile can
record zero executions for a CFG-reachable block. Giving such a block zero
spill cost can make the allocator choose a very different placement.

Clamp every accepted wave-derived frequency to at least one, as we already
do for nonzero counts that round down to zero. Unmeasured or rejected blocks
still use their existing MBFI frequency. Add a focused MIR test for a valid
zero-wave record.

This pattern arose in a profiled Composable Kernel convolution case. With
the separate spill correctness fixes and partial spilling enabled, the
zero-cost policy failed two CPU-reference checks; the positive floor passed
both. The test checks the cost directly; the application result was checked
separately on gfx950.

DeltaFile
+27-0llvm/test/CodeGen/AMDGPU/wave-profile-spill-zero-cost.mir
+3-2llvm/lib/CodeGen/SpillPlacement.cpp
+30-22 files

LLVM/project 5d86702 — llvm/lib/CodeGen SpillPlacement.cpp, llvm/test/CodeGen/AMDGPU wave-profile-spill.mir

[CodeGen] Use validated wave counts for AMDGPU spill costs

Lane-based block frequencies can understate the cost of a spill in
divergent GPU code: a wave still executes a block with only some lanes
active. Use measured block-wave counts to weight SpillPlacement's costs
relative to the original entry-wave count.

Opt in on AMDGPU only. Accept a measured count only when its IR block
maps uniquely to a machine block with matching predecessors and
successors. Keep the existing MBFI cost for unmeasured or rejected
blocks, and leave branch probabilities and general BFI unchanged.



DeltaFile
+251-0llvm/test/CodeGen/AMDGPU/wave-profile-spill.mir
+97-1llvm/lib/CodeGen/SpillPlacement.cpp
+348-12 files

LLVM/project 9a0c255 — llvm/include/llvm/ProfileData InstrProf.h, llvm/lib/ProfileData InstrProf.cpp

[InstrProf] Replace !PGOFuncName and !PGOName metadata with !guid (#214134)

!PGOFuncName and !PGOName metadata were attached to internal functions
and vtables during profile annotation to record their original
"<file>;<name>" PGO names before ThinLTO promoted and renamed them. In
post-link LTO passes, InstrProfSymtab read that metadata back so profile
records keyed by the original name's hash could still find the renamed
IR object.

Global objects now carry stable !guid metadata assigned before LTO
renaming, which records the MD5 hash of the original PGO name directly.
See: https://discourse.llvm.org/t/rfc-keep-globalvalue-guids-stable/84801

Use !guid instead of maintaining separate PGO name metadata:

- Stop emitting and reading !PGOFuncName and !PGOName in Clang and
  PGOInstrumentation, and remove the metadata helper functions
  (createPGOFuncNameMetadata, createPGONameMetadata,
  getPGOFuncNameMetadata, and the metadata name getters).

    [13 lines not shown]
DeltaFile
+29-83llvm/lib/ProfileData/InstrProf.cpp
+102-9llvm/unittests/ProfileData/InstrProfTest.cpp
+15-30llvm/include/llvm/ProfileData/InstrProf.h
+44-0llvm/test/tools/llvm-profdata/Inputs/guid-roundtrip.proftext
+35-0llvm/test/tools/llvm-profdata/guid-roundtrip.test
+28-0llvm/test/tools/llvm-profdata/Inputs/guid-roundtrip.c
+253-12217 files not shown
+291-17823 files

LLVM/project e521fe9 — llvm/lib/Transforms/Utils LowerSwitch.cpp, llvm/test/Transforms/LowerSwitch wave-profile.ll profile-weights.ll

[Transforms] Preserve wave profiles across CFG rewrites

HIP device PGO attaches measured wave counts to IR blocks. Later CFG
rewrites can drop counts from unchanged blocks or leave stale counts
on blocks that now execute differently. Either case makes the profile
unreliable for later optimizations.

Preserve counts through switch lowering, structurization, and loop
rotation only when a block still represents the same executions. Keep
unaffected counts and their IDs even when a loop header's count must be
invalidated. Transfer branch hints only for equivalent decisions, and
avoid assigning switch weights when default traffic cannot be traced
to one edge.




DeltaFile
+157-0llvm/test/Transforms/LowerSwitch/profile-weights.ll
+113-10llvm/lib/Transforms/Utils/LowerSwitch.cpp
+68-0llvm/unittests/Transforms/Utils/LoopRotationUtilsTest.cpp
+56-0llvm/test/Transforms/LowerSwitch/wave-profile.ll
+53-0llvm/test/Transforms/StructurizeCFG/wave-profile-loop-prefix.ll
+50-0llvm/test/Transforms/StructurizeCFG/wave-profile.ll
+497-108 files not shown
+629-5114 files

LLVM/project 5ef0a66 — llvm/lib/Transforms/Instrumentation PGOInstrumentation.cpp, llvm/test/Transforms/PGOProfile dense-wave-cfg.ll wave-profile-use.ll

[PGO] Load dense block wave counts from device profiles

Use the profile's dense layout flag to map appended wave-only slots after
the original block/select prefix. Reuse the producer's block selection so
eligible loop and reconvergence blocks retain their directly measured wave
frequencies. Keep select slots out of the block mapping.

Test dense and sparse profiles, measured zeros, generation and metadata
switches, and exact loop block identities after critical-edge splitting.
DeltaFile
+49-3llvm/test/Transforms/PGOProfile/wave-profile-use.ll
+15-1llvm/test/Transforms/PGOProfile/dense-wave-cfg.ll
+11-4llvm/lib/Transforms/Instrumentation/PGOInstrumentation.cpp
+75-83 files

LLVM/project dcef0bb — llvm/lib/Transforms/Instrumentation PGOInstrumentation.cpp, llvm/test/Transforms/PGOProfile wave-profile-use.ll

[PGO] Add a debugging switch for wave profile metadata

Add the hidden pgo-wave-metadata option, enabled by default, for debugging,
performance comparisons and disabling wave annotations when investigating
regressions without turning off ordinary PGO or uniformity hints.

Gate wave metadata emission while retaining the existing clearing of stale
function and block annotations during profile use. Profile collection and
ordinary count reconstruction are unchanged.

Test default/explicit enablement, disabling, retained counts and uniformity
hints, and replacement profiles with missing, mismatched or zero counts.

DeltaFile
+34-1llvm/test/Transforms/PGOProfile/wave-profile-use.ll
+6-1llvm/lib/Transforms/Instrumentation/PGOInstrumentation.cpp
+40-22 files

LLVM/project e655a71 — llvm/lib/Transforms/Instrumentation PGOInstrumentation.cpp, llvm/test/Transforms/PGOProfile wave-profile-use.ll

[PGO] Load GPU wave counts into IR metadata

GPU profiles contain wave counts alongside lane counts, but profile use
does not expose them to optimizations. Wave visits do not obey scalar
flow conservation, so unmeasured blocks cannot use counts reconstructed
from neighboring blocks or ordinary branch weights.

Map wave-counter indices to the blocks selected by PGO instrumentation,
after reproducing its critical-edge splits. Attach measured counts using
wave.profile metadata, retaining measured zeros and marking other blocks
unmeasured. Require a measured entry count for normalization and exclude
select-counter slots from the block mapping.

Validate the wave-counter layout against the accepted lane profile.
Clear old wave metadata when loading a replacement profile, including
when a function has no usable record. Do not emit wave metadata for
previously profiled functions: their branch weights may change counter
placement without changing the CFG hash. Keep ordinary lane-count
reconstruction and branch weights unchanged.

DeltaFile
+212-0llvm/test/Transforms/PGOProfile/wave-profile-use.ll
+44-0llvm/lib/Transforms/Instrumentation/PGOInstrumentation.cpp
+256-02 files

LLVM/project 37656e7 — llvm/docs LangRef.md, llvm/include/llvm/IR ProfDataUtils.h

[IR] Define GPU wave-profile metadata

Existing offload GPU profile counters measure lane executions, while
GPU instructions execute at wave granularity under an active-lane
mask. A block visited by every wave can therefore look cold when only
a few lanes are active. Scalar branch weights also cannot represent a
divergent wave visiting both successors before reconverging. These
profiles are a poor fit for optimizations that estimate work performed
by a wave.

Lane counters remain useful for measuring per-lane branch selectivity
and estimating work that scales with the number of active lanes.
Wave counters cannot replace them: a visit with one active lane and a
visit with all lanes active both count as one. The two profiles provide
complementary information about instruction execution and lane activity.

Introduce wave.profile and wave.profile.block metadata to represent
measured wave visits to IR blocks. Each dynamic visit with at least one
active lane contributes one to the count. A measured zero is distinct

    [18 lines not shown]
DeltaFile
+817-0llvm/unittests/IR/ProfDataUtilsTest.cpp
+396-0llvm/lib/IR/ProfDataUtils.cpp
+91-0llvm/test/Verifier/wave-profile.ll
+74-0llvm/include/llvm/IR/ProfDataUtils.h
+47-0llvm/docs/LangRef.md
+40-0llvm/test/Bitcode/wave-profile.ll
+1,465-05 files not shown
+1,525-011 files

LLVM/project 76c7824 — llvm/lib/Transforms/Instrumentation PGOInstrumentation.cpp, llvm/test/Transforms/PGOProfile dense-wave-instrumentation.ll dense-wave-profile-use.ll

[PGO] Collect dense AMDGPU block wave counts

Wave visits are not additive across divergent control flow. Sparse scalar
counter sites can leave repeated loop blocks unmeasured, so their wave
frequencies cannot be reconstructed from entry and edge counts.

Append zero-step instrumentation for eligible unmeasured AMDGPU blocks,
keeping the existing lane and select counter indices unchanged. Exclude the
appended slots from lane-flow reconstruction and uniformity annotation.

Identify the layout with a profile variant bit, preserve it through raw and
indexed readers/writers, and select it automatically during profile use.
Keep sparse profiles readable and reject incompatible merges, including
concatenated raw profiles. Ignore empty merge-worker contexts.

Enable dense collection for ordinary AMDGPU IR-PGO by default, with the
hidden -pgo-instrument-dense-wave-counts option for debugging. Leave
context-sensitive, coverage, and temporal instrumentation unchanged.


    [2 lines not shown]
DeltaFile
+108-0llvm/test/Transforms/PGOProfile/dense-wave-cfg.ll
+82-0llvm/test/Transforms/PGOProfile/dense-wave-profile-use.ll
+60-8llvm/lib/Transforms/Instrumentation/PGOInstrumentation.cpp
+59-0llvm/test/Transforms/PGOProfile/dense-wave-instrumentation.ll
+43-0llvm/test/tools/llvm-profdata/dense-wave-layout.test
+31-0llvm/unittests/ProfileData/InstrProfTest.cpp
+383-88 files not shown
+441-1414 files

LLVM/project 78f2beb — mlir/include/mlir/Dialect/SCF/IR SCFOps.td, mlir/lib/Dialect/SCF/IR SCF.cpp

[mlir][scf] Add unsignedCmp to scf.parallel and use it in the tiling in-bound check (#226130)

This change mirrors `scf.for` where `scf.parallel` gets an `unsignedCmp`
unit attribute, and every pass that rebuilds or lowers a `scf.parallel`
has to respect it:

- Propagate where a loop is rebuilt from another, such as parallel loop
tiling, parallel loop fusion (which also refuses to fuse loops of
different signedness), parallel-to-nested-fors and SCF-to-CF lowering.
- Decline where the bound arithmetic assumes signed values, such as
SCF-to-GPU, SCF-to-OpenMP, async-parallel-for and the
parallel-loop-collapsing test pass.

Fixes #223233
DeltaFile
+44-2mlir/test/Dialect/SCF/parallel-loop-tiling-inbound-check.mlir
+36-0mlir/test/Dialect/SCF/parallel-loop-fusion.mlir
+22-6mlir/include/mlir/Dialect/SCF/IR/SCFOps.td
+15-4mlir/lib/Dialect/SCF/IR/SCF.cpp
+16-0mlir/test/Dialect/Async/async-parallel-for-async-dispatch.mlir
+16-0mlir/test/Conversion/SCFToControlFlow/convert-to-cfg.mlir
+149-1213 files not shown
+243-2119 files

LLVM/project 7b5ec23 — llvm/test/tools/llvm-readobj/ELF/RISCV eflags-abi.test eflags.test

[RISCV][llvm-readobj] Generalize the eflags-abi test to cover all of the eflags. NFC (#227152)

Assisted-by: Claude
DeltaFile
+56-0llvm/test/tools/llvm-readobj/ELF/RISCV/eflags.test
+0-51llvm/test/tools/llvm-readobj/ELF/RISCV/eflags-abi.test
+56-512 files

LLVM/project 619028d — llvm/test lit.cfg.py, llvm/utils profcheck-xfail.txt

[ProfCheck] Exclude DirectX (#227157)

I don't think anyone runs profiling on DirectX, so exclude it for now.
Also move AMDGPU to the normal exclusion list given there are efforts
around PGO for AMDGPU currently.
DeltaFile
+12-4llvm/test/lit.cfg.py
+0-1llvm/utils/profcheck-xfail.txt
+12-52 files

LLVM/project b1faf8b — lldb/packages/Python/lldbsuite/test lldbplatform.py

[lldb/test] Register xros (visionOS) as a Darwin platform (#227142)

`lldbplatform.py` never registered `xros` at all: no enum value, no
`__name_lookup` entry, and it was absent from `__darwin_embedded`/
`darwin_all`. As a result, `platformIsDarwin()` returned `False` for
`xros`, and `finalize_build_dictionary`'s fallback branch, which indexes
`platform_name_to_uname` by the raw platform name, would `KeyError`
before any visionOS test could build.

Add `xros` alongside the other embedded Darwin platforms so it's
recognized the same way `ios`/`tvos`/`watchos`/`bridgeos` already are.

Signed-off-by: Med Ismail Bennani <ismail at bennani.ma>
DeltaFile
+10-3lldb/packages/Python/lldbsuite/test/lldbplatform.py
+10-31 files

LLVM/project c894e63 — .github/workflows libc-overlay-tests.yml libc-fullbuild-tests.yml

[Github] Temporary fix for #226230 (#227147)

We are seeing this issue in the libc++ runner sets and it can presumably
pop up in the libc runner sets as well in the case of an abnormally long
clone operation.
DeltaFile
+5-0.github/workflows/libc-overlay-tests.yml
+5-0.github/workflows/libc-fullbuild-tests.yml
+10-02 files

LLVM/project 0ec4a7f — lldb/packages/Python/lldbsuite/test lldbplatformutil.py

[lldb/test] Normalize appletvos to tvos in getPlatform() (#227138)

Mirror the existing `iphoneos` -> `ios` SDK-name normalization for
`tvOS`. Without it, an `--apple-sdk` value derived from "appletvos" made
`getPlatform()` return "appletvos", which isn't in
`lldbplatform.__name_lookup`, so `platformIsDarwin(`) returned `False`
and `finalize_build_dictionary` KeyError'd before any `remote-tvos` test
could build.

Signed-off-by: Med Ismail Bennani <ismail at bennani.ma>
DeltaFile
+4-0lldb/packages/Python/lldbsuite/test/lldbplatformutil.py
+4-01 files

LLVM/project 093f672 — llvm/lib/Target/AMDGPU GCNSubtarget.cpp

Clang format

Change-Id: Id1095f9eb7c6f81bc373a58d0bcc3584211f241e
DeltaFile
+2-1llvm/lib/Target/AMDGPU/GCNSubtarget.cpp
+2-11 files

LLVM/project 32fb337 — llvm/lib/Target/AMDGPU GCNSubtarget.cpp

Add comment on COPY latency calculation

Change-Id: I561f60aacf2180036a8567b4fa054b7d28a4f0df
DeltaFile
+10-0llvm/lib/Target/AMDGPU/GCNSubtarget.cpp
+10-01 files

LLVM/project fbac8e1 — llvm/lib/Target/SPIRV SPIRVPreLegalizer.cpp, llvm/test/CodeGen/SPIRV/legalization signed-narrow-int.ll

[SPIR-V] Sign extend narrow G_SMIN/G_SMAX operands and mask sign sensitive results (#226957)

This PR includes `G_SMIN` and `G_SMAX` to the list of sign sensitive
ops. It also masks the result of the sign sensitive ops to their
original widths with a bitwise AND because the later users assume the
upper bits are zero and without masking they read the wrong value.
DeltaFile
+53-1llvm/test/CodeGen/SPIRV/legalization/signed-narrow-int.ll
+26-6llvm/lib/Target/SPIRV/SPIRVPreLegalizer.cpp
+79-72 files

LLVM/project d718391 — flang/lib/Semantics check-acc-structure.h check-acc-structure.cpp, flang/test/Semantics/OpenACC acc-routine-loop-nesting.f90

[flang][acc] Warn when a loop is above its routine parallelism level (#227041)

OpenACC spec allows a routine to parent a loop at its own parallelism
level or below. A higher level is not allowed: gang inside a worker
routine, gang or worker inside a vector routine, gang, worker, or vector
inside a seq routine, and a gang dimension above the routine's gang
dimension. A device-specific routine clause counts the same way.

Flang accepted these loops with no diagnostic. Warn that the clause is
ignored. A loop at the same level or below is left alone, and illegal
nesting of loops remains an error.
DeltaFile
+355-50flang/lib/Semantics/check-acc-structure.cpp
+218-0flang/test/Semantics/OpenACC/acc-routine-loop-nesting.f90
+3-2flang/lib/Semantics/check-acc-structure.h
+576-523 files