LLVM/project cc2a16c — clang/include/clang/CIR/Dialect/Builder CIRBaseBuilder.h, clang/include/clang/CIR/Dialect/IR CIRAttrs.td

[CIR] Emit inbounds/inrange for vtable address points in initializers (#227520)

Classic codegen emits `inbounds inrange(...)` for vtable address points
in global initializers: VTT entries and the vptr of a
constant-initialized object. CIR never emitted any flags for
`#cir.global_view`. It only got `inbounds` and `nuw` when LLVM's
constant folder inferred them, which stopped for a while after
https://github.com/llvm/llvm-project/pull/226904 until #227651 restored
the folding. It never got `inrange`.

This patch emits the flags in CIR instead of relying on the folder:

- Add an `address_point` flag to `#cir.global_view`, spelled
`#cir.global_view<@vtable, [vtable index, slot], address_point>`.
- Set it where CIRGen builds VTT entries and constant vptrs, and keep it
when CIRGen and `CXXABILowering` rebuild the attribute.
- When lowering to LLVM, make an address point `inbounds`, with an
`inrange` covering the one vtable in the group that it points into. The
bounds come from the vtable global's type, so they also work when the

    [10 lines not shown]
DeltaFile
+35-0clang/test/CIR/IR/global-view-address-point.cir
+14-14clang/test/CIR/CodeGen/vtt.cpp
+24-3clang/lib/CIR/Lowering/DirectToLLVM/LowerToLLVM.cpp
+14-6clang/include/clang/CIR/Dialect/IR/CIRAttrs.td
+15-2clang/test/CIR/CodeGen/global-ptr-init.cpp
+6-4clang/include/clang/CIR/Dialect/Builder/CIRBaseBuilder.h
+108-296 files not shown
+123-4112 files

LLVM/project 5ee74a0 — llvm/lib/Target/AArch64 AArch64SVEInstrInfo.td SVEInstrFormats.td

[AArch64][SVE] Fix predicate type for quadword gather/scatter SDNodes (#228274)

GLD1Q_MERGE_ZERO and SST1Q_PRED use a single predicate bit per 128-bit
quadword (always nxv1i1), not one bit per data element.
SDT_AArch64_GATHER_VS and SDT_AArch64_SCATTER_VS incorrectly required
the predicate to have the same number of elements as the loaded/stored
vector via SDTCisSameNumEltsAs<0,1>, which only happened to hold for the
data width these nodes were originally used with.

Give these two nodes their own SDTypeProfile that fixes the predicate
type to nxv1i1, and update the corresponding patterns in
SVEInstrFormats.td to match.

Co-authored-by: Claude Sonnet 5.5 <noreply at anthropic.com>
DeltaFile
+16-16llvm/lib/Target/AArch64/SVEInstrFormats.td
+18-2llvm/lib/Target/AArch64/AArch64SVEInstrInfo.td
+34-182 files

LLVM/project 3fab81a — llvm/lib/Target/X86 X86ISelLowering.cpp

[X86] LowerSINT_TO_FP - pull out repeated CONCAT_VECTORS widening. NFC. (#230164)
DeltaFile
+5-7llvm/lib/Target/X86/X86ISelLowering.cpp
+5-71 files

LLVM/project 6d56e63 — llvm/utils/gn/secondary/clang-tools-extra/clangd/test BUILD.gn

[gn] port 9f3af63e886d (#230199)
DeltaFile
+9-0llvm/utils/gn/secondary/clang-tools-extra/clangd/test/BUILD.gn
+9-01 files

LLVM/project bc984a9 — llvm/lib/Target/AArch64 AArch64SchedPredExynos.td AArch64SchedPredicates.td, llvm/lib/Target/AArch64/MCTargetDesc AArch64AddressingModes.h

[AArch64] Remove stale packed extend/shift for memory ops(NFC) (#230068)

This patch removes stale extend/shift representation for memory ops.
They use packed representation which has been replaced by operands to
the instruction format.
DeltaFile
+0-29llvm/lib/Target/AArch64/MCTargetDesc/AArch64AddressingModes.h
+1-10llvm/lib/Target/AArch64/AArch64SchedPredicates.td
+1-4llvm/lib/Target/AArch64/AArch64SchedPredExynos.td
+2-433 files

LLVM/project f5ba409 — llvm/lib/Target/AMDGPU SISchedule.td, llvm/lib/Target/AMDGPU/Utils AMDGPUBaseInfo.h

[AMDGPU] Model gfx1250-strict WMMA latencies

gfx1250-strict shares the gfx1250 scheduling model, but some WMMA
instructions take longer on it:
- 16x16x64 FP8/BF8: 8 cycles instead of 4.
- f8f6f4: 16 cycles when any input is f8 and 8 cycles otherwise, instead of
  8 and 4.

The WMMA co-execution hazard category is derived from the latency, so this
also gives these instructions the required number of wait states on
gfx1250-strict.

Read the f8f6f4 matrix formats through named operands for both MCInst and
MachineInstr, so llvm-mca models the strict latencies too. This also fixes the
both-f4 check, which read fixed operand indices and could take other
immediates for f4 formats: inline constant scales of the scaled forms, which
gave too few hazard wait states, and the neg_lo/neg_hi operands of the
unscaled form in assembler input.
DeltaFile
+624-208llvm/test/CodeGen/AMDGPU/wmma-hazards-gfx1250-w32.mir
+340-127llvm/test/CodeGen/AMDGPU/wmma-coexecution-valu-hazards.mir
+72-0llvm/test/tools/llvm-mca/AMDGPU/gfx1250-strict-wmma-cycles.s
+23-12llvm/lib/Target/AMDGPU/SISchedule.td
+24-0llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.h
+16-7llvm/test/tools/llvm-mca/AMDGPU/gfx1250-wmma-cycles.s
+1,099-3546 files

LLVM/project 239c882 — flang/lib/Lower Bridge.cpp, flang/lib/Lower/OpenMP DataSharingProcessor.h DataSharingProcessor.cpp

[flang][OpenMP] Privatize loop IVs in the innermost parallel (#227486)

A sequential loop's iteration variable is predetermined private in the
innermost parallel, teams or task-generating construct that encloses
the loop. When the loop was nested in another construct, such as a
worksharing loop, lowering failed to privatize the variable in the
enclosing parallel region. Instead, it created a new local copy inside
the nested construct, so the parallel region's other references to the
variable used the shared host variable.

Fix this by deciding which construct privatizes a symbol based on the
scope that owns it. This also simplifies DataSharingProcessor: the
OMPConstructSymbolVisitor, which walked the parse tree to track where
symbols were defined, is no longer needed and has been removed.

I have noticed that metadirectives don't always own the symbols that
should be privatized in them, as semantics doesn't create a new scope.
I'm not very familiar with metadirectives, but it seems this causes some
privatization issues with non-explicitly specified DSAs, as

    [7 lines not shown]
DeltaFile
+206-0flang/test/Lower/OpenMP/predetermined-do-iv.f90
+70-130flang/lib/Lower/OpenMP/DataSharingProcessor.cpp
+0-73flang/lib/Lower/OpenMP/DataSharingProcessor.h
+6-7flang/test/Lower/OpenMP/unstructured.f90
+4-6flang/test/Lower/OpenMP/shared-loop.f90
+1-8flang/lib/Lower/Bridge.cpp
+287-2242 files not shown
+292-2278 files

LLVM/project b05db3d — llvm/lib/Transforms/InstCombine InstCombineMulDivRem.cpp, llvm/test/Transforms/InstCombine fmul.ll

[InstCombine] Fix profiles in foldMulSelectToNegate (#229865)

The condition is always the same as the original select, so we can just
propagate the metadata after using a matcher to ensure that we get the
actual SI out.
DeltaFile
+22-12llvm/lib/Transforms/InstCombine/InstCombineMulDivRem.cpp
+10-4llvm/test/Transforms/InstCombine/fmul.ll
+0-2llvm/utils/profcheck-xfail.txt
+32-183 files

LLVM/project 7658602 — llvm/test/CodeGen/AMDGPU fma.ll reorder-stores.ll

AMDGPU: Remove remaining uses of -amdgpu-scalarize-global-loads=false

Stop relying on the option to select vector loads from uniform kernel
argument pointers.

Index loads by the workitem id where the test is about memory operations.
Use volatile loads where the test depends on a uniform value held in
VGPRs. Use functions with inreg pointer arguments for the MUBUF encoding
tests, which preserves the existing encodings. Convert the early
if-conversion tests that only need operand values into functions.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+74-26llvm/test/CodeGen/AMDGPU/global-extload-i16.ll
+58-17llvm/test/CodeGen/AMDGPU/urem.ll
+41-29llvm/test/CodeGen/AMDGPU/load-global-i64.ll
+39-27llvm/test/CodeGen/AMDGPU/load-global-f64.ll
+36-23llvm/test/CodeGen/AMDGPU/reorder-stores.ll
+31-19llvm/test/CodeGen/AMDGPU/fma.ll
+279-14119 files not shown
+463-30425 files

LLVM/project ae7f7ef — .github/workflows pr-code-format.yml

[GitHub] Bump CI format container (#230192)

To pull in the bump to 23.1.3.

Fixes #230169
DeltaFile
+1-1.github/workflows/pr-code-format.yml
+1-11 files

LLVM/project ff995ee — llvm/test/CodeGen/AMDGPU llvm.amdgcn.trig.preop.ll select-vectors.ll

AMDGPU: Use functions in more operation tests instead of kernel loads

Stop relying on -amdgpu-scalarize-global-loads=false. Inputs are passed
as VGPR arguments, or inreg for SGPR operands. Kernels that check
multiple stores index their loads by workitem id.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+89-171llvm/test/CodeGen/AMDGPU/shift-i64-opts.ll
+73-181llvm/test/CodeGen/AMDGPU/v_mac_f16.ll
+80-119llvm/test/CodeGen/AMDGPU/s_movk_i32.ll
+20-91llvm/test/CodeGen/AMDGPU/v_mac.ll
+31-74llvm/test/CodeGen/AMDGPU/select-vectors.ll
+10-20llvm/test/CodeGen/AMDGPU/llvm.amdgcn.trig.preop.ll
+303-6566 files

LLVM/project 8c691cf — clang/lib/CodeGen CodeGenModule.cpp, clang/test/CodeGen/X86 ms-hotpatch.c

[CodeView] Set the S_COMPILE3 hotpatch flag from a module flag (#229868)

Previously, the `HotPatch` flag in CodeView `S_COMPILE3` was only
derived from `TargetOptions::Hotpatch`, which is set by clang's backend
setup but not by LTO code generation. Objects produced by LTO were
therefore never marked as hotpatchable, and lld's `/FUNCTIONPADMIN`,
which only pads chunks from hotpatchable objects, left their functions
unpadded, even though the `"patchable-function"` attribute still made
code generation pad the function entries.

After this PR, emit an `"ms-hotpatch"` module flag when compiling with
`/hotpatch`, and also set the CodeView flag from it. The flag uses the
`Min` merge behavior, so when LTO merges modules, it is reset to 0 if
any module was compiled without `/hotpatch`, and the result is never
wrongly marked as hotpatchable.

This came up while working on
https://github.com/llvm/llvm-project/pull/229477


    [4 lines not shown]
DeltaFile
+67-0llvm/test/DebugInfo/COFF/hotpatch-module-flag.ll
+21-0clang/test/CodeGen/X86/ms-hotpatch.c
+17-0llvm/docs/LangRef.md
+3-6llvm/include/llvm/Target/TargetOptions.h
+7-2llvm/lib/CodeGen/AsmPrinter/CodeViewDebug.cpp
+5-0clang/lib/CodeGen/CodeGenModule.cpp
+120-83 files not shown
+125-99 files

LLVM/project 2d4b179 — llvm/lib/Transforms/Utils LoopUnroll.cpp, llvm/test/Transforms/LoopUnroll runtime-unroll-remainder-deletes-loop-block.ll

[LoopUnroll] Snapshot loop blocks after runtime remainder generation (#229779)

Runtime remainder unrolling can simplify the enclosing loop nest and
delete blocks from the loop being unrolled. The early OriginalLoopBlocks
snapshot then contains dangling pointers, causing a crash during
dominator-tree updates.

Capture the snapshot after remainder generation and before cloning adds
new blocks to the loop. The included test aborts without the fix with
`cannot get DomTreeNode of block with different parent (exit 134)`, and
passes lit with the fix.

Assisted by AI tooling, to triage the bug and find an LLVM IR reproducer
for the crash.
DeltaFile
+47-0llvm/test/Transforms/LoopUnroll/runtime-unroll-remainder-deletes-loop-block.ll
+5-2llvm/lib/Transforms/Utils/LoopUnroll.cpp
+52-22 files

LLVM/project 3a8f293 — clang/include/clang/CIR/Dialect/Transforms CIRTransformUtils.h, clang/lib/CIR/Dialect/Transforms CIRTransformUtils.cpp

[CIR] Put try_throw's normal destination in the throw's region (#229915)

`replaceThrowWithTryThrow` creates the unreachable normal destination of
`cir.try_throw` at the end of the parent function.

This is not correct for coroutines. Doing so would create a reference to
a block outside their respective regions, leading to a verification
error.

This patch moves the unreachable normal destination of `cir.try_throw`
in `replaceThrowWithTryThrow` right below the throw's region.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+102-0clang/test/CIR/CodeGenCoroutines/coro-throw.cpp
+4-9clang/lib/CIR/Dialect/Transforms/CIRTransformUtils.cpp
+5-5clang/test/CIR/CodeGen/cleanup-throw-from-cleanup.cpp
+4-4clang/include/clang/CIR/Dialect/Transforms/CIRTransformUtils.h
+2-2clang/test/CIR/CodeGen/try-catch.cpp
+117-205 files

LLVM/project 23fae18 — llvm/utils/gn/secondary/llvm/unittests/Frontend BUILD.gn

[gn build] Port c4d9267a31897 (#230177)
DeltaFile
+1-0llvm/utils/gn/secondary/llvm/unittests/Frontend/BUILD.gn
+1-01 files

LLVM/project bd0d503 — llvm/lib/Target/RISCV RISCVISelLowering.cpp, llvm/test/CodeGen/RISCV clmulr.ll rv64zbc-intrinsic.ll

[RISCV] Combine (srl (clmul (and X, 0xffffffff), (and Y, 0xffffffff)), 31). (#228848)

into (srl (clmulr (shl X, 32), (shl Y, 32)), 32).
DeltaFile
+37-2llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+4-6llvm/test/CodeGen/RISCV/rv64zbc-intrinsic.ll
+2-2llvm/test/CodeGen/RISCV/clmulr.ll
+43-103 files

LLVM/project 111c4da — lld/MachO ConcatOutputSection.cpp

[lld][MachO] Write the contents of an output section in parallel (#229512)

This is part of the ld64.lld performance improvements tracked by #222068

Writer::writeSections() does actually run the output sections in
parallel, but each section's writeTo() is not done in parallel, so
there's still a lot of work done by one thread at a time.

We can parallelize the writeTo() calls pretty easily as each input
writes to its own range in the output buffer.

This commit is output-preserving.
DeltaFile
+11-15lld/MachO/ConcatOutputSection.cpp
+11-151 files

LLVM/project 7f82a6c — llvm/utils/gn/secondary/clang/lib/Lex BUILD.gn

[gn build] Port bd5327092e65a (#230176)
DeltaFile
+1-0llvm/utils/gn/secondary/clang/lib/Lex/BUILD.gn
+1-01 files

LLVM/project 34a16b4 — llvm/utils/gn/secondary/clang/lib/CodeGenUtils BUILD.gn

[gn build] Port b54c096369ab (#230174)
DeltaFile
+1-0llvm/utils/gn/secondary/clang/lib/CodeGenUtils/BUILD.gn
+1-01 files

LLVM/project 7a173eb — llvm/lib/Transforms/InstCombine InstCombineCalls.cpp, llvm/test/Transforms/InstCombine fpcast.ll

[InstCombine] Fold fpto[su]i.sat of NaN guarded select

fpto[su]i.sat returns 0 for both NaN and $\pm 0.0$, so a select that
replaces a NaN input with zero does not change the result, and the call
can take X directly.

```llvm
%not.nan = fcmp ord float %x, 0.0
%sel = select i1 %not.nan, float %x, float 0.0
%r = call i32 @llvm.fptosi.sat.i32.f32(float %sel)
  -->
%r = call i32 @llvm.fptosi.sat.i32.f32(float %x)
```

The inverted form `select (fcmp uno X, 0.0), 0.0, X` is folded too.
Only the canonical `fcmp ord/uno X, 0.0` is matched, which also covers
`oeq/une X, X` after fcmp canonicalization.
DeltaFile
+13-25llvm/test/Transforms/InstCombine/fpcast.ll
+16-1llvm/lib/Transforms/InstCombine/InstCombineCalls.cpp
+29-262 files

LLVM/project c1344d4 — llvm/test/Transforms/InstCombine fpcast.ll

[InstCombine] Add tests for fpto[su]i.sat of NaN guarded select. NFC

fpto[su]i.sat returns 0 for NaN and +/-0.0, so the select is redundant.
Cover ord/oeq/uno/une checks, scalar and vector, and negative cases.
DeltaFile
+179-0llvm/test/Transforms/InstCombine/fpcast.ll
+179-01 files

LLVM/project 7c60a6a — llvm/utils/gn/secondary/llvm/lib/Support BUILD.gn

[gn build] Port bd5327092e65 (#230175)
DeltaFile
+1-0llvm/utils/gn/secondary/llvm/lib/Support/BUILD.gn
+1-01 files

LLVM/project 9de0686 — flang/lib/Semantics check-omp-atomic.cpp check-omp-structure.h, flang/test/Semantics/OpenMP declare-target02.f90 flush02.f90

[flang][OpenMP] Switch clause verification to descriptor-based

Delete all the scattered pieces of clause verification that are now
replaced by the unified handling.
DeltaFile
+57-273flang/lib/Semantics/check-omp-structure.cpp
+300-0flang/lib/Semantics/check-omp-syntax.cpp
+22-10flang/lib/Semantics/check-omp-structure.h
+0-31flang/lib/Semantics/check-omp-atomic.cpp
+8-8flang/test/Semantics/OpenMP/flush02.f90
+15-0flang/test/Semantics/OpenMP/declare-target02.f90
+402-32236 files not shown
+516-38342 files

LLVM/project 8a8d1b2 — flang/lib/Semantics check-omp-syntax.cpp, flang/test/Semantics/OpenMP uses-allocators-version51.f90 uses-allocators-version50.f90

[flang][OpenMP] Improve diagnostics about modifier properties

Use OpenMPDeprecated and OpenMPFuture warning categories for modifier
diagnostics as well, analogously to how they are used for clauses.
DeltaFile
+66-30flang/lib/Semantics/check-omp-syntax.cpp
+4-4flang/test/Semantics/OpenMP/map-modifiers-v60.f90
+3-3flang/test/Semantics/OpenMP/uses-allocators-version50.f90
+3-3flang/test/Semantics/OpenMP/to-clause-v45.f90
+3-3flang/test/Semantics/OpenMP/from-clause-v45.f90
+2-2flang/test/Semantics/OpenMP/uses-allocators-version51.f90
+81-456 files not shown
+88-5212 files

LLVM/project 51d022a — flang/include/flang/Parser openmp-utils.h, flang/lib/Semantics check-omp-structure.h check-omp-syntax.cpp

[flang][OpenMP] Account for using different versions for modifier checks

When a modifier from a past/future version is accepted, use its properties
from the nearest version in which it is allowed.
DeltaFile
+118-75flang/lib/Semantics/check-omp-syntax.cpp
+3-1flang/include/flang/Parser/openmp-utils.h
+2-1llvm/include/llvm/Frontend/Directive/Spelling.h
+1-0flang/lib/Semantics/check-omp-structure.h
+124-774 files

LLVM/project c430d0f — flang/include/flang/Parser openmp-utils.h, flang/lib/Parser openmp-utils.cpp

[flang][OpenMP] Create generic interfaces for "container" entities

In short, clauses are containers of modifiers, and directives are
containers of clauses. With the abstractions in place, clauses and
modifiers look almost the same from the point of view of syntactic
properties.

Use these functions to implement generic GetAllowedElements and
GetElementVersionRange functions.
DeltaFile
+17-55flang/lib/Semantics/check-omp-syntax.cpp
+56-0flang/include/flang/Parser/openmp-utils.h
+31-0flang/lib/Parser/openmp-utils.cpp
+104-553 files

LLVM/project a6f6fff — flang/lib/Semantics check-omp-syntax.cpp

[flang][OpenMP] Properly use GetAllowedElements

Allowed elements are not just those explicitly listed, but also those
from the union of allowed sets.

For example, the set of modifiers allowed on a clause are those that
are listed on a given clause, plus the union of all modifier sets that
the clause allows.
DeltaFile
+25-13flang/lib/Semantics/check-omp-syntax.cpp
+25-131 files

LLVM/project 6442969 — flang/lib/Semantics check-omp-syntax.cpp, llvm/include/llvm/Frontend/OpenMP OMPDescriptors.h OMPDescriptors.h.inc

[OpenMP] Introduce descriptors for directives, clause groups and sets

Reuse Association and Category enum definitions from OMP.h.inc.
DeltaFile
+1,110-1llvm/lib/Frontend/OpenMP/OMPDescriptors.inc
+110-1llvm/include/llvm/Frontend/OpenMP/OMPDescriptors.h.inc
+108-0llvm/lib/Frontend/OpenMP/OMPDescriptors.cpp
+63-0llvm/include/llvm/Frontend/OpenMP/OMPDescriptors.h
+1-1flang/lib/Semantics/check-omp-syntax.cpp
+1,392-35 files

LLVM/project 7a62476 — .github/workflows/containers/github-action-ci-tooling Dockerfile

[GitHub] Bump CI Tooling Container to 23.1.3 (#230178)

To pull in 5340f7cc8814b89c01a786e85e1736e10159a9f0, which is impacting
clang-format results.

Part of fixing #230169.
DeltaFile
+1-1.github/workflows/containers/github-action-ci-tooling/Dockerfile
+1-11 files

LLVM/project 87069df — llvm/lib/CodeGen ScheduleDAGInstrs.cpp, llvm/test/CodeGen/AArch64 unanalyzable-store-sequencing.ll

[MISched] Implement `UnanalyzableFrontier`-based DAG construction algorithm (#227375)

Introduces a new algorithm for constructing the control dependencies in
the schedule DAG. The new algorithm is gated behind a (temporary)
`cl::opt` (`-enable-unanalyzable-store-sequencing`), currently off by
default. With the option disabled, the change is an effective NFC: edge
insertion order changes slightly, but the DAGs do not change materially.

The difference between this algorithm and the existing algorithm is
that, rather than maintaining all unanalyzable memory operations and
repeatedly querying all of them, we instead maintain a frontier of loads
and stores since a 'sequencing store'. A sequencing store is either (a)
a store that writes to different set of base objects to the last seen
sequencing store or (b) any store if there has been a load from another
base object since the last sequencing store. We treat the sequencing
stores carefully to ensure that all previously seen unanalyzable memory
operations that do not belong to the frontier necessarily transitively
succeed the current sequencing store. This makes it sound to sequence
preceding memory operations against the sequencing store and only those

    [9 lines not shown]
DeltaFile
+335-0llvm/test/CodeGen/AArch64/unanalyzable-store-sequencing.ll
+199-15llvm/lib/CodeGen/ScheduleDAGInstrs.cpp
+534-152 files