[SandboxVectorizer] Dispatch LoadStoreVec::runOnRegion on seed kind
Rebuilt on top of the direction-agnostic packOperands() redesign and
the fresh-Scheduler-per-sub-run fix (previously this same feature
existed on the now-superseded vectorize-loads branch, stacked on the
classifyStoreOperands-gated version of partition-store-bundles).
runOnRegion() previously assumed its seed slice was always a store
chain, unconditionally casting Bndl[0] to StoreInst. This crashed
(assertion in areConsecutive<StoreInst>) whenever -sbvec-collect-seeds
included "loads", since a load-seeded region's Aux holds LoadInsts.
Add createVectorLoad(), a standalone builder used only by the new
vectorizeLoads() (packOperands() replaced its old role of also being
vectorizeStores()'s all-loads fast path -- that path no longer exists,
per the direction-agnostic redesign).
Add a symmetric top-level path for load-kind seed slices:
- vectorizeLoads() builds the vector load via createVectorLoad(), then
[39 lines not shown]
[SandboxVectorizer] Vectorize partial store sub-bundles in LoadStoreVec
Rebuilt on top of the direction-agnostic packOperands() redesign
(previously this same feature existed on the now-superseded
partition-store-bundles branch, gated by classifyStoreOperands).
runOnRegion() previously required the entire seed chain to vectorize as
one unit. A seed slice can legitimately fail that as a whole while a
sub-run within it is still fine, e.g. because it spans an address gap
(SeedBundle::getSlice sorts by address but doesn't guarantee
contiguity) or a sub-range fails to schedule. Add
findLegalStoreRun()/isLegalStoreRun() to search for the longest
vectorizable run starting at a given position, and have runOnRegion()
call vectorizeStores() once per such run instead of once for the whole
chain. isLegalStoreRun() is purely an address/scheduling check now --
no operand-eligibility check is needed since packOperands() accepts any
operand kind.
The search only ever shrinks a candidate length, never grows one:
[35 lines not shown]
[SandboxVectorizer] Hoist getInsertPointAfterInstrs into VecUtils
Move BottomUpVec.cpp's file-local getInsertPointAfterInstrs() into
VecUtils, next to the getLowest()/getLastPHIOrSelf() primitives it's
built from. It has no BottomUpVec-specific state; the next commit adds a
second caller in LoadStoreVec.
Not hoisting BottomUpVec::createPack() itself here: it asserts a single
common scalar type (VecUtils::getCommonScalarType), which doesn't fit
LoadStoreVec's mixed-type ("enable-diff-types") requirement. That needs
its own extended packer, kept local to LoadStoreVec.cpp rather than
force-fitting the shared version.
No functional change: check-llvm Transforms/SandboxVectorizer passes
(28/28).
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
[SandboxVectorizer] Make LoadStoreVec::vectorizeStores direction-agnostic
Remove classifyStoreOperands()/isFoldableLoadOperand(): vectorizeStores()
no longer gates on whether a store chain's value operands are all loads,
all constants, or neither. Instead it always builds the vector value via
a new packOperands(), which packs any mix of loads, constants, or
arbitrary SSA values via extractelement/insertelement -- direction-
agnostic in the sense that it doesn't care what kind of operand it's
given, unlike the load-specific and constant-specific paths it replaces.
packOperands() combines operands at the granularity of their narrowest
common scalar element type (the same rule getCombinedVectorTypeFor()
uses), splitting a wider operand into multiple lanes via a bitcast. A
plain bitcast can't convert between pointer and non-pointer types, and
inttoptr requires an integer source, so reinterpretSameWidth() picks
bitcast, ptrtoint, or inttoptr as needed, routing a non-integer,
non-pointer operand (e.g. double) through an intermediate same-width
integer when the target granularity is a pointer.
[38 lines not shown]
[lldb][RISCV] Construct CSR information dynamically (#203234)
Custom RISC-V extensions may define control and status registers (CSRs)
with overlapping addresses. Therefore, when performing postmortem debug
of 32-bit RISC-V core dump images, dynamically construct CSR information
based on the set of enabled extensions.
Assisted-by: OpenAI GPT-5
[HLSL] Move `lerp` implementation to header files (#215692)
Closes #213097.
This PR replaces the previous implementation of `lerp` with a new one
inside the header files. It also cleans up the tests to use 3 distinct
parameters (x, y, s) for consistency with other similar tests.
The SPIRV intrinsic (`int_spv_lerp`) and its lowering are intentionally
kept, since a follow-up will pattern match `X + S * (Y - X)` back to the
extended instruction and needs the SPIRV intrinsic to do so.
Assisted-by: Claude Opus 4.8
[SandboxVectorizer] Extract LoadStoreVec::vectorizeStores from runOnRegion
Move runOnRegion()'s body -- the store-chain legality checks, operand
classification, vector value construction, and profitability decision
-- into a new vectorizeStores(Bndl, Rgn, Sched, A) method. runOnRegion()
now only builds the initial bundle from the region's Aux and calls
vectorizeStores() once. Also extract the cost-check tail into
acceptIfProfitable(), a small helper worth having on its own. NFC.
[HLSL] Move `radians` implementation to header files (#215379)
Closes #213095.
This PR replaces the previous implementation of `radians` with a new one
inside the header files.
The SPIRV intrinsic (`int_spv_radians`) and its lowering are
intentionally kept, since a follow-up will pattern-match `Val *
(pi/180)` back to the extended instruction and needs the SPIRV intrinsic
to do so.
Assisted-by: Claude Opus 4.8
DAG: Gracefully diagnose missing lrint/lround float-operand libcalls (#215064)
I believe this is my 9000th commit, merged during the solar eclipse
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
[AMDGPU][GlobalISel] Fix wide multiply with known-zero parts
Materialize a zero accumulator when all partial products for a destination part are skipped, avoiding an invalid null register during legalization.
Change-Id: I05294cf68ddd4363ccb20df7159981280e7c9c4e
Exclude more libc math C++ builtins from compiler-rt Bazel overlay (#214779)
Building on work in https://github.com/llvm/llvm-project/pull/208861,
adds recently-added additional libc math C++ builtins from the
compiler-rt Bazel overlay.
[flang][OpenMP] Support the FULL clause on the UNROLL construct (#214115)
`!$omp unroll full` is parsed and then aborts in lowering with
`not yet implemented: Unhandled clause FULL in UNROLL construct`. This
implements it. Fixes #214114.
`partial` landed in #206642 and bare `unroll` in #144785, so `full` is
the remaining clause on the
construct. clang already supports `#pragma omp unroll full`, and
`OpenMPIRBuilder::unrollLoopFull` already exists, so the work is the
MLIR op, its translation, and
the flang lowering.
### MLIR
Adds `omp.unroll_full`, mirroring `omp.unroll_heuristic`: one applyee,
no generatee, since no loop
remains after full unrolling. The name is the one already used as an
example in `CanonicalLoopOp`'s
[66 lines not shown]
Verify 2-D iterator dimension associations
Bind each two-dimensional iterator range and trace its induction
variable to the corresponding map bound and array extent in target update
and declare mapper lowering tests.
This makes range, induction-variable, bound, and dimension swaps observable.
[MIR] Quote basic block names that are not plain identifiers (#214054)
`MachineBasicBlock::printName` writes the IR basic-block name after
`bb.<N>.` unquoted, and the
MIR lexer reads it back with `isIdentifierChar` only. If the name
contains anything else -- a
comma, say -- the name is cut short on the way in and the remainder is
treated as syntax, so `llc`
cannot parse the MIR that `llc` just wrote:
```console
$ llc -mcpu=gfx90a -stop-before=greedy bb-comma-name.ll -o out.mir
$ llc -mcpu=gfx90a -x mir -start-before=greedy out.mir -o /dev/null
error: out.mir:170:8: expected ':'
```
because the printer emitted
```
[46 lines not shown]
[AssumptionCache] Remove incorrect assertion from `removeAffectedValues()` (#214524)
The assertion in `removeAffectedValues()` is trying to enforce that,
when we come to remove the affected values of a live assume call, each
affected value has a corresponding line in the cache. This may not be
the case if the assume call has been modified via a call to
`Use::set()`, as our value handles are only notified on deletion and
RAUW. We're not worried about this, though, as the cache is
conservative.
Remove this assertion and add a comment explaining why we may fail to
find.
Handle unsupported iterator locators without crashing
Analyze iterator map locators without using getBaseObject, whose coarray
case is intentionally unreachable. Share the analysis with map-info
generation so rootedness and lowering agree on supported shapes.
Preserve specific diagnostics for external coarrays and array-element
members, with regressions for both cases.
[libc++] Collapse `optional<T&>` inheritance hierarchy (#215286)
Resolves #215185
- Since `optional<T>` and `optional<T&>` are decoupled, there wasn't
much reason to have a base class for `optional<T&>`, so we can simplify
it by inlining all of the base members and private functions into it
directly.
- This leaves us with the iterator base which now only exposes the
`iterator` type, since we can also directly inline `begin()` and
`end()`.
---------
Co-authored-by: Nikolas Klauser <nikolasklauser at berlin.de>
[Clang] Include libc wrappers for LLVM environment CUDA / HIP (#208084)
Summary:
These wrappers (although they are empty right now) are used to inform
the offloading runtime of the supported libc / libm functions available
on the device. We do not include these for the standard HIP path because
they conflict with the alreaedy present utilities, but with the
LLVM-only route we should be able to use them.
[modulemap] Exclude the z/OS string.h wrapper from LLVM_Utils (#215800)
8fce476c8122 (https://github.com/llvm/llvm-project/pull/167703) replaced
`llvm/Support/SystemZ/zOSSupport.h` with
`llvm/Support/SystemZ/zos_wrappers/string.h`, a wrapper that pulls in
the system header via `#include_next` and then redeclares `strsignal`
and `strnlen` with `asm` labels. It is only meant to be reachable on
z/OS, and `llvm/CMakeLists.txt` adds the directory to the include path
solely when `CMAKE_SYSTEM_NAME` matches OS390.
The header was never excluded from the module map, though.
`LLVM_Utils.Support` is an umbrella over `llvm/Support`, so building
that module textually includes the wrapper on every host. This surfaced
building the Swift compiler on Windows, where `<string.h>` resolves to
the UCRT header that already declares `strnlen` as `_ACRTIMP`, i.e.
`__declspec(dllimport)`:
```
error: cannot apply asm label to function after its first use
[11 lines not shown]
[MLIR][ROCm] Export runtime wrappers on Windows (#213046)
Export the ROCm runtime wrapper entry points when building
`mlir_rocm_runtime` as a Windows DLL.
Use `__attribute__((visibility("default")))` on other platforms.
## Motivation
`extern "C"` prevents C++ name mangling, but it does not add symbols to
a
Windows DLL export table. Consequently, `mlir-runner` can load
`mlir_rocm_runtime.dll`, but ORC cannot resolve its `mgpu*` entry
points.
CUDA, SYCL, Vulkan, and SPIR-V runtime wrappers already use explicit
Windows
export annotations. This applies the same approach to the ROCm runtime.
[24 lines not shown]
Revert "[HLSL] Generate semantic signature metadata" (#215844)
Reverts llvm/llvm-project#212892
Build dependency for `DXILResource.h` was not updated. I will reland
with the corrected dependency.
[libc++] Guard container benchmarks on library availability (#215408)
Instead of using TEST_STD_VER, use FTMs or requires clauses to enable
benchmarks for some recent features like `append_range`. This is needed
since older versions of the library don't provide these features, so the
container benchmarks as a whole would fail to compile instead of just a
few methods being disabled.