[SCEV] Rename SCEVNoWrapFlags to SCEVFlags (NFC) (#225179)
In preparation to extend the flags that ScalarEvolution can represent,
rename SCEVNoWrapFlags to the more general SCEVFlags. In particular, we
preserve the (get|set)NoWrapFlags interface, which simply masks the
no-wrap flags from the newly-introduced (get|set)Flags.
See also: #225065.
[libc++] Fix "leak" of PMR resource in benchmark (#216114)
The benchmark was not releasing the monotonic buffer resource, which
means we'd grow a larger and larger pool as Google Benchmark would go
through more and more iterations of the benchmark.
[Offload] Lazily initialize platforms in the Offloading API (#227084)
Summary:
The Offloading library wraps around the underlying plugins. The problem
is that we currently initialize all plugins we find, even if they are
not needed for the program. This is very expensive for trivial uses, as
fully heterogenous usage is quite rare. In practice this means that you
will always pay a 200 ms penalty for having CUDA installed.
This patch changes the behavior to provide accessors into the plugins
and devices that allows them to be initialized lazily. We use a
once_flag, this should properly take a fast-path check while still
blocking on concurrent use.
Making full use of this will require a way to filter platforms more
specifically. I'm thinking of what this would look like as an API.
I'm thinking that we either have an extra iterate function that takes a
callback on the platform, or we just provide a helper to find all the
devices that can run a given image. Maybe both?
[2 lines not shown]
[Offload] Try loading dynamic libraries at global scope first (#227020)
Summary:
The intention of the dynamic path is to provide basic linking in cases
where the SDK is not avaialble at the user's build time. Previously this
did a distinct `dlopen` call on the library.
However, in cases where an existing HSA implementation exists at global
scope this can cause multiple instances to be open at the same time.
Specifically this is problematic if mixing with another implementation
(Like HIP or CUDA) or trying to intercept functions (like compiler-rt).
The compiler-rt interceptors currently carry interceptors into `dlsym`
when we should instead just load these at global scope, this makes them
behave more similar to a direct link in these scenarios, which was the
intention.
The fallback remains as the current behavior, should only trigger in
cases where there is an existing HSA in the process.
[flang-rt] Install the runtime libraries under the flang-rt component (#227073)
Summary:
The Flang-RT libraries were installed without a component, so they were
placed in `Unspecified` and there was no `install-flang-rt` target. The
runtimes build forwards `install-flang-rt` for every enabled runtime, so
requesting flang-rt in a distribution failed on the missing target, and
component installs skipped the libraries. Add the component and its
install targets, matching the existing `flang-rt-headers` and
`flang-rt-mod` components.
[InlineSpiller][AMDGPU] Implement subreg reload during RA spill
Currently, when a virtual register is partially used, the
entire tuple is restored from the spilled location, even if
only a subset of its sub-registers is needed. This patch
introduces support for partial reloads by analyzing actual
register usage and restoring only the required sub-registers.
This improvement enhances register allocation efficiency,
particularly for cases involving tuple virtual registers.
For AMDGPU, this change brings considerable improvements
in workloads that involve matrix operations, large vectors,
and complex control flows.
[Offload] Fix installing offload through its install components (#227036)
Summary:
Distribution builds install each runtime through its component's install
target rather than the global `install` target. Several pieces of
offload
did not work with this:
- `install-offload` only depended on `omptarget` and `LLVMOffload`, but
the component also installs `LLVMOffloadKernel`. The install failed
unless something else had already built it.
- The CUDA, HIP, and kernel language headers were installed without a
component, so component installs skipped them.
- The offload tools used `add_llvm_tool` which did not build and/or
install. We already removed `add_llvm_library` from the libraries. Now
these use standard `add_execulable`.
devel/libftdi: resurrect and fix broken tree
There are several ports still depend on this one.
Fixes: b1f10ca6b2ffbdad2ba2bd309584941063fcb7b3
Reported by: many
[compiler-rt] Give every installed file an install component (#227103)
Summary:
LLVM tries to expose install components to allow distributions to choose
the layers to build. Not all of the compiler-rt artifacts were installed
under this configuration. This PR makes this a consistent rule by
updating the existing helper function to require a parent target for the
component.
Adds things like
```
ninja install-dfsan
```
and ensures this installs everything
```
ninja install-compiler-rt-x86_64-unknown-linux-gnu
```
[runtimes] Disable sanitizers for configuration checks with --unwindlib=none (#227092)
Summary:
in a sanitized runtimes build, every configuration check is linked with
both -fsanitize=... and --unwindlib=none and fails, because the
sanitizer runtimes need the unwinder. This completes an existing TODO by
suppressing this for the compiler flag checks in this configuration.
This shouldn't affect any existing users. Motivation is compiling
`+asan` multilibs for existing OpenMP/Offloading configurations.
[OpenMP] Install the Fortran module files with libomp-mod (#227116)
Summary:
The `libomp-mod` install component only contains `omp_lib.h`. The module
files are only installed incidentally as part of the Fortran runtime or
directory copy. Just do this directly. These were intentinoally dropped
in https://github.com/llvm/llvm-project/pull/171515, in favor of the
directory copy. However, this means these are ignored when done through
the specific component target.
Shouldn't affect plain installs.
[clang][AST] Make Result parameter of `isCXX11ConstantExpr` mandatory (#226493)
We want to discourage people from just checking if something is a
constant expression without using the value, which we always compute
anyways.
There are only three call sites, all in clang, and all pass a value
already. Change the parameter type to a reference to enforce this.
[TableGen] Reuse `UnitMaskIdx` lambda in `computeRegUnitLaneMasks` (NFC) (#227249)
This simplifies the code at the expense of losing an assert that checks
that one `SUI` matches more than one `RU`s.
[flang-rt][Offload] Fix builds using the new amdgpu triple spelling (#227110)
Summary:
The offload cache files, such as `FlangOffload.cmake`, build the AMDGPU
runtimes for `amdgpu-amd-amdhsa`, but flang-rt and offload only
recognize AMDGPU triples beginning with `amdgcn`. Just check both like
we do elsewhere. Eventually this can all be dropped.
[LLVM] Infer compression format from zlib and zstd headers (#222773)
Summary:
This PR adds two routines that allow users to decompress bitstreams
without explicitly specifying the format. This is inferred from the
bytes
of the object, returning an error if it could not be identified. The
Zstd
standard exposes proper magic bytes, while Zlib needs to be inferred
through the two header bytes.
Because this is a raw bytestream, there is a chance that someone
could get bytes that pass this check, giving no 'reason', but since
the decompression routine requires the expected size there is no
chance it will falsely succeed.
[InstCombine] Use context in select-to-umin nonzero check (#225381)
foldSelectICmpMinMax() queries whether W is nonzero without an
instruction context, missing facts from a preceding llvm.assume or
dominating branch. Querying at the comparison enables the existing
select-to-umin fold.
Fix #225345
TwoAddressInstructions: Use a range loop to move the copy chain
Cleanup some messy iterator work.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
[libsycl] Add context to exception's throw. (#224668)
sycl::exception class can contain associated context (see 4.13.2.
Exception class interface).
Specification doesn't declare rules about when context must be
associated and when not.
Passing context to exception ctor in all places where context is
available.
Assisted-by: Claude Code.
---------
Signed-off-by: Tikhomirova, Kseniya <kseniya.tikhomirova at intel.com>
X86: Respect the exception model module flag in X86LFIRewritePass
This is preparation for removing the TargetOptions ExceptionModel field.
Currently this doesn't show an observable behavior change because the
codegen pass pipeline is driven by this field. Add the module flag based
check so in the future, if the pass runs on a module not using sjlj, it
will skip the sjlj specific handling.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
[Github] Temporary fix for #226230 (#227144)
It seems that with actions/checkout we run into a 5 minute timeout
somewhere (presumably in the websocket that communicates between the
runner and workflow pods), which causes actions/checkout to succeed but
without actually writing anything out, causing the rest of the workflow
to fail. This patch fixes that by forcing verbose output, which should
ensure there is no 5 minute period without any output.
This issue doesn't seem to impact other parts of the workflow due to
progress indicators.
[ARM] Invalidate LiveRegPos when erasing current position (#226937)
The load store optimizer could erase an instruction that is currently
holding the position of the LiveRegPos. Make sure we recalculate
LiveRegs if this happens.
Fixes #223630
TwoAddressInstructions: Move the rescheduled copy chain back to front
rescheduleMIBelowKill sinks an instruction below the kill of its tied
source, along with the run of copies that follows it. With LiveIntervals
the copies are spliced one at a time so handleMove sees a well-formed
block, but they were visited front to back and each inserted before the
previously moved one. This reversed them, and transiently moved a copy
below its use, asserting in handleMoveDown.
Walk them back to front instead, which also preserves the original order.
Exposed by #225174, which made LiveIntervals available here by default.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>