[lldb-dap] Convert test to use the require decorator (#213462)
This is a follow up to https://github.com/llvm/llvm-project/pull/212753
to convert the `lldb-dap` tests to use the `@require` decorators.
[OpenMP] Fix race in setupIndirectCallTable (#215274)
The IndirectCallTable variable where the table is constructed is a local
variable. The copy to the device is right now not synchronized, so the
frame where the table resides can be overwritten before the GPU copy has
accessed the host data which corrupts the device side table.
Fix it by make it a synchronous operation.
[OFFLOAD][OMP] Move OpenMP kernel argument processing from plugins (#213867)
This PR moves part of the OpenMP specific code that is inside the common
code of the plugins that converts the kernel arguments from the OpenMP
ABI to the expected format by the plugins. Also it handles the
additional argument for the KernelLaunch environment.
It also decouples the plugin interface structs from the OpenMP specific
ABI (KernelArgsTy) so changes to this interface do not require changes
to the OpenMP ABI anymore. This also allows to merge what was
KernelArgsTy and LaunchParamsTy into a single struct with all the
information. Because of this there's a number of small changes scattered
through the plugin infrastructure.
There is still some more OpenMP specific code that should be moved out
eventually (and because of this we had to retain some of that
informationt in the new KernelLaunchArgsTy for now) but I didn't want to
complicate this PR further.
Assisted by Claude-Sonnet-5.
[Support] Use block-scope statics for cl::opt registration globals (#215216)
Every static cl::opt constructor references GlobalParser and the
TopLevelSubCommand/AllSubCommands ManagedStatics, namespace-scope
globals in another translation unit. Although these are constant-
initialized in practice, Coverity's GLOBAL_INIT_ORDER checker cannot
prove it and reports one 'Initialization or destruction ordering is
unspecified' finding per option -- 15,356 of the 34,356 outstanding
defects (44.7%) in the LLVM Coverity project.
Move the three registries into block-scope statics behind accessor
functions. Block-scope statics are initialized on first use, so the
cross-TU initialization-order hazard pattern disappears while the
ManagedStatic semantics (lazy construction, destruction via
llvm_shutdown) are preserved unchanged.
Fix AArch64 ISel for unpacked types
Using DUP instructions only really works for packed types, as it will
not introduce the spacing between elements that is required for unpacked
types like nxv2f16.
For those, we'll first broadcast to their corresponding packed type, and
then extract the lo lanes into an unpacked type, effectively introducing
the required spacing.
[MLIR][OpenMP] Support calls added between MarkDeclareTarget runs
Currently, if there are multiple executions of the `MarkDeclareTarget`
pass in a compiler pipeline and somewhere between both runs function
calls get added to a non-declare_target function that was marked as such
implicitly, potential changes to the `device_type` won't get propagated.
This is because we can't distinguish between a user-specified
`declare_target` function attribute and one added by that pass. This
patch addresses this by adding a new parameter to `DeclareTargetAttr`
that is used by that pass to know whether new `declare_target`
information could be propagated to it.
The `automap` and `implicit` parameters are given default values to
simplify the representation of these attributes.
[Flang][OpenMP] Improve implicit declare_target propagation
After starting to run the `MarkDeclareTarget` pass later in the
pipeline, some limitations of its original implementation started to be
hit; specifically, some calls being missed could result in an overly
restrictive marking that would cause the `HostOpFiltering` pass to
remove reachable device code.
This patch aims to address these problems by making the following
changes:
- It makes sure to mark functions in every `RecipeInterface` op pointed
to by OpenMP operations.
- It recursively propagates and combines declare_target information from
target regions and explicitly set declare_target functions to unmarked
functions, but it never modifies explicitly marked functions.
- External and public functions can now only be marked with
`device_type(any)`. Before, marking them as `nohost` or `host` was
possible, but without the ability to see all users we can't give such
guarantees.
[7 lines not shown]
[Flang][OpenMP] Remove the DeleteUnreachableTargets pass
The `DeleteUnreachableTargets` pass was initially added to work around
a problem caused by the interaction between the function filtering and
host op filtering passes and generic MLIR optimizations.
Specifically, by running function and host op filtering before these
optimizations, we would lose the ability to detect unreachable
`omp.target` operations when compiling for the device. This would cause
GPU kernels being created for them, for which no host counterpart
existed.
By now having moved both passes towards the end of the compilation
pipeline, after FIR to LLVM lowering, removal of unreachable host and
device code is no longer impacted by them. This makes the
`DeleteUnreachableTargets` pass redundant.
[Flang][MLIR][OpenMP] Move function filtering to the omp dialect
The `FunctionFilteringPass`, which removes host-only functions when
compiling for an OpenMP target device, is currently defined only for
Flang. However, it implements logic that would generally be useful for
other frontends that can generate OpenMP offloading code. This patch
makes that transition by implementing the following changes:
- It rewrites the `FunctionFilteringPass` to work on lower level non-FIR
MLIR modules.
- It splits off preexisting logic to check for not-yet-implemented target
device features during function filtering to its own pass and it improves
context detection to properly diagnostic only device code.
- It delays function filtering to run at the end of the pipeline, as well
as the `MarkDeclareTargetPass`. This will enable the latter to run
only once in the pipeline after the `UnimplementedDeviceCheckPass`
becomes no longer necessary.
One side effect of these changes is that delaying the filtering will
cause all following passes to process host functions that are eventually
[4 lines not shown]
[AIX][SystemZ][Support] Check if file is dir on open instead of read (#214815)
See
https://github.ibm.com/compiler/llvm-project/commit/678f19f08296fec299438130cf5943714c590b7e
for the original change.
This original change would run fstat() on the file at every read(). In
the non-error situation that is a lot of redundant checking. Moving the
fstat() check to openNativeFileForRead() will reduce the checks to a
minimum and still produce the same error if someone tries to open a
directory.
workflows/release-task: Stop uploading lit to test.pypi.org (#214979)
The gh-action-pypi-publish action only supports being run once per job.
Running it twice results in the second upload always failing. Rather
than trying to create a complicated job structure to support uploading
to test.pypi.org and pypi.org, we just remove the test.pypi.org upload
for now.
release-tasks: Disable lit publishing for release candidates (#214972)
There is no rc in the lit version string, so release candidates get
published using the non-rc version number.
[mlir][acc] Fold present() clauses on device values (#212815)
The compiler must emit acc.device_ptr mapping for device values,
however, an existing present clause prevents that. A present on a device
value always holds, so fold it away to allow implicit data handling to
generate device_ptr mapping.
[Hexagon] Fix unusable SCS reg, make it selectable (#213820)
SCS hardcoded r19 as the shadow call stack pointer and required
-ffixed-r19. That was the wrong register to pick: r19 is precisely the
one the intended consumers cannot give up, so the feature was unusable
in practice.
* The Hexagon Linux kernel already reserves r19 for its thread-info
pointer (arch/hexagon/Makefile: "TIR_NAME := r19", documented there as
not configurable because it is hard-coded in several files).
* hexagon-hypervisor reserves r20-r28 (kernel/CMakeLists.txt), with r28
bound to a register global (H2K_gp).
That leaves h2 only r16-r19, so no single hardcoded choice can serve
both consumers.
Intersecting that with the callee-saved regs leaves r1{6,7,8}. So the
new default is r18.
[5 lines not shown]
AMDGPU/GlobalISel: RegBankLegalize rules for uniform i16 extending loads
Extending loads, i8 to i16, are legal on targets with true16.
Note: there is potentially a missing rule for uniform P4 when target
usesTrue16 and hasSMRDSmall but MMO does not satisfy isUL. Wasn't able
to construct an LLVM-IR test for this case, leaving it unsupported.
[libc++] Add tools for gathering historical benchmark data (#212775)
Benchmarking every commit of libc++ is prohibitively expensive: a single
run of the benchmark suite takes hours, and the data has to be
regenerated from scratch whenever the compiler, the OS or the benchmark
machines change. These tools instead sample the history at a coarse
granularity and drive libcxx-benchmark-commit.yml to fill in what is
missing.
Three tools cooperate, meant to be run periodically:
select-anchor-commits picks one commit per calendar bucket from Git
plan-benchmarks diffs that against what LNT already holds
dispatch-benchmarks requests the corresponding workflow runs
They keep no state of their own. They recompute the current and target
states from LNT and the GitHub Actions API, which allows running them in
a CRON. The dispatching of workflows is done using a budget, to avoid
launching tens of jobs and competing with other uses of the CI
[3 lines not shown]
[analyzer] Fix -analyzer-output=html assert on reversed and macro ranges
HTMLDiagnostics::HighlightRange guarded against a reversed range by
comparing line numbers, so a same-line reversal - which is what the piece for
an implicit copy constructor carries - reached html::HighlightRange.
Its scan walks from begin to end, ran off the end of the buffer, and asserted:
https://godbolt.org/z/sTb5qfjjd
Invalid position to insert! (RewriteRope.h)
It also added the end token's length itself and then passed a token range to
html::HighlightRange, which measured the token again, this time from the
interior. For most tokens the two cancel, but where the tail re-lexes longer
the highlight reached past the end of the range, e.g. over a trailing ';'.
Use getExpansionRangeInFile(), which rejects reversed and cross-file ranges,
then convert once and tell html::HighlightRange the range is already
char-granular.
[4 lines not shown]
[analyzer] Fix -analyzer-output=sarif crash on macro-expanded ranges
A path piece whose range ends inside a macro expansion aborted the whole
document: https://godbolt.org/z/61vWYcsWj
Cannot create a physicalLocation from invalid SourceRange!
convertTokenRangeToCharRange() built the end with
Lexer::getLocForEndOfToken(), which returns an invalid location for a macro
ID that is not at the end of its expansion, and used it unchecked. The
analyzer's own test corpus hits this in nine files; text and plist output
were unaffected because both already map such ranges to the expansion.
- Use getExpansionRangeInFile(), so the region covers the macro use like the
other two outputs.
- Fall back to a caret when the range is unusable. A thread flow needs a
location per piece, so dropping one would truncate the reported path. This
also stops reversed ranges producing regions with endColumn < startColumn.
[4 lines not shown]
[clang] Reject ranges getExpansionRangeInFile cannot represent
getExpansionRangeInFile was extracted verbatim and inherited two shortcomings
of the original loop, fixed here before the analyzer's SARIF and HTML consumers
depend on it:
- It mapped the end with getExpansionRange(SourceLocation), which always
reports a token range, so a char-range input was widened by a whole token.
Now using the getExpansionRange(CharSourceRange) overload, which keeps the flag.
- It passed reversed ranges through. Consumers walk begin->end; now returning
nullopt for those, as Lexer::makeFileCharRange already does.
Separate from the extraction so that stays NFC, and out of the consumer fixes
because it changes the shared helper's contract rather than one output.
Both contract changes, plus the invalid- and cross-file-range guards, are
covered by a GetExpansionRangeInFile unit test in
clang/unittests/Frontend/TextDiagnosticTest.cpp.
Assisted-By: claude
[clang][NFC] Extract getExpansionRangeInFile out of the diagnostic renderers (#214460)
Prep for the following commits, which fix crashes in the analyzer's
SARIF and HTML output on ranges that end inside a macro expansion.
Fixing them means mapping such a range into the reported file - the
normalization the frontend text and SARIF renderers already do, and that
the two analyzer consumers each do differently and incorrectly.
Hoist that logic into getExpansionRangeInFile, beside the DiagnosticRenderer
base both frontend renderers derive from, so the fixes reuse one
implementation instead of adding two more copies. TextDiagnostic and
SARIFDiagnostic move onto it here with no behavior change; the analyzer
consumers follow in later commits.
getFileID() replaces SARIFDiagnostic's getDecomposedLoc(...).first - equivalent
here, and what TextDiagnostic has used since c113cbb51005.
Assisted-By: claude
[InstCombine] Fold uadd.sat(X, C) - C to umin(X, ~C) (#215130)
`uadd.sat(X, C) - C --> umin(X, ~C)` for nonzero `C`.
The saturating add gives `X + C` or `UMAX`, so subtracting `C` leaves
`X` or
`UMAX - C`, which is the unsigned minimum. `UMAX - C == ~C`, so the
constant
is just the inverted `C`.
There was already a test documenting this miss in saturating-add-sub.ll
(`test_scalar_uadd_sub_const`) - it folds now.
https://alive2.llvm.org/ce/z/kLFWy7
Fixes #215103
[libc++] Simplify detection of win32-broken-utf8-wchar-ctype (#214797)
Instead of querying `_LIBCPP_HAS_LOCALIZATION` from Python, do it from
the source program. This fixes a bug where if `_LIBCPP_HAS_LOCALIZATION`
was not defined at all (which is the case for older versions of libc++),
the feature would then be defined immediately, regardless of the
platform we're on. That's because `and` has higher precedence than `or`
in Python, so we'd end up skipping the `_WIN32` check entirely.
[lldb][test] Remove Python <= 3.6 workaround (#215262)
re.Pattern was added in 3.7 and our minimum
is now 3.8.
Python 3.6.15:
>>> import re
>>> re.Pattern
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
AttributeError: module 're' has no attribute 'Pattern'
Python 3.7.17:
>>> import re
>>> re.Pattern
<class 're.Pattern'>
Python 3.8.20:
>>> import re
[3 lines not shown]