[flang-rt] Add headers as a dependency to the runtime (#228210)
Summary:
The flang-rt runtime implicitly depends on the flang compiler's header
gen. We need to depend on the target for the component.
Stray space change noise Leftover dead code
Stray space change noise Leftover dead code
flangast experiments
clang-format
Stray space change noise
Leftover dead code
clang-format
Post-merge fixes
clang-format
[OpenMP][offload] Fix single-thread SPMD cross-team reductions
If a kernel gets SPMDized by OpenMPOpt but does not have a parallel
region, it actually is only executed by a single thread.
forceSingleThreadPerWorkgroupHelper() is responsible for that and had
been introduced by https://reviews.llvm.org/D133120. The cross-team
reduction trusted that SPMD mode means that we're dealing with all
threads. Now, it checks explicitly.
Claude assisted with this patch.
[BoundsSafety][test] Add late-parsed counted_by type-attribute coverage
New tests exercising the late-parse fill-in mechanism:
- Sema/attr-counted-by-weird-type-positions{,-late-parsed}.c: counted_by
in assorted type positions, nested pointers, and rejection cases.
- Sema/attr-bounds-safety-function-ptr-param.c: attributes on
function-pointer-typed members.
- Modules/ and PCH/ bounds-safety-attributed-type-late-parsed: the
resolved type round-trips through serialization.
- Sema/attr-counted-by-late-parsed-regressions.c: guards against the
double-free on a nested-record decl-spec attribute and the null-count
escape on a free-function parameter.
[BoundsSafety] Create incomplete counted_by types and wire up the refill
Activate late parsing for the counted_by family (counted_by / sized_by and
their _or_null variants) in type-attribute position, under
-fexperimental-late-parse-attributes, on top of the type-attribute handling,
the validation helper and the refill machinery added in the previous commits.
When such an attribute is seen during type construction and its argument
can't be resolved yet, build the CountAttributedType immediately with
getIncompleteCountAttributedType and record it against the enclosing record;
its count expression is filled in at the closing brace via the refill logic.
Because enclosing types refer to the node by pointer, completing it in place
leaves the type chain untouched -- no rebuild, no TypeLoc re-emission.
- Sema::ActOnLateParsedTypeAttr builds the incomplete node; the parser
callback stores it on the LateParsedTypeAttribute and records the
attribute in the record currently being parsed.
Parser::CompleteLateParsedTypeAttributes drains that list at the closing
brace; a nested anonymous record hands its pending attributes up to the
[9 lines not shown]
[BoundsSafety] Handle the counted_by family as a type attribute
counted_by / sized_by (and their _or_null variants) were handled in only one
way: a declaration-position attribute went through handleCountedByAttrField,
which validated it and then patched the field afterwards with
FieldDecl::setType. There was no type-position handling at all.
Build the type during type construction instead, from a single handler that
serves both positions:
- Add HandleCountedByAttrOnType and dispatch the counted_by family to it
from processTypeAttrs, going through the shared
validateBoundsAttrTypeForTypePosition leaf.
- Remove handleCountedByAttrField. Its FieldDecl-based type-shape checks in
Sema::CheckCountedByAttrOnField are superseded by
Sema::ValidateBoundsAttrTypeShape, added in the previous commit and now
reached from the type path, and are deleted; no diagnostic is dropped. The
checks that genuinely need the FieldDecl (union member, non-flexible
array, cross-struct count) stay in CheckCountedByAttrOnField and run from
[11 lines not shown]
[BoundsSafety][NFC] Thread a late-parsed attribute list through declarators
A late-parsed type attribute is written in the middle of a declarator, so the
list it lands in has to travel with the declarator pieces until the enclosing
record can supply its argument. Add that storage and plumbing, with nothing
producing or consuming it yet:
- Move CachedTokens and LateParsedAttrList earlier in DeclSpec.h so DeclSpec,
Declarator and DeclaratorChunk can hold one.
- Give DeclSpec, Declarator and DeclaratorChunk a LateParsedAttrList, and let
Declarator::AddTypeInfo carry one onto the chunk it appends.
- Give ParseSpecifierQualifierList and ParseTypeQualifierListOpt an optional
LateParsedAttrList parameter, passed down to ParseDeclarationSpecifiers.
No functional change: the lists stay empty and no caller passes one. The next
commit populates them.
[BoundsSafety][NFC] Add the Sema/Parser bridge for late-parsed type attributes
A late-parsed bounds attribute has to build its type when the attribute is
seen, but its argument isn't parseable until the enclosing record is complete.
Building that type needs the Parser (which owns the cached tokens) and Sema
(which owns type construction) to meet:
- Sema::ActOnLateParsedTypeAttr validates a counted_by-family attribute for
the type position and, if valid, wraps the type in a CountAttributedType
whose count is not yet known, handing the node back for completion.
- Parser::ProcessLateParsedTypeAttrCallback is the Parser-side entry point,
registered on Sema so Sema can call back without including Parser.h (the
same pattern as LateTemplateParserCallback). It reuses an already-built
node so several declarators sharing one attribute share one type.
No functional change: nothing records late-parsed type attributes yet, so the
callback is never invoked. The next commit wires it up.
[offload][omp] Move reading _kernel_environment to libomptarget (#222606)
The xxxx__kernel_environment are only generated for OpenMP kernels. Move
reading them to libomptarget. We still pass the information needed for
launching kernels to the plugins.
Assisted by Claude.
Add Vassil's responsibilities to the CODEOWNERS file. (#228167)
I'd like to get more reliable github-based routing when pull requests
arrive in the areas I am maintaining.
For reference, see clang/Maintainers.md
[SystemZ][z/OS] Fix GOFF exception table (LSDA) section generation (#228097)
Fixes #226804
The z/OS Binder rejects LSDA exception tables with `IEW2353E ... ERROR
CODE IS 25000E` when emitted as renamable PRs under an initial-load
`C_WSA64` ED parented to the root code section.
Following the fix suggested by @mms-it-ch in #226804, emit the LSDA
similarly to static WSA data:
- Under its own Section Definition (`SD`) named `GCC_except.<func>` with
`ESD_BSC_Section`.
- With a `C_WSA64` Element Definition (`ED`) using deferred load
(`GOFF::ESD_LB_Deferred`) and doubleword alignment.
- As a non-renamable Part Reference (`PR`).
Co-authored-by: Yusra Syeda <yusra.syeda at ibm.com>
[AMDGPU] Validate scale_sel in v_cvt_scale_*
These instructions can be block16 or block32 depending on the target
and scale_sel bits. Block16 is not supported in strict mode.
Re-enable the rest of the instructions in the strict mode but validate
the scale selector.
[compiler-rt] Fix -shared-libsan usage with internal symbolizer (#227717)
Summary:
https://github.com/llvm/llvm-project/pull/226551 exposed a preexisting
issue when using `-shared-libsan` nad the internal symbolizer builds.
The symbolizer will try to look up the symbol via `RTLD_NEXT`, which
goes in load order. If the sanitizer library is loaded *after* `libc.so`
then this `dlsym` call will miss it.
The solution is to have a fallback that checks using `RTLD_DEFAULT`.
This
should allow us to find the symbol in these exceptional cases. I don't
expect much fallout as this covers a case that currently returns `NULL`.
Loading it in means it could grab a reference ahead of the runtime that
was intended to be intercepted, but for this use-case I don't think this
will apply.
[clang] Fix crash constant-evaluating the construction of huge arrays (#226899)
Fixes #173728
The constant evaluator runs on ordinary code all the time (range checks
for `-W` warnings, `isEvaluatable` in codegen), and it
default-constructs an array by building one `APValue` per element. The
element count got truncated to `unsigned` first and nothing checked its
size, so a local like `struct T {} s[0xFFFFFFFF][0]` ran it out of
memory; with the issue's reproducer, where the count wraps to 2^64 - 4,
an assertions build hits the "bounds check failed" assertion in
`adjustIndex` first. Copying such an array, e.g. into a lambda capture,
had the same problem. This goes back to at least Clang 3.4. The bytecode
interpreter has its own version: it emits code for every element, so a
constructor with a member like `T a[0xFFFFFFFF]` runs it out of memory
too.
Array default construction and `ArrayInitLoopExpr` now go through the
existing `CheckArraySize` guard, the same one `new` already uses, so
[6 lines not shown]
[LLDB] Acquire the module mutex at the start of SetLoadAddress (#227149)
Fixes a potential dead-lock from parallel module loading.
I received a quick-stack of LLDB hung loading a core with parallel
module loading enabled, where two threads were trying to mutate a given
module and an object file, but having acquired the module mutex first in
one case, and the object file's section mutex first in the second case,
which each trying to subsequently acquire the other lock.
In the update case, [SetLoadAddress acquires the section list mutex and
then tries to acquire the module
mutex](https://github.com/llvm/llvm-project/blob/60f717946cb5ca911b6be22169b9bd646225c59f/lldb/source/Symbol/ObjectFile.cpp#L614)
```
SectionList *ObjectFile::GetSectionList(bool update_module_section_list) {
std::lock_guard<std::recursive_mutex> guard(m_sections_mutex);
if (m_sections_up)
return m_sections_up.get();
[26 lines not shown]
[clang][OpenMP] Don't use fused dist schedule for teams loop emitted as distribute (#228129)
Fix teams loop reductions lowered as 'distribute' lose their loop.
Claude assisted with this patch.
[AMDGPU] Canonicalize num_records to its actual width in InstCombine
llvm.amdgcn.make.buffer.rsrc is overloaded on the type of its
num_records argument, but the hardware field it ends up in has a fixed
width (32 bits, or 45 bits on gfx1250 and up). Rewrite the intrinsic to
use that width, zero-extending or truncating num_records as needed, so
that IR-level optimizations can see that the extra bits of, for example,
the i64 that Clang emits are not demanded.
Targets that aren't concrete enough for the buffer resource layout to be
known are left alone.
AI disclosure: This was my idea but Claude wrote the code (and I've
tried to tighten up the comments)
[AMDGPU] Pre-commit tests for num_records canonicalizations (#217067)
Add tests for having InstCombine canonicalize the num_records argument
of llvm.amdgcn.make.buffer.rsrc to the width it will ultimately have,
which lets later passes see that, for example, the high bits of the i64
that Clang emits aren't used.
AI disclosure: Claude generated these and I've looked at them
[Clang][RISCV][P-ext] Add packed Q-format widening accumulate intrinsics (#228009)
Add support for the Packed "Q-format" Multiply with Widening Accumulate
intrinsics:
- `__riscv_pmqwacc_i32x2`
- `__riscv_pmqrwacc_i32x2`
RV32 selects the direct instructions, while RV64 lowers to the
spec-listed `zip16p` and packed Q-format accumulate sequences.
[CIR][CodeGen][NFC] Share hasExtraNeonArgument
Deduplicates `hasExtraNeonArgument` between CIR and classic CodeGen into
`TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share isEmptyFieldForLayout and isEmptyRecordForLayout
Deduplicates `isEmptyFieldForLayout` and `isEmptyRecordForLayout` between CIR
and classic CodeGen into a new `RecordLayoutUtils.h`. `ABIInfoImpl.h` and CIR's
`TargetInfo.h` re-export them with using-declarations, so the ~30 unqualified
callers are untouched.
Assisted-by: Claude Code (Claude Fable 5.1).