Add Vassil's responsibilities to the CODEOWNERS file. (#228167)
I'd like to get more reliable github-based routing when pull requests
arrive in the areas I am maintaining.
For reference, see clang/Maintainers.md
[SystemZ][z/OS] Fix GOFF exception table (LSDA) section generation (#228097)
Fixes #226804
The z/OS Binder rejects LSDA exception tables with `IEW2353E ... ERROR
CODE IS 25000E` when emitted as renamable PRs under an initial-load
`C_WSA64` ED parented to the root code section.
Following the fix suggested by @mms-it-ch in #226804, emit the LSDA
similarly to static WSA data:
- Under its own Section Definition (`SD`) named `GCC_except.<func>` with
`ESD_BSC_Section`.
- With a `C_WSA64` Element Definition (`ED`) using deferred load
(`GOFF::ESD_LB_Deferred`) and doubleword alignment.
- As a non-renamable Part Reference (`PR`).
Co-authored-by: Yusra Syeda <yusra.syeda at ibm.com>
[AMDGPU] Validate scale_sel in v_cvt_scale_*
These instructions can be block16 or block32 depending on the target
and scale_sel bits. Block16 is not supported in strict mode.
Re-enable the rest of the instructions in the strict mode but validate
the scale selector.
[compiler-rt] Fix -shared-libsan usage with internal symbolizer (#227717)
Summary:
https://github.com/llvm/llvm-project/pull/226551 exposed a preexisting
issue when using `-shared-libsan` nad the internal symbolizer builds.
The symbolizer will try to look up the symbol via `RTLD_NEXT`, which
goes in load order. If the sanitizer library is loaded *after* `libc.so`
then this `dlsym` call will miss it.
The solution is to have a fallback that checks using `RTLD_DEFAULT`.
This
should allow us to find the symbol in these exceptional cases. I don't
expect much fallout as this covers a case that currently returns `NULL`.
Loading it in means it could grab a reference ahead of the runtime that
was intended to be intercepted, but for this use-case I don't think this
will apply.
[clang] Fix crash constant-evaluating the construction of huge arrays (#226899)
Fixes #173728
The constant evaluator runs on ordinary code all the time (range checks
for `-W` warnings, `isEvaluatable` in codegen), and it
default-constructs an array by building one `APValue` per element. The
element count got truncated to `unsigned` first and nothing checked its
size, so a local like `struct T {} s[0xFFFFFFFF][0]` ran it out of
memory; with the issue's reproducer, where the count wraps to 2^64 - 4,
an assertions build hits the "bounds check failed" assertion in
`adjustIndex` first. Copying such an array, e.g. into a lambda capture,
had the same problem. This goes back to at least Clang 3.4. The bytecode
interpreter has its own version: it emits code for every element, so a
constructor with a member like `T a[0xFFFFFFFF]` runs it out of memory
too.
Array default construction and `ArrayInitLoopExpr` now go through the
existing `CheckArraySize` guard, the same one `new` already uses, so
[6 lines not shown]
[LLDB] Acquire the module mutex at the start of SetLoadAddress (#227149)
Fixes a potential dead-lock from parallel module loading.
I received a quick-stack of LLDB hung loading a core with parallel
module loading enabled, where two threads were trying to mutate a given
module and an object file, but having acquired the module mutex first in
one case, and the object file's section mutex first in the second case,
which each trying to subsequently acquire the other lock.
In the update case, [SetLoadAddress acquires the section list mutex and
then tries to acquire the module
mutex](https://github.com/llvm/llvm-project/blob/60f717946cb5ca911b6be22169b9bd646225c59f/lldb/source/Symbol/ObjectFile.cpp#L614)
```
SectionList *ObjectFile::GetSectionList(bool update_module_section_list) {
std::lock_guard<std::recursive_mutex> guard(m_sections_mutex);
if (m_sections_up)
return m_sections_up.get();
[26 lines not shown]
[clang][OpenMP] Don't use fused dist schedule for teams loop emitted as distribute (#228129)
Fix teams loop reductions lowered as 'distribute' lose their loop.
Claude assisted with this patch.
[AMDGPU] Canonicalize num_records to its actual width in InstCombine
llvm.amdgcn.make.buffer.rsrc is overloaded on the type of its
num_records argument, but the hardware field it ends up in has a fixed
width (32 bits, or 45 bits on gfx1250 and up). Rewrite the intrinsic to
use that width, zero-extending or truncating num_records as needed, so
that IR-level optimizations can see that the extra bits of, for example,
the i64 that Clang emits are not demanded.
Targets that aren't concrete enough for the buffer resource layout to be
known are left alone.
AI disclosure: This was my idea but Claude wrote the code (and I've
tried to tighten up the comments)
[AMDGPU] Pre-commit tests for num_records canonicalizations (#217067)
Add tests for having InstCombine canonicalize the num_records argument
of llvm.amdgcn.make.buffer.rsrc to the width it will ultimately have,
which lets later passes see that, for example, the high bits of the i64
that Clang emits aren't used.
AI disclosure: Claude generated these and I've looked at them
[Clang][RISCV][P-ext] Add packed Q-format widening accumulate intrinsics (#228009)
Add support for the Packed "Q-format" Multiply with Widening Accumulate
intrinsics:
- `__riscv_pmqwacc_i32x2`
- `__riscv_pmqrwacc_i32x2`
RV32 selects the direct instructions, while RV64 lowers to the
spec-listed `zip16p` and packed Q-format accumulate sequences.
[CIR][CodeGen][NFC] Share hasExtraNeonArgument
Deduplicates `hasExtraNeonArgument` between CIR and classic CodeGen into
`TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share isEmptyFieldForLayout and isEmptyRecordForLayout
Deduplicates `isEmptyFieldForLayout` and `isEmptyRecordForLayout` between CIR
and classic CodeGen into a new `RecordLayoutUtils.h`. `ABIInfoImpl.h` and CIR's
`TargetInfo.h` re-export them with using-declarations, so the ~30 unqualified
callers are untouched.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share hasOwnStorage
Deduplicates `hasOwnStorage` between CIR and classic CodeGen into
`RecordLayoutUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share canUseSingleInheritance
Deduplicates `canUseSingleInheritance` between CIR and classic CodeGen into
`ItaniumCXXABIUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Itanium __vmi_class_type_info flags computation
Deduplicates the `__vmi_class_type_info` and `__base_class_type_info` flags and
`computeVMIClassTypeInfoFlags` between CIR and classic CodeGen into
`ItaniumCXXABIUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Retire the CodeGenUtils.h catch-all header
Moves the last helpers out of `CodeGenUtils.h` into `ClassUtils.h`,
`ModuleUtils.h`, `TargetUtils.h` and `FunctionUtils.h` (`checkTargetFeatures`,
since it came from CodeGenFunction.cpp) and deletes the header. Only moves code
already on main, so it can be dropped on its own.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share requiresAMDGPUProtectedVisibility
Deduplicates `requiresAMDGPUProtectedVisibility` between CIR and classic CodeGen
into `TargetUtils.h`. The shared version takes a bool for "currently hidden" in
place of the `llvm::GlobalValue` and `cir::VisibilityKind` the two callers
passed.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Arm SME inlinability check
Deduplicates `ArmSMEInlinability` and `getArmSMEInlinability` between CIR and
classic CodeGen into a new `TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the bit-field and vbase layout ABI predicates
Deduplicates `isDiscreteBitFieldABI` and `isOverlappingVBaseABI` between CIR and
classic CodeGen into `RecordLayoutUtils.h`, as free functions taking the
`ASTContext`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen] Share isStandardLibraryRTTIDescriptor
Deduplicates `isStandardLibraryRTTIDescriptor` between CIR and classic CodeGen
into `ItaniumCXXABIUtils.h`, taking the classic implementation. The two copies
have been equivalent since #227781 filled in the builtin types CIR was missing.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Itanium __pbase_type_info flags and predicates
Deduplicates the `__pbase_type_info` flags, `containsIncompleteClassType` and
`extractPBaseFlags` between CIR and classic CodeGen into `ItaniumCXXABIUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
Targets: Remove redundant TRI arguments from InstrInfo helpers
This is directly available in TargetInstrInfo.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
CodeGen: Remove TRI arguments from TargetInstrInfo hooks
TRI can now always directly be referenced from TargetInstrInfo
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>