[LLDB] Acquire the module mutex at the start of SetLoadAddress (#227149)
Fixes a potential dead-lock from parallel module loading.
I received a quick-stack of LLDB hung loading a core with parallel
module loading enabled, where two threads were trying to mutate a given
module and an object file, but having acquired the module mutex first in
one case, and the object file's section mutex first in the second case,
which each trying to subsequently acquire the other lock.
In the update case, [SetLoadAddress acquires the section list mutex and
then tries to acquire the module
mutex](https://github.com/llvm/llvm-project/blob/60f717946cb5ca911b6be22169b9bd646225c59f/lldb/source/Symbol/ObjectFile.cpp#L614)
```
SectionList *ObjectFile::GetSectionList(bool update_module_section_list) {
std::lock_guard<std::recursive_mutex> guard(m_sections_mutex);
if (m_sections_up)
return m_sections_up.get();
[26 lines not shown]
[clang][OpenMP] Don't use fused dist schedule for teams loop emitted as distribute (#228129)
Fix teams loop reductions lowered as 'distribute' lose their loop.
Claude assisted with this patch.
[AMDGPU] Canonicalize num_records to its actual width in InstCombine
llvm.amdgcn.make.buffer.rsrc is overloaded on the type of its
num_records argument, but the hardware field it ends up in has a fixed
width (32 bits, or 45 bits on gfx1250 and up). Rewrite the intrinsic to
use that width, zero-extending or truncating num_records as needed, so
that IR-level optimizations can see that the extra bits of, for example,
the i64 that Clang emits are not demanded.
Targets that aren't concrete enough for the buffer resource layout to be
known are left alone.
AI disclosure: This was my idea but Claude wrote the code (and I've
tried to tighten up the comments)
[AMDGPU] Pre-commit tests for num_records canonicalizations (#217067)
Add tests for having InstCombine canonicalize the num_records argument
of llvm.amdgcn.make.buffer.rsrc to the width it will ultimately have,
which lets later passes see that, for example, the high bits of the i64
that Clang emits aren't used.
AI disclosure: Claude generated these and I've looked at them
[Clang][RISCV][P-ext] Add packed Q-format widening accumulate intrinsics (#228009)
Add support for the Packed "Q-format" Multiply with Widening Accumulate
intrinsics:
- `__riscv_pmqwacc_i32x2`
- `__riscv_pmqrwacc_i32x2`
RV32 selects the direct instructions, while RV64 lowers to the
spec-listed `zip16p` and packed Q-format accumulate sequences.
[CIR][CodeGen][NFC] Share hasExtraNeonArgument
Deduplicates `hasExtraNeonArgument` between CIR and classic CodeGen into
`TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share isEmptyFieldForLayout and isEmptyRecordForLayout
Deduplicates `isEmptyFieldForLayout` and `isEmptyRecordForLayout` between CIR
and classic CodeGen into a new `RecordLayoutUtils.h`. `ABIInfoImpl.h` and CIR's
`TargetInfo.h` re-export them with using-declarations, so the ~30 unqualified
callers are untouched.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share hasOwnStorage
Deduplicates `hasOwnStorage` between CIR and classic CodeGen into
`RecordLayoutUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share canUseSingleInheritance
Deduplicates `canUseSingleInheritance` between CIR and classic CodeGen into
`ItaniumCXXABIUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Itanium __vmi_class_type_info flags computation
Deduplicates the `__vmi_class_type_info` and `__base_class_type_info` flags and
`computeVMIClassTypeInfoFlags` between CIR and classic CodeGen into
`ItaniumCXXABIUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Retire the CodeGenUtils.h catch-all header
Moves the last helpers out of `CodeGenUtils.h` into `ClassUtils.h`,
`ModuleUtils.h`, `TargetUtils.h` and `FunctionUtils.h` (`checkTargetFeatures`,
since it came from CodeGenFunction.cpp) and deletes the header. Only moves code
already on main, so it can be dropped on its own.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share requiresAMDGPUProtectedVisibility
Deduplicates `requiresAMDGPUProtectedVisibility` between CIR and classic CodeGen
into `TargetUtils.h`. The shared version takes a bool for "currently hidden" in
place of the `llvm::GlobalValue` and `cir::VisibilityKind` the two callers
passed.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Arm SME inlinability check
Deduplicates `ArmSMEInlinability` and `getArmSMEInlinability` between CIR and
classic CodeGen into a new `TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the bit-field and vbase layout ABI predicates
Deduplicates `isDiscreteBitFieldABI` and `isOverlappingVBaseABI` between CIR and
classic CodeGen into `RecordLayoutUtils.h`, as free functions taking the
`ASTContext`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen] Share isStandardLibraryRTTIDescriptor
Deduplicates `isStandardLibraryRTTIDescriptor` between CIR and classic CodeGen
into `ItaniumCXXABIUtils.h`, taking the classic implementation. The two copies
have been equivalent since #227781 filled in the builtin types CIR was missing.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Itanium __pbase_type_info flags and predicates
Deduplicates the `__pbase_type_info` flags, `containsIncompleteClassType` and
`extractPBaseFlags` between CIR and classic CodeGen into `ItaniumCXXABIUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
Targets: Remove redundant TRI arguments from InstrInfo helpers
This is directly available in TargetInstrInfo.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
CodeGen: Remove TRI arguments from TargetInstrInfo hooks
TRI can now always directly be referenced from TargetInstrInfo
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[X86][APX] Use EVEX BLSI/BLSMSK for i8 patterns with EGPR (#226796)
This is a follow-up to #204746 and #205093, which added the i8 BLSI and
BLSMSK patterns.
The i8 BLSI/BLSMSK patterns were defined outside of Bls_Pats and always
selected the VEX forms, so with EGPR the register allocator could not
assign r16-r31 to their operands. Move them into Bls_Pats so that the
_EVEX variants are selected when EGPR is available.
Assisted-by: Claude Code
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[InstCombine] Declare command line options in TableGen (#227948)
Follow-up to the framework #226087:
Move the cl::opts into InstCombineCLOptions.td, with `prefix =
"instcombine-"`: -instcombine-max-num-phis sets CLOpts.max_num_phis.
CLOpts also replaces InstCombiner::MaxArraySizeForCombine.
Options that were not cl::Hidden are now listed by -help-hidden only,
and -instcombine-lower-dbg-declare is now a bool.
Aided by Opus 5.5
[mlir][xegpu] Distribute extract_strided_slice over multiple dims (#227478)
SgToLaneVectorExtractStridedSlice only handled a single distributed
dimension, and this PR handles distributing multiple dimensions. It
scales each distributed dim by its own lane_layout entry. Both the size
and the offset along a dim shrink by the number of lanes that split it.
A dim the lanes do not split keeps its offset untouched, and may still
carry non-unit lane_data as before.
assisted-by-claude
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply at anthropic.com>
[IDF] Use BitVectors indexed by DFS number for visited sets (NFC). (#227832)
IDFCalculatorBase::calculate tracks visited dominator tree nodes in two
SmallPtrSets. The DFS numbers computed at the start of calculate are
unique and dense in [0, number of nodes), with the DFS out-number of the
root being the number of nodes, so use them to index BitVectors instead.
This reduces hashing overhead and improves compile-time, depending on
configuration/workload:
stage1-O3: -0.04%
stage1-ReleaseThinLTO: -0.05%
stage1-ReleaseLTO-g: -0.22%
stage1-aarch64-O3: -0.06%
stage2-O3: -0.04%
stage2-clang: -0.04%
https://llvm-compile-time-tracker.com/compare.php?from=b96b66160ace30c2b5eef1f4afdb14c95ecc26cd&to=5ffb13d6d1f87195bba8af13366a152a7de2bfe0&stat=instructions:u
[3 lines not shown]
[mlir][vector] Add an eliminate-vector-masks pass (#226517)
`eliminateVectorMasks` had no in-tree caller other than a test pass, so
no pipeline could use it. This exposes it as an opt-in
`-eliminate-vector-masks` pass and drops the test pass.
The pass rewrites `vector.create_mask` ops that can be proven all-true
into `vector.constant_mask`; canonicalization then folds those away, so
a masked transfer becomes an unmasked one. Mask sizes are bounded with
`ValueBoundsOpInterface`, so a loop-derived size like `%dim - %iv` can
be proven.
Scalable dimensions need a `vscale` range, given by
`vscale-min`/`vscale-max`. Leaving them unset means fixed-size reasoning
only; a half-specified or inverted range is rejected rather than
silently ignored.
`eliminate-masks.mlir` keeps its checks and only switches its two RUN
lines, which shows the pass behaves as the test pass did.
[6 lines not shown]
[SPIRV] Fix crash when two functions share an alias scope (#225017)
When two functions referenced the same alias scope metadata, they were
incorrectly sharing a virtual register. This caused either a crash or
broken SPIR-V output with undefined references. The fix ensures each
function creates its own virtual register for alias scope metadata
---------
Co-authored-by: Michal Paszkowski <michal at michalpaszkowski.com>
[llvm-profdata] Remove exitWithError and LSan leak workaround
With all subcommands propagating llvm::Error to main, exitWithError,
exitWithErrorCode, and the LSan leak suppression workaround are no
longer needed.
Assisted-by: Gemini
[NFCI][llvm-profdata] Propagate Error in loadInput and mergeWriterContexts
Propagate Error from loadInput and mergeWriterContexts in mergeInstrProfile,
supplementInstrProfile, and overlapInstrProfile. In mergeInstrProfile's
ThreadPool, catch errors from worker threads, stop scheduling new jobs,
and return the first encountered fatal error.
Ensure ~WriterContext() consumes any pending unhandled errors in
WriterContext::Errors upon destruction.
Not NFC as destructors are run on the stack and ThreadPool workers exit
earlier on error.
Assisted-by: Gemini