[NewPM] Add helper function to skip for opt-bisect
With the introduction of CodeGen passes to the NewPM infrastructure, we
get the concept of required passes that perform optimizations and want
to skip them when running under opt bisect. This occurs for passes like
SelectionDAG and StackColoring. Add a helper function to make it easy to
query, although make it specific to opt bisect unlike the LegacyPM
skipFunction. A helper function is slightly more convenient (i.e., no
need to forward the pass name explicitly).
Reviewers: nikic, arsenm, aeubanks
Pull Request: https://github.com/llvm/llvm-project/pull/223486
[CHERI][RISCV][AsmPrinter] Use pointer index size rather than pointer size in AsmPrinter constant lowering. (#220100)
This prevents a crash due to APInt width mismatches during
accumulateConstantOffset for CHERI targets.
clang: Fix exception model flag tests if webassembly isn't build (#223770)
The wasm and emscripten RUN lines require the WebAssembly backend due to
the use of -mllvm -wasm-enable-eh, until that flag is fully replaced with
the new exception model flag.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
[flang][cuda][NFC] Move CUFDeviceIsActive function to the right place (#223453)
The definition and declaration were done in different files (descriptor
and allocator). Move all to allocatable as this check is used for the
automatic deallocation.
[[mlir][linalg] Infer reduction-neutral padding values in rewriteAsPaddedOp (#216517)
**Problem**:
with no explicit options.paddingValues, rewriteAsPaddedOp padded every
operand with zero — silently corrupting reductions (a padded
maximumf/mulf element gets combined).
**Change**:
Added a new public linalg::inferPaddingValues(builder, toPad) to pick a
semantics-preserving value per operand:
- no reduction dim → zero;
- contraction → zero (0 annihilates through the multiply);
- other reductions → the combiner's neutral (-inf for maximumf, 1 for
mulf, …), found via matchReduction so the real accumulator combiner is
used.
rewriteAsPaddedOp calls it when options.paddingValues is empty. If no
value can be inferred, it returns failure() with the error "could not
infer a padding value"; the caller is then expected to determine the
[6 lines not shown]
[docs] Prefer portable Markdown documentation links
Document and enforce source-relative Markdown links for documents and
generated heading anchors. Diagnose nonportable project: and generated HTML
links, and convert the existing LLVM Markdown links to the preferred form.
[UTC] Strip only standalone positional %s in update_test_checks.py (#221153)
Previously, update_test_checks.py used tool_cmd_args.replace("%s", "")
to strip the input file when passing the IR via stdin. However, this
blindly removed "%s" from any option argument, leaving an empty value
(e.g., -lowertypetests-read-summary=).
Update the argument stripping to only remove standalone positional
"< %s" or "%s", preserving option arguments with "%s" values so they can
be expanded via common substitutions.
This enables update_test_checks.py to work for tests such as
llvm/test/Transforms/LowerTypeTests/cfi-jumptable-hotness-summary.ll,
which passes -lowertypetests-read-summary=%s on the RUN line.
Assisted-by: Gemini
[CMake] Prune dead and unnecessary try_compiles from config.h (#223765)
This relands the change directly on `main` after #223600 was
accidentally
merged into its stacked base branch. This reuses the approvals from
@aengelke and @MaskRay on #223600.
Several config-ix checks either have no consumer or can be performed
directly by the source that needs the feature. Remove six compiler
invocations from a typical Linux configure:
* HAVE_PTHREAD_MUTEX_LOCK has been unused since LLVM switched its mutex
implementation to std::recursive_mutex in 2019.
* The Linux magic-header results have never been propagated to config.h,
so CMake builds already use Path.inc's fallback constants.
* FE_ALL_EXCEPT and FE_INEXACT can be tested directly after including
cfenv.
* The Valgrind and CrashReporter headers can be detected with
__has_include in their only consuming translation units.
[8 lines not shown]
RuntimeLibcallsEmitter: Let a consumer's own library variant beat its exclusion (#221953)
A target can pull a shared library via LibraryRef<Lib, [impls]> to drop
some impls, then re-add its own versions through a same-name library variant
guarded on that target. Previously the exclusion's setUnavailable calls were
emitted at the end of setAvailableLibFuncs_<name>, after every variant, so they
clobbered the target's own re-adds.
Defer emitting a variant until after the exclusions when it re-adds an
impl its own consumer excludes (same predicates), so the target's re-add wins
while the exclusion still suppresses every other variant's contribution.
This is yet unused infrastructure for future changes.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
CodeGen: Remove TargetOptions::EABIVersion (#222672)
The field's only effect was gating the __aeabi_mem*[4|8] libcalls via
the IsEABI4/IsEABI5 predicates. That distinction is derivable from the
triple's environment, so replace the two predicates with a single
triple-derived IsEABIVersion and delete the field.
The clang -meabi option and clang::TargetOptions::EABIVersion are
retained (now codegen-inert); the llc/opt -meabi flag is removed. -meabi
now only takes effect on triples with a bare-EABI/GNU environment pair
(arm-none-eabi <-> gnueabi), which is the only case with a triple
representation.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
[CIR][EH] Add dialect elements for dynamic exception specification (#223503)
This change introduces the CIR dialect elements that will be needed to
implement dynamic exception specification handling. At this point
nothing generates any of these elements, and they are not handled in CFG
flattening or EH ABI lowering. Those will be implemented in follow-up
changes.
See clang/docs/CIR/CleanupAndEHDesign.md for the design of this feature.
Assisted-by: Cursor / various models
[Clang] Keep qualifiers outside of OverflowBehaviorType nodes (#222190)
`__typeof_unqual__` wasn't stripping `const` off of an overflow behavior
type:
```c
__typeof_unqual__(const __ob_trap int) x; // '__ob_trap const int'
```
The canonical type was always right, so `x` really is writable -- we
just printed it as const, which is the opposite of the truth.
`getOverflowBehaviorType()` hoisted qualifiers out for the canonical
type but stored the underlying type verbatim in the sugar node, so
`const __ob_trap int` became `OBT{const int}` with nothing qualified at
the `QualType` level. `getUnqualifiedType()` checks the canonical type
to decide there's something to strip, then strips it by desugaring one
step at a time. `OverflowBehaviorType` isn't sugar, so that walk
dead-ends on it and the qualifiers never come off.
[29 lines not shown]
clang: Fix exception model flag tests if webassembly isn't build
The wasm and emscripten RUN lines require the WebAssembly backend due to the
use of -mllvm -wasm-enable-eh, until that flag is fully replaced with the
new exception model flag.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
[X86] Require legal result types for f32 estimates
The single-precision paths in `getRecipEstimate` and `getSqrtEstimate`
match a fixed set of scalar and vector f32 types, then check the
corresponding subtarget features. Those checks do not imply that the
result type is legal. Soft-float makes these result types illegal. Scalar
f32 is also illegal on a 32-bit SSE1 target without SSE2 or x87. The hooks
can then build an X86-specific estimate node with an illegal result type
before type legalization. Type legalization has no rule for that node, so
compilation fails.
Require a legal result type before entering the single-precision estimate
paths. The f16 paths in these hooks already require a legal type. After
this patch, the f32 paths match that behavior. The hooks return no estimate
for an illegal result type, so the original division or square root stays
in the DAG and follows its normal lowering. Estimate code generation for
legal result types is unchanged.
Fixes https://github.com/llvm/llvm-project/issues/217801
Assisted-by: Claude Opus 5, GPT-5.6 Sol.
[SLP]Fix insert point for struct-call extractvalue entries with multi-block scalars
Emit the vector extractvalue right after the vectorized call it depends
on, so it dominates all uses.
Fixes #223726
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/223766
[OpenMP][docs] Finish MyST migration for OpenMP docs (#222194)
This was prepared with rst2myst plus LLM-assisted cleanup. The main
manual intervention was porting CommandLineArgumentReference.(rst|md) to
the semantic [`option`
directive](https://www.sphinx-doc.org/en/master/usage/domains/standard.html#directive-option),
which causes some rendering differences.
Tracking issue: #201242
See the [migration guide] for more information.
[migration guide]:
https://llvm.org/docs/SphinxQuickstartTemplate.html#markdown-migration-guidelines
This is a stacked PR based on #222191, which is a standalone commit that
renames *.rst -> *.md before this PR lands for history preservation
purposes.
Please spot check my work and approve if it looks good. You can use the
HTML links below to confirm it renders properly.
[179 lines not shown]
[OpenMP][docs] Rename reStructuredText files to Markdown (#222191)
Tracking issue: #201242
See the [migration guide] for more information.
[migration guide]:
https://llvm.org/docs/SphinxQuickstartTemplate.html#markdown-migration-guidelines
This is the initial straight rename commit. It will probably break the
docs build, but it has to be a separate PR for blame preservation
purposes.
Validation:
- Confirmed this commit contains exactly 34 `.rst` to `.md` renames and
no content changes.
- The complete stacked migration is validated in the follow-up PR.
[bazel] Use new rules_cc API to determine musl targeting (#223549)
This avoids requiring hermetic-llvm in the public BUILD file API.
Downstream users need to update to this version of rules_cc
[CMake] Factor check- target deps to reduce build.ninja size by 45%
Many LLVM developers are not aware, but our lit CMake logic creates a
check-* target for every eligible subdirectory inside a lit test suite.
There are thousands of these targets; in the benchmark configuration:
```
❯ ninja -C build -t targets all | grep '^check-' | wc -l
2667
```
Chris Bieneman originally added this feature in March 2015. The original
change suppressed these targets for Xcode and Visual Studio because the
target clutter hurt IDE usability; rL344555 later generalized that
behavior behind LLVM_ENABLE_IDE.
When using the Ninja generator, each of these check-${proj}-${subdir}
targets repeats its suite's transitive test dependencies, like so:
[52 lines not shown]
[CMake] Remove unused configuration probes
Several config-ix checks either have no consumer or can be performed directly
by the source that needs the feature. Remove six compiler invocations from a
typical Linux configure:
* HAVE_PTHREAD_MUTEX_LOCK has been unused since LLVM switched its mutex
implementation to std::recursive_mutex in 2019.
* The Linux magic-header results have never been propagated to config.h, so
CMake builds already use Path.inc's fallback constants.
* FE_ALL_EXCEPT and FE_INEXACT can be tested directly after including cfenv.
* The Valgrind and CrashReporter headers can be detected with __has_include in
their only consuming translation units.
Also remove the obsolete definitions from the GN and Bazel configurations.
In a fresh minimal LLVM configure, CMake profiling attributed 2067.1 ms to 45
config-ix probes before this change and 1846.9 ms to 39 probes after it.
[6 lines not shown]
[offload][omp] Move OpenMP KLE to libomptarget
Move preparations related to OpenMP KLE and
dynamicCGroupMem fallback out of the plugins and
into libomptarget.
Resructure Device::launch as it grew too large.
[DAG][X86] Do not narrow trunc(select) to a type undesirable for SELECT
DAGCombiner rewrites trunc(select c, a, b) as select c, (trunc a), (trunc b)
whenever truncating is free. On X86 that fires for i8, and where CMOV is
available there is no 8-bit form of it, so LowerSELECT widens the narrow select
back to i32 through ANY_EXTENDs, which lower to MOVZX. The narrowing then only
buys a byte ALU chain plus a MOVZX for each operand that became an i8 op:
orb $64, %al orl $64, %eax
movzbl %al, %eax cmovneq %rcx, %rax
cmovnel %ecx, %eax
Gate the fold on isTypeDesirableForOp(ISD::SELECT, VT) and let X86 report i8 as
undesirable when it can use CMOV. Without CMOV the select lowers to a branch
instead, where i8 operands are fine, so narrowing stays enabled there. The hook
defaults to isTypeLegal(VT), so guarded by isTypeLegal it is a no-op for every
target that does not override it. A select of two constants stays exempt: it
introduces no truncate of a computed value, so the narrow form is never worse.
Assisted-by: Claude Code
[X86][NFC] Pre-commit tests for trunc(select) narrowing
DAGCombiner narrows trunc(select c, a, b) to the truncated type. On X86 that
gives an i8 select, which is widened back to i32 wherever CMOV is available,
so the narrowing only adds a byte ALU chain and a MOVZX.
Assisted-by: Claude Code
[CIR] Support pointer-to-int global initializers (#220643)
This allows CIR to emit global initializers where a constant address is
cast to an integer.
This fixes compound literals like `unsigned long addr = (unsigned
long)(int[]){1, 2, 3}` and also handles global variable, array element
and function addresses cast to integers.
Fixes #216618
RuntimeLibcallsEmitter: Let a consumer's own library variant beat its exclusion
A target can pull a shared library via LibraryRef<Lib, [impls]> to drop some
impls, then re-add its own versions through a same-name library variant guarded
on that target. Previously the exclusion's setUnavailable calls were emitted at
the end of setAvailableLibFuncs_<name>, after every variant, so they clobbered
the target's own re-adds.
Defer emitting a variant until after the exclusions when it re-adds an impl its
own consumer excludes (same predicates), so the target's re-add wins while
the exclusion still suppresses every other variant's contribution.
This is yet unused infrastructure for future changes.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
RuntimeLibcallsEmitter: Add isolated LibcallLibrary variant
This is essentially a hack to not break the common core functions shared
by most targets. Most targets have essentially the same base set of
compiler-rt or libc/libm functions, but a few are so radically different
there is nothing in common (e.g., the GPU targets have a handful of functions,
arm64ec changes every single function). This bit will pull these out of the
library merge-by-name system, and emitted as its own special case. This allows
the single name to be universal across targets without duplicating large
tables in the emitted inc file.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
RuntimeLibcallsEmitter: Handle calling conv in per-library functions
Pull handling of the default calling convention into the
setAvailableLibFuncs_<name> functions, so the library logic will be
fully contained.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
[CIR] Support SubstNonTypeTemplateParmExpr for aggregates (#223603)
Lower `SubstNonTypeTemplateParmExpr` in `AggExprEmitter` by visiting its
replacement expression, matching classic Clang codegen (`CGExprAgg.cpp`)
and existing scalar (#146751), complex (#146755), and lvalue (#182920)
implementations.
The missing aggregate NTTP gap was analyzed with LLM assistance.
Fixes #223533