[bazel] Use new rules_cc API to determine musl targeting (#223549)
This avoids requiring hermetic-llvm in the public BUILD file API.
Downstream users need to update to this version of rules_cc
[offload][omp] Move OpenMP KLE to libomptarget
Move preparations related to OpenMP KLE and
dynamicCGroupMem fallback out of the plugins and
into libomptarget.
Resructure Device::launch as it grew too large.
[DAG][X86] Do not narrow trunc(select) to a type undesirable for SELECT
DAGCombiner rewrites trunc(select c, a, b) as select c, (trunc a), (trunc b)
whenever truncating is free. On X86 that fires for i8, and where CMOV is
available there is no 8-bit form of it, so LowerSELECT widens the narrow select
back to i32 through ANY_EXTENDs, which lower to MOVZX. The narrowing then only
buys a byte ALU chain plus a MOVZX for each operand that became an i8 op:
orb $64, %al orl $64, %eax
movzbl %al, %eax cmovneq %rcx, %rax
cmovnel %ecx, %eax
Gate the fold on isTypeDesirableForOp(ISD::SELECT, VT) and let X86 report i8 as
undesirable when it can use CMOV. Without CMOV the select lowers to a branch
instead, where i8 operands are fine, so narrowing stays enabled there. The hook
defaults to isTypeLegal(VT), so guarded by isTypeLegal it is a no-op for every
target that does not override it. A select of two constants stays exempt: it
introduces no truncate of a computed value, so the narrow form is never worse.
Assisted-by: Claude Code
[X86][NFC] Pre-commit tests for trunc(select) narrowing
DAGCombiner narrows trunc(select c, a, b) to the truncated type. On X86 that
gives an i8 select, which is widened back to i32 wherever CMOV is available,
so the narrowing only adds a byte ALU chain and a MOVZX.
Assisted-by: Claude Code
[CIR] Support pointer-to-int global initializers (#220643)
This allows CIR to emit global initializers where a constant address is
cast to an integer.
This fixes compound literals like `unsigned long addr = (unsigned
long)(int[]){1, 2, 3}` and also handles global variable, array element
and function addresses cast to integers.
Fixes #216618
RuntimeLibcallsEmitter: Let a consumer's own library variant beat its exclusion
A target can pull a shared library via LibraryRef<Lib, [impls]> to drop some
impls, then re-add its own versions through a same-name library variant guarded
on that target. Previously the exclusion's setUnavailable calls were emitted at
the end of setAvailableLibFuncs_<name>, after every variant, so they clobbered
the target's own re-adds.
Defer emitting a variant until after the exclusions when it re-adds an impl its
own consumer excludes (same predicates), so the target's re-add wins while
the exclusion still suppresses every other variant's contribution.
This is yet unused infrastructure for future changes.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
RuntimeLibcallsEmitter: Handle calling conv in per-library functions
Pull handling of the default calling convention into the
setAvailableLibFuncs_<name> functions, so the library logic will be
fully contained.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
[CIR] Support SubstNonTypeTemplateParmExpr for aggregates (#223603)
Lower `SubstNonTypeTemplateParmExpr` in `AggExprEmitter` by visiting its
replacement expression, matching classic Clang codegen (`CGExprAgg.cpp`)
and existing scalar (#146751), complex (#146755), and lvalue (#182920)
implementations.
The missing aggregate NTTP gap was analyzed with LLM assistance.
Fixes #223533
[lldb] Compare registers by LLDB register number (#223704)
Fix `GetIndexOfChildWithName`/`GetChildMemberWithName` on `Windows
x86_64` (Release builds), where comparing registers by RegisterInfo
pointer identity could spuriously fail to find sp/rsp in the GPR set.
Compare by `eRegisterKindLLDB` number instead.
Failing bot: https://ci-external.swift.org/job/lldb-windows/job/main
[mlir][sparse] Lower affine before SCF to CF in sparsifier (#219860)
Fix the lowering order in the sparsifier pipeline by running
`convert-scf-to-cf` after `lower-affine`.
Previously, `convert-scf-to-cf` ran while Affine control-flow operations
could still be present. For example, lowering an `scf.if` nested inside
an `affine.for` introduces multiple blocks into the loop body, violating
`affine.for`'s single-block region requirement and causing verification
to fail.
Running `lower-affine` first converts Affine control flow to SCF. The
subsequent `convert-scf-to-cf` pass can then lower both the existing SCF
operations and those introduced by Affine lowering to CFG-based control
flow.
Fixes #217895
[AggressiveInstCombine] Fix shared libraries build (#223756)
Broken by #223728 by depending upon a function in ProfileData. Add that
as a dependency to fix the build.
[lldb][Windows] Key lldb-server's loaded-module list by base address (#223445)
`NativeProcessWindows` keys `m_loaded_modules` by `FileSpec`, so a
second mapping of an already-loaded DLL overwrites the first one's base
address, and `OnUnloadDll` then erases by address. The wrong mapping
gets dropped.
This reproduces reliably in swiftlang where `swiftCore.dll` is loaded
twice.
This patch maps by base address instead (the identity `UNLOAD_DLL`
carries), so an unload resolves to the right file and only retires the
reported entry when it is that file's last mapping.
It introduces the `LoadedModuleList` class to key the modules correctly
and 9 regression unit tests.
---------
Co-authored-by: Nerixyz <nero.9 at hotmail.de>
RuntimeLibcalls: Update library-ref test for removed EABIVersion parameter
This test was added on main after the parent commit was written, so its
CHECK lines still expect EABIVersion in the generated dispatcher calls.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
[DAGCombiner] Port custom DAG combine for `rem` from NVPTX (#210344)
Part of #116695
Supersedes #167147
Port NVPTX combine for `rem` to DAGCombine. Folds `Num % Den -> Num -
(Num / Den) * Den` if `DIVREM` is not supported by the backend.
[ConstraintElim] Unify no-wrap queries. (#223537)
Unify no-wrap queries in new isKnownNoWrap helper, which checks both
wrap flags on instructions and query-based reasoning in a single place.
Checks signed/unsigned wrap depending on a flag.
Update both decompose and tryToStrengthenFlags to use shared helper.
This improves results in a number of cases, because some rules were only
implemented for tryToStrengthenFlags and others in decompose:
https://github.com/dtcxzyw/llvm-opt-benchmark-nightly/pull/1313
Note that there are 2-3 small regressions due added flags/removed
branches pessimizing other parts of the pipeline.
Compile-time impact in the noise for geomean. Notable individual changes
is mafft -0.05% in stage1-ReleaseThinLTO, and lencod +0.05% for
stage2-O3.
[2 lines not shown]
[Clang][Sema] Add Diagnostic for using matrix logical on non HLSL targets (#223252)
Make Clang emit an error message instead of crashing when using matrix
logical operations on non-HLSL targets.
Issue #222381
[CIR] Take the va_arg register demand from the classifier
rewriteVAArg rebuilt how many registers of each class the fetched type
needs by inspecting that type. ArgInfo and ArgClassification now carry
the demand the x86-64 classifier already computes.
The register arm also has to copy into a full-size temp whenever the
registers carry less than the whole result, not only when the value
starts at a byte offset. A record whose tail holds no field took the
neighbouring slot as part of its value.
Assisted-by: Cursor / claude-opus-5
[AggressiveInstCombine] Preserve Profile Info for [0,1] Memset Guard (#223728)
This showed up as a profcheck failure from #213240. If the memset has
profile information, we can use that to synthesize appropriate branch
weights.
RuntimeLibcalls: Remove the dead IsDefault emitter machinery (#219947)
The IsDefault bit on RuntimeLibcallImpl fed a LibCallToDefaultImpl map
in the TableGen backend that was populated but never read.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
[AMDGPU] Skip vector atomicrmw operands in AtomicOptimizer (#223309)
A uniform `<2 x i32>` `atomicrmw add` reaches a scalar-to-vector
lane-count cast and asserts in AMDGPUAtomicOptimizer.
The divergent-value path rejects types in `isLegalCrossLaneType`,
but the uniform-value path does not.
Returned vectors also enter scalar result construction.
Reject vector types in `visitAtomicRMWInst` and `visitIntrinsicInst`
before either value path. The regression covers unused
global `add`, returned LDS `and`,
and a uniform scalar `i16` control that must still optimize, using DPP
and iterative strategies on wave64 and wave32 subtargets.
[LowerAtomic] Mark cmpxchg select with unknown branch weights (#223736)
LowerAtomic in cmpxchg lowering creates a select to figure out what to
store. This is conditioned on whether or not the compare value is equal
to the value in memory, which we currently have no way of knowing the
probability of. So mark the branch weights explicitly unknown.
It seems like we were lacking coverage of this case before #223157,
which made this test pop up on the profcheck bot.
CodeGen: Remove TargetOptions::EABIVersion
The field's only effect was gating the __aeabi_mem*[4|8] libcalls via the
IsEABI4/IsEABI5 predicates. That distinction is derivable from the triple's
environment, so replace the two predicates with a single
triple-derived IsEABIVersion and delete the field.
The clang -meabi option and clang::TargetOptions::EABIVersion are
retained (now codegen-inert); the llc/opt -meabi flag is removed. -meabi
now only takes effect on triples with a bare-EABI/GNU environment pair
(arm-none-eabi <-> gnueabi), which is the only case with a triple
representation.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
RuntimeLibcalls: Add LibraryRef for dispatch-with-exclusion (#218869)
Let a SystemRuntimeLibrary dispatch a shared provider library while
dropping the impls the target replaces, since a library reference cannot nest
inside (sub ...). This is a compromise from the ideal of explicitly listing all calls,
but getting to that point is prooving to be difficult.
The opt-out is emitted inside setAvailableLibFuncs_<lib>, so the single library's
logic is self contained.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
Cache DeclContext-to-Decl conversions
Store the owning Decl pointer when constructing each DeclContext and route
parent traversal and generic casts through it. This avoids a lazy lookup and
does not imply support for concurrent AST traversal.
CTMark O0 (three alternating matched-build samples, CPU 6): 28.990700 s ->
28.702600 s (-0.9938%). Peak build RSS: 258016 -> 259276 KiB (+0.4883%).
All 632 normalized objects matched in each pair.
MLIR build-time medians (three alternating matched-build samples, CPU 6):
- `mlir/lib/RegisterAllDialects.cpp`: 2.8115% fewer retired instructions,
0.7353% less user CPU, 0.1972% less wall time, and 7144 KiB (+0.5257%)
peak RSS.
- `mlir/lib/Dialect/LLVMIR/IR/NVVMDialect.cpp`: 1.5841% fewer retired
instructions, 1.1120% less user CPU, 0.9893% less wall time, and 912 KiB
(+0.0795%) peak RSS.
All MLIR outputs matched after removing only `.comment`; each concurrency
[3 lines not shown]
[MachinePipeliner] Use VirtRegOrUnit instead of Register appropriately (#177535)
Resolve the FIXMEs in MachinePipeliner added in #167730, primarily by
replacing `Register` with `VirtRegOrUnit` where necessary to remove
invalid `static_cast`s. Based on my local testing, there was no
noticeable performance impact.
---------
Co-authored-by: Harsha Jagasia <harsha.jagasia at amd.com>
[clang][bytecode] Apply pointer casts to opaque pointers (#223607)
For `((char *)&sqlite3Prepare_sParse) + 4`, the final byte offset should
be `4`, not `4 * sizeof(sqlite3Prepare_sParse)`. To handle that, we need
to actually pass the cast along to the opaque pointer.