[libc] Delete CharConstant utility from printf_core/core_structs.h. (#229193)
This helper was added in #222988, but this was a silly
over-complication. ASCII *strings* are not valid UTF-16 or UTF-32, but
the integral values for individual characters map directly to all Unicode encodings.
[LV] Remove EpilogueLoopVectorizationInfo (NFC). (#229411)
EpilogueLoopVectorizationInfo only holds the main loop VF and UF and the
epilogue VF, which are already available in processLoop. Use them
directly and pass them to preparePlanForEpilogueVectorLoop instead.
clang: Do not pass an empty -syslibroot for --sysroot= on Darwin
-isysroot and DEFAULT_SYSROOT, but it checked only for the presence
of the option. An empty --sysroot= then produced -syslibroot "",
where previously it fell through to -isysroot. An empty --sysroot= is
the usual way to defeat DEFAULT_SYSROOT, so treat it as absent.
Use this to fix the darwin-static-lib tests when clang is built with
DEFAULT_SYSROOT or CLANG_USE_XCSELECT.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
clang: Fix nostdlib.c test with CLANG_USE_XCSELECT
With CLANG_USE_XCSELECT, the driver injects the host macOS SDK as the sysroot
for the i386-apple-darwin run line. The SDK does not support i386, so
-Wincompatible-sysroot fires. Pass --no-xcselect, as other darwin driver tests
do.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[SlotIndexes] Add queries for stale indexes
An erased instruction leaves its index list entry in place, making the
index indistinguishable from a block boundary entry. Add
isBlockBoundaryIndex() and isStaleIndex() to tell the two apart, and
canonicalizeIndex() to resolve a stale index to the closest preceding
instruction's register slot, or the block start if none survives.
NFC. No caller yet. LiveDebugVariables is next.
[LiveDebugVariables] Repair stale SlotIndexes
The analysis keeps its indexes from before the first register allocator
until DBG_VALUEs are emitted, by which point passes in between have
erased some of the instructions they point at. Resolve them at the
start of each allocator run and before emitting.
SlotIndexes can then reclaim the entries of erased instructions without
sparing the ones held here, which would have made generated code depend
on -g. Emitted locations are unchanged, except that intervals resolving
to one position now emit a single DBG_VALUE rather than identical
consecutive ones.
clang: Fix tests when built with DEFAULT_SYSROOT
Configuring with DEFAULT_SYSROOT applies the default sysroot to every target,
not just the host/default. This broke cross-target driver tests that expect the
sysroot to be derived from the GCC installation or the driver's install
directory. Pass an empty --sysroot= to defeat the configured default, following
the precedent of 4bc05627199
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
Revert "[SLP]Support copyable GEP/zext main ops and relax gather pointer compatibility" (#229397)
It caused "isa<> used on a null pointer" assertion failures. See comment
on the PR for a reproducer.
Reverts llvm/llvm-project#221599
[TargetLowering][X86] Prefer 'r' over 'm' for foldable "rm" inline asm operands
An "rm" (register-or-memory) inline asm operand has always resolved to
'm', because getConstraintPreferences() picks the most general
constraint present, and 'm' is more general than 'r'. That's safe, since
memory can't run out, but it forces a value that could stay in a
register through a stack slot even when there's no register pressure
(https://github.com/llvm/llvm-project/issues/20571).
Prefer 'r' instead where the register allocator can fold the register
back to a stack slot when it runs out of registers, and mark the
register operand foldable (InlineAsm::Flag::setRegMayBeFolded()) so it
does. Both allocators can: the greedy allocator folds an operand when it
spills its value, and the fast allocator folds operands up front when
the asm's register operands wouldn't fit.
ParseConstraints() sets AsmOperandInfo::MayFoldRegister for an operand
whose constraint codes are exactly {r, m}, above -O0, on a target that
opts in through the new supportsRegMemInlineAsmFolding() hook, which
[20 lines not shown]
[libcxx] Add 'DoNotOptimize' to a test that is missing it. (#228253)
While developing ClangIR, we realized that we emitted slightly better
code here that InstCombine figured out how to omit new/delete in this
test. As a result, this fails with -fclangir.
This usage of DoNotOptimize seems to be the blessed version (according
to the comment on it which perfectly captures this issue?).
[RegAllocFast] Fold foldable inline asm operands under register pressure
An inline asm register operand marked foldable (from an "rm" constraint)
may be replaced with a stack slot when the register allocator runs out
of registers. The greedy allocator does that when it spills the value.
The fast allocator can't: it assigns the operands of an instruction one
at a time, and folding replaces the instruction. So it reported
"inline assembly requires more registers than available" instead.
Before allocating a block, estimate from each inline asm's own operands
whether they fit in registers, and fold as many foldable registers as
needed to make them fit, so the common case without pressure still gets
a register. Values that only live across the asm don't count, since the
allocator spills them when it needs their registers. The estimate
follows allocateInstruction():
- First all defs get distinct registers, avoiding physreg defs; then all
uses, along with the defs still occupied while the uses are read
(early-clobber and tied defs, see isLiveThroughDef()), get distinct
[25 lines not shown]
[CodeGen] Report an error for a direct inline asm output in memory
An inline asm output returned by value has no memory to write to, yet a
constraint such as "=rm" picks memory, the most general constraint, as
does "=m". SelectionDAG asserted on that ("Can only indirectify direct
input operands!"), and GlobalISel dereferenced a null pointer. Clang
never emits such an output, since it passes the address of a memory
output, but other IR can. Report "cannot handle direct memory outputs
yet for constraint 'm'" instead, like the other inline asm errors there.
Assisted-by: Claude Opus 5.5
[TargetInstrInfo] Fix folding inline asm operands next to tied operands
foldInlineAsmMemOperand() swapped a register operand for the target's
memory operands with MachineInstr::removeOperand(), which asserts when a
later operand is tied, because moving it would break the tie. Inline asm
lists every input after every output, so folding any operand that comes
before another tied pair asserted, e.g. an "rm" input followed by the
input of a "+r" operand, or one of two "+rm" operands. Without
assertions, the moved operands kept stale tie indices. Untie the
operands, rebuild the operand list, and re-tie the remaining pairs at
their new positions.
It also gave up when the register appears in more than one operand, e.g.
one value passed to two "rm" operands, which left the greedy allocator
unable to spill that value at all. Fold every such operand into the
stack slot.
Finally, take MayLoad from the folded operands rather than from every
read of the register: a folded def whose tied use is another virtual
[8 lines not shown]
[MIR] Serialize the "foldable" inline asm register operand flag
The symbolic MIR syntax for inline asm operand flags dropped
InlineAsm::Flag's RegMayBeFolded bit, which marks a register operand
(from an "rm" constraint) that the register allocator may fold to a
stack slot. Printing MIR and parsing it back therefore silently changed
what the allocator was allowed to do, e.g. with -stop-before and
-run-pass, and a test could only set the bit through a raw numeric flag.
Print the bit as a trailing "foldable", as MachineInstr::print() already
does, and parse it back. A tied use stores its matched operand number in
the same bits, so it never prints the marker.
Assisted-by: Claude Opus 5.5
[CodeGen] Ignore debug instructions in EarlyIfConversion scan budget (#224495)
`EarlyIfConverter::hasCallOrLoopInRange()` counted raw debug
instructions
against `MaxRegionInstrs`. Enough `DBG_VALUE` instructions could
therefore
stop the data-dependent load-to-condition search and change ordinary
AArch64
code generation.
Use `instructionsWithoutDebug(..., /*SkipPseudoOp=*/false)` for all
ranges
examined by the search and cache the non-debug instruction count for
complete
blocks. The explicit `false` preserves the existing treatment of pseudo
probes. Add a MIR regression test at the scan-budget boundary.
Fixes #224482
AI disclosure: This patch was developed with OpenAI Codex 5.6 Sol and
co-reviewed by Claude and Yongqiang Tian.
[BitcodeReader] Avoid integer overflow in strtab bounds check (#228786)
`Record[0] + Record[1]` can wrap, since both are read from the file, so
an out-of-bounds offset passes the check; compare without adding
instead.
Prepared with AI assistance (Claude); I reviewed the change myself.
[LV] Add support for widening loads/stores to a UF x VF (#217670)
This patch adds support for widening loads and stores to UF x VF. It is
currently limited to unmasked operations, but we plan to extend it to
masked loops via the wide active lane mask.
For now, this is driven by a new TTI hook,
`hasMultiVectorLoadStore`. A small VPlan transform uses that hook to
replace `VPWidenMemoryRecipe` with `WideVectorLoad` or `WideVectorStore`
(and any required extracts/concats).
On AArch64 this is used to target multi-vector load/store instructions.
[CIR][CodeGen][NFC] Retire the CodeGenUtils.h catch-all header
Moves the last helpers out of `CodeGenUtils.h` into `ClassUtils.h`,
`ModuleUtils.h`, `TargetUtils.h` and `FunctionUtils.h` (`checkTargetFeatures`,
since it came from CodeGenFunction.cpp) and deletes the header. Only moves code
already on main, so it can be dropped on its own.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share hasOwnStorage
Deduplicates `hasOwnStorage` between CIR and classic CodeGen into
`RecordLayoutUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share isEmptyFieldForLayout and isEmptyRecordForLayout
Deduplicates `isEmptyFieldForLayout` and `isEmptyRecordForLayout` between CIR
and classic CodeGen into a new `RecordLayoutUtils.h`. The 25 callers now name
them as `CodeGenUtils::isEmptyFieldForLayout` and
`CodeGenUtils::isEmptyRecordForLayout`, like the other shared helpers.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the bit-field and vbase layout ABI predicates
Deduplicates `isDiscreteBitFieldABI` and `isOverlappingVBaseABI` between CIR and
classic CodeGen into `RecordLayoutUtils.h`, as free functions taking the
`ASTContext`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share hasOwnStorage
Deduplicates `hasOwnStorage` between CIR and classic CodeGen into
`RecordLayoutUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the bit-field and vbase layout ABI predicates
Deduplicates `isDiscreteBitFieldABI` and `isOverlappingVBaseABI` between CIR and
classic CodeGen into `RecordLayoutUtils.h`, as free functions taking the
`ASTContext`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share isEmptyFieldForLayout and isEmptyRecordForLayout
Deduplicates `isEmptyFieldForLayout` and `isEmptyRecordForLayout` between CIR
and classic CodeGen into a new `RecordLayoutUtils.h`. The 25 callers now name
them as `CodeGenUtils::isEmptyFieldForLayout` and
`CodeGenUtils::isEmptyRecordForLayout`, like the other shared helpers.
Assisted-by: Claude Code (Claude Fable 5.1).
[clang][OpenMP] Extend no-loop promotion to the split teams loop
Widen the no-loop promotion to apply not only to the fused target teams
distribute parallel for but include the same region written as separate
directives, other directives promotable to parallel for, and regions
carrying a schedule clause that asks for a mapping a no-loop kernel
already provides.
Use the SPMD_NO_LOOP kernel tag to carry the promotion decision, based
on whether the kernel can get fully promoted to a no-loop kernel. This
eliminates the previous discrepancy where a clause got the tag but not
the promotion.
[clang][OpenMP] Emit lastprivate final copies in no-loop kernels
Promotion rejected lastprivate because the no-loop branch leaves the
worksharing path before it privatizes or copies out, so the clause would
have been silently dropped.
Privatize the non-counter variables in the parallel region and copy them
out under the last iteration that applyWorkshareLoop publishes, forcing
the exit barrier the copy reads through. Loop counters stay at the
distribute level and reach EmitOMPSimdFinal.
[CIR][CodeGen][NFC] Retire the CodeGenUtils.h catch-all header
Moves the last helpers out of `CodeGenUtils.h` into `ClassUtils.h`,
`ModuleUtils.h`, `TargetUtils.h` and `FunctionUtils.h` (`checkTargetFeatures`,
since it came from CodeGenFunction.cpp) and deletes the header. Only moves code
already on main, so it can be dropped on its own.
Assisted-by: Claude Code (Claude Fable 5.1).