sys.mk: CTFMERGE: don't assume objfiles always have CTF sections
There are many instances, e.g. in the kernel, where
object files don't have CTF sections, but still need
to be merged into one that does have a CTF section.
This fixes cases like the following:
--------------------------------------------------------------
>>> stage 3.1: building everything
--------------------------------------------------------------
linking kernel.full
ctfmerge -t -L VERSION -g -o kernel.full ...
ERROR: ctfmerge: Input file force-dynamic-hack.pico was partially built from C sources, but no CTF data was present
Removing kernel.full
kernel.full ---
[kernel.full] Error code 1
[8 lines not shown]
[lldb][NativePDB] Use `%build` helper for thread locals test (#229195)
The test failed on ARM because it tries to link to an arm64
`msvcrtd.lib` instead of an x86_64 one. We don't really care about the
architecture here, so use the one from the host.
MachineLICM: Stop checking kill flags in register pressure estimate
This was checking kill flags, or hasOneNonDBGUse as a kill approximation. The
kill flag case appears to be of no practical use. InstrEmitter does not emit
the kill flag in the multiple user cases that would be required and I've only
managed to trigger a different hoisting decision with hand modified MIR.
The flag disagrees with hasOneNonDBGUse in 40 CodeGen tests (mostly
custom-inserter loops such as AMDGPU waterfalls and atomic expansions),
but removing it does not change the output of any CodeGen test.
[flang][test] Narrow dead SELECT CASE check (#229164)
Narrow the `CHECK-NOT` pattern in `select-case-statement.f90` from the
generic string `888` to `arith.constant 888`.
The original pattern could accidentally match `888` within generated
symbol names or source-file hashes, causing path-dependent test
failures. The
narrower pattern continues to verify that the dead `CASE DEFAULT` value
is not lowered while avoiding unrelated textual matches.
This is a test-only NFC change.
[SLP][AMDGPU][NFC] Precommit test for alternate node fmul cost (#229185)
The vector fmul of an alternate [fmul | fadd] node only feeds the
lane-select shuffle and cannot fuse with the fsub that uses the scalar
fmul, but it is currently priced as fused, so this is vectorized.
[AMDGPU] Return scalar results from uniform cvt tests
Tests with uniform inputs are meant to produce a scalar result. Return
an integer type so amdgpu_ps passes it back in SGPRs.
Change-Id: I833c140f46bb5bc347a6b9ad89c475b08bf75f3e
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
SystemZ: Don't convert AND to RISBG if CC is live (#229063)
The RISBG-type replacements either don't define CC or set it with
different semantics, so the conversion dropped or clobbered a CC value
that was still used.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[clang][deps][test] Filter modules-in-stable-dirs output with scan-deps-filter, NFC (#229155)
Pipe the output through %scan-deps-filter to capture only the expected
results. The CHECK lines are unchanged.
[RISCV][Disassembler] Reject GPR pairs with x16-x31 under RVE (#229166)
GPR pairs used by instructions like Zilsd ld/sd must use registers
accessible under RVE. Registers x16-x31 are not available with RVE, so
DecodeGPRPairRegisterClass should fail decoding when RVE is enabled and
the register number is in that range.
Co-authored-by: Claude Sonnet 5 <noreply at anthropic.com>
[ConstantFold] Don't let noipa block global address equality folding (#228186)
When deciding whether two globals may have the same address, globals
that may be replaced at link time are treated as unsafe.
`isInterposable()`
also returns true for `noipa` function definitions by default, but
`noipa` doesn't change which definition (and therefore which address)
a function resolves to. Query `isInterposable(/*CheckNoIPA=*/false)` so
comparisons involving `noipa` functions fold like any other function
with the same linkage.
This is one of a few refinements of noipa identified while working on
making optnone imply noipa - without these refinements,
optnone-implies-noipa, might substantially change clang -O0 codegen. I'm
open to discussing whether these refinements are the right direction,
though.
Assisted-By: Claude
[AMDGPU] Use common check prefixes in packed convert tests
SDAG and GlobalISel, and some subtargets, now produce identical output
for these tests. Merge their prefixes and drop the unused ones.
Change-Id: Id7ee0d1e7a3585c249ad2696d47d10a1ada7f105
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[Verifier] Don't treat noipa callees as non-inlinable in the !dbg check (#228241)
The verifier requires calls to inlinable functions (from functions with
debug info) to carry a !dbg location, and skips interposable callees
since they can't be inlined. `isInterposable()` also returns true for
`noipa` definitions by default, but `noipa` does not prevent inlining
(InlineCost already queries `isInterposable(/*CheckNoIPA=*/false)`).
Do the same here so calls to `noipa` callees are still checked.
This is one of a few refinements of noipa identified while working on
making optnone imply noipa - without these refinements,
optnone-implies-noipa, might substantially change clang -O0 codegen. I'm
open to discussing whether these refinements are the right direction,
though.
Assisted-By: Claude
[DWARF] Don't index variables whose address is only their value (#228574)
A local whose value is the address of a global, as in
long g;
void f(void) { long *p = &g; sink(p); }
can be described by DW_OP_addr(x) g, DW_OP_stack_value. The variable
does not live at that address, but the verifier and the DWARFLinkers
took any DW_OP_addr in a location to mean static storage: the verifier
demanded a .debug_names entry for such a variable, and the linkers added
one when the location was a single expression. Since DWARFLinker turns
DW_OP_addrx into DW_OP_addr, a dSYM with such a variable in a location
list failed dsymutil's output verification.
Add DWARFExpression::isMemoryLocation(), which tests whether a location
description, or one of its pieces, is a memory location description, and
only treat an address as static storage where it is. Use it in the
verifier and, through hasImplicitAddressLocation(), in both linkers.
Assisted-by: Claude
[AMDGPU] Use amdgpu_ps in packed convert tests
amdgpu_ps has no callable function prolog or epilog. Use it for the
former kernel tests and restore it for the gfx950 and gfx13 tests.
Return integer results as float so they stay in VGPRs.
gfx1250 tests keep the default calling convention: entry functions
there get a workaround sequence longer than that prolog.
Change-Id: I33ffefcb0c43b7c8e22805ae2429930d397917f8
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[mlir][OpenMP] Allow multi-block omp.iterator regions
Frontends lower a locator inside an omp.iterator region like any other
expression, so the region can contain control flow. For example, Flang
lowers LEN_TRIM in `depend(iterator(i=1:n), in: a(i+len_trim(s)))` to a
loop, and later passes inline SUM or COUNT as loops. After control-flow
conversion the region has several blocks, which the single-block
omp.iterator rejects:
%it = omp.iterator(%i: i64) = (%c1 to %n step %c1) {
llvm.br ^scan(%slen : i64)
^scan(%k: i64): // LEN_TRIM skips trailing blanks
...
llvm.cond_br %blank, ^scan(%km1 : i64), ^done
^done:
... // address of a(i+len_trim(s))
omp.yield(%addr : !llvm.ptr)
} -> !omp.iterated<!llvm.ptr>
[9 lines not shown]