[CIR] Lower language address spaces nested in record members (#228934)
TargetLowering now derives from RecordRewritingTypeConverter, so it
rebuilds records whose members reach a language address space and keeps
the rest. This fixes "member type mismatch" for OpenCL and SYCL records
with address-space qualified pointer members.
[AMDGPU] Split the true16 fcopysign pattern by uniformity (#226898)
A uniform 16-bit value already lives in the low half of an SGPR, so it can go straight to V_BFI_B32_e64. Only a divergent VGPR_16 operand needs the widening REG_SEQUENCE, which was ill-formed for the uniform case (operand wider than the lo16 slot it names) and confused DetectDeadLanes.
Fixes: ROCM-31212
[CIR][SSCP] Return null callable region for function declarations (#228641)
Declarations and aliases have empty bodies; returning them made SCCP
treat external calls as never returning and fold their results to
constants.
[lldb] For Mach-O re-export symbols, save target binary FileSpec (#228266)
Mach-O re-export symbols specify the name of the target function to call
AND the binary name to find it in. I didn't correctly save the binary
name in the DataFileCache representation. This change does that, and
adds a new test (only runs on Darwin) which does a `target modules dump
symtab` on a system dylib that has many re-export symbols, then creates
a DataFileCache of that binary, loads a new target using the DFC and
re-executes `target modules dump symtab`, and checks that the output is
identical between them.
rdar://187739510
[clang] Precommit test for auto-init of small struct fields (#229225)
Currently when using the -ftrivial-auto-var-init-max-size flag, record
types are either fully initialized if they are smaller than the max
size, or left fully uninitialized even if they have individual fields
smaller than the max size. This change extends existing tests for
auto-init max size to cover the case where a struct contains a struct
containing a small field, in preparation for changes to automatically
initialize small fields inside structs.
[CIR][OpenMP] Add support for host_eval so that SPMD kernels can be used
This patch adds support for host_eval so that SPMD kernels and in the future
num_threads etc. can be implemented correctly.
Assisted-by: Cursor / Claude Sonnet 5 High
[sanitizer] Clear pending dlerror on thread exit before unregistering (#228930)
Starting with glibc 2.34
(https://sourceware.org/git/?p=glibc.git;h=fada9018199c), failed dlfcn
calls store a per-thread error struct in TLS (__libc_dlerror_result)
instead of a pthread_key_t, and free it in __libc_thread_freeres() after
all pthread TSD destructors have completed.
Because ASan, LSan, and HWASan unregister the exiting thread in their
pthread TSD destructors during __nptl_deallocate_tsd(), LSan stops
scanning that thread's TLS before __libc_thread_freeres() frees
__libc_dlerror_result. If a leak check runs in that window (for example,
when the main thread exits while a detached worker thread is finishing)
and libc debug symbols are not installed to match the default *dlerror*
suppression on _dlerror_run, LSan reports a false positive leak from
dlsym/dlopen.
This manifests as flaky (~3% failure rate) LSan leaks on
`LLVM :: ExecutionEngine/JITLink/x86-64/` tests on
[9 lines not shown]
RISCV: Do not add a live VL def to inline asm that clobbers it
Inline asm that already clobbers VL or VTYPE was given an additional
live implicit def of the same register, contradicting the dead clobber
def.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[AMDGPU] Reserve ENABLE_WAVEFRONT_SIZE32 on gfx125x (#229162)
Wave32-only targets (gfx1250, gfx1250-strict, gfx1251, gfx12-5-generic)
reserve the kernel descriptor's ENABLE_WAVEFRONT_SIZE32, so it must be
0.
Also update the documentation regarding which targets allow selecting
the wave size through this bit.
[AMDGPU] op_sel[0] must be zero for v_cvt_sr_fp8/bf8_f16 (#227413)
For v_cvt_sr_fp8_f16 and v_cvt_sr_bf8_f16, current hardware ignores the
op_sel[0] bit. SP3 docs have been updated to specify op_sel[0] must be zero.
This work makes the change accordingly for both MC and instruction
selection. In true16 mode the src0.h form is thus no longer encodable.
These instructions use op_sel[2:3] to select a byte in vdst and require
op_sel[1:0] to be zero, so set HasOpSel = 0 on their profiles: there is no
op_sel operand at all and Inst{11-12} are hardwired to zero. src0 is then
restricted to the low half of a VGPR through the true16 profile, using new
definitions:
- VGPR_16_LO16 -- VGPR0_LO16 through VGPR255_LO16, the most a 9-bit VOP3
source can encode;
- VS_16_LO16 -- VGPR_16_LO16, SReg_32 and LDS_DIRECT, as the operand's
register class;
- VSrcT_f16_LO16 and FPT16_LO16InputMods for Src0RC64/Src0Mod (VOP3);
- VGPROp_16_LO16 and FPT16_LO16VRegInputMods for Src0VOP3DPP/Src0ModVOP3DPP
[9 lines not shown]
[DWARFLinker] Add missing DebugInfoDWARFLowLevel dependency (#229222)
Since afca127bbb62, hasImplicitAddressLocation() in LLVMDWARFLinker uses
DWARFExpression from DebugInfoDWARFLowLevel. Builds with
BUILD_SHARED_LIBS=ON failed to link libLLVMDWARFLinker with undefined
references to DWARFExpression. Add the component to the CMake library,
and the matching dependency to the GN build.
Assisted-by: Claude
DAG: Fix error message casing to match style policy (#229212)
The developer guidelines suggests that diagnostic messages should start
with a lowercase letter and not end in a period
Co-authored-by: Claude <noreply at anthropic.com>
[AArch64][Win] Account for the fixed object area when classifying CSR (#222111)
In Windows frames, the fixed object area sits above the callee-saved
register area:
+---------------+
| Fixed objects |
+---------------+
| Callee-saved |
+---------------+
| Locals |
+---------------+
Include the fixed object area in the callee-saved threshold used by
`resolveFrameOffsetReference()`. Without this, objects in the
callee-saved area can be incorrectly addressed through the base
pointer in stack-realigned frames.
This became visible after #147421 moved catch objects into the fixed
[3 lines not shown]
[MLGO] Make regalloc eviction input tensor shapes runtime variables (#224598)
The column count was hardcoded to 33 for X86
It now comes from `-mlregalloc-num-allocatable-regs` (default 32, plus
one column for the candidate), compiled model whose shapes do not match
falls back to the default policy
[lldb] Use file(MAKE_DIRECTORY) to create header staging dir (#228606)
This simplifies the build graph for modifying lldb headers for
installation. It also fixes the `clean` target by not making a target
responsible for creating this directory. CMake will create it as needed
during configuration time instead of at build time.
rdar://161109746
[SPIR-V] Mark elementwise intrinsics IntrTriviallyScalarizable (#227286)
This mirrors the DirectX intrinsics
The SPIR-V legalizer will use it to split elementwise intrinsics with
illegal vector widths
Required for https://github.com/llvm/llvm-project/pull/227287
[mlir][OpenACC] Skip implicit routine marking for host-only calls (#227758)
Calls nested in a host-only branch of `acc.on_device` do not run on the
device. Do not try to attach implicit acc routine information to them.
[CI] Exclude CIR from Windows premerge testing (#229208)
After #227957, the check-clang-cir target is only defined when
CLANG_ENABLE_CIR is ON. The Windows premerge build never sets that
option (only monolithic-linux.sh receives enable_cir), but
compute_projects.py still selects check-clang-cir for CIR changes on
Windows. Every PR touching CIR now fails the Windows job before any test
runs:
ninja: error: unknown target 'check-clang-cir'
Before #227957 the target existed unconditionally, and the CIR tests
were all unsupported on Windows because CIR was disabled, so excluding
CIR there loses no coverage. It also stops building mlir for CIR-only
changes on Windows.
Enabling real CIR testing on Windows (passing enable_cir through to
monolithic-windows.sh) can be done separately.
[2 lines not shown]
[DWARFLinker] Relocate DW_OP_addrx by the delta of its own symbol (#228583)
When rewriting DW_OP_addrx or DW_OP_constx into a relocated address,
DWARFLinker applied the adjustment of whatever owned the expression.
For a location list that is the enclosing function, but the .debug_addr
entry may name a data symbol, which the linker moves by a different
amount. For a variable holding the address of _g at 0x100004008 this
produced:
[0x100000418, 0x100000424): DW_OP_addr 0x100000790, DW_OP_stack_value
Look up the relocation of each operand's own .debug_addr slot instead,
the way a variable's single location is already handled, and fall back
to the owner's adjustment only where there is none. Do this in both the
classic and the parallel linker.
rdar://188852150
Assisted-by: Claude
[CIR][OpenMP] Add support for the OpenMP 'for' directive
This patch adds support for wsloop in ClangIR: the `for` directive and its
combined forms `parallel for` and `target parallel for`. This is lowered to an
omp.wsloop + omp.loop_nest, nested utilizing the existing queue-based
decomposition.
Assisted-by: Cursor / Claude Sonnet 5 High
Add braces to multi-line if in asm parser
Change-Id: I20476ccf51eb2901a1181aa038d15e6c1e99e8cc
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>