DAG: Fix error message casing to match style policy (#229212)
The developer guidelines suggests that diagnostic messages should start
with a lowercase letter and not end in a period
Co-authored-by: Claude <noreply at anthropic.com>
[AArch64][Win] Account for the fixed object area when classifying CSR (#222111)
In Windows frames, the fixed object area sits above the callee-saved
register area:
+---------------+
| Fixed objects |
+---------------+
| Callee-saved |
+---------------+
| Locals |
+---------------+
Include the fixed object area in the callee-saved threshold used by
`resolveFrameOffsetReference()`. Without this, objects in the
callee-saved area can be incorrectly addressed through the base
pointer in stack-realigned frames.
This became visible after #147421 moved catch objects into the fixed
[3 lines not shown]
[MLGO] Make regalloc eviction input tensor shapes runtime variables (#224598)
The column count was hardcoded to 33 for X86
It now comes from `-mlregalloc-num-allocatable-regs` (default 32, plus
one column for the candidate), compiled model whose shapes do not match
falls back to the default policy
[lldb] Use file(MAKE_DIRECTORY) to create header staging dir (#228606)
This simplifies the build graph for modifying lldb headers for
installation. It also fixes the `clean` target by not making a target
responsible for creating this directory. CMake will create it as needed
during configuration time instead of at build time.
rdar://161109746
[SPIR-V] Mark elementwise intrinsics IntrTriviallyScalarizable (#227286)
This mirrors the DirectX intrinsics
The SPIR-V legalizer will use it to split elementwise intrinsics with
illegal vector widths
Required for https://github.com/llvm/llvm-project/pull/227287
[mlir][OpenACC] Skip implicit routine marking for host-only calls (#227758)
Calls nested in a host-only branch of `acc.on_device` do not run on the
device. Do not try to attach implicit acc routine information to them.
[CI] Exclude CIR from Windows premerge testing (#229208)
After #227957, the check-clang-cir target is only defined when
CLANG_ENABLE_CIR is ON. The Windows premerge build never sets that
option (only monolithic-linux.sh receives enable_cir), but
compute_projects.py still selects check-clang-cir for CIR changes on
Windows. Every PR touching CIR now fails the Windows job before any test
runs:
ninja: error: unknown target 'check-clang-cir'
Before #227957 the target existed unconditionally, and the CIR tests
were all unsupported on Windows because CIR was disabled, so excluding
CIR there loses no coverage. It also stops building mlir for CIR-only
changes on Windows.
Enabling real CIR testing on Windows (passing enable_cir through to
monolithic-windows.sh) can be done separately.
[2 lines not shown]
[DWARFLinker] Relocate DW_OP_addrx by the delta of its own symbol (#228583)
When rewriting DW_OP_addrx or DW_OP_constx into a relocated address,
DWARFLinker applied the adjustment of whatever owned the expression.
For a location list that is the enclosing function, but the .debug_addr
entry may name a data symbol, which the linker moves by a different
amount. For a variable holding the address of _g at 0x100004008 this
produced:
[0x100000418, 0x100000424): DW_OP_addr 0x100000790, DW_OP_stack_value
Look up the relocation of each operand's own .debug_addr slot instead,
the way a variable's single location is already handled, and fall back
to the owner's adjustment only where there is none. Do this in both the
classic and the parallel linker.
rdar://188852150
Assisted-by: Claude
Add braces to multi-line if in asm parser
Change-Id: I20476ccf51eb2901a1181aa038d15e6c1e99e8cc
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[VPlan] Handle div/rem replicate recipes outside regions in cost. (#229203)
VPReplicateRecipe::computeCost unconditionally dereferenced getRegion()
for div/rem recipes. A non-single-scalar replicate recipe may be hoisted
out of the loop region into the vector preheader, where getRegion()
returns nullptr, causing a crash. Treat recipes outside any region as
unpredicated, matching the existing handling for loads and stores.
Fixes https://github.com/llvm/llvm-project/issues/229124.
Fix per-locator iterator lowering regressions
Reuse ordinary locator lowering inside depend and affinity iterators to
preserve component and vector-subscript dependence support. Keep locator
evaluation and temporary cleanup inside the iterator region so empty
selected ranges suppress evaluation. Preserve affinity section lengths.
Retain bound and step widths through widened trip-count calculation and
reconstruct induction values in their declared types. Reject iterated
dependences on unsupported target-data operations during translation.
Bind source-occurrence checks to each locator's ranges and yielded value.
Add locator, wide-range, empty-range, and target-data diagnostic coverage,
plus locators whose lowering adds control flow to the iterator region
(LEN_TRIM, allocatable function results, and SUM inlined at -O1).
[AMDGPU] Reserve ENABLE_WAVEFRONT_SIZE32 on gfx125
Wave32-only targets (gfx1250, gfx1250-strict, gfx1251, gfx12-5-generic)
reserve the kernel descriptor's ENABLE_WAVEFRONT_SIZE32; it must be 0.
The compiler set it to 1.
The wave size is selectable only on targets with both supports-wave32
and supports-wave64. Elsewhere, clear the bit, stop printing
.amdhsa_wavefront_size32, and reject the directive in the assembler.
The disassembler now rejects a set bit on every target without a
selectable wave size. This replaces the gfx9-only check and also covers
GFX6-GFX8, matching the documentation.
Document which targets select the wave size through this bit.
Change-Id: I0c3942272f03ff5aba04a15ca063b7a02e6fed06
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[mlir][flang][OpenMP] Select depend and affinity iterators per locator
Flang expands each iterator-dependent depend or affinity locator over all
ranges in the clause, including iterators absent from that locator. For
example:
depend(iterator(i=1:2, j=3:m), in: a(i))
When m < 3, the unused empty j range suppresses both required dependences
on a. This can lose task ordering. A nonempty unused range instead repeats
entries unnecessarily; affinity locators have the same expansion problem.
Select ranges according to the iterator identifiers appearing in each
locator's source, preserving declaration order. Making this change safely
also requires preserving source occurrences and handling empty ranges:
- Record resolved iterator references before using the folded expression
for lowering. In a(j/d + 0*i), folding removes 0*i, but an empty i range
must still suppress the locator and prevent evaluation of j/d.
[16 lines not shown]
DAG: Fix error message casing to match style policy
The developer guidelines suggests that diagnostic messages should start with a
lowercase letter and not end in a period
Co-authored-by: Claude <noreply at anthropic.com>
[offload][omp] Load and resolve device binaries through liboffload
Migrate DeviceTy::loadBinary and global/kernel symbol resolution off
GenericPluginTy::load_binary/get_global/get_function onto liboffload's
Program/Symbol API, encapsulated in a new ProgramTy abstraction that wraps
an ol_program_handle_t. Kernel symbol resolution still needs the plugin's
opaque GenericKernelTy* handle for the legacy launch path, obtained via a
temporary __ol_tgt_GetKernelFromSymbol helper rather than new public
liboffload API surface. Removes the now-dead __tgt_device_binary type and
the corresponding GenericPluginTy methods and exports entries.
[offload][omp] Query device info directly through liboffload
Route DeviceTy::getInfo through olGetDeviceInfo instead of the
plugin's obtain_device_info, and drop the now-unused
GenericPluginTy::obtain_device_info wrapper and its liboffload
export.
[offload][omp] Manage memory allocation through liboffload
Migrate DeviceTy::allocData/deleteData off GenericPluginTy::data_alloc/
data_delete onto liboffload's olMemAlloc*/olMemFree, migrate
targetLockExplicit/targetUnlockExplicit off data_lock/data_unlock onto
olMemRegister/olMemUnregister (fixing a latent bug where these passed
the OpenMP-visible device number instead of the plugin device id), and
migrate DeviceTy::isAccessiblePtr onto a new olMemIsAccessible API
(added with a unit test) since liboffload had no equivalent for
querying accessibility of arbitrary, not-necessarily-liboffload-
allocated pointers. Removes the now-dead GenericPluginTy::data_alloc/
data_delete/data_lock/data_unlock/is_accessible_ptr wrappers and their
exports entries.
[offload][omp] Remove data_fence
All supported backends execute enqueued work on a given queue in
submission order (CUDA, AMDGPU, and Host queues are always in-order;
Level Zero's default and non-default in-order/synchronous command
modes are as well), so the explicit data-fence used to order a
pointer-attachment after prior data transfers is unnecessary. Remove
the DeviceTy/GenericDeviceTy/GenericPluginTy dataFence chain and the
now-unused olQueueBarrier liboffload API added to support it.
[clangd] Handle template template parameter pack indexes correctly in TargetFinder::VisitTemplateSpecializationType (#228996)
The code was already handling template template parameters, but it used
TemplateName::getAsTemplateDecl() which does not work for a template
template parameter referenced inside a pack index.
Fixes https://github.com/clangd/clangd/issues/2712