[lld][MachO] Add `--warn-missing-subsections-via-symbols` flag (#221464)
This change introduces:
* `--warn-missing-subsections-via-symbols`: Warns when an input object
file with non-empty sections is missing `MH_SUBSECTIONS_VIA_SYMBOLS`.
* `--no-warn-missing-subsections-via-symbols`: Disables the warning
(default).
Also adds documentation and lit tests.
Object files missing `MH_SUBSECTIONS_VIA_SYMBOLS` can prevent
dead-stripping and subsection splitting.
We have seen a number of cases where we are compiling assembly, and this
allows us to track them down.
[mlir] Outline registered operation-model allocation (NFC) (#222803)
Allocate operation-model storage through an out-of-line helper so every
registration TU does not instantiate the allocation machinery.
Across three controlled -j16 rebuilds of 44 registration TUs, median
instructions fell from 2.284T to 2.202T (-3.594%) and median wall time
fell from 40.44s to 38.54s (-4.698%).
Assisted-by: Codex
[mlir][AMDGPU] Take an `arch` target ID instead of triple/chip/features
`features` was a general `-mattr` string, which needed a general feature
parser and let callers ask for arbitrary combinations we have no interest
in supporting. In practice the only things anyone sets are the wavefront
size and the xnack/sramecc settings that come off a device query.
Replace `triple`, `chip` and `features` with a single `arch` option that
names the target the way Clang does, parsed by `llvm::AMDGPU::TargetID`
rather than by hand. It accepts
- a processor, with optional target-ID modifiers: `gfx942`,
`gfx942:xnack+`, `gfx9-4-generic`;
- a triple: `amdgpu9.42-amd-amdhsa`;
- a full target ID: `amdcgn-amd-amdhsa--gfx90a:sramecc+:xnack-`, which
is what `rocminfo` prints for a device's ISA, so that output can be
pasted straight in.
Since `chipset=gfx942` becomes `arch=gfx942`, migration is a rename.
[22 lines not shown]
[mlir][ROCDL] Carry `arch`'s xnack/sramecc onto the module
`rocdl-attach-target` rejected a target ID that pinned xnack or sramecc,
because `#rocdl.target` feeds a TargetMachine and the backend no longer
accepts those two as subtarget features. Now that the module attributes
exist, migrate them instead of refusing: `TargetInfo` gains
`migrateArchFeaturesToModuleFlags`, which records the settings the target
ID pinned as `rocdl.xnack` / `rocdl.sramecc` on a module, and
`rocdl-attach-target` calls it on each module it attaches to.
A setting the target ID leaves open, or that the GPU does not support, is
left alone rather than written as false: an absent flag means "either",
so writing false would be a different request. That also means an
attribute already on the module survives an `arch` that says nothing
about the feature, while an `arch` that does pin it wins as the more
specific request.
[mlir][AMDGPU] Keep `chipset` as a deprecated alias for `arch`
Renaming the option meant every existing invocation of these passes had
to be updated in lockstep. Accept the old spelling instead: `chipset` on
`convert-amdgpu-to-rocdl`, `convert-gpu-to-rocdl`, `convert-arith-to-amdgpu`,
`convert-math-to-rocdl` and `amdgpu-emulate-atomics`, and `chip` on
`gpu-lower-to-rocdl-pipeline`, which is what each of them was called
before the rename.
`arch` wins whenever it names a target; the alias is consulted only when
`arch` is still at the sentinel that means "no target given", so with
neither given the error still names the unusable default rather than an
empty string, and a stale alias value is reported as itself.
[mlir] Migrate AMDGPU/ROCDL to targets, not chipset versions
**migration tl;dr:** `chipset=` becomes `triple=`, migrate off of
`amdgpu::Chipset` to `ROCDL::TargetInfo`, and eventually change
`gfxXYZ` to `amdgpuX.YZ-amd-amdhsa` in that `triple` argument.
`amdgpu::Chipset` was an awkward hack that was hard to keep up to date
with changes in the compiler/new architectures, and didn't properly
support generic targets (and has been strongly disfavored by the
compiler team).
This PR replaces `amdgpu::Chipset` with `ROCDL::TargetInfo`, a
structure that uses LLVM's TargetParser and the underlying LLVM
features tables to get the real nature of the target being compiled
for.
This also helps MLIR move to
new-style (`-mtriple=amdgpuX.YZ-amd-amdhsa`) over "old
style" (`-mtriple=amdgcn-amd-amdhsa -mcpu=gfxXYZ`) triples.
[40 lines not shown]
[mlir][AMDGPU][NFC] Pre-commit tests for incorrect version checks
There'll be a refactoring from `amdgpu::Chipset` to
`ROCDL::TargetInfo`, thus also moving from chip version checks to
features checks. This commit adds tests for incorrect lowerings that
were allowed by the current code.
- gfx90c is >= gfx90a but stil needs atomic emulation (it doesn't
have buffer fmax and so on).
- gfx90c is also >= gfx90a but has no barrier back-off, so it needs
the inline asm workaround around `s_barrier` that it isn't getting
- gfx908 doesn't have a packed fp16 atomic add but we thought it did
- gfx950 is mistakenly allowing xf32 MFMAs
- gfx1200 is allowing permlane_swap instructions that it doesn't have
- gfx11.7 should be allowing OCP FP8 conversions but isn't on the list
This also cleans up some redundant tests with a --check-prefixes
AI disclosure: Claude found these and wrote the tests.
[2 lines not shown]
[mlir][ROCDL] Add `rocdl.xnack` and `rocdl.sramecc` module attributes
Since 27eeb7370281, the AMDGPU backend takes the xnack
and sramecc target-ID settings from the `amdgpu.xnack` and
`amdgpu.sramecc` module flags instead subtarget features, making the
old usage a hard error.
This commit adds `rocdl.xnack` and `rocdl.sramecc` module attributes
to the discardable attribute list the ROCDL dialect defines in order
to represent these flags and adds translations for them.
Omitting them means to leave these modifiers at
their default "either" state, which isn't the same as setting them to
false.
AI disclosure: Claude wrote this code and I reviewed it and tried to
reword the comments to something better.
[AMDGPU] Expose buffer resource num_records width in TargetParser
This also fixes the conflict in gfx12.
Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
[AMDGPU][NFC] Account for the LDS bank count column in the GPU table test
AMDGPUTargetDefSubArchSpelling.td spells out every column of the emitted
GPUInfo rows, so it has to be updated whenever one is added. The
num_records width and getLDSBankCount landed independently, and each
CHECK line only grew by one, leaving them a column short.
Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
[mlir] Use pointers for static operation hooks (NFC)
Return raw function pointers from static operation-hook accessors to avoid
instantiating unique-function construction machinery for every operation.
Across three controlled -j16 rebuilds of 44 registration TUs, median
instructions fell from 2.284T to 2.201T (-3.646%) and median wall time fell
from 41.04s to 38.80s (-5.458%).
Assisted-by: Codex
[LLVMABI] Add support for SVE types in the LLVM ABI library (#221375)
This change extends the LLVM ABI library's VectorType to be able to
describe SVE types and updates Clang's QualTypeMapper to map them.
The AArch64 target info class in the ABI library is still a work in
progress. It will continue to report "not yet implemented" for function
signatures involving SVE types. This change is a neceasary prerequisite
for correctly handling them or correctly deferring handling of these
specific types.
Assisted-by: Cursor / claude-opus-5
[mlir] Outline registered operation-model allocation (NFC)
Allocate operation-model storage through an out-of-line helper so every
registration TU does not instantiate the allocation machinery.
Across three controlled -j16 rebuilds of 44 registration TUs, median
instructions fell from 2.284T to 2.202T (-3.594%) and median wall time fell
from 40.44s to 38.54s (-4.698%).
Assisted-by: Codex
[mlir][MemRef] Split the dialect declaration to improve build time (NFC) (#222801)
Introduce a self-contained dialect declaration and use it in thirteen
consumers that need registration but not the generated operation
umbrella.
Across three controlled rebuilds of the affected TUs, median
instructions fell from 296.237B to 278.383B (-6.027%) and median wall
time fell from 47.57s to 44.50s (-6.454%).
Assisted-by: Codex
[mlir][SCF] Split the dialect declaration (NFC)
Introduce a self-contained dialect declaration and use it in nine configured
consumers that do not need the generated operation umbrella.
Across three controlled -j16 rebuilds of the affected TUs, median instructions
fell from 162.029B to 150.032B (-7.404%) and median wall time fell from 8.82s
to 8.73s (-1.020%).
Assisted-by: Codex
[mlir] Outline registered operation-model allocation (NFC)
Allocate operation-model storage through an out-of-line owner so every
registration TU does not instantiate the allocation and deletion machinery.
Across three controlled -j16 rebuilds of 44 registration TUs, median
instructions fell from 2.284T to 2.203T (-3.519%) and median wall time fell
from 40.60s to 38.53s (-5.099%).
Assisted-by: Codex
[mlir][Vector] Split the dialect declaration to improve build time (NFC) (#222799)
Introduce a self-contained dialect declaration and use it in ten
consumers that need registration but not the generated operation
umbrella.
Across three controlled rebuilds of the affected TUs, median
instructions fell from 209.277B to 192.946B (-7.804%) and median wall
time fell from 31.94s to 29.26s (-8.391%).
Assisted-by: Codex
[mlir][GPU] Split the dialect declaration to improve build time (NFC) (#222798)
Introduce a self-contained dialect declaration and use it in two
consumers that need registration but not the generated operation
umbrella.
Across three controlled rebuilds of the affected TUs, median
instructions fell from 82.473B to 75.865B (-8.012%) and median wall time
fell from 12.51s to 11.34s (-9.353%).
Assisted-by: Codex
[mlir][ROCDL] Split the dialect declaration to improve build time (NFC) (#222794)
Introduce a self-contained dialect declaration and use it in three
consumers that need registration but not the generated operation
umbrella.
Across three controlled rebuilds of the affected TUs, median
instructions fell from 129.372B to 111.587B (-13.748%) and median wall
time fell from 20.62s to 17.72s (-14.064%).
Assisted-by: Codex
[mlir][MemRef] Split the dialect declaration (NFC)
Introduce a self-contained dialect declaration and use it in thirteen consumers
that need registration but not the generated operation umbrella.
Across three controlled rebuilds of the affected TUs, median instructions fell
from 296.237B to 278.383B (-6.027%) and median wall time fell from 47.57s to
44.50s (-6.454%).
Assisted-by: Codex
[mlir][Linalg] Split the dialect declaration to improve build time (NFC) (#222795)
Introduce a self-contained dialect declaration and use it in five
consumers that need registration but not the generated operation
umbrella.
Across three controlled rebuilds of the affected TUs, median
instructions fell from 130.162B to 114.154B (-12.298%) and median wall
time fell from 20.30s to 17.85s (-12.069%).
Assisted-by: Codex
[mlir][Affine] Split out the dialect declaration to improve build time (NFC) (#222766)
Move AffineDialect to a self-contained declaration header and narrow 23
callers that only register or reference the dialect class.
Median instructions fell from 700.601B to 681.315B (-2.753%) and wall
time from 107.36s to 104.17s (-2.971%).
Assisted-by: Codex
[mlir][Vector] Split the dialect declaration (NFC)
Introduce a self-contained dialect declaration and use it in ten consumers
that need registration but not the generated operation umbrella.
Across three controlled rebuilds of the affected TUs, median instructions fell
from 209.277B to 192.946B (-7.804%) and median wall time fell from 31.94s to
29.26s (-8.391%).
Assisted-by: Codex