[llvm-profdata] Remove exitWithError and LSan leak workaround
With all subcommands propagating llvm::Error to main, exitWithError,
exitWithErrorCode, and the LSan leak suppression workaround are no
longer needed.
Assisted-by: Gemini
[NFCI][llvm-profdata] Propagate Error in loadInput and mergeWriterContexts
Propagate Error from loadInput and mergeWriterContexts in mergeInstrProfile,
supplementInstrProfile, and overlapInstrProfile. In mergeInstrProfile's
ThreadPool, catch errors from worker threads, stop scheduling new jobs,
and return the first encountered fatal error.
Ensure ~WriterContext() consumes any pending unhandled errors in
WriterContext::Errors upon destruction.
Not NFC as destructors are run on the stack and ThreadPool workers exit
earlier on error.
Assisted-by: Gemini
[NFCI][llvm-profdata] Propagate Error in merge subcommand
Change merge_main and its helpers to return Error and handle it
with reportError in main.
Not NFC as destructors are run on the stack.
Assisted-by: Gemini
[NFCI][llvm-profdata] Propagate Error in show subcommand
Change show_main and its helpers to return Error and handle it
with reportError in main.
Not NFC as destructors are run on the stack.
Assisted-by: Gemini
[NFCI][llvm-profdata] Propagate Error in overlap subcommand
Change overlap_main and its helpers to return Error and handle it
with reportError in main.
Not NFC as destructors are run on the stack.
Assisted-by: Gemini
[NFCI][llvm-profdata] Propagate Error in order subcommand
Change order_main to return Error and handle it with reportError
in main.
Not NFC as destructors are run on the stack.
Assisted-by: Gemini
[NFC][llvm-profdata] Introduce ProfdataError, makeError, and reportError (#228153)
Define ProfdataError, makeError, and reportError, and adapt
exitWithError and exitWithErrorCode to use reportError(makeError(...)).
This establishes the error infrastructure for propagating Error
up to main across incremental commits.
Assisted-by: Gemini
[clang][Serialization] Fix latent bug in adjustFilenameForRelocatableAST (#227910)
This patch fixes a latent bug in `adjustFilenameForRelocatableAST`. It
relies on a null terminator, but `PreparePathForOutput` passes in a
`StringRef::data()` that has no null-terminator. Today this is benign as
nothing passes in a path that would hit the attempt at checking for a
null terminator, but future patches will.
The fix both makes the code safer by removing the C string usage, and
handles passing in the base path itself by representing it as `.` and
resolving that back to the (possibly relocated) base path on read.
No test on the writer part as it's not triggerable via an integration or
unit test.
Assisted-by: Claude Code: opus-5.5
[lit] Preload ASan runtime for ld64 tests on arm64 too (#227873)
get_asan_rtlib() only preloaded the runtime on x86 hosts. On arm64 this
silently skipped the preload, so ASan-instrumented libLTO.dylib aborts
when ld64 dlopen()s it.
rdar://188411900
llc: Verify MIR outputs by default
MIR is validated by the machine verifier on read, but by default wasn't
validated on output. This differs from opt, which runs the verifier on
output unless explicitly disabled. -verify-machineinstrs is frequently used as
a much more expensive way of getting the verifier run, since that runs between
every pass.
Run the machine verifier at the end of the pipeline whenever it stops before code
emission. Add -disable-mir-output-verify to suppress this. It is separate from
-disable-verify, which still only controls verification of the IR input. The
extra verifier is skipped when the verifier already runs after every machine
pass, with -verify-machineinstrs in the legacy pass manager or -verify-each in
the new pass manager.
Some tests had to force disabling the verifier in a few tests which already fail
the verifier.
Unlike opt, the driver can't simply verify after the pass manager finishes.
[5 lines not shown]
[libc] Implement scandir and its unit tests (#223198)
scandir is a POSIX function from <dirent.h> header that scans the directory and returns all files
contained within, optionally applying user-provided filtering and comparator functions.
Provide implementation which allows substituting the `Dir` class to allow for dependency
injection to cover various error cases.
[flang][cuda] Resolve scalar constant address to device copy for reads (#227802)
Example test.cuf:
```
module m
integer, constant :: int_0_d = 2
end module
program test
use m
integer :: e(10)
e = 0
e = int_0_d
print *, e ! expect 2 (x10)
end program
```
During lowering, a data transfer from device to host is generated
because reads from a constant scalar consults the device copy. But the
address passed to the cudaMemcpy is the host shadow address. Depending
[11 lines not shown]
[mlir] Migrate AMDGPU/ROCDL to targets, not chipset versions
**migration tl;dr:** Replace usages of `amdgpu::Chipset` with `ROCDL::TargetInfo`, ideally move from `chipset=` to `arch=`. If you don't use upstream pipelines, call 'TargetInfo::migrateArchFeaturesToModuleFlags` at the appropriate location.
Further note: if you've got a build pipeline that's getting a `gfxXXX` name from something like `rocm_agent_enumerator`, using a full triple name like the ones you get from `rocminfo` is preferred.
`amdgpu::Chipset` was an awkward hack that was hard to keep up to date
with changes in the compiler/new architectures, and didn't properly
support generic targets (and has been strongly disfavored by the
compiler team).
This PR replaces `amdgpu::Chipset` with `ROCDL::TargetInfo`, a
structure that uses LLVM's TargetParser and the underlying LLVM
features tables to get the real nature of the target being compiled
for.
This also helps MLIR move to
new-style (`-mtriple=amdgpuX.YZ-amd-amdhsa`) over "old
style" (`-mtriple=amdgcn-amd-amdhsa -mcpu=gfxXYZ`) triples.
[36 lines not shown]
[mlir][ROCDL] Add TargetInfo to replace Chipset, allow features queries (#223562)
Add a new ROCDL::TargetInfo struct that parses AMDGPU triples and
target names using the same logic that Clang and LLVM
use (TargetParser) and maintains the set of features available on a
given GPU.
This is an improvement over the old `amdgpu::Chipset` struct since
that was just a version number and often became stale compared to the
knowledge exposed by LLVM, such as gfx1170 having OCP FP8 support even
though other gfx11 chips don't have it.
This struct also allows for moving to new-style
triples (amdgpu9.42-amd-amdhsa vs amdgcn-amd-amdhsa--gfx942, for
example), which is an ongoing migration in other parts of the compiler
that this PR lets us follow.
It also enables compiling for generic targets, like `gfx11-generic`,
which can be run on all chips in a generation.
[14 lines not shown]
[AMDGPU] Canonicalize num_records to its actual width in InstCombine (#217068)
This PR adds code to InstCombineIntrinsic to change the width of the
num_recods field (by extension or truncation) to the correct width for
the target triple (if a concrete enough target triple has been set) so
that LLVM IR-level optimizations can see the lack of demand for the
high bits, for example.
Assisted by Claude, which also found those buffer lowering edge cases
Docs: remove gdb startup text from WritingAnLLVMPass. (#228236)
It is simply noise for readers of the doc.
Also, it is (incorrectly) flagged by license-compliance scanners as
indicating GPL-licensed content. Simplest to just drop it.