[NFC][TSan] Make Context::nreported atomic (#228602)
Make Context::nreported an atomic_uint32_t so that Finalize and
__tsan_report_count can read it without relying on report_mtx.
Assisted-by: Gemini
[mlir] Set WrapNamespaceBodyWithEmptyLines: Never (#221421)
MLIR predominantly writes namespace bodies without a wrapping empty
line: in mlir/lib and mlir/include 35% of bodies start with an empty
line and 32% end with one.
```
namespace ns { // no empty line below
struct A {
};
} // namespace ns // no empty live above
```
New code drifts the other way, partly related to LLM preference. Pin the
majority style using a clang-format 20 feature to stop further damage.
[flang] Add a pass-pipeline config callback hook for plugins
The HLFIR-to-FIR pipeline extension points live on MLIRToLLVMPassPipelineConfig,
which the frontend builds as a local of CodeGenAction, out of reach of a plugin.
Add a process-global registry of callbacks that run on the config before the
pipeline is built. A plugin registers one from a static initializer, so it is in
place before any compilation begins, as FrontendPluginRegistry does for plugin
actions. Both code generation entry points invoke the callbacks, lowerHLFIRToFIR
for -emit-fir and generateLLVMIR for -emit-llvm/-emit-obj, and are mutually
exclusive for a given compilation, so a plugin sees the same behaviour whichever
output was asked for.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
[flang] Address review comments on the pipeline config callbacks
- Adopt the suggested doc comment wording and name the callback type
(PassPipelineConfigCallback) instead of spelling std::function / auto.
- Run the callbacks in generateLLVMIR once the config is fully set up, so a
plugin sees the OpenMP/OpenACC, vscale and other settings as on -emit-fir.
- Document that -emit-fir only reaches the HLFIR extension points, that
callbacks must not register callbacks, and that bbc/tco/fir-opt do not run
the registry.
- Tests: drop the redundant epRan flag and use a plain global recorder.
Assisted-by: Claude Code (Opus 5.5)
[mlir][bufferization] Fix placement walk crashes and invalid captures (#228050)
`findPlacementBlock`, which is shared by `buffer-hoisting` and
`buffer-loop-hoisting`, can move an allocation outside the operation
used to build the liveness analysis. For example, when
`transform.bufferization.buffer_loop_hoisting` targets a `scf.while`,
moving the allocation into the enclosing function block causes the
liveness lookup to crash.
The walk can also cross an `IsIsolatedFromAbove` operation and introduce
an illegal capture, or reach an unreachable enclosing block and
dereference a missing dominator node. The initial reachability check on
the allocation's block doesn't cover that second case.
This PR keeps allocations inside the analysis target and isolated
regions, and stops the walk before querying a dominator node for an
unreachable block. Movement between blocks within the same region
remains allowed.
[15 lines not shown]
[Dexter] Add timeout and finishing guard for startup
A process can be terminated without start. In that case, the DAP.py
will spin over there and causes test loops forever. Add a timeout and
finish guard in here to at least stop the test.
[RISCV] Don't truncate MaxAlignment to int when realigning the stack (#228225)
emitPrologue casts MaxAlignment to int before checking whether the mask
fits in an ANDI immediate. With an alloca of align 4294967296 the cast
gives 0, the isInt<12> check passes, and we emit `andi sp, sp, 0`.
Casting to int64_t instead makes this case use the srli/slli sequence.
This is the patch Mitch Briles posted on the issue, but with int64_t
instead of long, since long is only 32 bits on Windows hosts.
The test is RV64 only. On RV32 this alignment already fails with "Should
only materialize 32-bit constants for RV32" while setting up the frame
(before and after this change), so it couldn't go into
stack-realignment.ll alongside the RV32 run lines.
Fixes #228084
Co-authored-by: Mitch Briles <mitchbriles at gmail.com>
[ORC] Replace JIT-dispatch handlers with call-controller handlers (#228628)
This is the executor-to-controller counterpart to Proxy: where Proxy
provides a uniform way to call functions in the executor,
call-controller handlers provide a uniform way for the executor to call
handlers in the controller.
Handlers are now registered as CallControllerHandlerBindings, which
decouples ExecutionSession's handler registration from
SimplePackedSerialization. The new bindCallControllerHandlerSPS utility
makes it easy to introduce a new SPS handler with a one-liner, using
either a lambda:
ES.registerCallControllerHandlers(
JD, bindCallControllerHandlerSPS<int32_t(int32_t, int32_t)>(
SymbolNameSpec::c("add_tag"),
[](unique_function<void(int32_t)> Return,
int32_t X, int32_t Y) {
Return(X + Y);
[39 lines not shown]
[Clang][HIP] Skip internalization for non-LTO device links (#225859)
A non-LTO HIP device link can combine native relocatable objects with
bitcode libraries. This can happen with both RDC and non-RDC
compilation. LLD must resolve undefined symbols in the native objects
with definitions from the bitcode.
The LTO step processing the bitcode cannot see references from native
objects. Internalizing non-kernel functions can therefore hide
definitions that those objects need. Do not enable AMDGPU
internalization for these mixed non-LTO links.
LTO mode does not have this problem. It links the bitcode modules before
internalization, so the references are visible when symbols are
resolved. Keep internalization for LTO links and when compiling non-RDC
main modules.
[Clang][HIP] Skip internalization for non-LTO device links (#225859)
A non-LTO HIP device link can combine native relocatable objects with
bitcode libraries. This can happen with both RDC and non-RDC
compilation. LLD must resolve undefined symbols in the native objects
with definitions from the bitcode.
The LTO step processing the bitcode cannot see references from native
objects. Internalizing non-kernel functions can therefore hide
definitions that those objects need. Do not enable AMDGPU
internalization for these mixed non-LTO links.
LTO mode does not have this problem. It links the bitcode modules before
internalization, so the references are visible when symbols are
resolved. Keep internalization for LTO links and when compiling non-RDC
main modules.
[mlir][scf] Fix scf.for value bounds for empty and unsigned loops (#226761)
Value bounds assumes that an `scf.for` result is `init + ceildiv(ub -
lb, step) * (yield - iter_arg)`. This is only true if the loop runs at
least once. When the loop does not run, `ceildiv(ub - lb, step)` is
negative instead of 0. For unsigned loops, `ub - lb` is computed as if
the bounds were signed, which also gives a wrong trip count. The
induction variable bounds of unsigned loops have the same signed
problem.
Example:
```mlir
%r = scf.for %iv = %c5 to %c2 step %c1 iter_args(%arg = %c7) -> index {
%n = arith.addi %arg, %c1 : index
scf.yield %n : index
}
%m = affine.min affine_map<(d0) -> (d0, 5)>(%r)
```
The loop does not run, so `%r` is 7 and `%m` should be 5. Value bounds
[13 lines not shown]
[orc-rt] Use the Error matchers in NativeDylibManagerSPSCITest (#228629)
Use the Error matchers introduced in 4c8a437d0487 to clean up error
checks in NativeDylibManagerSPSCITest.
[CodeGen] Avoid sinking EH pad blocks (#224812)
EH pad blocks should not be sinkable since the control flow comes from
the unwinder instead of the predecessor.
This exclusion also matches the behavior of other sinkers.
Fixes #224802.
[lldb] Don't run code for CanJIT() while creating the ObjC runtime (#227555)
When libobjc lacks class_getMethodImplementation, the ObjC trampoline
handler calls CanJIT() to decide whether to warn. Without the _M
packet, CanJIT() allocated memory by calling mmap in the inferior.
Setting up that call unwinds the stack, which asks for the ObjC
runtime. The runtime isn't registered yet, so LLDB creates another one
and recurses until the stack overflows.
Ask the process plugin whether it can allocate memory instead of
trying it. ProcessGDBRemote probes the _M packet and otherwise checks
for an mmap symbol, neither of which runs code in the inferior.
rdar://188335027
[lldb][FreeBSDKernel] Find .debug files using pseudo sysroot (#224845)
Unlike userspace processes, it is common to obtain kernel dump from a
machine (e.g. QEMU) which different from the machine debugging the dump.
In this case users need to run `target symbols add foo.debug` for the
kernel and each kernel object, which becomes quite inconvenient when the
machine had dozens of kernel modules loaded during the dump.
This patch adds functionality to the dynamic loader so that it loads
symbol files automatically when kernel or kenrel modules are loaded. It
assumes a pseudo sysroot. When kernel is located at
`/foo/boot/kernel/kernel`, it will look for
`foo/usr/lib/debug/boot/kernel/kernel.debug` and same for kernel
modules. In this case, the only thing users need to do is copying the
dumped machine's `/usr/lib/debug/boot` relative to `/boot` that is being
debugged.
Assisted-by: GPT
[CASPlugin] Move CASPluginTest next to the CAS unit tests (#227546)
Set up the CAS test plugin like CGTestPlugin: move it from
llvm/tools/libCASPluginTest to llvm/unittests/CAS/CASPluginTest and
build it as CASPluginTest${LLVM_PLUGIN_EXT}, without the lib prefix or
a version, so its name is the same on all platforms. The lit tests now
find it via %llvmshlibdir/CASPluginTest%pluginext instead of a
configured LLVM_CAS_PLUGIN_TEST_PATH, and CASTests gets its path from
the build system through the CAS_PLUGIN_PATH definition.
[llvm-objcopy][MachO] Fix use-after-free when stripping (#228607)
We must preserve symbols referenced by the indirect symbol table even
with strip all, and we must preserve symbols referenced by relocations
without strip all.
Fixes: https://github.com/llvm/llvm-project/issues/228595
Assisted-by: codex