LiveVariables: Remove dead live-in handling and Defs plumbing
No physical register is tracked at the start of a block, so handling
the block live-ins was a no-op. The Defs list was only appended for
instruction defs, which runOnInstr already collects.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
LiveVariables: Only visit tracked physical registers
Keep a bitvector of physical registers with a recorded def or use in
the current block. Register mask handling, the end of block scan, and
the per-block reset now only visit those registers instead of every
register. This is significant for targets with many registers, such as
AMDGPU.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
CodeGen: Strip LiveVariables down to dead flag computation
This analysis is dead and there are no more explicit uses. There are still
passes implicitly relying on adjustments of dead flags. Missing dead
flags are added, and implicit-def operands are added for partially dead
physical registers.
The whole pass should be deleted, but it's taking a while to get all the dead
flag changes through the rest of the compiler. As a stop-gap to try to recover
some compile time regression, and avoiding new users appearing, strip the pass
down to only commputing the dead flags.
The main side effect of this is kill flags are no longer made accurate, which
is the source of the test churn.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[CIR] Register a static's extended temporaries where they are built (#229896)
Temporaries whose lifetime a static extends were all destroyed together,
in one batch registered after the static's whole initializer ran. Now
each one is registered for destruction right after it is built, the way
classic codegen does it. That fixes a verifier failure when a static
extends more than one temporary, a variable's own destructor going
missing, objects destroyed in the wrong order, and a temporary destroyed
even when it was never built.
Assisted-by: Claude Code / claude-opus-5-5
[nfc][mlir][acc] Improve acc.on_device documentation (#230143)
This updates the operation documentation because as per OpenACC spec, it
is not just a runtime call, but can be evaluated at compile time. Also
add the spec section.
[CodeGen] Honor -function-splitting=none in BasicBlockSections (#227880)
With -function-splitting=none, functions with a basic block sections
profile are still laid out according to the profile, but all their basic
blocks are now emitted in a single section, instead of the profile's
clusters and a cold section.
[mlir][LLVM] Verify llvm.invoke callees like llvm.call (#229954)
Generalize CallOp's symbol-use verification into a template shared with
InvokeOp, which now implements SymbolUserOpInterface. Direct invokes get
the same checking calls have: callee resolution, operand and result
types, vararg attributes, and the debug-location rule.
This newly rejects malformed direct invokes that used to be accepted and
translate. `mlir/test/Dialect/LLVMIR/roundtrip.mlir` had such an invoke
and now uses a matching callee.
Assisted-by: Muse Code (Muse).
---------
Co-authored-by: Bruno Cardoso Lopes <bruno.cardosolopes at gmail.com>
[mlir][LLVMIR] Bound import diagnostics and share one slot tracker (#229956)
Import diagnostics render LLVM entities unbounded: a global prints its
entire initializer (a merged vtable reaches megabytes), and every
rendering builds its own ModuleSlotTracker, whose module walk makes a
per-instruction diagnostic quadratic in module size. To the scale we
link and use the LLVM IR dialect, this has bitten us many times
downstream, usually leading to OOM's and other silly behaviors from
something that isn't supported anyways. This solution has helped us
finding and fixing a lot of missing attributes, which will be
contributed soon.
Cap each rendering at 256 characters and print globals by name only,
diagnostics only need to identify the entity. Share one lazily created
tracker across the import instead: printing through it numbers exactly
like the per-call tracker, the module is never mutated during import so
slots stay valid, and each print establishes its own function context.
Thread the tracker through the module-flag converters so every
diagnostic uses the same numbering.
[5 lines not shown]
[CIR] Enable bytecode emission through the driver and cc1 (#229607)
Teach the driver to serialize CIR directly to bytecode: new EmitCIRBC
frontend action kind, `TY_CIRBC` (.cirbc) driver type, and
EmitCIRBCAction writing the in-memory module with
`mlir::writeBytecodeToFile` instead of printing text. `.cirbc` inputs
map to the CIR language like .cir.
Tests cover cc1 emission with magic-byte and read-back pins, .cir-input
and .cirbc-input round trips, and driver forwarding, phase selection,
default output name, and flag-implied pipeline selection.
I started proper bytecode encoding support in #229586, so this is a step
into using that through the driver.
[compiler-rt] Remove dlsym interceptor and support `-shared-libsan` for CSan
Summary:
Follow the UBSan offload runtime. Offload now resolves HSA through the
global scope, so the `dlsym` interceptor is no longer needed. The real
HSA entry points are still taken from the loaded HSA library rather than
`RTLD_NEXT`, since every DSO with a static runtime exports the same
wrappers and they would otherwise chain back into each other.
Build `libclang_rt.csan.so` with the offload objects folded in. The
exported HSA wrappers report failure when HSA is absent and warn when HSA
was loaded ahead of the runtime. The preinit hook moves to a separate
`csan_offload-preinit` archive for executables.
[compiler-rt] Add 'csan' library for the concurrency sanitizer
Summary:
Adds the runtime for the concurrency sanitizer, both CPU and GPU.
Fundamentally, this works using the following pseudocode:
```c
static u64 watchpoints[N]; // Hash-indexed, zero is empty.
// Emitted before the access, so we never trip on our own write.
void check_access(volatile void *addr, u32 size, u32 type) {
// Every access probes. A read conflicts only with a watched write, a
// write conflicts with either.
if (u64 *wp = find_watchpoint(addr, size, type))
consume(wp, this_pc()); // Hand our location to the owner.
if (!should_sample()) // Wave-uniform, 1-in-N chance.
return;
[17 lines not shown]
[Clang] Support `-shared-libsan` for offload CSan
Summary:
Follow the UBSan handling. The shared runtime embeds the HSA
interceptors, so `csan_offload` is only linked with static runtimes and
executables pull in `csan_offload-preinit` to initialize early.
[Clang] Add support for the `-fsanitize=concurrency` runtime
Summary:
Add the frontend sanitizer kind, function attributes, pass pipeline
integration, predefined macro, driver handling, and documentation for
ConcurrencySanitizer.
[Offload] Add OFFLOAD_FORCE_BLOCKING force-synchronization escape hatch (#222635)
- Add OF_ForceSyncOps BoolEnvar and forceSyncOps() on GenericDeviceTy
with actual spelling: OFFLOAD_FORCE_BOCKING
- Add shouldForceSync predicate and drain external async info objects in
AsyncInfoWrapperTy::finalize() when the flag is set
Assisted-by: Claude Code
[AMDGPU] Limit the fmul fusion discount to a matching context type
The fmul is free when its context instruction feeds a fusable fadd or
fsub. A vector fmul priced with a scalar lane as the context inherits
that fusion only when it is emitted lane by lane. On a packed type the
scalar user does not show that the vector fmul feeds a vector fadd, and
products extracted into a scalar fadd chain keep the packed fmul while
the fma is lost.
[NFC][SLP] Precommit tests for the phantom load saving (#228901)
The idea of what the tests show. On AMDGPU, consecutive scalar loads
coalesce in the backend, so the saving SLP counts for vectorizing them
is phantom.
Assisted-By: Claude Code Opus 5