Refactor subreg spilling logic.
Avoiding Lanebitemask manipulations.
Moved most of generic the calculations to use TargetRegisterInfo APIs.
Added a new API to get the covering subreg index given a lanebitmask.
[clang][bytecode] Reject matrix lvalue-to-rvalue casts in non-HLSL (#227305)
To fix test/CodeGen/AArch64/abi-classify-return-types.c.
The code path in the existing tests in
`test/SemaHLSL/Types/BuiltinMatrix/MatrixConstantExpr.hlsl` don't go
though `CheckLiteralType()`, so aren't rejected.
[AMDGPU] Test precommit for subreg reload
This test currently fails due to insufficient
registers during allocation. Once the subreg
reload is implemented, it will begin to pass
as the partial reload help mitigate register
pressure.
[InlineSpiller][AMDGPU] Implement subreg reload during RA spill
Currently, when a virtual register is partially used, the
entire tuple is restored from the spilled location, even if
only a subset of its sub-registers is needed. This patch
introduces support for partial reloads by analyzing actual
register usage and restoring only the required sub-registers.
This improvement enhances register allocation efficiency,
particularly for cases involving tuple virtual registers.
For AMDGPU, this change brings considerable improvements
in workloads that involve matrix operations, large vectors,
and complex control flows.
[MLIR][LLVMIR] Restore constant folding for global initializer GEPs
Before #226904, a GEP in a global initializer went through
IRBuilder::CreateGEP, and MLIR's IRBuilder<TargetFolder> folded the result
with ConstantFoldConstant. That combines nested GEPs, folds null and integer
bases, and infers inbounds and nuw when the offset stays within the global.
Building the constant expression directly skipped the folder, so the output
lost those flags.
Fold the constant again. This also applies to inrange GEPs, which used to
bypass the folder; the only visible difference there is that an inbounds
GEP with a non-negative offset now also gets nuw.
CIR's vtable, VTT and constant pointer tests check for the inferred flags
(for example CIR/CodeGen/vtt.cpp) and have failed since #226904.
Assisted-by: Claude Code (Claude Fable 5.1).
[MLIR][Remark] Make remark reporting thread-safe and order final remarks by source position
Passes nested under the multithreaded pass manager report into the same
RemarkEngine from worker threads, but nothing in the engine or the
policies was synchronized. RemarkEmittingPolicyFinal inserted into its
map from several threads at once, and under RemarkEmittingPolicyAll the
streamer, including the LLVM remark serializer, was called from several
threads at once. ThreadSanitizer reports both races in the new unit tests.
RemarkEngine now holds a lock around every call into the policy. It is a
recursive llvm::sys::SmartMutex, the same type as the DiagnosticEngine's
lock. report() takes it, and so does a new finalizePolicy(), which the
engine destructor and mlir-opt now use instead of calling finalize() on
the policy directly. For the All policy the streamer and the diagnostic
printer also run under the lock, so custom policies and streamers need no
lock of their own.
With the race fixed, the final policy's creation order is still not
deterministic: which thread reports first depends on scheduling.
[16 lines not shown]
[clang][bytecode] Fix initializer-list edge cases (#227587)
Namely, initializer lists for primitive types and discarding the result
of an initializer list.
[MLIR][Remark] Make remark reporting thread-safe and order final remarks by source position
Passes nested under the multithreaded pass manager report into the same
RemarkEngine from worker threads, but nothing in the engine or the
policies was synchronized. RemarkEmittingPolicyFinal inserted into its
map from several threads at once, and under RemarkEmittingPolicyAll the
streamer, including the LLVM remark serializer, was called from several
threads at once. ThreadSanitizer reports both races in the new unit tests.
RemarkEngine now holds a lock around every call into the policy. It is a
recursive llvm::sys::SmartMutex, the same type as the DiagnosticEngine's
lock. report() takes it, and so does a new finalizePolicy(), which the
engine destructor and mlir-opt now use instead of calling finalize() on
the policy directly. For the All policy the streamer and the diagnostic
printer also run under the lock, so custom policies and streamers need no
lock of their own.
With the race fixed, the final policy's creation order is still not
deterministic: which thread reports first depends on scheduling.
[16 lines not shown]
[MLIR][Remark] Emit final-policy remarks in deterministic order
RemarkEmittingPolicyFinal stores remarks in a DenseSet whose hash covers the
location pointer and the hash seed, so finalize() emitted them in bucket
order, which depends on the build and on where things landed in memory. That
is why mlir/test/Pass/remark-final.mlir used CHECK-DAG.
The engine assigns every remark a RemarkId from a monotonic counter when it is
created, and the set already keeps the newer of two remarks with the same
identity. finalize() now sorts the drained remarks by that ID before emitting,
so remarks come out in creation order and a replaced identity takes the
position of its last report. Linked remarks still follow their parent.
Order only. The identity, DenseMapInfo<Remark> and the header are unchanged.
Only remarks handed to the policy outside the engine have no ID; the unit
tests that do so assert unordered or single results.
Assisted-by: Claude Code (Claude Fable 5.1)
[lldb][Windows] Lock the thread list when lldb-server accesses it (#226974)
`NativeProcessProtocol::Threads()` returns a `LockingAdaptedIterable`
that holds `m_threads_mutex` for the lifetime of the iteration, so
callers may hold references into `m_threads` across the loop body. On
Windows, this is only done when removing threads:
```
// OnCreateThread
m_threads.push_back(std::move(thread));
...
// OnExitThread
std::lock_guard<std::recursive_mutex> guard(m_threads_mutex);
llvm::erase_if(m_threads, ...);
```
`m_threads` is a `std::vector<std::unique_ptr<NativeThreadProtocol>>`.
When the debug-event thread appends and the vector reallocates, every
reference held by a concurrent iteration on the main loop dangles. This
[17 lines not shown]
Refactor subreg spilling logic.
Avoiding Lanebitemask manipulations.
Moved most of generic the calculations to use TargetRegisterInfo APIs.
Added a new API to get the covering subreg index given a lanebitmask.
[Remarks] Fix YAML remark round-trip for values that need escaping
The serializer writes argument values with more than one newline as
literal block scalars. A block scalar has no escapes, so a value that
also contained a control character other than tab or newline was
written with the raw byte, which strict YAML readers reject. Such
values now use the double-quoted form.
The parser took the raw scalar text and stripped only single quotes, so
double-quoted values came back with their quotes and escapes, and ''
inside single quotes was not unescaped. Decode scalars with
ScalarNode::getValue instead and report escape errors. Unescaped values
and block scalar values are copied into storage owned by the parser.
Block scalar values used to point into the YAML document, which next()
frees before returning the remark.
Assisted-by: Claude
[MLIR][LLVMIR] Restore constant folding for global initializer GEPs
Before #226904, a GEP in a global initializer went through
IRBuilder::CreateGEP, and MLIR's IRBuilder<TargetFolder> folded the result
with ConstantFoldConstant. That combines nested GEPs, folds null and integer
bases, and infers inbounds and nuw when the offset stays within the global.
Building the constant expression directly skipped the folder, so the output
lost those flags.
Fold the constant again for the non-inrange case. The inrange path never
went through the folder and is unchanged.
CIR's vtable, VTT and constant pointer tests check for the inferred flags
(for example CIR/CodeGen/vtt.cpp) and have failed since #226904.
Assisted-by: Claude Code (Claude Fable 5.1).
[lldb][Windows] Strip the extended-length prefix from host and process paths (#227373)
This patch introduces a helper function to strip the extended-length
path prefix from windows paths.
ARM: Preserve the dead flag when activating the optional CPSR def
This did not preserve the original dead flag, so it would be recomputed
later by LiveVariables or RegAllocFast.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[clangd] Extract to Function: mark unmodified captured parameters const
Every captured variable was always passed by non-const reference, even when the extracted code never modifies it, resulting in misleading function signatures.
Determine, for each captured variable, whether it's ever (possibly) mutated within the extraction zone, and add `const` to the parameter's type when it isn't. Still passed by reference either way, to avoid a copy.
The check is folded directly into the existing zone traversal (`captureZoneInfo`'s `ExtractionZoneVisitor`), so its cost stays proportional to the size of the code being extracted. Direct mutations (assignment, increment/decrement, non-const method calls, explicit casts to non-const reference, non-const-reference call arguments, ...) are recognized precisely from each occurrence's immediate syntactic context. Anything that aliases a captured variable (a reference bound to it, its address taken, capture by reference in a lambda, a forwarding-reference call argument, a non-const-ref range-for loop variable) is conservatively treated as a possible mutation, without tracing whether the alias itself is later mutated -- trading a little precision in alias-heavy code for a simple, linear-cost check.
Array-typed captures are never made const, since array-to-pointer decay and array-element mutation have enough edge cases that it wasn't worth special-casing for a rare pattern.
Assisted-by: Claude
[Remarks] Escape control characters in multi-line YAML remark arguments
Argument values with more than one newline are written as YAML literal
block scalars. A block scalar has no escapes, so a value that also
contains a control character (other than tab and newline) was written
with the raw byte, which strict YAML readers reject. Use the
double-quoted form for such values instead.
Assisted-by: Claude
[lldb] Share the address lookup of thread until and StepOverUntil (#226978)
`thread until` and `SBThread::StepOverUntil` turned their targets into
addresses using different code. This commit introduces a helper function
`GetStepUntilAddresses` to unify behavior and the error message.
It resolves lines as `thread until` did: a line without line table
entries resolves to the nearest following line with entries. For
StepOverUntil, a line that is only the call site of an inlined function
now resolves that way too.
The wording of errors is that of `thread until`.
[flang][codegen] Report a shape or slice cg-rewrite cannot read
cg-rewrite folds a fir.shape, fir.shape_shift, fir.shift or fir.slice
into the code-gen form by reading it through its defining op. A value
that has none cannot be folded. For a slice this went unreported: the
rewrite dropped it and produced a descriptor for the whole array rather
than the section it names. For a shape it reached a cast on a null
defining op.
Report it instead, and say which operand. Rebuilding the value covers a
block argument, but not every case: a slice chosen by an arith.select
has no single value to take apart.
[flang][codegen] Rebuild compile-time-only block arguments in cg-rewrite
A fir.shape, fir.shape_shift, fir.shift or fir.slice only describes an
array at compile time. cg-rewrite reads one through its defining op and
folds it into the code-gen form, so none of these types has an LLVM
lowering and none is expected to reach codegen.
A pass that merges two blocks differing only in such a value passes it
as a block argument instead, and a block argument has no defining op.
The fold then finds nothing to read: a slice is dropped, leaving a
descriptor for the whole array rather than the section, and a shape
reaches a cast on a null defining op.
Rebuild the value in the block. The operands behind it are integers,
which can be block arguments, so take those as arguments, forward them
along each branch, and build the value from them at the top of the
block. The merge is kept and no block is duplicated.