[lldb] Step through inlined code when stepping out of inline frame N > 0 (#226910)
To step out of an inline frame N > 0, ThreadPlanStepOut first steps out
to frame N, then queues a plan that steps through the inlined block. Due
to a bug, ShouldStop marks the plan complete right after it queues that
plan, so the thread stops as soon as it reaches frame N, with no stop
reason.
In fact, the comment next to that code describes the intended behavior
while implementing something different: only call the StepOut plan done
if the "step out from line" plan cannot be queued.
For example, consider this backtrace:
```
frame #0: 0x00000001000003c0 deep`sink(x=81) at deep.c:4:6 [opt]
frame #1: 0x0000000100000418 deep`level3(a=48) at deep.c:11:3 [opt] [inlined]
frame #2: 0x0000000100000404 deep`level2(b=37) at deep.c:20:3 [opt] [inlined]
frame #3: 0x00000001000003f0 deep`level1(c=42) at deep.c:28:3 [opt] [inlined]
frame #4: 0x00000001000003dc deep`main at deep.c:33:3 [opt]
[8 lines not shown]
[RISCV] Eliminate redundant materializations after register comparison (#227673)
`RISCVRedundantCopyElimination` can eliminate redundant immediate
materializations when a branch establishes a known immediate value, but
does not handle register-register `BEQ`/`BNE` when one operand is a
non-zero immediate materialized in a register.
Extend the pass to recognize `ADDI` from `X0` and `QC_LI`
materializations before the branch. On the equality edge, reuse that
known immediate so an identical materialization in the successor can be
removed. `X0` is excluded as a target register because writes to it are
discarded.
Tests cover `BEQ`/`BNE`, both operand orders, clobbers, `X0`, RV32/RV64,
and the relevant Xqci and XAndes variants.
AI tool usage: OpenAI Codex assisted with this contribution. I reviewed
and take responsibility for the final changes.
Assisted-by: OpenAI Codex
[clang-repl] Evaluate a void expression that has no trailing semicolon (#228334)
An expression statement without a semicolon at the end of an input is
replaced by a call to `__clang_Interpreter_SetValueNoAlloc` that
captures its value. For an expression of type `void`, the expression was
not an argument of that call and was dropped, so `f()` did not call `f`.
Build `(E, __clang_Interpreter_SetValueNoAlloc(...))` instead.
Fixes #219800.
Assisted-by: Claude Opus 5.5
[orc-rt] Use the Error matchers in SimpleNativeMemoryMapTest (#228335)
Use the Error matchers introduced in 4c8a437d0487 to clean up error
checks in SimpleNativeMemoryMapTest.
Revert "[ORC] Use __unw_add_dynamic_eh_frame_section/__unw_remove_dyn… (#227894)
…amic_eh_ frame_section in RegisterEHFrames.cpp (#212260)"
Revert commit d49626c811c304c047e028a80cedc0d377f49dfb due to link
failures on some platforms (see discussion in PR).
[AMDGPU] Add getCmpSelInstrCost override to re-enable SimplifyCFG speculation for vector types (#208043)
Commit ["[CostModel] Handle all cost kinds in
getCmpSelInstrCost"](https://github.com/llvm/llvm-project/commit/0967957d7a94e1b5c749c6e963bdca25f3c6d749)
changed the base cost model to consider the type in the costs for
non-throughput cost kinds. This caused SimplifyCFG to stop folding
branches to selects on AMDGPU, since
vector selects now report higher costs that exceed the folding
threshold.
Add an getCmpSelInstrCost override for AMDGPU to partially restore the
old behavior by returning a constant unit cost for cost kind
`TCK_SizeAndLatency`. This affects speculative execution in SimplifyCFG
and SpeculativeExecution.
[orc-rt] Add VettedPeer, require it for the socket transport (#228336)
The controller at the other end of a channel can make the executor run
arbitrary code, so attaching to it is a trust decision.
VettedPeer<ChannelT> makes that decision explicit: it can only be made
through one of three named factories -- inherited, checked or unchecked
-- each naming the reason the peer is trusted.
createSimpleRemoteCAOverSocket now takes a VettedPeer<SocketHandle>
rather than a bare SocketHandle.
VettedPeer neither verifies nor records the reason. The factories exist
so that the decision can't be skipped by omission, and so that every
choice is visible in the source.
The socket:adopt connector trusts its socket as inherited, and documents
the resulting precondition on its callers.
Assisted-by: Claude
[Fuchsia] Escape ';' in STAGE2_ CMake variables (#228272)
When forwarding STAGE2_* variables to EXTRA_ARGS in Fuchsia.cmake and
Fuchsia-stage2-instrumented.cmake, semicolons in list values must be
replaced with '|' to match the LIST_SEPARATOR of ExternalProject_Add.
In 70cf616b331c, 'list(APPEND EXTRA_ARGS "-D${variableName}=...")' was
added using the raw '${${variableName}}' before replacing ';' with '|',
and Fuchsia-stage2-instrumented.cmake similarly omitted the ';' to '|'
replacement.
When a list variable like STAGE2_CROSS_TOOLCHAIN_FLAGS_NATIVE contains
semicolon-separated '-D...' flags (such as CMAKE_EXE_LINKER_FLAGS),
appending it without escaping splits it into separate top-level CMake
arguments for stage2. Because STAGE2_CROSS_TOOLCHAIN_FLAGS_NATIVE sorts
alphabetically after STAGE2_CMAKE_*_LINKER_FLAGS, its split flags
overwrite the stage2 linker flags with the stage0 toolchain's libc++.a,
causing stage2 link failures on macOS.
[llvm-profdata] Propagate Error in loadInput and mergeWriterContexts (#228158)
[llvm-profdata] Propagate Error in loadInput and mergeWriterContexts
Propagate Error from loadInput and mergeWriterContexts in
mergeInstrProfile,
supplementInstrProfile, and overlapInstrProfile. In mergeInstrProfile's
ThreadPool, catch errors from worker threads, stop scheduling new jobs,
and return the first encountered fatal error.
Ensure ~WriterContext() consumes any pending unhandled errors in
WriterContext::Errors upon destruction.
Not NFC as destructors are run on the stack and ThreadPool workers exit
earlier on error.
With all subcommands propagating llvm::Error to main, exitWithError,
exitWithErrorCode, and the LSan leak suppression workaround are no
longer needed.
Assisted-by: Gemini
[SandboxVec][LoadStoreVec] Support constant vectors of mixed types
createConstantVector() previously packed the constant store operands
as-is, which only worked when every store had the same element type.
Take the lane type from VecUtils::getCombinedVectorTypeFor() instead and
reinterpret each constant's bits as that type, going through an integer
of matching width via ptrtoint/inttoptr/bitcast. Constants wider than a
lane (e.g. an i64 in an <N x i32>) are split across several lanes in
memory order. Bail out when a constant cannot be reinterpreted, such as
a non-integral pointer or a relocatable address that needs splitting.
Also flatten vector-typed ConstantPointerNull into per-lane nulls, and
bail out on the remaining vector constants such as poison rather than
packing them into the result.
Co-authored-by: Cursor <cursoragent at cursor.com>
[MC] Declare command line options in TableGen (#228321)
Move the cl::opts into MCCLOptions.td, except the -dx-* options, which
the DirectX backend reads and writes. The name avoids MCOptions, a
common variable name for MCTargetOptions.
New OptionsStruct member kinds replace cl:: features the port needs:
* OptionalBoolField, a std::optional<bool>, replaces cl::boolOrDefault
(-use-leb128-directives)
* EnumField replaces cl::values, showing its values as the -help-hidden
metavar
* DefaultOnOffField, an EnumField of std::optional<bool>, replaces the
file-local Default/Enable/Disable enum (-dwarf-extended-loc), which
DwarfDebug.cpp also defines for four options
Aided by Opus 5.5
[offload][omp] Move strict threads & groups computation to libomptarget (#222607)
Move all computation related to strict threads and strict groups to
libomptarget as it's specific to OpenMP. This also removes all refrences
to ExecutionModes from the Plugin Interface.
With this, the data from the OpenMP kernel environment being used is
ReductionDataSize used to create the KLE, and maxNumThreads used by
RecordReplay. They're are purposely left for a future PR.
Assisted by Claude.
[SandboxVec][LoadStoreVec] Support constant vectors of mixed types
createConstantVector() previously packed the constant store operands
as-is, which only worked when every store had the same element type.
Take the lane type from VecUtils::getCombinedVectorTypeFor() instead and
reinterpret each constant's bits as that type, going through an integer
of matching width via ptrtoint/inttoptr/bitcast. Constants wider than a
lane (e.g. an i64 in an <N x i32>) are split across several lanes in
memory order. Bail out when a constant cannot be reinterpreted, such as
a non-integral pointer or a relocatable address that needs splitting.
Also flatten vector-typed ConstantPointerNull into per-lane nulls, and
bail out on the remaining vector constants such as poison rather than
packing them into the result.
Co-authored-by: Cursor <cursoragent at cursor.com>
[IRBuilder] Propagate fast-math flags to folding in CreateBinOpFMF (#227543)
Fast-math flags are used by InstSimplifyFolder, so we need to pass them
to the folding function in CreateBinOpFMF.
[orc-rt] Use the Error matchers in NativeDylibManagerTest (#228323)
Use the Error matchers introduced in 4c8a437d0487 to clean up error
checks in NativeDylibManagerTest.
[compiler-rt][msan] Ignore addrinfo padding in getaddrinfo interceptor (#219124)
Avoid checking padding bytes in `struct addrinfo` in the `getaddrinfo`
interceptor.
The interceptor used a single `COMMON_INTERCEPTOR_READ_RANGE` for the
whole structure, which caused MSan to report uninitialized padding even
when all
named members were initialized.
Check each member separately instead.
Testing
- Added `getaddrinfo-padding.cpp` to cover the padding case
- Verified the new test fails before the fix and passes after the change
- Verified `getaddrinfo-positive.cpp` still reports uninitialized fields
- Ran `ninja -C build/runtimes/runtimes-bins check-msan` on AArch64
Linux
Fixes #162216
[clang][bytecode] Skip `emitInitScope` for `EvalEmitter` (#228092)
The opcode ultimately doesn't do anything in `EvalEmitter`, so try to
skip it as early as possible.
[CIR] Upstream missing support for floating point unary operator (#193215)
- Add Support for `__bf16/Float16/bfloat16/float128` scalar pre/post
increment and decrement.
- Add support for float vector increment and decrement, and add test
cases for both float vector and integer vector.
- Add test case for __bf16/Float16/bfloat16/double/long double/
float128 scalar unary operator with target `x86_64`.
- Part of task https://github.com/llvm/llvm-project/issues/192316
[llvm-profdata] Propagate Error in loadInput and mergeWriterContexts
Propagate Error from loadInput and mergeWriterContexts in mergeInstrProfile,
supplementInstrProfile, and overlapInstrProfile. In mergeInstrProfile's
ThreadPool, catch errors from worker threads, stop scheduling new jobs,
and return the first encountered fatal error.
Ensure ~WriterContext() consumes any pending unhandled errors in
WriterContext::Errors upon destruction.
Not NFC as destructors are run on the stack and ThreadPool workers exit
earlier on error.
With all subcommands propagating llvm::Error to main, exitWithError,
exitWithErrorCode, and the LSan leak suppression workaround are no
longer needed.
Assisted-by: Gemini
[NFCI][llvm-profdata] Propagate Error in merge subcommand (#228157)
Change merge_main and its helpers to return Error and handle it
with reportError in main.
Not NFC as destructors are run on the stack.
Assisted-by: Gemini
[compiler-rt] Remove dlsym interceptor and support `-shared-libsan` for CSan
Summary:
Follow the UBSan offload runtime. Offload now resolves HSA through the
global scope, so the `dlsym` interceptor is no longer needed. The real
HSA entry points are still taken from the loaded HSA library rather than
`RTLD_NEXT`, since every DSO with a static runtime exports the same
wrappers and they would otherwise chain back into each other.
Build `libclang_rt.csan.so` with the offload objects folded in. The
exported HSA wrappers report failure when HSA is absent and warn when HSA
was loaded ahead of the runtime. The preinit hook moves to a separate
`csan_offload-preinit` archive for executables.
[compiler-rt] Add 'csan' library for the concurrency sanitizer
Summary:
Adds the runtime for the concurrency sanitizer, both CPU and GPU.
Fundamentally, this works using the following pseudocode:
```c
static u64 watchpoints[N]; // Hash-indexed, zero is empty.
// Emitted before the access, so we never trip on our own write.
void check_access(volatile void *addr, u32 size, u32 type) {
// Every access probes. A read conflicts only with a watched write, a
// write conflicts with either.
if (u64 *wp = find_watchpoint(addr, size, type))
consume(wp, this_pc()); // Hand our location to the owner.
if (!should_sample()) // Wave-uniform, 1-in-N chance.
return;
[17 lines not shown]
[Clang] Support `-shared-libsan` for offload CSan
Summary:
Follow the UBSan handling. The shared runtime embeds the HSA
interceptors, so `csan_offload` is only linked with static runtimes and
executables pull in `csan_offload-preinit` to initialize early.