[lldb][test] Fix TestMinidumpSizeOfImage.py with glibc >= 2.41 (#212972)
Test added by #188363.
In glibc 2.41:
> * dlopen and dlmopen no longer make the stack executable if a shared
> library requires it, either implicitly because of a missing GNU_STACK
> ELF header (and default ABI permission having the executable bit set)
> or explicitly because of the executable bit in GNU_STACK, and the
> stack is not already executable. Instead, loading such objects will
> fail.
https://lists.gnu.org/archive/html/info-gnu/2025-01/msg00014.html
The test program used for this test provides all the PHDRS itself, but
did not include a PT_GNU_STACK entry.
So the loader would assume it wanted exectuable stack. The main program
was not using an executable stack and so the loader refused to change
[8 lines not shown]
[LV] Only use legacy scalarization costs with replicate regions. (#212738)
The scalarization costs in InstsToScalarize are based on the assumption
that the instructions are scalarized and predicated. There are a number
of VPlan transformations that can simplify/remove replicate regions. If
there are no replicate regions in a plan, nothing is predicated and
scalarized, so the costs in InstsToScalarize will be inaccurate.
Skip the fallback in those cases, using the more accurate VPlan-based
cost info.
PR: https://github.com/llvm/llvm-project/pull/212738
[flang][CodeGen] Replace fir.select* FIR-to-LLVM patterns with stubs
`fir.select`, `fir.select_case`, `fir.select_rank`, and `fir.select_type`
are lowered to cf.* earlier in the pipeline (`--fir-select-ops-conversion`
and `--fir-polymorphic-op`). Their FIR-to-LLVM conversion patterns are
dead in a correct pipeline. Replace them with a single templated stub
`SelectShouldHaveBeenConvertedStub<OP>` that emits `"'fir.<op>' op should
have already been converted"` and fails legalization, so running
`--fir-to-llvm-ir` standalone on stale IR reports a clear diagnostic
instead of "unable to legalize".
`Fir/convert-to-llvm.fir`'s six select* test blocks are removed (the
lowering no longer runs; CF-level coverage lives in
`Fir/SelectOpsConversion/`). `Fir/convert-to-llvm-invalid.fir` gains a
stub-error test per op. `Fir/Todo/select_case_with_character.fir` is
retargeted to check the equivalent diagnostic now emitted by
`--fir-select-ops-conversion`.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply at anthropic.com>
[flang][Transforms] Add SelectOpsConversion pass
Introduces `--fir-select-ops-conversion`, which lowers `fir.select`,
`fir.select_case`, and `fir.select_rank` to the control-flow dialect
(`cf.switch` / `cf.cond_br` / `cf.br`) while preserving the CFG shape.
`fir.select_case` becomes an if-then-else ladder of `arith.cmpi` +
`cf.cond_br`; Fortran `UNSIGNED` selectors use `ule`. Signed / unsigned
FIR integer values are normalized to signless via `fir.convert` first.
`fir.select_type` is not handled here — it is already lowered by
`--fir-polymorphic-op` (`PolymorphicOpConversion`).
The pass runs in the default FIR optimizer pipeline right after
`PolymorphicOpConversion`. Pipeline-check tests are updated to expect
`SelectOpsConversion` in the sequence; `Fir/select.fir` and
`Lower/volatile3.f90` are relaxed to accept the newly-canonicalized form
of the lowered output.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply at anthropic.com>
CSKY: Consume "float-abi" module flag
Start respecting float-abi, and fall back on the TargetOptions
field if not present.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
[flang][AllocationPlacement] prevent promotion of mock result to allocmem (#212825)
Under the new experimental pass that can move automatic arrays to the
heap, "mock" result storage may be promoted to allocmem before
AbstractResult removes them and replace them by a hidden result.
Add a fir.must_be_stack attribute to these mock alloca so they are never
promoted. So that passes dealing with array function results ABI can
expect to find an fir.alloca and remove it.
AMDGPU: Export the TargetParser feature bitset
Previously this bitset was only used to populate the feature
name string map used by clang. Eventually this will replace
the current bitmask integer. AArch64 already has a similar
interface.
Co-authored-by: Claude (Claude-Opus-4.8)
AMDGPU: Tablegenerate TargetParser feature sets
Traditionally we maintained 2 parallel feature mechanisms,
one in clang (later moved to TargetParser), with largely
mirrored subtarget features defined in the backend. Start
directly taking feature information from the backend and putting
it into TargetParser. This is still in a compromise mid-migration
state. We still have both the legacy "ArchAttr" bitfield integer,
plus a new AMDGPUFeatureBitset field stored in the table, which
isn't yet exported.
For the moment, the new bitset is only used to populate the
feature string name map, which is the big maintainability win.
This also lists an explicit subset of exported features to
avoid churn.
Co-authored-by: Claude (Claude-Opus-4.8)
[flang][OpenMP] Lower allocate align modifier on parallel
Lower the align modifier on OpenMP allocate clauses for parallel constructs.
Carry per-item alignments through the OpenMP dialect and select
__kmpc_aligned_alloc for aligned private storage while retaining the existing
allocation and cleanup behavior for unaligned items.
Add source, verifier, preservation, LLVM IR, i386 ABI, and overflow coverage.
Keep unsupported construct kinds and device lowering unchanged.
Assisted-by: GitHub Copilot
[libc++][NFC] Rename the streambuf members (#212277)
This refactors `streambuf` to contain a `_GetArea` and a `_PutArea`.
This makes the code significantly easier to read, since the pointers
belonging together are bundled in a struct.
[VPlan] Make simplifyRecipe more like InstCombine
Most combines in simplifyRecipe RAUW a value, but not all of them erase the old recipe.
Unify them and bring it in line with InstCombine by having it return a VPValue, which simplifyRecipes can then call RAUW with, and automatically erase the old recipe.
Similarly to InstCombine, combines that modify a recipe should return the same recipe.
[libc++] Remove SFINAE checks in tuple which are always true (#212765)
We have a specialization for `tuple` with no arguments, so checking
`sizeof...(_Tp) >= 1` in the primary template will never be false.
[lldb] Fix stale L1 memory cache read after memory write (#208347)
A `memory write` can leave stale bytes in the L1 memory cache, so a later
`memory read` of the address that was just written returns the old value.
The L1 cache (`m_L1_cache`) is a map keyed by each chunk's start address, and
chunks can overlap: a read larger than an L2 cache line
(`target.memory-cache-line-size`, 512 by default) bypasses L2 and is stored
whole in L1, so two large reads can produce two chunks that both cover the same
address.
`Flush()` invalidates the L1 cache on a write. It started at the chunk at or
below the flushed address and walked forward, stopping at the first chunk that
did not intersect. It therefore never inspected a chunk that starts below the
flushed address but is long enough to reach into it, leaving that chunk behind
with the stale byte. A later read fully contained in that chunk is served from
the cache and returns the old value.
Fix `Flush()` to walk the whole L1 cache and erase every chunk that intersects
[7 lines not shown]
[SPARC] Support the .seg directive (#209001)
Support the legacy .seg directive used by SunOS SPARC assembly. Map
"text", "data", "data1", and "bss" to their corresponding MC sections,
using subsection 1 for "data1".
[WebAssembly][TTI] Avoid crash when costing scalable vector shifts (#212759)
I see the following
```
anutosh491 at Anutoshs-MacBook-Air llvm-project % cat /private/tmp/wasm-scalable-shift-cost.ll
define <vscale x 4 x i32> @shift(<vscale x 4 x i32> %x,
<vscale x 4 x i32> %amount) {
%result = shl <vscale x 4 x i32> %x, %amount
ret <vscale x 4 x i32> %result
}
anutosh491 at Anutoshs-MacBook-Air llvm-project % build-assert/bin/opt \
-mtriple=wasm32-unknown-unknown \
-mattr=+simd128 \
-passes='print<cost-model>' \
-disable-output \
/private/tmp/wasm-scalable-shift-cost.ll
[27 lines not shown]
IR: Introduce "float-abi" module flag (#210821)
This is intended to eliminate the FloatABIType TargetOptions field.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
[RISCV] Fix crash in convertToVLMAX when AVL is an ADDI of a frame index (#212923)
`getConstant()` recognizes an AVL that is an immediate by checking
the opcode is `ADDI`, then reading the second operand and checking
if it is `X0`. However the second operand may be a frame index and
`getReg()` asserts with "This is not a register operand!".
We should guard the register access with `isReg()` before comparing
against `X0`.
Fixes #212797.
[Docs][AMDGPU] Explain completion of async operations (#212756)
This improves the somewhat hand-wavey "memory model" currently described
for async operations. While this version is also not complete, it
prepares for the more complete memory model being written down.
[RISCV] Split and rename WriteVSlideI/WriteVISlide1X/WriteVFSlide1F (#212184)
Split each of these SchedWrites into separate slide-up and slide-down
variants:
- WriteVSlideI -> WriteVSlideUpI, WriteVSlideDownI
- WriteVISlide1X -> WriteVISlide1Up, WriteVISlide1Down
- WriteVFSlide1F -> WriteVFSlide1Up, WriteVFSlide1Down
SpacemiT X100 and A100 have different latencies and/or throughput for
slide up vs. slide down operations, so they need separate SchedWrites to
model that difference.
clang/AMDGPU: Forward xnack/sramecc mode to the assembler
When assembling a .s file with no target ID directive, the requested
xnack/sramecc mode has no module flag to carry it. Forward the mode
requested via -mxnack/-msramecc (or the -mcpu target ID modifiers) to the
assembler as a target feature so it is recorded in the object's e_flags.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
CSKY: Avoid constructing a temporary CSKYSubtarget in AsmPrinter (#212944)
This was reconstructing a subtarget equivalent to the global
subtarget instead of just using the available one.