[orc-rt] Make split-file a required regression test tool. (#227510)
split-file is an LLVM utility, like FileCheck and not, so it's available
wherever they are. Require it in the same way, rather than gating tests
on a split-file lit feature.
[MC] Add baseline test for MCContext::getSubtargetCopy HwMode loss
No change intended here, just adding test coverage showing that
MCContext::getSubtargetCopy currently slices away the target's
<Target>GenMCSubtargetInfo subclass and causes getHwMode() and
getHwModeSet() to return 0.
This commit was created with the help of AI tools
[llvm-objdump][RISC-V] Do not leak file-level features into $x<ISA> regions
Previously, RISCVISATargetCache::get initialized per-region features from
Base.SubtargetInfo->getFeatureString(), which already included the
file-level Tag_RISCV_arch and --mattr flags. Because RISCVISAInfo::toFeatures()
only emits "+ext" for extensions present in the "$x<ISA>" mapping symbol
(and never "-ext"), any extension enabled in Tag_RISCV_arch or via a
conflicting --mattr remained enabled even in regions whose "$x<ISA>"
mapping symbol disabled it.
Start from an empty SubtargetFeatures instead and populate it from the
parsed "$x<ISA>" mapping symbol (plus non-conflicting --mattr flags).
This commit was created with the help of AI tools
[RISC-V][MC] Reject x0 as the temporary register for store pseudos
The zero register as the temporary for the store pseudo is illegal since
it is used to synthesize the store address and using zero would mean
that the auipc result is ignored and we store to an invalid location.
[BOLT] Ensure runtime is built with uncompressed debug symbols (#218126)
Currently, if the runtime is built with compressed debug symbols it
causes an out-of-bounds read and segfault when instrumenting a binary.
This is due to JITLink currently not supporting SHF_COMPRESSED symbols
as it would need to decompress the symbols first to update the reloc
symbols
As some distributions are starting to enable compressed debug sections
in their toolchains by default, ensure the runtime is built without them
until JITLink supports updating SHF_COMPRESSED relocs.
[TableGen] Support HwMode registers in CompressInstEmitter
Previously, CompressInstEmitter only handled plain Register and
RegisterClass records when matching and validating operands in
CompressPat definitions. Attempting to use a RegisterByHwMode (such as a
mode-dependent stack pointer) or a RegClassByHwMode with mode-dependent
registers failed because CompressInstEmitter could not resolve the
underlying register or register class for a given subtarget configuration.
Introduce HwModePredicates in CodeGenHwModes to resolve HwModeSelect
records using the subtarget features required by a CompressPat (with
predicates implying the absence of all non-default modes resolving to
DefaultMode). This allows using a single instruction definition for
compressed instructions like RISC-V C_ADDI4SPN/C_ADDI16SP across HwModes
instead of duplicating the instruction definitions for each mode.
This commit was created with the help of AI tools
[orc-rt] Remove duplicate section from regression test README. NFC. (#227503)
The "Tests with more than one source file" section appeared twice.
Remove the older copy.
[mlir][xegpu] Check uArch block shapes in VectorToXeGPU transfer lowering (#217179)
vector.transfer_read/transfer_write were lowered to xegpu.load_nd /
xegpu.store_nd whenever the target chip was pvc, bmg or cri and the
transfer had rank >= 2, without asking whether the target's subgroup 2D
block instructions can access the requested tile at all.
Changing it to query the uArch 2D block load/store description up front
and keep the block path only for tiles whose two innermost dims are each
a multiple of a supported block size- the property the later layout
propagation and blocking passes rely on when they split a tile into
hardware-sized blocks. Everything else takes the scattered
load_gather/store_scatter fallback that both patterns already have.
Deriving 2D block support from the uArch instruction registry also
replaces the hardcoded chip list, resolving the TODO left there; the
three chips that registry covers are the same three that were listed.
assisted by claude
[lldb-dap] End the session before sending the disconnect response (#227106)
Previously the `disconnect` request asks the debug adapter to disconnect
from the debuggee, ending the debug session, and then to shut down.
lldb-dap responded right after killing or detaching the process. So the
`exited` event, terminatedCommands and exitedCommands can be sent after
`disconnect` could reaches the client.
`DisconnectRequestHandler` now stores its response in
`DAP::on_session_end`. `DAP::Loop()` sends it once the session ended.
The teardown/cleanup is now:
- The event threads stop, after they report the exit of a killed
process.
- `terminated` Event is sent, if the event thread didn't send it.
- The transport thread stops.
- If there is a pending `configuration_done` we reply.
- The requests queued behind `disconnect` request are cancelled.
- The `disconnect` response is sent last.
- The debugger is destroyed.
[17 lines not shown]
[Clang] Do not assume existing substition for template specialization
Mangled symbol is not always available when having template
specialization. For example, a template alias mangled the type when
actually doing substitution. In previous code, B in A<B>::foo() is
always mangled when reaching A<B>. When doing substition insides A<T>,
B is always available. Alias is not the case here, B is mangled when
reaching A<T> inside.
As now, TemplateName::SubstTemplateTemplateParm might needs to be
resolved further as B is not always defined, we remove it from the
switch case and find recursively like the origianl path.
Assisted-by: Claude # Tests, ReleaseNotes
[BOLT] Emit DW_AT_high_pc matching the class of its form (#227066)
DW_AT_high_pc holding an offset from DW_AT_low_pc is a DWARF 4 addition;
in DWARF 2 and 3 the attribute is always class address.
updateLowPCHighPC() always wrote "HighPC - LowPC" while reusing whatever
form the input DIE had, so a DWARF 2 CU using DW_FORM_addr ended up
storing a length where an end address is expected.
BOLT should write the end address when the form is DW_FORM_addr, and
the length otherwise. The default form for a new attribute is now
DW_FORM_addr below DWARF 4, and size needs to be widened to 64 bits so
DW_FORM_data8 won't be truncated. The FunctionRanges.empty() code path
no longer treats a raw DW_FORM_addr high_pc as a size.
Assited-by: opus
[libc][aarch64] Use inline_memcpy_aligned_access_64bit under -mstrict-align (#227428)
Under `-mstrict-align`, clang cannot determine both `src` and `dst` are
aligned. The end of `inline_memcpy_aarch64` aligns `src` but it cannot
confirm `dst` is aligned. As a result,
`builtin::Memcpy<64>::loop_and_tail(dst, src, count)` lowers to
`__builtin_memcpy_inline(..., 64)` but clang emits many many single-byte
load/store instructions to ensure correctness. This can be very slow, so
if unaligned accesses are not supported, opt for adjusting both `src`
and `dst` such that we can use 8-byte loads/stores.
`inline_memcpy_aligned_access_64bit` already does this so we can just
reuse that.
[Hexagon,test] Enable +audio for intrinsics-v67.ll (#227362)
The test uses llvm.hexagon.M7.vdmpy{,.acc}, which lower to
M7_dcmpyrwc{,_acc}. Those instructions require the audio feature, so add
-mattr=+audio to the RUN line to avoid a predicate-check crash in the
asm printer.
[clang][docs] Use doc links to HLSL/ResourceTypes in AttrDocs.td (#227407)
#222504 made the docs build warn about absolute links to pages of the
same Sphinx project, and the docs build turns warnings into errors. The
HLSL resource attribute docs added in #213346 linked to
https://clang.llvm.org/docs/HLSL/ResourceTypes.html, which after #226729
broke the clang docs build (docs-clang-html, docs-clang-man). Use the
{doc} role instead, like the rest of AttrDocs.td.
[lldb] Add label, title and divider ANSI settings (#226606)
Commands that print structured listings hard-coded their colors, but
terminal themes render ANSI colors differently, so users couldn't adjust
them the way they can with the prompt, progress or autosuggestion
colors.
Add `label-ansi-prefix`, `title-ansi-prefix` and `divider-ansi-prefix`
settings, each with a matching suffix, for the field labels, entry
titles and dividers such listings are made of. Their defaults are the
colors `scripting extension list` uses today, and it now reads them from
these settings.
Signed-off-by: Med Ismail Bennani <ismail at bennani.ma>
[flang][openacc] Add ACCEraseUnusedKernelAllocations pass (#227481)
fir.declare carries a debug-memory effect and fir.freemem is a real
free, so an unused fir.allocmem inside acc.compute_region stays live
through ordinary dead-code elimination and is lowered to a checked
device malloc.
Delete that allocation when it has no uses, or when every use is
fir.freemem, a ViewLikeOpInterface view such as fir.convert, or
fir.declare of that storage. A load, store, or other memory use keeps
it. An unused private-recipe allocation is removed the same way as an
unused source array.
[clang][CIR] Match CIR return value with OGCG using mem2reg (#227241)
This fixes the test failure introduced by
[226534](https://github.com/llvm/llvm-project/pull/226534)
Currently, the `sve/len.c` test fails because of this difference between
OGCG and CIR:
```
; classic codegen ; CIR (-fclangir)
%2 = mul nuw i64 %1, 16 %6 = mul nuw i64 %5, 16
ret i64 %2 store i64 %6, ptr %3 ; retval slot
%7 = load i64, ptr %3
ret i64 %7
```
So, adding the return value checks as they are breaks the test.
This adds `mem2reg`, so that the CIR output matches OGCG and the same
return value check works for both.
[RISCV][P-ext] Add scalar multiply high intrinsics (#225561)
Add intrinsics, Clang builtins and SelectionDAG support for the scalar
multiply high operations.
Support
[multiply-high](https://github.com/riscv/riscv-p-spec/blob/master/P-ext-intrinsics.adoc#multiply-high).
On RV32 the non-rounding forms map to the M-extension mulh/mulhu/mulhsu
instructions and the rounding forms to the P-extension
mulhr/mulhru/mulhrsu instructions.
On RV64 the scalar spellings reuse the packed pmulh.w family on the low
words via the existing PMULH*_W patterns.