[CodeGen] Drop dead SlotIndexes before allocation
SlotIndexes keeps the index list entry of an erased instruction and only
clears its instruction pointer. Live range sizes are measured in slot
indexes and greedy ranks ranges by size, so the leftovers inflate some
ranges more than others and reorder allocation, spilling heavily on
register-starved functions.
Add SlotIndexes::compactIndexes() to erase them. Erased entries are
unlinked, so LiveIntervals, LiveStacks and LiveDebugVariables first
report the indexes they hold via appendReferencedIndexes(); a missed one
trips an assert instead of dangling.
Off by default behind -greedy-compact-slot-indexes, since it changes
allocation across much of the test suite.
[NFC] Use logical && instead of bitwise & with a logical NOT operand (#223798)
## Summary
Two spots computed `x & !y`, mixing a bitwise AND with a logical NOT:
* CodeGen: `MachineOperand::isKill()` — `IsDeadOrKill & !IsDef`
* OpenMP/Attributor: `AAExecutionDomainFunction::updateImpl` —
`StoredED.IsReachingAlignedBarrierOnly &
!IsEndAndNotReachingAlignedBarriersOnly`
In both cases all operands are booleans (1-bit bitfields / `bool`), so
the
computed value is unchanged. Switch to the logical AND (`&&`) to express
the
intended logical conjunction and restore short-circuiting.
NFC. Resolves CodeQL `cpp/incorrect-not-operator-usage` reports on these
lines.
[6 lines not shown]
[flang][cuda] Implicitly attribute ALLOCATABLE/POINTER components as managed
Under -gpu=mem:managed, resolve-names implicitly attributes allocatables and
pointers declared in an ordinary scope as managed.
Apply the same attribution to components in Post(ComponentDecl). An explicitly
attributed component keeps its own attribute, and a translation unit without
CUDA Fortran enabled is left alone.
Record the attribution in ObjectEntityDetails::cudaDataAttrIsImplicit, since
an attribute the compiler applied is not a user requirement: a memory space
the user did ask for on an enclosing object takes precedence over it.
[GlobalISel] Add computeKnownBits for G_CLMUL opcode (#223980)
Port the computeKnownBits support for ISD::CLMUL from SelectionDAG path
to the GlobalISel G_CLMUL.
[Clang] Fix padding clearing logic for packed boolean vectors in big endian
The memory layout of packed boolean vectors in big endian mode is quite
involved. This patch adds support for determining the occupied bits of
this type in big endian mode, so that the correct bits are cleared as
padding.
[clang] Fix mangling of constrained auto referring to a parameter when the return type has an ABI tag (#223121)
Fixes #204178
Mangling `template<typename T> auto f(T t, c<decltype(t)> auto) -> s;`
asserts with `ParmVarDecl is not visible in current parameter
environment` when `s` carries an ABI tag, as `std::string` does under
libstdc++. The invented template parameter's constraint refers to the
function parameter `t`, and the mangler encodes that reference by its
nesting depth, so the function's parameter scope has to be entered
before the name is mangled. `mangleFunctionEncoding` does that in the
common case, but when the return type has a tag it first mangles the
name with a temporary mangler to collect the tags in use, and the scope
was pushed on the outer mangler after the temporary one had already
copied its depth state. The temporary mangler saw depth zero. Without
assertions this didn't crash but produced `fp_` instead of `fL0p_`, a
symbol GCC doesn't emit and that differs from clang's own mangling of
the same declaration without the tag.
[4 lines not shown]
[SLP][X86][NFC] Add pre-commit test for width-3 shared-weight reduction (#222243)
Pre-commit test for three independent reductions over contiguous bytes
sharing a single broadcast weight. It captures current codegen: even
with -slp-vectorize-non-power-of-2=true the width-3 tree is discarded to
scalar because the non-power-of-2 <3 x i8> load is over-priced (the
folded vpinsr*(mem) is double-counted). A follow-up X86 cost fix flips
the NPOT run to the width-3 vector form.
cc @alexey-bataev
---------
Co-authored-by: Cursor <cursoragent at cursor.com>
[flang][OpenMP][NFC] Add missing lowering test coverage (#221079)
Adds three tests for behavior that already works but is untested:
- target-generic-spmd.f90: whether host_eval is emitted for separate,
non-combined target/teams/distribute directive nests, covering the
generic versus SPMD kernel classification. The existing host-eval.f90
covers only combined constructs.
- hlfir-to-fir-conv-omp.mlir: alloca block selection for OpenMP regions
during HLFIR-to-FIR conversion.
- reduction-target-spmd.f90: a reduction on target teams distribute
parallel do is placed on omp.teams.
verbs/mlx5: Add GRE and MPLS flow specification filter
[PATCH 30/31] FreeBSD OFED support for DPDK MLX5 PMD
a) Allow verbs applications packet steering of GRE tunneled traffic.
Adding GRE flow specification based on RFC 2890.
GRE consists of flags, protocol and key fields.
IPv4 protocol 47 (IPPROTO_GRE) can be used when GRE packets are
encapsulated in IPv4.
b) verbs: Add MPLS flow specification filter
Add MPLS flow specification based on RFC 3032.
MPLS spec defined with label field which includes the
label value and additional parameters such as: BoS, TC and TTL.
MPLS allows stacking multiple labels in sequence.
In addition, the MPLS header can be encapsulated on top of different
layers, e.g.: ETH, IP (rfc4023), UDP (rfc7510), GRE (rfc4023).
Therefore, when using the flow creation verb, the application should
[14 lines not shown]
IB/mlx5: Expose GRE flow spec to user-kernel ABI header
[PATCH 29/31] FreeBSD OFED support for DPDK MLX5 PMD
a) Add ib_uverbs_flow_spec_gre to define the rule to match GRE
encapsulation protocol.
The spec includes the generic specs header, type, size and reserved
fields while the filter itself is defined as ib_uverbs_flow_gre_filter
and includes:
- Checksum present bit, key present bit and version bits in a single
16bit field.
- Protocol type field - Indicates the ether protocol type of the
encapsulated payload.
- Key field - present if key bit is set and contains an application
specific key value.
b) IB/uverbs: Expose MPLS flow spec to user-kernel ABI header
[24 lines not shown]
fixup! verbs/mlxs: mlx5: Add tunnel offloads support in direct verbs
Fix is based on upstream rdma-core commit ee54f9d5348f ("mlx5: Add
loopback flags to QP creation").
Validate create_flags against a supported-flags mask instead of rejecting
anything that is not MLX5DV_QP_CREATE_TUNNEL_OFFLOADS. A create_flags of
0 is now accepted, unknown bits are rejected, and the vendor flag is OR-ed
in rather than assigned. The loopback flags that commit also adds do not
exist in this tree, so only the validation is taken.
Sponsored by: NVidia networking
MFC after: 1 month
fixup! IB/mlx5: Expose GRE flow spec to user-kernel ABI header
Import Linux upstream commit a93b632c4531 ("IB/mlx5: Fix GRE flow
specification").
Honour the user-supplied GRE protocol mask instead of hard-coding 0xffff,
so wildcard and partial masks work.
Also give flow_spec_data[] 8-byte alignment, so this userspace copy of
struct ib_uverbs_flow_spec_hdr matches the kernel UAPI on 32-bit builds.
The kernel header has always used __aligned_u64 here and only our copy
diverged, so there is no upstream commit for that half. The attribute is
spelled out instead of introducing an __aligned_u64 macro, which
<infiniband/types.h> does not provide and which libbnxtre already defines
for itself.
Sponsored by: NVidia networking
MFC after: 1 month
fixup! verbs/mlx5: Add GRE and MPLS flow specification filter
Import rdma-core upstream commit ff01da2c5ac2 ("verbs: Fix typo in copying
IBV_FLOW_SPEC_UDP/TCP 'val'").
The TCP/UDP filter value was copied using the size of the IPv4 filter,
overrunning the destination subobject.
Also accept inner MPLS flow specs, matching 6d6f29721eea ("verbs: Allow
creation of inner MPLS flow spec"). Upstream deliberately has no inner
case for GRE - it is a tunnel header, matched only in the outer stack - so
that part of the review comment is not applied.
Sponsored by: NVidia networking
MFC after: 1 month
fixup! verbs/mlx5: Expose tag matching capabilities
Reject TM-SRQ creation when tm_cap.max_ops is 0, so the internal command
QP's send queue is not sized to zero work-queue entries.
Sponsored by: NVidia networking
MFC after: 1 month
fixup! IB/mlx5: Add tunneling offloads support
Import Linux upstream commit 4e2b53a5cb5a ("IB/mlx5: Report inner RSS
capability").
Define MLX5_RX_HASH_INNER as (1UL << 31). Shifting 1 into the sign bit of
a plain int is undefined behaviour, and user space compiles this header
too. The main.c half of that commit is already carried by the preceding
fixup.
Sponsored by: NVidia networking
MFC after: 1 month
fixup! IB/mlx5: Expose multi-packet RQ capabilities
Import Linux upstream commit ccc870879027 ("IB/mlx5: Allow creation of a
multi-packet RQ").
Only the create_rq() part was missing: the striding RQ parameters we
validate and store on the WQ never reached the firmware. Program
two_byte_shift_en, single_stride_log_num_of_bytes and
single_wqe_log_num_of_strides, with the two log values encoded relative to
their MLX5_MIN_* bases so the 3-bit firmware fields do not truncate them.
Upstream renamed these WQ context fields later on, so the names used here
are the ones this tree's mlx5_ifc.h still carries.
Sponsored by: NVidia networking
MFC after: 1 month
verbs/mlxs: mlx5: Add tunnel offloads support in direct verbs
[PATCH 28/31] FreeBSD OFED support for DPDK MLX5 PMD
a) In order to enable offloading such as checksum and LRO for incoming
tunneling traffic, the QP should be created with tunnel offloads flag -
MLX5DV_QP_CREATE_TUNNEL_OFFLOAD.
b) Reports capability of which tunneling type supports
the tunneling offloads.
c) verbs: Add support in RSS of the inner packet
Some user space application would like to do RSS on the inner packet
fields instead of the outer. When user will set the IBV_RX_HASH_INNER
bit with one of the other hash fields, then the RSS will be on the inner
packet.
Differential revision: https://reviews.freebsd.org/D32203
MFC after: 1 month
verbs/mlx5: Allow creation of a Multi-Packet RQ using direct verbs
[PATCH 27/31] FreeBSD OFED support for DPDK MLX5 PMD
a) Add needed definitions to allow creation of a Multi-Packet RQ
using the mlx5 direct verbs interface.
In order to create a Multi-Packet RQ, one needs to provide a
mlx5dv_wq_init_attr containing the following information in its
striding_rq_attrs struct:
- single_stride_log_num_of_bytes: log of size of each stride
- single_wqe_log_num_of_strides: log of number of strides per WQE
- two_byte_shift_en: When enabled, hardware pads 2 bytes of zeros
before writing the message to memory (e.g. for IP alignment).
b) Add a helper function to verify 64 bit comp mask
The common check for a mask is as follows:
if (comp_mask & ~COMP_MASK_SUPPORTED_VALUES)
return EINVAL;
[11 lines not shown]
libmlx5: Report Multi-Packet RQ capabilities through mlx5 direct verbs
[PATCH 26/31] FreeBSD OFED support for DPDK MLX5 PMD
A Multi-Packet RQ is a receive queue where multiple packets are
written to the same WQE. Each message starts in the beginning of a
stride. The total size of the scatter elements of each WQE is
determined upon RQ creation and all the posted WQEs should meet the
determined size.
A Multi-Packet RQ reduces the number of needed post-recv operations
thus increasing performance.
It reduces memory footprint by allowing each packet to consume a
different number of strides instead of the whole WR.
Differential revision: https://reviews.freebsd.org/D32201
MFC after: 1 month
IB/mlx5: Fix ABI alignment to 64 bit
[PATCH 25/31] FreeBSD OFED support for DPDK MLX5 PMD
Struct mlx5_ib_striding_rq_caps was not aligned to 64 bit as
it should have been. Add a 32 bit reserved field.
Differential revision: https://reviews.freebsd.org/D32200
MFC after: 1 month
verbs/mlx5: Expose tag matching capabilities
[PATCH 23/31] FreeBSD OFED support for DPDK MLX5 PMD
a) Expose tag matching capabilities and show them in ibv_devinfo.
b) verbs: Introduce tag matching SRQ
Introducing tag matching SRQ (TM-SRQ), which retains basic semantic of
regular SRQ, reports completions to own CQ, and has additional tag
based message receiving mechanism.
Detailed description for the TM-SRQ usage was added into
Documentation/tag_matching.md
c) mlx5: Add support to tag matching SRQ type
Create command QP for a tag-matching SRQ. Command QP used for
inserting/removing entries from the tag matching list. This command QP
is hidden from the users in the mlx5_srq structure.
New verb ibv_post_srq_ops() will be added in next patch to use it.
[2 lines not shown]
IB/mlx5: Add tunneling offloads support
[PATCH 22/31] FreeBSD OFED support for DPDK MLX5 PMD
a) The device can support receive Stateless Offloads for the inner
packet's fields only when the packet is processed by TIR which is
enabled to support tunneling. Otherwise, the device treats the
packet as an ordinary non-tunneling packet and receive offloads
can be done only for the outer packet's field.
In order to enable receive Stateless Offloading support for incoming
tunneling traffic the TIR should be created with tunneled_offload_en.
Tunneling offloads is supported only be raw ethernet QP.
This patch includes:
- New QP creation flag for tunneling offloads.
- Reports device capabilities.
b) IB/mlx5: Add support for RSS on the inner packet
Some user space application would like to do RSS on the inner
[6 lines not shown]
fixup! libmlx5: Report SW parsing capabilities through mlx5 direct verbs
Document the full struct mlx5dv_context (cqe_comp_caps, sw_parsing_caps)
and the mlx5dv_context_comp_mask enum in mlx5dv_query_device.3, and note
that these caps are only valid when the matching comp_mask bit is asked
for on input and returned on output.
Sponsored by: NVidia networking
MFC after: 1 month
libmlx5: Report SW parsing capabilities through mlx5 direct verbs
[PATCH 24/31] FreeBSD OFED support for DPDK MLX5 PMD
Software parsing (SWP) is a feature that can be used to instruct the
device to stop using its internal parser and to parse packets on the
transmit path according to offsets set for each packet.
Through this feature, the device allows the handling of checksum and LSO
by the hardware according to the location of IP and TCP/UDP headers.
Report various SW parsing capabilities and supported QP types through
mlx5 direct verbs interface.
Differential revision: https://reviews.freebsd.org/D32199
MFC after: 1 month
fixup! IB/mlx5: Add tunneling offloads support
Drop the unused MLX5_IB_QP_TUNNEL_OFFLOAD flag. Tunnel offload is driven
by qp->tunnel_offload_en, so the enum bit was dead and misleading.
Also advertise MLX5_RX_HASH_INNER in query_device()'s
rss_caps.rx_hash_fields_mask, so user space can discover inner RSS.
Sponsored by: NVidia networking
MFC after: 1 month
IB/mlx5: Expose multi-packet RQ capabilities
[PATCH 21/31] FreeBSD OFED support for DPDK MLX5 PMD
a) This patch reports the device's striding RQ capabilities to
the user-space:
- min/max_single_stride_log_num_of_bytes: Log of min/max number of
bytes in a single stride.
- min/max_single_wqe_log_num_of_strides: Log of min/max number of
strides in a single WQE.
- supported_qpts: A bit mask to know which QP types support multi-
packet RQ, for now only Raw Packet QPs.
b) Allow creation of a multi-packet receive queue.
In order to create a multi-packet RQ, the following fields in
the mlx5_ib_rwq should be set:
- log_num_strides: Log of number of strides per WQE
- single_stride_log_num_of_bytes: Log of a single stride size
- two_byte_shift_en: When enabled, hardware pads 2 bytes of zeros
[4 lines not shown]
fixup! IB/mlx5: Expose software parsing for Raw Ethernet QP
Report sw_parsing_caps using the same condition that enables software
parsing on the SQ (both eth_net_offloads and swp), so what we advertise
matches what the driver actually does.
Sponsored by: NVidia networking
MFC after: 1 month
libmlx5: Report if kernel allows using MPW in SQ
PATCH 20/31] FreeBSD OFED support for DPDK MLX5 PMD
a) Use flag MLX5DV_CONTEXT_FLAGS_MPW_ALLOWED to indicate hardware
supports multi packet WQE and it's enabled in SQ context.
Flag MLX5DV_CONTEXT_FLAGS_MPW is deprecated, shall not be used
in new applications.
b) Report if enhanced multi packet send WQE is supported through mlx5
direct verbs.
Differential revision: https://reviews.freebsd.org/D32194
MFC after: 1 month