fix(AMDGPU): fold AND of any-extended booleans
Demanded-bits simplification can turn a boolean sign extension into an
any extension before the target AND combine. Choose sign extension for
the undefined high bits so the AND still folds to a select.
Cover reduced demand, swapped operands, and zero extension. Restore the
fold in setcc-multiple-use after symmetric demanded-bits simplification.
[Mach-O] Parallelize ICF's section sort. NFC (#230269)
Port of #223216 (ELF) and #229849 (COFF) to Mach-O. Replace the
single-threaded llvm::stable_sort of the ICF inputs with a parallelSort
over packed 64-bit keys, then gather the inputs in key order.
The key is icfEqClass[0] in the high 32 bits and the input index in the
low bits. With --icf=safe_thunks, bit 31 is set for inputs that are not
keepUnique, so keepUnique inputs still come first within each class. The
index breaks remaining ties by original position, so the resulting order
is exactly what stable_sort produced and the output is unchanged.
Merge tag 'pmdomain-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/ulfh/linux-pm
Pull pmdomain provider fixes from Ulf Hansson:
- imx: Serialize power on/off across sibling domains for imx8m-blk-ctrl
- rockchip: Fix a couple of errors during probe
* tag 'pmdomain-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/ulfh/linux-pm:
pmdomain: rockchip: don't ignore clock lookup errors on attach
pmdomain: rockchip: fix clock leak on domain probe failure
pmdomain: rockchip: propagate subdomain add errors
pmdomain: imx8m-blk-ctrl: Serialize power on/off across sibling domains
[SelectionDAG] avoid rounding exact f32->bf16 (STRICT_)FP_ROUND
For f32-to-bf16, DAGTypeLegalizer::SoftPromoteHalfRes_FP_ROUND currently
ignores the Trunc flag and unconditionally lowers ISD::FP_ROUND to
ISD::FP_TO_BF16 as follows:
- On X86 targets with +avx512bf16 or +avxneconvert, ISD::FP_TO_BF16
lowers to vcvtneps2bf16, which is lossy (it flushes subnormals to
zero and quietens sNaNs.)
- On targets without hardware bf16 conversion instructions, it performs
rounding via a runtime libcall (e.g. __truncsfbf2) or a software
rounding sequence (which can also be lossy, e.g. quieting sNaNs).
The above is unnecessary and lossy; bitcasting to i32 and extracting
the upper 16 bits is cheaper and lossless. Do that.
Tested:
[11 lines not shown]
fix(SelectionDAG): simplify commuted demanded bits
AND/OR demanded-bit simplification uses RHS known bits to simplify the
LHS, but does not retry the RHS using LHS known bits, making
optimizations depend on operand order.
Retry the RHS when the LHS reduces its demanded bits. Add AArch64
and AMDGPU codegen coverage.
Merge tag 'dma-mapping-7.3-2026-10-09' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux
Pull dma-mapping fixes from Marek Szyprowski:
"Two more fixes for the corner cases in the DMA-mapping SWIOTLB code
(Peng Fan and Marek Szyprowski)"
* tag 'dma-mapping-7.3-2026-10-09' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux:
swiotlb: fix default_swiotlb_limit() for non-growable default pool
iommu/dma: skip swiotlb bounce for DMA_ATTR_MMIO in iommu_dma_map_phys
[X86] Keep upper-16-bit f32 extractions in XMM for f16/v8i16
Teach combineBitcast and combineVectorInsert to keep this upper-16-bit
f32 extraction in XMM registers rather than bouncing through a GPR.
This paves the way for the following change in SelectionDAG:
https://github.com/llvm/llvm-project/pull/230557
update to py3-setuptools-84.0.0, ok tb@ kmos@
various changes to follow, linking py-standard-pkg-resources to
the build. and using it to replace or add to py-setuptools deps
for ports still needing pkg_resources.
thanks tba for running a bulk build and fixing a bunch of ports
[CIR] Emit the frame-pointer function attribute and module flag
CIR now honors -mframe-pointer= the way classic CodeGen does. Functions
and the module carry a new #cir.frame_pointer attribute, which lowers to
the "frame-pointer" function attribute and module flag.
Assisted-by: Cursor / claude-opus-5.5
import ports/sysutils/py-standard-pkg-resources, ok tb@ kmos@
Redistribution of the pkg_resources library previously provided by
setuptools. This was removed in setuptools 82, but quite a range of
Python software still needs it.
[SelectionDAG] Add extension attribute on Memset libcall arg. (#229883)
The Src value previously had no extension, but with this it is first
any-extended to i32 and then further extended if needed for the target.
A new ArgListEntry constructor is added that may set either IsSExt or IsZExt
given the third Attribute argument.
Add middleware support for LIO ALUA HA
Wire up the middleware side of LIO ALUA high-availability: load
lio_ha.ko with per-node addresses on service start, manage ALUA
state across failover events, clean up STANDBY configfs on pool
export, and add pre-flight validation that targets have static
initiator ACLs before ALUA can be enabled.
For each target, create a portal-less phantom TPG carrying the peer
node's controller group so that a single RTPG response from any
connected port lists both ALUA groups. Write tpgt_N/rtpi explicitly
before enable so that relative target port IDs in RTPG match the
tag formula (portal.tag on Node A, portal.tag + 32000 on Node B)
rather than being auto-assigned sequentially by the kernel.
ALUA group states are driven by role and ha_state:
MASTER + synced local=OPTIMIZED remote=NONOPTIMIZED
MASTER + connected local=OPTIMIZED remote=TRANSITIONING
[4 lines not shown]
Add context_params to iscsi_scsi_connect()
Pass an optional context_params dict through to python-scsi's
init_device(), so a test can set pre-login iscsi.Context options such
as an explicit ISID. context_params_supported() lets callers skip when
the installed python-scsi predates it.
Also ruff.
[ARM] Add VECTOR_REG_CAST to VBSP operands in PerformORCombine. (#230604)
The SDTypeProfile says the operand types should match the result.
Assisted-by: Claude
fix: add missing define in asan_mapping_sparc64.h (#230686)
#217530 added a new define to `asan_mapping.h`. SPARC has its own file
that provides the same defines, and I forgot to define the new macro
there.
SPARC takes the same default as all other non-darwin platforms.
[VPlan] Bail out on FindLast reductions with extra phi users. (#230676)
handleFindLastReductions uses the mask of the single non-phi incoming
value of the find-last blend as the condition for updating the
reduction. If the data value itself is the reduction phi on some paths,
e.g. via a nested blend for an inner if, the mask is also true for lanes
that keep the old value, and extract-last-active may then pick a stale
lane.
Conservatively bail out if the reduction phi has users other than the
find-last select/blend and the header mask select.
linux: Exposes renderD nodes and chardev in sysfs
To allow normal users to render through the render device, we expose the
renderD node. This enables Wayland applications to use hardware
acceleration when running under the Linux emulator.
Additionally, libdrm and Mesa need to look up
/sys/dev/char/<major>:<minor> and <pcidev>/drm to identify the
corresponding renderer device (e.g., a renderD device). We expose this
path as well so that libdrm can locate the renderer.
Differential Revision: https://reviews.freebsd.org/D59190
(cherry picked from commit 102adf88e6e8f83a9ab9769732d168c501f8eb6c)
libthr: Support disable spinloop
Like yieldloops, we shoulde be able to set _thr_spinloops to zero.
Originally, it makes us to enformce default spin time even if we try to
disable it. Make MUTEX_ADAPTIVE_SPINS a one time initialization now.
Reviewed by: kib
MFC after: 2 weeks
Differential Revision: https://reviews.freebsd.org/D60486