[mlir][scf] Fold Self-addition of the induction variable into the loop range (#228725)
scf-for-loop-range-folding only folded an arith op into the loop range
if the induction variable had a single use (hasOneUse()). That counts
operand slots rather than users, so `arith.addi %i, %i`, a single
operation that uses the induction variable twice, was rejected.
now we check whether the induction variable is used in a single
operation.
Fixes #228474
Assisted-by: Claude
textproc/py-phonemizer: New port: Simple text to phones converter for multiple languages
PR: 298595
Approved by: Baptiste Daroussin <bapt at FreeBSD.org> (on behalf of portmgr@)
devel/py-loky: New port: Robust implementation of concurrent.futures.ProcessPoolExecutor
PR: 298680
Approved by: Baptiste Daroussin <bapt at FreeBSD.org> (on behalf of portmgr@)
devel/py-cpsat-logutils: New port: Parse CP-SAT (OR-Tools) solver logs into pydantic models
PR: 298874
Approved by: Baptiste Daroussin <bapt at FreeBSD.org> (on behalf of portmgr@)
[RegAllocFast] Lower tied operands, absorbing TwoAddressInstructionPass (#225316)
TwoAddressInstructionPass inserts a copy for every tied operand it
cannot rewrite, which register allocation then tries to fold away. Teach
RegAllocFast to lower tied operands itself so that the pipeline can skip
the pass:
* A tied use that dies at the instruction takes over the tied def's
register when its class contains it, and is otherwise copied into it,
with the instruction's other reads of the value following the copy.
* Copy hints follow ties as chain links, so argument copies feeding
two-address chains still fold.
* REG_SEQUENCE and INSERT_SUBREG expand to subregister COPYs; an undef
REG_SEQUENCE source needs one only where a use reads its lane.
PHIElimination still runs, so the allocator's input is not SSA. The
lowering keys on the TiedOpsRewritten property, which the allocator now
sets itself, so partial pipelines (-run-pass, -start-before) follow the
MIR they are given.
[10 lines not shown]
[AMDGPU] Fix missed WMMA C-operand co-exec hazard
The gfx1250 WMMA co-execution hazard check treats only A, B and the
SWMMAC index as registers the in-flight MMA still reads. C (src2 of a
non-SWMMAC WMMA) is missing, so a VALU scheduled into the MMA's shadow
can clobber C and the MMA consumes the new value.
This is latent while C is tied to vdst, since the existing D check then
covers it. It miscompiles where the tie does not hold: for
v_wmma_bf16f32_16x16x32_bf16, whose D is narrower than C, and for the
_threeaddr form of any WMMA.
drm/amdgpu: hold a runtime PM reference for P2P dma-buf attachments
From Mike Lothian
aeaa7bdc0ea2c725331a423575bb7f93c040f8fb in linux-6.18.y/6.18.54
636139603b99d2e3a18a46cf3f8d39313ce8042e in mainline linux
drm/amdgpu: lock bo before calling amdgpu_vm_bo_update_shared
From Pierre-Eric Pelloux-Prayer
801d8647dcb0d34f0654b11f932e4ed365092c5e in linux-6.18.y/6.18.54
36ffc58b8a8704e690a0ce679db26baa5759256f in mainline linux
drm/amdgpu: fix rmmio iounmap skipped on device removal
From Chengjun Yao
cd55dde2b63789a3dafe982c89844f537501d096 in linux-6.18.y/6.18.54
5155002b03b24ba3ef91c5c313b8cf0171b24904 in mainline linux
18471 simnet rejects its maximum MTU and reports an inaccurate MTU range
Reviewed by: Robert Mustacchi <rm at fingolfin.org>
Approved by: Dan McDonald <danmcd at oxide.computer>
drm/amdgpu: check ras and obj before dereference
From Dmitriy Chumachenko
37583946d8751f8e285c467d770ac0b609b82a23 in linux-6.18.y/6.18.54
723d4dc628d764b19cf9efca14b82cca5ff020c9 in mainline linux
drm: Fix drm_pending_vblank_event leak in error path for out_fence_ptr
From Thadeu Lima de Souza Cascardo
daefd7ff159b0d1fb8c9e64430dcbe2aca1bdd09 in linux-6.18.y/6.18.54
9eb1a393c89a79c4210230d23e7d88d239c61d7b in mainline linux
[tsan]: fix Go race syso build on s390x with GCC (#225217)
Native GCC on s390x defines __GCC_HAVE_SYNC_COMPARE_AND_SWAP_16, which
causes the SpinMutex func_cas overload for a128 to be skipped. The
generic template emits a __sync_val_compare_and_swap_16 libcall instead
of inlining CDSG (GCC cannot prove alignment) and that symbol
is not available without libatomic, which the Go race syso does not
link.
Extend the SpinMutex condition to also cover SANITIZER_GO builds,
regardless of __GCC_HAVE_SYNC_COMPARE_AND_SWAP_16. This is safe since
all atomic accesses in Go go through the TSan trampolines.
Verified with s390x-linux-gnu-g++ (GCC 12): without the fix the syso
contains an unresolved __sync_val_compare_and_swap_16 reference; with
the fix all three __tsan_go_atomic128_* symbols are present and no
libatomic references remain.