[RISCV] Correct violations of SDTCisSameNumEltsAs. (#230612)
We missed a conversion to scalable vector in
lowerVectorMaskVecReduction.
lowerINSERT_SUBVECTOR used an all 1s mask with the wrong type. This made
a call to getDefaultVLOps unnecessary as both of its Mask and VL are
overridden.
Assisted-by: Claude
CodeGen: Compute live-outs of a split critical edge while updating LiveIntervals (#230499)
Fill in the new block's live-out set in the same loop that updates the
live intervals, instead of walking the original block's set a second
time with another liveAt query per register. A register is live out of
the new block if it is a PHI source or is live into Succ, which the
update already checks.
Instructions retired in phi-node-elimination, x86_64 -O3, on a generated
chain of N compare blocks branching to a shared PHI block:
N before after after/before
1k 94,119,121 88,904,465 0.94
2k 353,559,942 330,642,394 0.94
4k 1,201,233,649 1,136,359,710 0.95
8k 4,305,218,922 4,136,184,186 0.96
16k 16,076,293,247 15,539,552,966 0.97
gcc-c-torture compile/20001226-1.c (liveintervals,phi-node-elimination
[2 lines not shown]
[flang][CUDA] Safely verify registered kernel symbols (#230588)
`cuf.register_kernel` resolved GPU module and kernel symbols during
nested function verification, potentially racing with concurrent changes
to the sibling GPU module. Add `SymbolUserOpInterface` and move these
lookups to `verifySymbolUses()`, where they run as part of module-level
symbol verification while the sibling IR is stable.
bnxt_en: dcb: stop zeroing ETS and PFC config in bnxt_dcb_init()
bnxt_dcb_init() currently pushes all-zero ETS and PFC settings to
the firmware at every attach. This reserves bandwidth allocation
for RoCE traffic, which prevents L2 traffic from reaching line rate.
These settings should only be updated with valid values when the
RoCE driver is loaded. Therefore, we should stop programming them
during bnxt_dcb_init().
Signed-off-by: Andy Gospodarek <gospo at broadcom.com>
Reviewed by: chandrakanth.patil_broadcom.com, gallatin
MFC after: 2 weeks
Sponsored by: Broadcom Inc.
Differential Revision: https://reviews.freebsd.org/D60516
[flang][cuda] initialize the CUDA module after all units register managed variables (#230350)
With relocatable device code, every unit registers its managed variables
on the same CUDA module, and the runtime does not fill variables
registered after the module is initialized. Move CUFInitModule to a
second constructor with the next priority value, so that it runs after
the registration constructors of every unit in the executable or shared
library.
This will fix segmentation fault in example like:
```
module m1
integer, managed :: x1
contains
attributes(global) subroutine k1()
x1 = 1
end subroutine
end module
[20 lines not shown]
[SeparateConstOffsetFromGEP] Track cast state during offset extraction (#229839)
Keep the cast state used while searching for a constant offset and reuse
it when rebuilding the GEP index. This lets constants be extended or
truncated at the point where they are found, instead of carrying
separate sign/zero extension flags and then redistributing casts in a
second pass.
This also allows the RHS-of-sub zero-extension case because the constant
is zero-extended before it is negated.
devel/lfcbase: 1.23.6 -> 1.23.7, databases/cego: 2.54.39 -> 2.54.41
lfcbase:
- in BigDecimal::scaleTo, rounding mode HALFUP was not treated correctly
cego:
- cgclt
show systemspace
now aggregates space type and indicate usage in percent for each space
- Added flag _stdNullComp to support standard SQL Null value
comparison behaviour ( comparison returns false in any case )
As default, cego specific comparison is enabled, where null values
are the lowest values ( e.g. -1 > null returns true )
Note : This ordering is also used for btree value sorting,
where a well defined ordering for null values is needed.
This SQL standard behaviour can be enabled be setting
STDNULLCOMP=ON in the database xml file
arc: fix race between arc_release() and arc_read_done()
Consider the following scenario:
1. arc_release() is called on hdr with one buf, but which has
IO_IN_PROGRESS (reading more raw data in encypted pool while
keeping decrypted data in buf).
2. arc_release() moves hdr to anon state and discards its identity.
3. arc_read_done() is called, adds the 2nd buf to hdr, increasing
b_refcnt to 2.
Now we have hdr in anon state with two bufs and without identity.
Or here's a racing scenario:
1. arc_release() checked that hdr is not in anon state, but before
taking hash_lock
2. arc_read_done() takes hash_lock, moves hdr to anon state, in
case of an error.
[28 lines not shown]
L2ARC: Reorder header destruction for in-flight L2 writes
With multiple L2ARC devices, headers can be destroyed asynchronously
(e.g., during zpool sync) while L2_WRITING is set. The original code
destroyed L2HDR before L1HDR, causing ABDs to lose their device
association (b_l2hdr.b_dev) when arc_hdr_free_abd() is called.
This caused ABDs to be added to the global free-on-write list without
device information. When any L2ARC device completed its write and
attempted to free these orphaned ABDs, it would panic on
ASSERT(!list_link_active(&abd->abd_gang_link)) because the ABD was
still part of another device's vdev_queue I/O aggregation gang.
Fix by extending l2ad_mtx lock scope to cover L1HDR destruction and
reordering to destroy L1HDR before L2HDR when L2_WRITING is set. This
ensures arc_hdr_free_abd() can access b_l2hdr.b_dev to properly tag
ABDs with their device for deferred cleanup.
Reviewed-by: Alexander Motin <alexander.motin at TrueNAS.com>
[5 lines not shown]
L2ARC: Preserve L2HDR in arc_release() for in-flight writes
When arc_release() is called on a header with a single buffer and
L2_WRITING set, the L2HDR must be preserved for ABD cleanup (similar
to the arc_hdr_destroy() case). If we destroy the L2HDR here, later
arc_write() will allocate a new ABD and call arc_hdr_free_abd(),
which needs b_l2hdr.b_dev to properly defer ABD cleanup, causing
VERIFY(HDR_HAS_L2HDR(hdr)) to fail.
Allocate a new header for the buffer in the single_buf_l2writing
case (single buffer + L2_WRITING), leaving the original header with
L2HDR intact. The original header becomes an "orphan" (no buffers, no
b_pabd) but retains device association for ABD cleanup when
l2arc_write_done() completes.
The shared buffer case (HDR_SHARED_DATA) is excluded because L2ARC
makes its own transformed copy via l2arc_apply_transforms(), so the
original ABD is not used by the L2 write. The header can be safely
reused without allocating a new one.
[14 lines not shown]
RuntimeLibcalls: Require system library members to be libraries (#230306)
Every SystemRuntimeLibrary now lists only LibcallLibrary and LibraryRef
members, so the inline path that expanded unhomed RuntimeLibcallImpl
members directly into the system's block is dead. Remove it, and error
on any member that is not a library.
This drops the SystemAvailableImpls bitset and the predicate groups
from the system setup function. Conditional and calling-convention
groups now only appear inside a LibcallLibrary, whose own setup
function emits them. The generated RuntimeLibcalls.inc is unchanged.
Co-authored-by: Claude Opus <noreply at anthropic.com>
RuntimeLibcalls: Organize functions into libraries (#229562)
Associate runtime functions with the library which provides them. Add
shared LibcallLibrary defs, like compiler-rt, libm and libc, along with
OS and target specific variants, and replace each target's hand-listed
SystemRuntimeLibrary body with references to them. The resulting libcall
sets are unchanged.
Targets which differ from the shared libraries opt out of individual
functions, e.g. AVR has no sin, cos or sincos, x86 has no fp128 sincosl,
and PPC replaces some f128 compiler-rt helpers. Libraries which should
not merge with the generic variants are Isolated, like arm64ec's
compiler-rt and SPIRV's libc.
Co-authored-by: Claude Opus <noreply at anthropic.com>
[AMDGPU] Split negative global offsets into saddr and voffset on gfx1250 (#224527)
gfx1250 sign extends the saddr vector offset, so a big negative constant
fits in a VGPR and needs no 64-bit address add
FreeBSD: do not clear dirty bits outside a sub-block write
page_busy() shrinks the written range to DEV_BSIZE boundaries. A write
that starts and ends inside one block leaves nbytes at -DEV_BSIZE, and
vm_page_clear_dirty() then clears that block and every one above it. A
page dirtied through mmap goes clean and the store is never written.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Alexander Motin <alexander.motin at TrueNAS.com>
Signed-off-by: Nick Price <nprice at FreeBSD.org>
Closes #19268
ctf*: exit with error upon terminate()
The initial port of the CTF tools had a FreeBSD-specific patch to print
the termination message but exit with a 0 status, with a goal of getting
as much to build as possible and silently ignoring any issues.
We're now past the point where silently ignoring failures makes sense.
Any future issues need to be found and addressed.
PR: 276826
PR: 276930 [exp-run]
Reviewed by: markj
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D43743
[AMDGPU] Let SIOptimizeVGPRLiveRange handle AV-class virtual registers (#226527)
The pass filters candidates with `isVectorRegister()`, which is false
for AV classes, so most vector virtual registers on gfx90a+ have been
skipped since #166482 and #166483.
Accept virtual registers whose class has vector registers and no SGPRs.
Similar to #211560.
Co-Authored-By: Claude Opus 5.5 <noreply at anthropic.com>
---------
Co-authored-by: Matt Arsenault <Matthew.Arsenault at amd.com>