[flang][OpenMP] Privatize loop IVs in the innermost parallel (#227486)
A sequential loop's iteration variable is predetermined private in the
innermost parallel, teams or task-generating construct that encloses
the loop. When the loop was nested in another construct, such as a
worksharing loop, lowering failed to privatize the variable in the
enclosing parallel region. Instead, it created a new local copy inside
the nested construct, so the parallel region's other references to the
variable used the shared host variable.
Fix this by deciding which construct privatizes a symbol based on the
scope that owns it. This also simplifies DataSharingProcessor: the
OMPConstructSymbolVisitor, which walked the parse tree to track where
symbols were defined, is no longer needed and has been removed.
I have noticed that metadirectives don't always own the symbols that
should be privatized in them, as semantics doesn't create a new scope.
I'm not very familiar with metadirectives, but it seems this causes some
privatization issues with non-explicitly specified DSAs, as
[7 lines not shown]
[InstCombine] Fix profiles in foldMulSelectToNegate (#229865)
The condition is always the same as the original select, so we can just
propagate the metadata after using a matcher to ensure that we get the
actual SI out.
AMDGPU: Remove remaining uses of -amdgpu-scalarize-global-loads=false
Stop relying on the option to select vector loads from uniform kernel
argument pointers.
Index loads by the workitem id where the test is about memory operations.
Use volatile loads where the test depends on a uniform value held in
VGPRs. Use functions with inreg pointer arguments for the MUBUF encoding
tests, which preserves the existing encodings. Convert the early
if-conversion tests that only need operand values into functions.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
sysconf: add a poudriere target
poudriere.conf is handed to sysrc(8), including the jail, tree, and
set overlays poudriere(8) consults. make, src, and src-env are
sub-targets over the matching make(1) files under poudriere.d. -l
and -L list one path per line; -L includes overlays that do not exist
yet. A bare -- ends the sub-target so the next word is a variable.
Bump sysconf(8) to 2.2.
Reviewed by: fuz, kfv, bcr
Differential Revision: https://reviews.freebsd.org/D59943
sysconf: accept a unique prefix of a target keyword
Exact match still wins. An ambiguous prefix is rejected with the
list of matches. Move keyword selection and the rc pass-through
into sysconf-targets.c. Bump sysconf(8) to 2.1.
Reviewed by: kfv
Differential Revision: https://reviews.freebsd.org/D59924
ZIL: avoid deadlock in xattr owner check
zfs_xattr_owner_unlinked() runs with an assigned transaction. A final
zrele() can run inode or vnode cleanup inline. During Linux writeback,
cleanup may wait for the caller's own I_SYNC state. Cleanup may also
open another transaction while the current transaction remains assigned.
The caller cannot return to commit its transaction, so the txg remains
open and cannot sync.
Treat the input znode as borrowed on all platforms. Release only parents
acquired by zfs_zget() with zfs_zrele_async(). This removes the Linux
zhold()/zrele() pair and the platform split. Check the current walk node
in the assertion.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Jaromir Hamala <jaromir.hamala at gmail.com>
Closes #19253
CI: wait for SSH after restarting sshd on FreeBSD
FreeBSD 16-current CI failed while transferring src.txz immediately
after restarting sshd: scp received Connection refused and the build
VM initialization aborted before any tests ran.
Poll SSH readiness for up to thirty attempts before transferring the
archive. Use the existing one-second connection timeout. A server
that remains unavailable still fails at the transfer.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
ZTS: wait for pool suspension and the blocked writer
The ten-poll suspension wait can expire before failed writes suspend
the pool on a busy system. Observing SUSPENDED also does not guarantee
that the background writer has reached dmu_tx_try_assign() yet.
Wait for both pool suspension and an increase in dmu_tx_suspended,
with a shared sixty-poll budget. Sample the counter before starting
the writer so earlier tests cannot satisfy the check. Detect a writer
that exits during the wait and report state and counters on timeout.
Keep the checks that the writer blocks and the pool can be resumed.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
ZTS: use IPv4 loopback explicitly in HTTPS key tests
The Python HTTPS fixture binds an IPv4 socket, while clients resolve
localhost independently and can select IPv6. Use the same explicit
IPv4 loopback address for the server and keylocation URLs. Add an IP
subject alternative name to the test certificate so verification
continues to validate the endpoint.
Debian 13 CI reset every HTTPS key request without logging a request
at the fixture. This change removes address-family ambiguity; the
original reset's precise cause is not established by the CI logs.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
ZTS: retry busy pool export in the zvol replay test
The zvol can remain briefly open for device probing after its ext4
filesystem has been unmounted. Pool export then fails with pool is
busy before the test can import the pool and verify log replay.
Use the existing bounded busy retry for export. Keep the frozen pool,
intent-log inspection, import, and checksum verification unchanged.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
AMDGPU: Use functions in more operation tests instead of kernel loads
Stop relying on -amdgpu-scalarize-global-loads=false. Inputs are passed
as VGPR arguments, or inreg for SGPR operands. Kernels that check
multiple stores index their loads by workitem id.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
ZTS: sync the resilver finishing txg before checking waiter exit
Pool status reports a finished resilver before its finishing txg has
synced. zpool wait also waits for that txg, whose config and label
writes are subject to the test's injected I/O delay. Starting the
two-second exit grace period at the status change can fail spuriously.
Sync the pool when the activity check observes completion, then check
that the waiter exits within the existing grace period. Keep the
checks for premature exit and nonzero status for both wait commands.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
ZTS: check TRIM rate limits without assuming minimum throughput
Rate limiting caps maximum throughput, but a busy vdev can trim more
slowly. The fixed sleep and cumulative progress checks failed with
32% trimmed when the test expected at least 35%.
Wait for each progress target with a bounded timeout and verify that
it did not arrive faster than the configured limit permits. Check
resumption without specifying a new rate, a higher rate, and complete
trimming at the maximum rate. Retain suspension and completion checks.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
ZTS: wait for each livelist condense before overwriting again
Condensing runs in a background thread and commits in a later txg.
Rapid overwrites can skip condense opportunities while that thread
is busy, leaving seven entries instead of the expected six.
Wait for the expected livelist length after each overwrite, with a
bounded retry and diagnostics on failure. Inspect only the test clone.
Keep the exact entry-count checks and the deactivation checks.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
[CodeView] Set the S_COMPILE3 hotpatch flag from a module flag (#229868)
Previously, the `HotPatch` flag in CodeView `S_COMPILE3` was only
derived from `TargetOptions::Hotpatch`, which is set by clang's backend
setup but not by LTO code generation. Objects produced by LTO were
therefore never marked as hotpatchable, and lld's `/FUNCTIONPADMIN`,
which only pads chunks from hotpatchable objects, left their functions
unpadded, even though the `"patchable-function"` attribute still made
code generation pad the function entries.
After this PR, emit an `"ms-hotpatch"` module flag when compiling with
`/hotpatch`, and also set the CodeView flag from it. The flag uses the
`Min` merge behavior, so when LTO merges modules, it is reset to 0 if
any module was compiled without `/hotpatch`, and the result is never
wrongly marked as hotpatchable.
This came up while working on
https://github.com/llvm/llvm-project/pull/229477
[4 lines not shown]
[LoopUnroll] Snapshot loop blocks after runtime remainder generation (#229779)
Runtime remainder unrolling can simplify the enclosing loop nest and
delete blocks from the loop being unrolled. The early OriginalLoopBlocks
snapshot then contains dangling pointers, causing a crash during
dominator-tree updates.
Capture the snapshot after remainder generation and before cloning adds
new blocks to the loop. The included test aborts without the fix with
`cannot get DomTreeNode of block with different parent (exit 134)`, and
passes lit with the fix.
Assisted by AI tooling, to triage the bug and find an LLVM IR reproducer
for the crash.
science/linux-ai-ml-env: Remove dependency on nvidia-driver
Instead, ask user to install it via pkg-message. This allows choosing between
nvidia-driver and nvidia-driver-devel, until we get provides/requires.
PR: 291139
[CIR] Put try_throw's normal destination in the throw's region (#229915)
`replaceThrowWithTryThrow` creates the unreachable normal destination of
`cir.try_throw` at the end of the parent function.
This is not correct for coroutines. Doing so would create a reference to
a block outside their respective regions, leading to a verification
error.
This patch moves the unreachable normal destination of `cir.try_throw`
in `replaceThrowWithTryThrow` right below the throw's region.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
ZTS: stabilize interior dnode reallocation test
send_realloc_dnode_interior can exhaust all retries without reaching
an interior slot of the freed dnode. In the FreeBSD CI failure, the
freed dnode was object 128 while every retry reused objects 10-17.
Deleting the filler files lets each retry revisit the same lower slots.
Keep filler files between attempts and allow enough allocations to
fill the slots below the freed dnode. Stop at the first interior slot,
then remove the filler files before taking the incremental snapshot.
Preserve the receive and directory comparison checks.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19259
[lld][MachO] Write the contents of an output section in parallel (#229512)
This is part of the ld64.lld performance improvements tracked by #222068
Writer::writeSections() does actually run the output sections in
parallel, but each section's writeTo() is not done in parallel, so
there's still a lot of work done by one thread at a time.
We can parallelize the writeTo() calls pretty easily as each input
writes to its own range in the output buffer.
This commit is output-preserving.
[InstCombine] Fold fpto[su]i.sat of NaN guarded select
fpto[su]i.sat returns 0 for both NaN and $\pm 0.0$, so a select that
replaces a NaN input with zero does not change the result, and the call
can take X directly.
```llvm
%not.nan = fcmp ord float %x, 0.0
%sel = select i1 %not.nan, float %x, float 0.0
%r = call i32 @llvm.fptosi.sat.i32.f32(float %sel)
-->
%r = call i32 @llvm.fptosi.sat.i32.f32(float %x)
```
The inverted form `select (fcmp uno X, 0.0), 0.0, X` is folded too.
Only the canonical `fcmp ord/uno X, 0.0` is matched, which also covers
`oeq/une X, X` after fcmp canonicalization.