[CIR] Emit inbounds/inrange for vtable address points in initializers (#227520)
Classic codegen emits `inbounds inrange(...)` for vtable address points
in global initializers: VTT entries and the vptr of a
constant-initialized object. CIR never emitted any flags for
`#cir.global_view`. It only got `inbounds` and `nuw` when LLVM's
constant folder inferred them, which stopped for a while after
https://github.com/llvm/llvm-project/pull/226904 until #227651 restored
the folding. It never got `inrange`.
This patch emits the flags in CIR instead of relying on the folder:
- Add an `address_point` flag to `#cir.global_view`, spelled
`#cir.global_view<@vtable, [vtable index, slot], address_point>`.
- Set it where CIRGen builds VTT entries and constant vptrs, and keep it
when CIRGen and `CXXABILowering` rebuild the attribute.
- When lowering to LLVM, make an address point `inbounds`, with an
`inrange` covering the one vtable in the group that it points into. The
bounds come from the vtable global's type, so they also work when the
[10 lines not shown]
[AArch64][SVE] Fix predicate type for quadword gather/scatter SDNodes (#228274)
GLD1Q_MERGE_ZERO and SST1Q_PRED use a single predicate bit per 128-bit
quadword (always nxv1i1), not one bit per data element.
SDT_AArch64_GATHER_VS and SDT_AArch64_SCATTER_VS incorrectly required
the predicate to have the same number of elements as the loaded/stored
vector via SDTCisSameNumEltsAs<0,1>, which only happened to hold for the
data width these nodes were originally used with.
Give these two nodes their own SDTypeProfile that fixes the predicate
type to nxv1i1, and update the corresponding patterns in
SVEInstrFormats.td to match.
Co-authored-by: Claude Sonnet 5.5 <noreply at anthropic.com>
[AArch64] Remove stale packed extend/shift for memory ops(NFC) (#230068)
This patch removes stale extend/shift representation for memory ops.
They use packed representation which has been replaced by operands to
the instruction format.
[AMDGPU] Model gfx1250-strict WMMA latencies
gfx1250-strict shares the gfx1250 scheduling model, but some WMMA
instructions take longer on it:
- 16x16x64 FP8/BF8: 8 cycles instead of 4.
- f8f6f4: 16 cycles when any input is f8 and 8 cycles otherwise, instead of
8 and 4.
The WMMA co-execution hazard category is derived from the latency, so this
also gives these instructions the required number of wait states on
gfx1250-strict.
Read the f8f6f4 matrix formats through named operands for both MCInst and
MachineInstr, so llvm-mca models the strict latencies too. This also fixes the
both-f4 check, which read fixed operand indices and could take other
immediates for f4 formats: inline constant scales of the scaled forms, which
gave too few hazard wait states, and the neg_lo/neg_hi operands of the
unscaled form in assembler input.
[flang][OpenMP] Privatize loop IVs in the innermost parallel (#227486)
A sequential loop's iteration variable is predetermined private in the
innermost parallel, teams or task-generating construct that encloses
the loop. When the loop was nested in another construct, such as a
worksharing loop, lowering failed to privatize the variable in the
enclosing parallel region. Instead, it created a new local copy inside
the nested construct, so the parallel region's other references to the
variable used the shared host variable.
Fix this by deciding which construct privatizes a symbol based on the
scope that owns it. This also simplifies DataSharingProcessor: the
OMPConstructSymbolVisitor, which walked the parse tree to track where
symbols were defined, is no longer needed and has been removed.
I have noticed that metadirectives don't always own the symbols that
should be privatized in them, as semantics doesn't create a new scope.
I'm not very familiar with metadirectives, but it seems this causes some
privatization issues with non-explicitly specified DSAs, as
[7 lines not shown]
[InstCombine] Fix profiles in foldMulSelectToNegate (#229865)
The condition is always the same as the original select, so we can just
propagate the metadata after using a matcher to ensure that we get the
actual SI out.
AMDGPU: Remove remaining uses of -amdgpu-scalarize-global-loads=false
Stop relying on the option to select vector loads from uniform kernel
argument pointers.
Index loads by the workitem id where the test is about memory operations.
Use volatile loads where the test depends on a uniform value held in
VGPRs. Use functions with inreg pointer arguments for the MUBUF encoding
tests, which preserves the existing encodings. Convert the early
if-conversion tests that only need operand values into functions.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
sysconf: add a poudriere target
poudriere.conf is handed to sysrc(8), including the jail, tree, and
set overlays poudriere(8) consults. make, src, and src-env are
sub-targets over the matching make(1) files under poudriere.d. -l
and -L list one path per line; -L includes overlays that do not exist
yet. A bare -- ends the sub-target so the next word is a variable.
Bump sysconf(8) to 2.2.
Reviewed by: fuz, kfv, bcr
Differential Revision: https://reviews.freebsd.org/D59943
sysconf: accept a unique prefix of a target keyword
Exact match still wins. An ambiguous prefix is rejected with the
list of matches. Move keyword selection and the rc pass-through
into sysconf-targets.c. Bump sysconf(8) to 2.1.
Reviewed by: kfv
Differential Revision: https://reviews.freebsd.org/D59924
ZIL: avoid deadlock in xattr owner check
zfs_xattr_owner_unlinked() runs with an assigned transaction. A final
zrele() can run inode or vnode cleanup inline. During Linux writeback,
cleanup may wait for the caller's own I_SYNC state. Cleanup may also
open another transaction while the current transaction remains assigned.
The caller cannot return to commit its transaction, so the txg remains
open and cannot sync.
Treat the input znode as borrowed on all platforms. Release only parents
acquired by zfs_zget() with zfs_zrele_async(). This removes the Linux
zhold()/zrele() pair and the platform split. Check the current walk node
in the assertion.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Jaromir Hamala <jaromir.hamala at gmail.com>
Closes #19253
CI: wait for SSH after restarting sshd on FreeBSD
FreeBSD 16-current CI failed while transferring src.txz immediately
after restarting sshd: scp received Connection refused and the build
VM initialization aborted before any tests ran.
Poll SSH readiness for up to thirty attempts before transferring the
archive. Use the existing one-second connection timeout. A server
that remains unavailable still fails at the transfer.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
ZTS: wait for pool suspension and the blocked writer
The ten-poll suspension wait can expire before failed writes suspend
the pool on a busy system. Observing SUSPENDED also does not guarantee
that the background writer has reached dmu_tx_try_assign() yet.
Wait for both pool suspension and an increase in dmu_tx_suspended,
with a shared sixty-poll budget. Sample the counter before starting
the writer so earlier tests cannot satisfy the check. Detect a writer
that exits during the wait and report state and counters on timeout.
Keep the checks that the writer blocks and the pool can be resumed.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
ZTS: use IPv4 loopback explicitly in HTTPS key tests
The Python HTTPS fixture binds an IPv4 socket, while clients resolve
localhost independently and can select IPv6. Use the same explicit
IPv4 loopback address for the server and keylocation URLs. Add an IP
subject alternative name to the test certificate so verification
continues to validate the endpoint.
Debian 13 CI reset every HTTPS key request without logging a request
at the fixture. This change removes address-family ambiguity; the
original reset's precise cause is not established by the CI logs.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
ZTS: retry busy pool export in the zvol replay test
The zvol can remain briefly open for device probing after its ext4
filesystem has been unmounted. Pool export then fails with pool is
busy before the test can import the pool and verify log replay.
Use the existing bounded busy retry for export. Keep the frozen pool,
intent-log inspection, import, and checksum verification unchanged.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
AMDGPU: Use functions in more operation tests instead of kernel loads
Stop relying on -amdgpu-scalarize-global-loads=false. Inputs are passed
as VGPR arguments, or inreg for SGPR operands. Kernels that check
multiple stores index their loads by workitem id.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
ZTS: sync the resilver finishing txg before checking waiter exit
Pool status reports a finished resilver before its finishing txg has
synced. zpool wait also waits for that txg, whose config and label
writes are subject to the test's injected I/O delay. Starting the
two-second exit grace period at the status change can fail spuriously.
Sync the pool when the activity check observes completion, then check
that the waiter exits within the existing grace period. Keep the
checks for premature exit and nonzero status for both wait commands.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
ZTS: check TRIM rate limits without assuming minimum throughput
Rate limiting caps maximum throughput, but a busy vdev can trim more
slowly. The fixed sleep and cumulative progress checks failed with
32% trimmed when the test expected at least 35%.
Wait for each progress target with a bounded timeout and verify that
it did not arrive faster than the configured limit permits. Check
resumption without specifying a new rate, a higher rate, and complete
trimming at the maximum rate. Retain suspension and completion checks.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
ZTS: wait for each livelist condense before overwriting again
Condensing runs in a background thread and commits in a later txg.
Rapid overwrites can skip condense opportunities while that thread
is busy, leaving seven entries instead of the expected six.
Wait for the expected livelist length after each overwrite, with a
bounded retry and diagnostics on failure. Inspect only the test clone.
Keep the exact entry-count checks and the deactivation checks.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19265
[CodeView] Set the S_COMPILE3 hotpatch flag from a module flag (#229868)
Previously, the `HotPatch` flag in CodeView `S_COMPILE3` was only
derived from `TargetOptions::Hotpatch`, which is set by clang's backend
setup but not by LTO code generation. Objects produced by LTO were
therefore never marked as hotpatchable, and lld's `/FUNCTIONPADMIN`,
which only pads chunks from hotpatchable objects, left their functions
unpadded, even though the `"patchable-function"` attribute still made
code generation pad the function entries.
After this PR, emit an `"ms-hotpatch"` module flag when compiling with
`/hotpatch`, and also set the CodeView flag from it. The flag uses the
`Min` merge behavior, so when LTO merges modules, it is reset to 0 if
any module was compiled without `/hotpatch`, and the result is never
wrongly marked as hotpatchable.
This came up while working on
https://github.com/llvm/llvm-project/pull/229477
[4 lines not shown]
[LoopUnroll] Snapshot loop blocks after runtime remainder generation (#229779)
Runtime remainder unrolling can simplify the enclosing loop nest and
delete blocks from the loop being unrolled. The early OriginalLoopBlocks
snapshot then contains dangling pointers, causing a crash during
dominator-tree updates.
Capture the snapshot after remainder generation and before cloning adds
new blocks to the loop. The included test aborts without the fix with
`cannot get DomTreeNode of block with different parent (exit 134)`, and
passes lit with the fix.
Assisted by AI tooling, to triage the bug and find an LLVM IR reproducer
for the crash.
science/linux-ai-ml-env: Remove dependency on nvidia-driver
Instead, ask user to install it via pkg-message. This allows choosing between
nvidia-driver and nvidia-driver-devel, until we get provides/requires.
PR: 291139