science/linux-ai-ml-env: Remove dependency on nvidia-driver
Instead, ask user to install it via pkg-message. This allows choosing between
nvidia-driver and nvidia-driver-devel, until we get provides/requires.
PR: 291139
ZTS: stabilize interior dnode reallocation test
send_realloc_dnode_interior can exhaust all retries without reaching
an interior slot of the freed dnode. In the FreeBSD CI failure, the
freed dnode was object 128 while every retry reused objects 10-17.
Deleting the filler files lets each retry revisit the same lower slots.
Keep filler files between attempts and allow enough allocations to
fill the slots below the freed dnode. Stop at the first interior slot,
then remove the filler files before taking the incremental snapshot.
Preserve the receive and directory comparison checks.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Tony Hutter <hutter2 at llnl.gov>
Signed-off-by: George Melikov <mail at gmelikov.ru>
Closes #19259
tailscale: update to 1.104.1
All:
- WireGuard packet buffering uses less memory while waiting for peer
handshakes.
- WireGuard packet queue depth scales with CPU core count, reducing
peak memory usage under heavy traffic.
Linux:
- Client-side netmap caching improves start up speed and connectivity
when the Tailscale control plane is unreachable.
- WireGuard allocates a single packet buffer for batched I/O,
improving throughput and reducing memory usage.
- An issue that could prevent traffic over IPv6 link-local addresses
from flowing through direct and Tailscale Peer Relay connections is resolved.
- An issue preventing Linux desktop detection from working on Wayland
environments is resolved.
Darwin:
- Client-side netmap caching improves start up speed and connectivity
[16 lines not shown]
[MISched] Implement `UnanalyzableFrontier`-based DAG construction algorithm (#227375)
Introduces a new algorithm for constructing the control dependencies in
the schedule DAG. The new algorithm is gated behind a (temporary)
`cl::opt` (`-enable-unanalyzable-store-sequencing`), currently off by
default. With the option disabled, the change is an effective NFC: edge
insertion order changes slightly, but the DAGs do not change materially.
The difference between this algorithm and the existing algorithm is
that, rather than maintaining all unanalyzable memory operations and
repeatedly querying all of them, we instead maintain a frontier of loads
and stores since a 'sequencing store'. A sequencing store is either (a)
a store that writes to different set of base objects to the last seen
sequencing store or (b) any store if there has been a load from another
base object since the last sequencing store. We treat the sequencing
stores carefully to ensure that all previously seen unanalyzable memory
operations that do not belong to the frontier necessarily transitively
succeed the current sequencing store. This makes it sound to sequence
preceding memory operations against the sequencing store and only those
[9 lines not shown]
[AMDGPU] Form VOPD dot2 pairs with a literal in src1
A V_DOT2 with a register in src0 and a literal in src1 is matched as if
commuted, and commuted when the VOPD pair is built. This recovers the
pairs lost once MachineCSE stopped leaving those literals in src0.
Co-Authored-By: Claude <noreply at anthropic.com>
[Darwin][TSan][Test-only] Make flaky norace-objcxx-run-time.mm test diagnosable (#229931)
This test has been failing occasionally with only the hint 'exit 66' -
it may be that the FileCheck is passing due to a TSan report being
before the 'Done' log (and the CHECK-NOT being after it).
This patch replaces the CHECK-NOT with the --implicit-check-not lit
shell flag, which should allow us to see the TSan error report if it
occurs again. The patch also adjusts the barrier mechanisms.
Assisted by: claude
rdar://189381572
[CIR][CodeGen] Share the EH personality selection logic
Deduplicates the EH personality selection (`getEHPersonality`,
`getCXXEHPersonality`) between CIR and classic CodeGen, taking the classic
implementation. CIR's copy lacked the z/OS, Wasm and GNUstep-on-CygMing cases;
none are reachable in CIR today, so no test changes.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Retire the CodeGenUtils.h catch-all header
Moves the last helpers out of `CodeGenUtils.h` into `ClassUtils.h`,
`ModuleUtils.h`, `TargetUtils.h` and `FunctionUtils.h` (`checkTargetFeatures`,
since it came from CodeGenFunction.cpp) and deletes the header. Only moves code
already on main, so it can be dropped on its own.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen] Share isStandardLibraryRTTIDescriptor
Deduplicates `isStandardLibraryRTTIDescriptor` between CIR and classic CodeGen
into `ItaniumCXXABIUtils.h`, taking the classic implementation. The two copies
have been equivalent since #227781 filled in the builtin types CIR was missing.
Assisted-by: Claude Code (Claude Fable 5.1).
CodeGen: Strip LiveVariables down to dead flag computation
This analysis is dead and there are no more explicit uses. There are still
passes implicitly relying on adjustments of dead flags. Missing dead
flags are added, and implicit-def operands are added for partially dead
physical registers.
The whole pass should be deleted, but it's taking a while to get all the dead
flag changes through the rest of the compiler. As a stop-gap to try to recover
some compile time regression, and avoiding new users appearing, strip the pass
down to only commputing the dead flags.
The main side effect of this is kill flags are no longer made accurate, which
is the source of the test churn.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
LiveVariables: Only visit tracked physical registers
Keep a bitvector of physical registers with a recorded def or use in
the current block. Register mask handling, the end of block scan, and
the per-block reset now only visit those registers instead of every
register. This is significant for targets with many registers, such as
AMDGPU.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
LiveVariables: Remove dead live-in handling and Defs plumbing
No physical register is tracked at the start of a block, so handling
the block live-ins was a no-op. The Defs list was only appended for
instruction defs, which runOnInstr already collects.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[CIR][CodeGen][NFC] Share requiresAMDGPUProtectedVisibility (#227261)
Deduplicates `requiresAMDGPUProtectedVisibility` between CIR and classic CodeGen
into `TargetUtils.h`. The shared version takes a bool for "currently hidden" in
place of the `llvm::GlobalValue` and `cir::VisibilityKind` the two callers
passed.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR] Split ComplexBinOps for float and int (#226725)
Split the Complex binary operations into float and int versions and
remove the unnecessary range kind from div and mul int ops
[Offload] Fix: Target JIT registers no backends when LLVM_LINK_LLVM_DYLIB=ON (#227370)
fixes: https://github.com/llvm/llvm-project/issues/221871
related: https://github.com/llvm/llvm-project/pull/104647
#104647 fixed the `'attributor-manifest-internal' registered more than
once` issue when LLVM_LINK_LLVM_DYLIB=ON, which is caused by linking
both libLLVM.so and various archives that are found via
llvm_map_components_to_libnames for jit support, by wrapping the whole
JIT loop in `if (NOT LLVM_LINK_LLVM_DYLIB)`.
However, it also disabled LIBOMPTARGET_JIT_${target}, which led to the
`failure to jit IR image: Unable to find target for this triple (no
targets are registered)` error as the issue #221871 describes.
So, I put `target_compile_definitions(PluginCommon PRIVATE
"LIBOMPTARGET_JIT_${target}")` outside of `if (NOT
LLVM_LINK_LLVM_DYLIB)` so the JIT registers backends via the
`LLVMInitialize*` calls in JIT.cpp:
[76 lines not shown]
[CostModel][X86] Update AVX1/AVX2 uitofp v2i64 -> v2f32 costs based off worst case llvm-mca numbers (#230098)
Since #206518, SLP uses fdiv as a vectorization seed. A pair of uitofp
i64 -> float feeding fdivs by a common divisor (e.g. a complex number
with unsigned 64-bit parts divided by a float, as in Julia's
`Complex{UInt64} / Float32`) is now fully vectorized on AVX1/AVX2
targets. Without AVX512DQ, uitofp v2i64 -> v2f32 has no native
instruction: it halves large elements, converts each lane through a GPR
with cvtsi2ss and fixes the result up with blends (14 instructions).
This made the code slower than before #206518 (3.1 -> 3.5 ns per call on
Kaby Lake; 1.25x in Julia's BaseBenchmarks).
SLP picked that sequence because the AVX table costs it at 10, below
what it measures. Use the worst case llvm-mca throughput from
check_cost_tables.py across btver2, sandybridge, haswell, broadwell,
skylake, alderlake and znver1-3, which is 11 (haswell/broadwell). With
that, SLP keeps the conversions scalar and only vectorizes the fdiv, as
it already does on SSE4.1 targets.
[mlir][tosa] Extend allow-non-finites to max_pool2d (#229869)
`TosaToLinalgNamed` seeds float max pooling with `APFloat::getLargest`,
i.e. `-3.40282347E+38` for f32. The spec asks for `-infinity`:
MAX_POOL2D (§2.3.8) initializes its accumulator with
`minimum_s<in_out_t>()`, and Table 5 (§1.9) defines that as `-infinity`
for the floating-point types.
This adds an `allow-non-finites` option to `tosa-to-linalg-named` that
selects the conforming seed, mirroring the option added to
`tosa-to-linalg` in #222140. It is off by default, so nothing changes
unless you ask for it — the finite seed stays available for targets that
can't represent infinities, which includes TOSA's own `fp8e4m3_t`.
This is the follow-up I mentioned on #222140, where `max_pool2d` was
left out to keep that review small.
The `nan_mode = IGNORE` arm is deliberately untouched. The spec requires
a NaN seed there and #225744 already implements it, so between the two
[22 lines not shown]
[AArch64][llvm] Fix incorrect diagnostic (index/immediate in simm9)
Improve the error message in the simm9 diagnostics and use the word
'immediate' instead of 'index'. This was noticed when creating the
CFLT instructions for Armv9.8-A, but would have caused a lot of
unrelated churn if merged as part of that change.
Co-authored-by: Martin Wehking <martin.wehking at arm.com>
[AArch64][llvm] Fix constant expressions in CMPBR immediate aliases
Use the standard immediate parser for CMPBR aliases so that parenthesized
expressions and explicit unary plus are accepted. This also simplifies
the amount of code required, since we can remove `tryParseAdjImm0_63`,
and reuse code in `AdjImmAsmOperand` which does the same job.
Check immediates in their original ranges and apply the adjustment
when rendering MC operands. Add coverage for constant expressions in
`cbge`, `cbhs`, `cble` and `cbls`.