Always write INFO/version when storing an SMB share ACL
store_share_acl only wrote the INFO/version key into share_info.tdb when
the file did not exist. This fixes a regression introduced when we added
stateful failover for SMB to truenas 26 and adds test coverage.
[CIR][NFC] Add missing NYI handling for some x86 builtins (#230609)
While doing a recent code review, I noticed that there were some x86
builtins that were incorrectly falling through to code that handles
builtins below them in a switch. This change adds an errorNYI diagnostic
rather than falling through.
AMDGPU: Index loads by workitem id in tests shared with r600 (#230667)
These tests relied on -amdgpu-scalarize-global-loads=false to select
vector loads from uniform pointer arguments. They also have r600 run
lines, so keep the kernels and index the input pointers by the workitem
id instead. Also fix shl_v2i16 not using its computed pointers, and
v_shl_32_i64 using the workgroup id. Also add some uniform variants
of some cases.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
fix(AMDGPU): guard OR folds with shared conditions
A shared uniform condition still needs materializing after folding a
divergent OR to a select, and sharing can introduce an extra SCC
conversion. The extension's single-use check does not prevent this
code-size regression.
Check the condition's uses and divergence as well. Add regression and
control cases, and consolidate the boolean OR tests in or.ll.
[CIR] Pass x87 long double vectors on x86_64 (#230302)
Vectors of x87 long double now go through x86_64 calling-convention
lowering instead of hitting NYI. The ABI library sizes their elements at
128 bits like clang does, so the signatures match classic codegen.
Unions are still moved as a value of their storage type, and a long
double stores only 10 of its 16 bytes. So a union holding an x87 value
next to another member stays NYI unless it's a plain long double and the
other members fit in those 10 bytes. That also stops a silent miscompile
of unions like `union { long double ld; char c[16]; }`.
Assisted-by: Cursor / Claude Opus 5.5
[RISCV] Reassociate add (add X, (ext Y)), (ext Z) -> add (add (ext Y), (ext Z)), X (#230657)
If Y and Z are extended from the same bitwidth, then we can reassociate
it so they are in the same inner add.
This then allows combineBinOpOfExt to kick in and narrow the inner add
to vwadd.vv. The motivating case is the sad_* kernels in x264.
Reassociating them allows more arithmetic to stay in a smaller LMUL and
takes up to 40% less cycles on sad_16x16 on the K3. This gives a ~3.1%
speedup on 525.x264_r overall.