CodeGen: Run LiveIntervals before PHIElimination and drop LiveVariables from it (#228618)
Move LiveIntervals to run before PHIElimination in the optimized register allocation
pipeline, and make PHIElimination maintain LiveIntervals only.
This removes the last explicit use of LiveVariables. The actual analysis is no longer used.
There are implicit dependencies on the side effects of running the analysis due to
adjustments of dead flags, so further work is still needed to complete the removal.
This perturbs register allocation in a number of tests. The same codegen result
can be achieved by not preserving the analysis and recomputing fresh. Greedy is
just sensitive to the exact slot index and value numbering with identical MIR.
Measured across every affected test the emitted instruction count goes from
145124 to 145196, +0.050%, with changes in both directions. The largest regression
is AArch64/phi.ll, where the GlobalISel output gains about 30
instructions and no longer matches the SelectionDAG output; the largest
improvements are ARM/fpclamptosat.ll and PowerPC/common-chain.ll.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[mlir][acc] Report independent loops not parallelized due to unstructured control flow
An independent acc loop whose body has unstructured control flow (for
example a backward GOTO) cannot be represented as a structured loop, so it
is executed sequentially. This happened silently: no remark was emitted for
the loop, and the user only saw a serial kernel without a reason.
Emit a remark for such loops stating that they are not parallelized because
their control flow could not be represented as a structured loop. This lets
users identify loops that are asserted parallel but run sequentially.
Co-Authored-By: Claude
Move 2FA force option to the current API models
## Problem
The HA synchronization change added a required `force` argument to `auth.twofactor.update`, `user.renew_2fa_secret` and `user.unset_2fa_secret`. That change was intended for 28.0, but it was written before the 27 -> 28 API version bump and its `force` field ended up in the v27 pydantic models while the v28 ones were never updated. Current API is v28, so validation never supplies the default and every call fails with `missing 1 required positional argument: 'force'`, which breaks all the 2FA integration tests.
## Solution
Moved the `force` field (defaulting to `False`) from the v27 argument models to the v28 ones, since v27 has already shipped. v27 clients keep working as their calls are adapted to the current models. Also let the `audit_extended` lambdas accept `force`, otherwise a caller passing it explicitly would get an audit entry without the username.
NAS-144203 / 28.0.0-BETA.1 / Standalone zettarepl daemon (#19941)
zettarepl is our only `multiprocessing` user within middleware, and
multiprocessing-related bugs are still being discovered (i.e.
https://github.com/truenas/middleware/pull/19837)
This PR replaces zettarepl process being managed by middleware and
communicated with using multiprocessing data structures with zettarepl
being managed by systemd and communicated with using middleware events
and method calls.
Now, zettarepl daemon is started by middleware as any other system
service. When started, it connects to the middleware and listens to
private `zettarepl.command` event that sends it task definition, task
execution commands, etc. On the other hand, zettarepl daemon can call
`zettarepl.notify` method to notify middleware of a task completion
progress or an error.
The paradigm is the same, but previously we had pickled data structures
[4 lines not shown]
[AMDGPU] Price vector f32 to f16 fptrunc by its packing form (#229950)
The base cost scalarizes the conversion and charges 4, 10, 22 and 46
for 2, 4, 8 and 16 lanes. The backend rounds every lane with
v_cvt_f16_f32 and packs the halves in pairs, which takes N + N/2
instructions. A packed conversion rounds a pair per instruction and
true16 writes a lane into either half of a register, which gives
ceil(N/2) and N.
Co-authored-by: Michael Selehov <michael.selehov at amd.com>
Assisted-By: Claude Code Opus 5
---
<sub>Stack created with <a
href="https://github.com/github/gh-stack">GitHub Stacks CLI</a> • <a
href="https://gh.io/stacks-feedback">Give Feedback 💬</a></sub>
Co-authored-by: Michael Selehov <michael.selehov at amd.com>
[NFC][AMDGPU] Add tests for the cost of vector f32 to f16 fptrunc (#229949)
Assisted-By: Claude Code Opus 5
---
<sub>Stack created with <a
href="https://github.com/github/gh-stack">GitHub Stacks CLI</a> • <a
href="https://gh.io/stacks-feedback">Give Feedback 💬</a></sub>
Add force option to current 2FA API models
## Problem
The HA synchronization change added a required `force` argument to `auth.twofactor.update`, `user.renew_2fa_secret` and `user.unset_2fa_secret`, but the matching `force` field only landed in the v27 API models. Current API is v28, so validation never supplies the default and every call fails with `missing 1 required positional argument: 'force'`, which breaks all the 2FA integration tests.
## Solution
Added the same `force` field (defaulting to `False`) to the v28 argument models. Also let the `audit_extended` lambdas accept `force`, otherwise a caller passing it explicitly would get an audit entry without the username.
[CIR][AMDGPU] Add support for AMDGCN s_sendmsg_rtn builtins (#223226)
Adds codegen for the following AMDGCN s_sendmsg_rtn builtins:
- __builtin_amdgcn_s_sendmsg_rtn
- __builtin_amdgcn_s_sendmsg_rtnl
These are lowered to the `llvm.amdgcn.s.sendmsg.rtn` intrinsic.
Assisted by: Claude Opus 5
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[mlir][acc] Report independent loops not parallelized due to unstructured control flow
An independent acc loop whose body has unstructured control flow (for
example a backward GOTO) cannot be represented as a structured loop, so it
is executed sequentially. This happened silently: no remark was emitted for
the loop, and the user only saw a serial kernel without a reason.
Emit a remark for such loops stating that they are not parallelized because
their control flow could not be represented as a structured loop. This lets
users identify loops that are asserted parallel but run sequentially.
Co-Authored-By: Claude
[Clang] Mark scoped_atomics with !noalias.addrspace(private)
The HIP specification marks atomics on thread private memory as UB.
Scoped atomics used within a HIP context are also considered UB,
unless explicitley specified via a command line argument.
These are now annotated with !noalias.addrspace(5) for amdgpus,
to avoid an expensive runtime check.
ports.7: Document how to use mdo(1) for SU_CMD
Reviewed by: olce
MFC after: 1 week
Sponsored by: fme AG
Differential Revision: https://reviews.freebsd.org/D60443
ports.7: Document how to use mdo(1) for SU_CMD
Reviewed by: olce
MFC after: 1 week
Sponsored by: fme AG
Differential Revision: https://reviews.freebsd.org/D60443
Merge tag 'pwrseq-fixes-for-v7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux
Pull power sequencing fixes from Bartosz Golaszewski:
- add missing PCI device IDs for Thinkpad T14s gen6 which should have
been part of commit a39ac4651e3b ("power: sequencing: pcie-m2: Match
WCN6855 and WCN7851 UART BT variants by subdevice ID") in
pwrseq-pcie-m2
- fix memory leak in pwrseq-pcie-m2 (leaking the array returned by
of_regulator_bulk_get_all())
* tag 'pwrseq-fixes-for-v7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux:
power: sequencing: pcie-m2: Fix leaking array from of_regulator_bulk_get_all()
power: sequencing: pcie-m2: Add Lenovo ThinkPad T14s gen6 WCN7850 subsystem PCI ids
Merge tag 'gpio-fixes-for-v7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux
Pull gpio fixes from Bartosz Golaszewski:
- fix runtime PM leak in error path in gpio-xilinx
- fix race when arming the IRQ poll worker in gpio-mpsse
- fix devres cleanup path on probe error in gpio-exar
* tag 'gpio-fixes-for-v7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux:
gpio: mpsse: fix race when arming the IRQ poll worker
gpio: exar: initialize the ID before registering its cleanup
gpio: xilinx: fix runtime PM leak on request error path
[Clang][Driver] Fix inverted diagnostic condition (#225640)
The diagnostic string for warn_drv_preprocessed_input_file_unused and
warn_drv_input_file_unused require %1/%2 to be true iff the causing
option is *un*available. Swap the option. Also use `getSpelling()` to
include the dash in the printed input.
warn_drv_input_file_unused was already part of #218802 which was
reverted.
preprocessed-input-file-unused.c test case generated by AI
[AMDGPU] Extend new hazard CFG walk for wmma instruction support (#229428)
Replaces `getWaitStatesSinceVALU` with `getMaxVALUWindowDeficit` and
aligns `checkWMMACoexecutionHazards` implementation with
`checkMAIHazards90A`. Now elides the visited set CFG walk and uses the
BFS CFG walk instead.
Fixes ROCM-32079
AI Assisted
[OpenMP] Fix debug locations in GPU reductions (#228622)
GPU reduction codegen temporarily switches the IRBuilder to AllocaIP to
create reduction storage. In Clang, AllocaIP points before the unlocated
allocapt marker, which clears the current debug location. Restoring the
insertion point does not restore the debug location, so the following
inlinable runtime calls lack !dbg.
Use InsertPointGuard for temporary alloca insertion regions, preserving
both the caller insertion point and debug location.
Remove redundant saveIP/restoreIP pairs around helper emitters that
already use InsertPointGuard internally to preserve the builder state.
Add an OpenMPIRBuilder unit test that models Clang's allocapt insertion
point and verifies the expected runtime call has a valid debug location.
Assisted by gpt-5.6.
[2 lines not shown]