[lldb] Fix offload bundle unit tests on 32-bit systems (#226922)
DataExtractor::GetByteSize returns uint64_t for reasons I don't yet
understand. I might change it to size_t but let's unblock the build for
now.
Fixes #222362 / 08ee0ef5da10d27d5776df1e1ff81c46a38d9b5c.
[X86] Select standalone ~(-1 << n) masks as BZHI (#226158)
InstCombine canonicalizes (1 << n) - 1 to xor (shl -1, n), -1. When that
mask feeds an AND, X86DAGToDAGISel::matchBitExtract folds the whole
expression into BZHI, and the add form is selected as BZHI even on its
own because Select enters matchBitExtract from ISD::ADD. The xor form on
its own is not: a mask that is returned, stored, or consumed by anything
but AND is selected as mov -1; shlx; not.
Enter matchBitExtract from ISD::XOR too under BMI2, so the standalone
mask becomes mov -1; bzhi. That is one instruction shorter and avoids
the shlx+not dependency chain. BMI1-only targets keep the shift form,
since BEXTR would need the count moved into bits 15:8 first.
Assisted-by: Claude Code
ports-mgmt/pkg-be-plugin: new port
pkg-be-plugin is a pkg(8) plugin that automatically creates a ZFS
boot environment before each install, upgrade, or deinstall
transaction. If a transaction leaves the system in a broken state,
the pre-transaction BE provides a clean rollback point.
The plugin uses libbe(3) directly with no subprocess invocations of
bectl(8) or zfs(8). Boot environments are pruned automatically based
on a configurable keep count and minimum age.
Co-authored-by: Michael Osipov <michaelo at FreeBSD.org>
PR: 295305
[VPlan] Remove (X && Y) | (X && !Y) -> X combine. NFC (#219368)
We have smaller combines that can take care of this now that we process
recipes in a worklist after #213900
[GlobalISel][NFC] Restore the extended LLT flag to its saved value in tests (#224897)
AArch64GISelMITest's setUp() constructs an AArch64TargetMachine, which
unconditionally enables the process-global extended LLT flag. Ending the
ExtLLT tests with setUseExtended(false) therefore disables the flag for
every test that runs afterwards in the same process.
Save the flag's value on entry and restore it on exit instead.
[lld][LoongArch] Prevent relaxation oscillation for LA.PCRel and CALL
Relaxation of pcalau12i+addi (relaxPCHi20Lo12, isInt<22>) and
call36/call30 (relaxMediumCall, isInt<28>) can oscillate: shrinking
one section moves a symbol, which flips isInt<N> for other sites and
changes bytesDropped again.
Follow the same approach as RISCV::relaxCall: after a few passes, do
not allow remove to increase beyond the previous pass's value
(cur - delta). Pass that cap as prevRemove into the two helpers;
range checks may still clear remove (0) when the target goes out of
range.
[CIR][AMDGPU] Implement __builtin_amdgcn_*_dpp* builtins (#226469)
This commit implements the `__builtin_amdgcn_update_dpp`,
`__builtin_amdgcn_mov_dpp`, and `__builtin_amdgcn_mov_dpp8` builtins in
CIR, closely matching the implementation in
CodeGenFunction::EmitAMDGPUBuiltinExpr from OGCG.
Assisted-by: Claude Sonnet 5
Signed-off-by: Steffen Holst Larsen <sholstla at amd.com>
Reland [flang] Support scoped LICM and OpenACC capture provenance (#225401) (#226909)
Follow compute-region capture operands when checking whether scalar and
scalar-descriptor loads are safe to speculate. Preserve the existing
optional, array-element, and loop-modification safety checks.
Add an optional only-inside operation-name selector while retaining
function-scoped alias analysis. Resolve the name once per function and
compare interned operation names during ancestor traversal. The default
continues to select all loops. An explicit name selects loops with a
matching ancestor within the function, including the function itself.
Cover capture safety, host exclusion, non-OpenACC and nested scopes,
loop boundaries, unmatched names, and the function boundary.
The motivation for this is to allow running LICM only on device relevant
loop at O0.
Reland #225401 with CMakeFiles.txt change to fix shared library builds.
[CodeExtractor][Verifier] Fix OoB read when a DIExpression is used multiple times (#226857)
#224360 made fixupDebugInfoPostExtraction reuse the existing
DIExpression, but this is not sound if the expression is referenced
multiple times, as occurs with cold/hot code splitting.
This PR is the trivial fix of restricting this change to only apply if
there is a single user of the expression.
I've also added an additional verifier guard to capture these failures.
Fixes #226848
AI usage: Claude used to find a way to construct
dbg-value-arg-index-out-of-range.ll so that I could add a verifier guard
that bypassed the other existing verifier guards.
AMDGPU: Make rewrite-vgpr-mfma-to-agpr-spill-multi-store.ll less allocator sensitive
This test is sensitive to the exact split and spills which occur, and disappeared
under a future upstream improvement. Use basic RA with a fixed occupancy since it more
stably produces the spill pattern.
Also add a codegen reference test for the same kernel, so future codegen improvements are
visible.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>