AMDGPU: Use LiveIntervals instead of LiveVariables in SIOptimizeVGPRLiveRange
Drop the LiveVariables dependency and the hand-written VarInfo
maintenance; the pass already knew how to recompute the affected
intervals with LiveIntervals, so make that the only path.
The legacy pass manager cannot schedule a pass requiring both
LiveIntervals and LiveVariables here, since LiveVariables (and its
UnreachableMachineBlockElim dependency) invalidates the LiveIntervals
just computed for it. In the AMDGPU pipeline, anchor the pass after
MachineLoopInfo instead of PHIElimination so LiveIntervals is computed
before PHIElimination, which is required anyway: the pass needs SSA and
introduces new PHIs. Preserving SlotIndexes keeps the transitive
last-user chain intact.
Since LiveVariables is no longer maintained past this point, PHIElimination
now splits critical edges using LiveIntervals, which accounts for most of
the test churn: LiveIntervalCalc drops kill flags and adds dead flags.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
Gate SMB fast path on the Veeam entitlement
This commit adds changes to drop the separate SMB_FASTPATH license feature and gate the ZFS block cloning / integrity streams smb.conf parameters on SMB_VEEAM instead, since Veeam Fast Clone depends on them and the two were never meant to be licensed independently.
Mips: Mark the $gp setup copy dead when erasing the call's $gp use (#227216)
MipsOptimizePICCall drops the implicit $gp operand from a call when the
lazy binding stub for the callee has already run. The copy that set $gp
up for that call then has no reader left, but nothing flagged it, so the
MIR carried a live def until a later liveness recomputation cleaned it
up.
The pass already walks each block in order, so track the reaching
definition of $gp as it goes and mark it dead when the use is erased.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[libc++] Encode the standard version in the ABI tag (#218527)
This prevents ODR mismatch issues from biting us across standard
versions. The order of elements in the ABI tag is now: libc++ version,
hardening mode, assertion semantic, exceptions, standard version.
Fixes #218524
[AArch64] ISel support for nxv1i1 vector_reverse (#226939)
This ensures llvm.vector.reverse can be selected for nxv1i1 types. These intrinsics can be generated by LoopVectorizer for `VF = vscale x 1` and loops iterating in reverse order.
[AMDGPU] Use S_CMP to lower a copy of a lane mask to SCC (#221445)
SIFixSGPRCopies lowers a copy of a lane mask to SCC as an AND with EXEC,
whose destination register is created by the pass and is always dead.
The AND with EXEC is only needed because SCC has to be "any active lane
is set". When the lane mask already has 0 in the bits of all inactive
lanes, that is just "the mask is non-zero", which S_CMP computes without
needing a destination register. In some cases the S_CMP can be optimized
away later by SIInstrInfo::optimizeCompareInstr.
Co-authored-by: Claude Opus 5 (1M context) <noreply at anthropic.com>
[LifetimeSafety] Fix off-by-one crash in `lifetime_capture_by` argument indexing (#227231)
Fixes an off-by-one crash in the lifetime_capture_by attribute handling.
The CapturingArgIdx is already an index into the full Args array, but
the code was incorrectly creating a CallArgs array (dropping the first
argument for instance methods) and then indexing into it with
CapturingArgIdx, causing incorrect argument access and potential
out-of-bounds crashes.
Currently crashes at head: https://godbolt.org/z/7rYczK1z6
[GVN] Move `GVNPass` and `GVNLeaderMap` out of `GVN.h` and into 'GVN.cpp` (NFC)
Rename `GVNPass` to `GVNPassImpl`. Leave in the header only
the class for interfacing with the pass manager.
X86: Fix missing ... separator between functions in mir test (#227228)
The update_mir_test_checks output is incomplete so the later functions
here were really manually checked.
[BFI] Use MapVector in combineWeightsByHashing. (#227027)
combineWeightsByHashing iterated over a DenseMap, which depends on the
hash. Use MapVector to get consistent results, independent of index
type/values.
This is mainly to ensure consistency between users that use different
index numbers, i.e. used for IR BFI and VPlan's use.
PR: https://github.com/llvm/llvm-project/pull/227027
[C++20] [Modules] Load friends for classes in ADL (#219094)
Close https://github.com/llvm/llvm-project/issues/218228
The root cause of the problem is the corresponding friend is not loaded
at the point of ADL.
This patch tries to fix this simply by loading the friends at the point
of ADL. Note that this may be best efficient if there are a lot of
friends. We just think it is rare. If it is really possible, we can
change the structure of friends from a list to a name lookup table.
(cherry picked from commit 01aedf3b34325ad2f74a77325a2da9a7e36ae3d5)
[ld64.lld, llvm-otool] Minimal arm64e.x1 support (#222721)
Just enough for `llvm-otool -hv` to dump the cpusubtype, and for
ld64.lld to not reject .tbd files that have an arm64e.x1 slice.
This is needed to link mac binaries against the macOS 27 SDK.
(cherry picked from commit b8007a8e4020b8bca2b12e941660e10bf5bf6716)
[Windows] Don't export clang::interp symbols for plugins (#221295)
With `LLVM_EXPORT_SYMBOLS_FOR_PLUGINS=ON`, `clang.exe` built from the
23.x release branch exports 66679 symbols on x86_64 Windows, over the
65535 limit of the PE export table, so it no longer links:
```
lld-link: error: too many exported symbols (got 66679, max 65535)
```
For comparison, the same configuration on 22.1.8 exports 65465 symbols,
so the headroom was already almost gone.
`clang::interp`, the constant expression bytecode interpreter, accounts
for roughly 4500 of the exported symbols (counted with `llvm-readobj
--coff-exports` on a linked aarch64 23.1.0 `clang.exe`). Its headers
live in `clang/lib/AST/ByteCode` and are not installed; the only mention
in a public header is the forward declaration of `interp::Context` in
`ASTContext.h`. A plugin cannot call into it, so there is no reason to
[21 lines not shown]
[SelectionDAG] Fix result index and vector width in unrollExpandedOp (#225886)
Fixes #224127.
In `DAGTypeLegalizer::WidenVectorResult`, `unrollExpandedOp` computes
the unroll count and widened vector type from the result being legalized
(`ResNo`) rather than unconditionally using result 0. For multi-result
nodes where result types differ (e.g. `ISD::FFREXP`), this prevents
mismatched vector widths and preserves the correct result index from
`DAG.UnrollVectorOp`.
Assisted-by: Claude
---------
Co-authored-by: Demetrios Chiuratto Agourakis <agourakis82 at gmail.com>
(cherry picked from commit 5cab2963e13496c272a482145b11796e9e26eeb4)
[CI] Bump python dependencies
Bump CI deps for 23.x specifically. Bumping everything like we did in
main causes some test failures in MLIR due to some changes that happened
there after the 23.x branch was cut.
[VPlan] Use BlockFrequencyInfo's full mass for execution frequencies. (#227006)
Use BlockFrequencyInfo's full mass instead of custom 2^63 for "always
execute".
This updates VPlan's block frequencies to match BFI.
PR: https://github.com/llvm/llvm-project/pull/227006
net/samba424: fix sys_proc_fd_path() buffer type in vfs_freebsd.c
At all 5 call sites, vfs_freebsd.c declared a plain "char buf[PATH_MAX]"
and passed its address to sys_proc_fd_path(fd, &buf). Samba 4.24
changed that function's signature to
"char *sys_proc_fd_path(int fd, struct sys_proc_fd_path_buf *buf)",
so "&buf" here has type "char (*)[PATH_MAX]", not the expected
"struct sys_proc_fd_path_buf *" - a type mismatch left over from an
incomplete migration of this patch to the new API (the sibling
patch-source3_modules_vfs__zfsacl.c was already updated correctly).
Change all 5 declarations to "struct sys_proc_fd_path_buf buf;" to
match the current signature.
PR: 298884
Co-Authored-By: Claude Sonnet 5 <noreply at anthropic.com>
Approved by: samba (kiwi)
[Openmp] Add OMPT device tracing support tests
These tests concern functionality used by parts of the OMPT device
tracing implementation.
Assisted-by: Claude Code
[Offload][OMPT] Add tracing orchestration and libomptarget integration
Wire the OMPT device tracing subsystem into libomptarget, completing the tracing pipeline from record production through buffer management to tool delivery. Include the synchronous fallback and initialization fixes in the integration that requires them.
Assisted-by: Claude Code
[Offload][OMPT] Route device callbacks through OmptProfilerTy
Compile and activate the complete OMPT profiler backend now that its tracing dependencies are available, and remove the superseded direct callback dispatch from PluginInterface in the same transition.
Assisted-by: Claude Code