AMDGPU: Fix SILoadStoreOptimizer dropping gds bit on DS merges
When forming read2/write2 from a pair of DS_READ/DS_WRITE
instructions, the gds operand was unconditionally set to 0, turning
GDS accesses into LDS accesses. Preserve the gds bit, refuse to pair
an LDS access with a GDS access, and select the M0-reading opcode
variant for GDS on targets that otherwise use the _gfx9 forms.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[lldb][NativePDB] Require native for thread locals (#229480)
Skips the test for builds that are not native and intends to fix the
still broken bot after #229340.
The test shouldn't require `native`, but, as I mentioned in the comment,
it seems like the linking in `build.py` doesn't set the correct
architecture or the build doesn't happen on the correct target. Either
way, I can't debug this locally, so this configuration is skipped.
hwpmc: handle delayed IBS NMIs on Zen 6
On Zen 6, an extra IBS NMI can arrive after later samples. Keep the
credit until the empty NMI arrives, and handle fetch and op samples when
both are ready.
Reviewed by: mhorne
Fixes: 34b00ed041a4 ("hwpmc: fix IBS fetch and op NMI handling")
Fixes: e51ef8ae490f ("hwpmc: Initial support for AMD IBS")
Sponsored by: AMD
Differential Revision: https://reviews.freebsd.org/D60367
pmc.h: bump PMC_VERSION_MINOR
Bump for the addition of PMC_OP_GETCAPS and the recently added Intel
CPUs.
Sponsored by: The FreeBSD Foundation
(cherry picked from commit e39d3a6b32331437da6c13a4aeb67e5bcca67625)
libpmc: Query hwpmc for caps
This change allows for fine-grained capabilities per counter index. This
is particularly useful for AMD where subclasses are not exposed to the
general PMC code, but other architectures also have asymmetric behaviors
when it comes to specific counter indices.
A new PMC_OP_GETCAPS op is added to the hwpmc(4) ioctl interface.
Reviewed by: mhorne
Sponsored by: Netflix
Pull Request: https://github.com/freebsd/freebsd-src/pull/2058
(cherry picked from commit 44a983d249d05d932b6cff333f130baf70febc22)
[CIR] Pass only 'this' to an inherited ctor from a virtual base
The base-object variant of an inheriting constructor whose inherited
constructor lives in a virtual base takes no arguments, since it doesn't
construct that base. We were hitting an NYI here. This change passes
just 'this', as classic codegen does.
Assisted-by: Claude Code / Claude Opus 5.5
lua-language-server: update to 3.19.1.
Un-BROKENs this package, patches were merged.
Over two years of development.
Provided by Sunil Nimmagadda in PR 60857.
[NVPTX] Use register-or-immediate operands for fns (#229224)
The PTX ISA is very abstract and high level, it supports immediate or
register operands in almost any instruction.
Prototype simplifying the NVPTX MIR opcodes and instruction selection by
no longer discriminating between registers and immediates.
[clang][deps] Track directory dependencies of modules (#222202)
In some cases a module depends on the contents of a directory in addition to
individual files. This happens for umbrella modules and framework modules, where
adding a header changes the module without changing any of its input files.
This records those directories in the module file, and adds
`-fmodules-validate-directory-dependencies` to treat an implicitly built module
as out of date when one of them changed after the module was built. It is off
by default.
Assisted-by: Claude Code: opus-5.5
[AMDGPU] Omit hardwired-on SRAMECC ELF mode
Mirror #227740 for SRAMECC. Without on/off modes, SRAMECC is implied
by EF_AMDGPU_MACH; encoding a mode makes consumers infer an
unsupported :sramecc+ modifier.
Depends on #225540.
Change-Id: I7bf76efbb86c7030eae36dc06ca7da0dd64cce87
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
Add alloc factor logic
RAIDz vdevs always perform allocations in multiples of nparity + 1
sectors. This is to prevent fragmentation from leaving small scattered
records everywhere that cannot be allocated and just pollute the
spacemaps. This rule was always preserved until the introduction of the
allocation range code, which allows allocation of not just specific
sizes, but anywhere in a range. The sizes that end up being allocated
may not be a multiple of nparity + 1 sectors in certain cases (primarily
when the allocation is at the end of a metaslab). If that happens, when
converting from the asize to the psize and back, we can end up with two
different psize values, which can theoretically cause issues in a few
different ways.
This PR adds logic to the metaslab code to check if there is a factor
that the allocations must be a multiple of, and if we're doing a
dynamically sized allocation, ensures that the allocation meets that
requirement.
[5 lines not shown]
devel/py-cysignals: use USE_PYTHON=autoplist
Fixes packaging with free-threaded Python.
Also specify USE_PYTHON=pytest instead of USES=pytest
PR: 299133
Approved by: thierry (maintainer)
(cherry picked from commit 578ce00b23cfeb275b0c1872f0f70f64a5784ad7)
pmc.h: bump PMC_VERSION_MINOR
Bump for the addition of PMC_OP_GETCAPS and the recently added Intel
CPUs.
Sponsored by: The FreeBSD Foundation
(cherry picked from commit e39d3a6b32331437da6c13a4aeb67e5bcca67625)
libpmc: Query hwpmc for caps
This change allows for fine-grained capabilities per counter index. This
is particularly useful for AMD where subclasses are not exposed to the
general PMC code, but other architectures also have asymmetric behaviors
when it comes to specific counter indices.
A new PMC_OP_GETCAPS op is added to the hwpmc(4) ioctl interface.
Reviewed by: mhorne
Sponsored by: Netflix
Pull Request: https://github.com/freebsd/freebsd-src/pull/2058
(cherry picked from commit 44a983d249d05d932b6cff333f130baf70febc22)
[CIR][NFC] Share LLVM checks with classic in inherited-ctors.cpp
The OGCG checks repeated the LLVM checks except for one call and one
vtable store in VirtualDelegatingCtor. This change runs the LLVM prefix
on both outputs and keeps LLVMCIR/OGCG only for those two lines.
Assisted-by: Claude Code / Claude Opus 5.5
[DebugInfo] Add DILayerLoc/DILayerLocList and DILocation irlayers operand (#215666)
Programs lowered through intermediate IRs lose their position in those
IRs by the time they reach a backend: a DILocation records only the
original source coordinate. Add an optional `irlayers` operand carrying
one coordinate per intermediate level, so a profiler can show the IR
text a program was actually compiled from.
---------
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply at anthropic.com>
[flang][openacc] Accept enclosing labeled DO loops sharing a label with an ACC loop (#229502)
When an `!$acc loop` or combined construct is associated with a labeled
DO,
AccNonBlockDoConstruct builds a DoConstruct for it during parsing. If an
enclosing labeled DO without a directive shares the same terminating
label,
that outer loop remains a LabelDoStmt until CanonicalizeDo, while its
terminator is now nested inside the DoConstruct. AnalyzeLabels runs
before
CanonicalizeDo and treats every DoConstruct as a new scope, so it
rejected
the outer loop with "Label 'N' is not in DO loop scope":
do 10 k = 1, m
!$acc loop
do 10 i = 1, n
10 a(i, k) = 0.
[8 lines not shown]
AMDGPU: Improve optimization remarks for atomic lowering (#229352)
Improve the phrasing when the atomicrmw legalization emits the hardware
instruction. Replace the unhelpful "due to an unsafe request" wording in
the optimization remark with the actual reason the native instruction was
legal to use.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[AMDGPU] Omit ELF XNACK modes for hardwired-on targets (#227740)
Since #212792, hardwired-on XNACK targets such as gfx1250 emit
`XNACK_ON`, causing consumers to infer an unsupported `:xnack+`
modifier.
Restore the previous zero ELF XNACK field for these targets in code
object V4 and later, preserving replay-safe code generation and
encodings for targets with selectable modes. Clarify the ABI
documentation and add regressions.
Assisted by: Codex (GPT-6)
[msan][test] Add tests for llvm.vector.partial.reduce.add/fadd (#229196)
Mostly, the first and second arguments are the same (e.g., `<4 x i32>
@llvm.vector.partial.reduce.add(<4 x i32>, <4 x i32>)`) which can be
handled heuristically, but sometimes the second argument is a larger
multiple of the first argument (e.g., `<4 x i32>
@llvm.vector.partial.reduce.add(<4 x i32>, <8 x i32>)`), which is
handled strictly.
N.B. although these intrinsics are LLVM cross-platform, we keep the
tests in the AArch64 directory to maintain the correspondence with the
directory (llvm/test/CodeGen/AArch64/) that the tests are forked from.
ztest: do not use a detached vdev in online_vdev()
online_vdev() drops the config lock to online a device. If the
configuration changed meanwhile, the vdev may have been detached or
replaced and freed, which the function detects before returning. Its
verbose message then still read the old top-level vdev's state, a
use-after-free that AddressSanitizer reports during zloop runs, and
both abort paths returned the possibly freed vdev to stop the walk.
Log the saved GUID and expected generation together with the current
configuration generation, without dereferencing the old vdev. Stop
the walk with the root vdev, which stays valid under the reacquired
lock.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Kamil Monicz <kamil at monicz.dev>
Closes #19224