tests: keep base virtio_gpu(4) attached on the arm64 lanes
Sets hw.virtio_gpu_drm.takeover=0 at the loader. On qemu virt, vtgpu0 is
the live console this harness reads and types at, and VirtIOGraphics'
takeover detaches it; the next console write then faults with
esr 0x96000047. That is nextbsd-kernel#170, filed 2026-09-01, and it
predates this harness -- it was hidden only because the old test had a
passwordless root shell and powered the VM off about a second after the
kext load was requested, before the ~7s graphics chain finished.
At the loader rather than in the image, because the default must stay on:
under UTM and Virtualization.framework base vtgpu displays nothing, so
video is blind from the bootloader until the kext loads. Shipping the
takeover off would leave those machines blind permanently. CI is the
environment that cannot tolerate it, so CI is what opts out, and no shipped
image changes.
The cost is stated in the comment rather than left implicit: the arm64 lanes
no longer exercise the DRM handoff at all. #170 option 3 is what would let
[3 lines not shown]
[RISCV] Remove getInstSizeVerifyMode override (#226241)
Now that we generate the correct instruction sizes for the compressed
instructions we can remove this function override.
[AMDGPU] Apply occupancy-aware register allocation anti-hints (#218074)
This PR overrides applyRegAllocationAntiHints in the SIRegisterInfo so
anti-hints
try to protect against occupancy regression. Changing of the allocation
order is
confined to the vgpr budget of the current occupancy. Reordering is
skipped when
the allocation is near the budget (80% of the target-occupancy vgpr
limit or 95%
of the current occupancy vgpr limit). Below this, the anti-hinted
registers are
moved behind the non-anti-hinted ones within the budget.
## Stack
PR **3/4**. Depends on #218073. Next: Add anti-hints in
GCNPreRAOptimizatioins #218075.
[GlobalISel] Fix segment size when narrowing G_INSERT (#226275)
When an inserted value starts inside a destination piece, bound the
overlap by the size of the inserted value. Including the offset from the
start of the piece can produce an extract wider than its source.
---------
Co-authored-by: Hongyu Chen <hongchen at nvidia.com>
[SandboxVec][LoadStoreVec][NFC] Tighten bundle element types
Use Instruction*/Constant* for LoadStoreVec APIs where that is what
callers hold, and template getCombinedVectorTypeFor so BndlRef is not
forced through a non-covariant Value* conversion.
Co-authored-by: Cursor <cursoragent at cursor.com>
tests: report a panic after login as a panic, and correct the last message
Correcting myself. The previous commit said the stale-prompt race was why
ROOT-NOATIME failed on arm64. It was not. Both arm64 lanes were dying on
the VirtIOGraphics takeover panic
(nextbsd-kernel-extensions#82): img-test (arm64) took the same
esr 0x96000047 data abort three seconds after the mount command went out,
so mount printed nothing because the kernel was gone.
The stale-prompt race is real -- I reproduced it locally, where a leftover
prompt satisfies `-re {[#%$] $}` before a slow command emits anything --
and the sentinel is what turned a misleading "/ is mounted without
noatime" into an accurate "mount printed no line", which is how the panic
was found. But it did not cause that failure and the commit message should
not have said it did.
The gap that hid this: only the stage-1 login block watched for a panic, so
a panic after login surfaced as absent output rather than as a crash. Both
post-login blocks now match it and say so.
[2 lines not shown]
[Support] Remove cl::Grouping (#224958)
This feature emulates POSIX's grouped short options in a non-perfect
way. https://reviews.llvm.org/D61270 made every single-character
cl::option implicitly group, leading to weird error message for `opt
-foo=bar`: `opt: for the -o option: may not occur within a group!`
Every tool (primarily binutils-style tools) has since migrated to
OptTable,
with llvm-cov gcov the last (#224955).
Delete this feature, which would block TableGen based representation.
LLM-aided
tests: read the mount output to a sentinel, not to a stale prompt
ROOT-NOATIME failed on arm64 while amd64 passed, on an image that was in
fact mounted correctly -- the serial log shows launchd's own
"root-rw: / remounted read-write,noatime" right there.
The cause is a race, not the image. The preceding block matches only the
text NB-SHELL-READY, which leaves that command's trailing shell prompt
sitting in expect's buffer. The mount block then offered `-re {[#%$] $}`
as an alternative, so the leftover prompt satisfied it immediately and
the block returned before `mount` had printed anything. The "on / (...)"
line therefore never reached the transcript that the verdict greps. On
amd64 the output usually won the race; on the slower arm64 guest the
stale prompt did.
Both mount blocks now bracket the output with a sentinel and read until
it arrives, and both fail loudly instead of warning -- a WARN here was
worse than useless, because the verdict then hard-failed anyway with a
misleading message about noatime.
[6 lines not shown]
tests: fix the third boot test the same way
iso-boot-test.sh had both bugs the other two had: it sent "root" with an
empty password, which root's "*" field rejects since nextbsd-overlays
f9dcd5b (#278), and it waited for "login:" in one block before deciding in
another -- so on an image where automatic login works there was no prompt,
the first block spent its eight minutes, and the second never ran.
Three copies of the same login sequence, fixed three times across three
commits, because each time I fixed the file that produced the error in front
of me instead of looking for the others. There are exactly three and this is
the last: boot-test.sh, img-boot-test.sh, iso-boot-test.sh, confirmed by
grepping every test for a login assumption rather than waiting for CI to
name the next one.
halt goes through sudo here too, since admin is not root.
[BOLT,test] Add missing pipe in dwarf5-debug-names-skip-forward-decl.s (#226338)
With cl::Grouping, --check-prefix=POSTCHECK parses as -c -h..., so
llvm-dwarfdump prints --help and exits 0, and the checks never run.
[libc++] Run the PR benchmarking tooling from main (#226278)
Previously, we'd use the workflow file and machines.json from the main
branch, but the rest of the tools (e.g. build-at-commit) would be taken
from the PR head. This patch switches to using the tools from main and
only using the PR head's content for the actual code and benchmarks.
This should make it easier to evolve the tools and the workflow files
without breaking PR benchmarking for people who submit PRs from slightly
outdated bases.
[RISCV] Support frame pointers in SiFive CLIC preemptible handlers (#221318)
This commit adds frame-pointer support for SiFive CLIC preemptible
handlers by using t0 to save mcause and mepc to dedicated stack slots in
both cases.
tests: wait for a shell in one block, not two
The previous commit fixed the wrong half. Both tests had a stage that
waited for "login:" and then a stage that handled either a prompt or an
automatic login. I replaced the second -- the one that printed the error I
had seen -- and left the first demanding a prompt.
On an image where automatic login works there is no prompt, so the first
block waited out its full eight minutes and the second never ran. That is
why all four lanes failed on the run against fresh packages, with
FAIL: 'login:' prompt not seen within 8 minutes
rather than the login rejection from before. The packages were the fix for
the earlier failure -- they carry autologin-user, so admin is now logged in
automatically -- and the test could not cope with its own success.
Splitting the wait from the decision was the mistake: an earlier block
could contradict a later one about what boot looks like. One block now does
both, and the panic check folds into it.
[CIR] Update the fastmath contract lowering test
cir-to-llvm now requires a module triple, and the LLVM dialect prints
fastmath flags as fastmath<contract> rather than an attribute dictionary.
Co-authored-by: Cursor <cursoragent at cursor.com>
[AMDGPU][InstCombine] Fold zero dot operands to accumulator
Fold AMDGPU dot intrinsics when either operand is zero.
`dot(a, 0) = 0` and `dot(0, b) = 0`, so replace the intrinsic with its accumulator.
This avoids unrelated clamp and add/sub reassociation cases.
[libc++] Take the benchmark testing configuration from the test suite (#226267)
Previously, we'd use the testing configuration (the .cfg.in file) from
the checked out test suite for the PR benchmarking, but we'd use the one
from `main` for the historical benchmarking. Since the config file is
logically part of the test suite (just like the rest of the Lit setup),
it makes more sense to take it from the checked-out test suite version.
[SelectionDAG][AMDGPU] Fold mul24 with an AND operand whose low bits are zero
Use SimplifyMultipleUseDemandedBits to simplify AND operands based on the
low 24 bits consumed by mul24.
Fold the multiply to zero when the simplified operand is zero.
This folds cases such as:
mul24(x & 0xff000000, y) -> 0
[AMDGPU] Fold 24 bit multiply with zero low bits
Fold `MUL_I24` and `MUL_U24` to zero when either operand has known zero
low 24 bits.
For example:
```
llvm.amdgcn.mul.i24(x, 0x01000000) -> 0
```
[SelectionDAG] Handle constants in SimplifyMultipleUseDemandedBits
Replace a non-zero constant with zero when none of its set bits are
demanded.
This allows users of `SimplifyMultipleUseDemandedBits` to eliminate
irrelevant constant bits while preserving the convention that a null
SDValue indicates no simplification.
[AMDGPU] Fold a constant add/sub into the sudot4/sudot8 accumulator (#225322)
Fold a constant add into the accumulator operand of sudot4 and sudot8
when
clamping is disabled:
```
sudot(a, b, C1, false) + C2 -> sudot(a, b, C1 + C2, false)
```
Subtraction by a constant is canonicalized to addition of its negation.
[HLSL] Add support for dynamic resources (#221103)
Adds support for dynamic resources, also known as bindless or directly
indexed resources. The design is described in
https://github.com/llvm/wg-hlsl/issues/204.
The change adds:
- The `ResourceDescriptorHeap` and `SamplerDescriptorHeap` global
symbols, which expose the Shader Model 6.6 descriptor heaps.
- Internal `hlsl::__detail::heap_resource_info` and
`hlsl::__detail::heap_sampler_info` types, which are produced by
indexing the corresponding descriptor heap. HLSL resource and sampler
types can be implicitly constructed from these values.
- Constructors that accept the corresponding internal heap info type for
all currently implemented resource classes.
- Clang builtins `__builtin_hlsl_resource_handlefromheap` and
`__builtin_hlsl_resource_counterhandlefromheap` for creating resource
and counter handles from descriptor heap indices.
- Sema validation and return-type handling for the new builtins.
[8 lines not shown]
[MIR][AMDGPU] Serialize register allocation anti-hints (#218073)
This PR adds MIR print and parse support for the register allocation
anti-hints through
a per vreg anti-hints: [ .. ] field in the registers: section. The
printer only emits
the anti-hint registers only when the it is non-empty to keep existing
tests unchanged.
## Stack
PR **2/4**. Depends on #218071. Next: Apply Occupancy-Aware allocation
anti-hints #218074.
[AMDGPU] Track allocator-reserved register pressure separately
GCNRegPressure counts a tuple at its live lane count, but the excess and
critical limits the scheduler compares against are expressed in
allocatable registers. A tuple is reserved at its full class width for
its whole live range, so lanes that are dead at a given point are still
unavailable to any other value, and the live lane count underestimates
what the allocator has to set aside.
Track a third set of per-kind values recording the reserved width, and
use it for the pressure the GCN trackers hand to the scheduling
heuristics.
Also add -amdgpu-trackers-compare-rp to print the GCN tracker and the
generic tracker pressure side by side for every candidate, which is what
the discrepancy was found with.
WIP
AI-Assisted