T4: pin boot tests to shared harness v0.2.3 (banner→prompt login) (#513)
* T4: boot gates run on the shared nextbsd-ci harness @v0.2.1 (login-only)
img-boot-test.sh + iso-boot-test.sh become thin entry points onto the shared
harness (NB_LOGIN_ONLY; the live ISO adds NB_MEDIA=cd): they extract the
zipped artifact, check out nextbsd/nextbsd-ci at v0.2.1 (pinned tag = no drift),
and run harness/boot-test.sh. The loader un-mute dance, the arch-aware qemu
argv, login detection and teardown now come from the shared harness (one place,
every arch, kept green by its selftest) instead of the in-repo copies.
Delete the 1512-line tests/boot-test.sh monolith (nothing in CI invoked it) and
the in-repo tests/loader.exp.inc + tests/qemu-arch.sh (superseded by the
harness's contract.exp.inc + qemu-arch.sh). The jobs' serial-log dumps point at
the harness transcript. Both gates stay NON-GATING as before (not in release's
needs); the gate is the harness exit class.
* Bump the shared-harness pin to v0.2.2 in the thin img/iso boot wrappers
[4 lines not shown]
T4: boot gates run on the shared nextbsd-ci harness @v0.2.1 (login-only)
img-boot-test.sh + iso-boot-test.sh become thin entry points onto the shared
harness (NB_LOGIN_ONLY; the live ISO adds NB_MEDIA=cd): they extract the
zipped artifact, check out nextbsd/nextbsd-ci at v0.2.1 (pinned tag = no drift),
and run harness/boot-test.sh. The loader un-mute dance, the arch-aware qemu
argv, login detection and teardown now come from the shared harness (one place,
every arch, kept green by its selftest) instead of the in-repo copies.
Delete the 1512-line tests/boot-test.sh monolith (nothing in CI invoked it) and
the in-repo tests/loader.exp.inc + tests/qemu-arch.sh (superseded by the
harness's contract.exp.inc + qemu-arch.sh). The jobs' serial-log dumps point at
the harness transcript. Both gates stay NON-GATING as before (not in release's
needs); the gate is the harness exit class.
Bump the shared-harness pin to v0.2.2 in the thin img/iso boot wrappers
v0.2.2 = v0.2.1 (v0.2.0 + login-only) plus a capped post-banner login-wait
window (180s) that probes for a shell quickly instead of waiting out the whole
global budget on an autologin image (the arm64 30-min 'hang').
Bump the shared-harness pin to v0.2.2 in the thin img/iso boot wrappers
v0.2.2 = v0.2.1 (v0.2.0 + login-only) plus a capped post-banner login-wait
window (180s) that probes for a shell quickly instead of waiting out the whole
global budget on an autologin image (the arm64 30-min 'hang').
Bump the shared-harness pin to v0.2.2 in the thin img/iso boot wrappers
v0.2.2 = v0.2.1 (v0.2.0 + login-only) plus a capped post-banner login-wait
window (180s) that probes for a shell quickly instead of waiting out the whole
global budget on an autologin image (the arm64 30-min 'hang').
T4: boot gates run on the shared nextbsd-ci harness @v0.2.1 (login-only)
img-boot-test.sh + iso-boot-test.sh become thin entry points onto the shared
harness (NB_LOGIN_ONLY; the live ISO adds NB_MEDIA=cd): they extract the
zipped artifact, check out nextbsd/nextbsd-ci at v0.2.1 (pinned tag = no drift),
and run harness/boot-test.sh. The loader un-mute dance, the arch-aware qemu
argv, login detection and teardown now come from the shared harness (one place,
every arch, kept green by its selftest) instead of the in-repo copies.
Delete the 1512-line tests/boot-test.sh monolith (nothing in CI invoked it) and
the in-repo tests/loader.exp.inc + tests/qemu-arch.sh (superseded by the
harness's contract.exp.inc + qemu-arch.sh). The jobs' serial-log dumps point at
the harness transcript. Both gates stay NON-GATING as before (not in release's
needs); the gate is the harness exit class.
T4: boot gates run on the shared nextbsd-ci harness @v0.2.1 (login-only)
img-boot-test.sh + iso-boot-test.sh become thin entry points onto the shared
harness (NB_LOGIN_ONLY; the live ISO adds NB_MEDIA=cd): they extract the
zipped artifact, check out nextbsd/nextbsd-ci at v0.2.1 (pinned tag = no drift),
and run harness/boot-test.sh. The loader un-mute dance, the arch-aware qemu
argv, login detection and teardown now come from the shared harness (one place,
every arch, kept green by its selftest) instead of the in-repo copies.
Delete the 1512-line tests/boot-test.sh monolith (nothing in CI invoked it) and
the in-repo tests/loader.exp.inc + tests/qemu-arch.sh (superseded by the
harness's contract.exp.inc + qemu-arch.sh). The jobs' serial-log dumps point at
the harness transcript. Both gates stay NON-GATING as before (not in release's
needs); the gate is the harness exit class.
rpi: boot normally, not verbose (#506)
* rpi: boot normally, not verbose
The Pi boot partition shipped `FreeBSD: -v` in cmdline.txt, from when the
board was being brought up and its serial log was the only instrument
there was. That is no longer the right default, and leaving it was a trap
rather than a nicety.
The Pi rootfs carries nextbsd-overlays' loader.conf.d, so it already gets
boot_mutemsgs="YES" (#363). RB_VERBOSE is exactly what that mute exempts,
so a -v that arrived would turn the quiet console back off on this board
alone, while every other machine stayed quiet.
It does not arrive today: parse_fdt_bootargs() only parses when
fdt_get_chosen_bootargs() succeeds, and a tryboot with that line produced
a boot that was not verbose, so the firmware appears not to be writing
/chosen/bootargs at all (nextbsd-kernel#93). So the flag was inert -- and
would have started working silently, on every Pi image, the day that was
[198 lines not shown]
rpi: boot normally, not verbose (#506)
* rpi: boot normally, not verbose
The Pi boot partition shipped `FreeBSD: -v` in cmdline.txt, from when the
board was being brought up and its serial log was the only instrument
there was. That is no longer the right default, and leaving it was a trap
rather than a nicety.
The Pi rootfs carries nextbsd-overlays' loader.conf.d, so it already gets
boot_mutemsgs="YES" (#363). RB_VERBOSE is exactly what that mute exempts,
so a -v that arrived would turn the quiet console back off on this board
alone, while every other machine stayed quiet.
It does not arrive today: parse_fdt_bootargs() only parses when
fdt_get_chosen_bootargs() succeeds, and a tryboot with that line produced
a boot that was not verbose, so the firmware appears not to be writing
/chosen/bootargs at all (nextbsd-kernel#93). So the flag was inert -- and
would have started working silently, on every Pi image, the day that was
[198 lines not shown]
tests: accept a bare shell prompt as evidence of login
The arm64 live ISO now boots, pivots and autologins correctly, and the test
failed it anyway. The serial ends at a working zsh prompt:
admin at virt-8-2 ~ %
FAIL: neither a login prompt nor an automatic login in 8 minutes
Neither accepted form reaches the serial there: getty autologins, so no
"login:" prompt is printed, and the "login on console as admin" marker does
not appear on the ISO either -- the next thing on the wire after getty's
banner is the shell. img-test passed only because that marker does appear on
the installed image.
Both stage-1 expect blocks and both shell-level verdicts now accept the
prompt as a third form.
The prompt carries SGR escapes inside it -- "\033[32madmin\033[39m@..." --
so "admin@" never appears literally, hence admin[^@]{0,12}@. I first wrote
[6 lines not shown]
tests: drop the takeover tunable; base virtio_gpu is gone from the kernel
hw.virtio_gpu_drm.takeover came from nextbsd-kernel-extensions#83, which is
closed -- gating the eviction had been tried before and broke the UTM kext.
So these lines set a kernel environment variable that nothing will ever read,
above a comment explaining a panic, which is worse than no line at all: it
reads as though the arm64 lanes are protected when they are not.
nextbsd-kernel#249 removes base virtio_gpu(4) from the kernel instead, so
there is no console for the DRM kext to evict and nothing for a tunable to
gate. These lanes then exercise the DRM bind again rather than opting out of
it, which is the coverage the tunable would have cost.
Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
tests: keep base virtio_gpu(4) attached on the arm64 lanes
Sets hw.virtio_gpu_drm.takeover=0 at the loader. On qemu virt, vtgpu0 is
the live console this harness reads and types at, and VirtIOGraphics'
takeover detaches it; the next console write then faults with
esr 0x96000047. That is nextbsd-kernel#170, filed 2026-09-01, and it
predates this harness -- it was hidden only because the old test had a
passwordless root shell and powered the VM off about a second after the
kext load was requested, before the ~7s graphics chain finished.
At the loader rather than in the image, because the default must stay on:
under UTM and Virtualization.framework base vtgpu displays nothing, so
video is blind from the bootloader until the kext loads. Shipping the
takeover off would leave those machines blind permanently. CI is the
environment that cannot tolerate it, so CI is what opts out, and no shipped
image changes.
The cost is stated in the comment rather than left implicit: the arm64 lanes
no longer exercise the DRM handoff at all. #170 option 3 is what would let
[3 lines not shown]
tests: report a panic after login as a panic, and correct the last message
Correcting myself. The previous commit said the stale-prompt race was why
ROOT-NOATIME failed on arm64. It was not. Both arm64 lanes were dying on
the VirtIOGraphics takeover panic
(nextbsd-kernel-extensions#82): img-test (arm64) took the same
esr 0x96000047 data abort three seconds after the mount command went out,
so mount printed nothing because the kernel was gone.
The stale-prompt race is real -- I reproduced it locally, where a leftover
prompt satisfies `-re {[#%$] $}` before a slow command emits anything --
and the sentinel is what turned a misleading "/ is mounted without
noatime" into an accurate "mount printed no line", which is how the panic
was found. But it did not cause that failure and the commit message should
not have said it did.
The gap that hid this: only the stage-1 login block watched for a panic, so
a panic after login surfaced as absent output rather than as a crash. Both
post-login blocks now match it and say so.
[2 lines not shown]
tests: read the mount output to a sentinel, not to a stale prompt
ROOT-NOATIME failed on arm64 while amd64 passed, on an image that was in
fact mounted correctly -- the serial log shows launchd's own
"root-rw: / remounted read-write,noatime" right there.
The cause is a race, not the image. The preceding block matches only the
text NB-SHELL-READY, which leaves that command's trailing shell prompt
sitting in expect's buffer. The mount block then offered `-re {[#%$] $}`
as an alternative, so the leftover prompt satisfied it immediately and
the block returned before `mount` had printed anything. The "on / (...)"
line therefore never reached the transcript that the verdict greps. On
amd64 the output usually won the race; on the slower arm64 guest the
stale prompt did.
Both mount blocks now bracket the output with a sentinel and read until
it arrives, and both fail loudly instead of warning -- a WARN here was
worse than useless, because the verdict then hard-failed anyway with a
misleading message about noatime.
[6 lines not shown]
tests: fix the third boot test the same way
iso-boot-test.sh had both bugs the other two had: it sent "root" with an
empty password, which root's "*" field rejects since nextbsd-overlays
f9dcd5b (#278), and it waited for "login:" in one block before deciding in
another -- so on an image where automatic login works there was no prompt,
the first block spent its eight minutes, and the second never ran.
Three copies of the same login sequence, fixed three times across three
commits, because each time I fixed the file that produced the error in front
of me instead of looking for the others. There are exactly three and this is
the last: boot-test.sh, img-boot-test.sh, iso-boot-test.sh, confirmed by
grepping every test for a login assumption rather than waiting for CI to
name the next one.
halt goes through sudo here too, since admin is not root.
tests: wait for a shell in one block, not two
The previous commit fixed the wrong half. Both tests had a stage that
waited for "login:" and then a stage that handled either a prompt or an
automatic login. I replaced the second -- the one that printed the error I
had seen -- and left the first demanding a prompt.
On an image where automatic login works there is no prompt, so the first
block waited out its full eight minutes and the second never ran. That is
why all four lanes failed on the run against fresh packages, with
FAIL: 'login:' prompt not seen within 8 minutes
rather than the login rejection from before. The packages were the fix for
the earlier failure -- they carry autologin-user, so admin is now logged in
automatically -- and the test could not cope with its own success.
Splitting the wait from the decision was the mistake: an earlier block
could contradict a later one about what boot looks like. One block now does
both, and the panic check folds into it.
tests: log in as admin, because root is disabled now
Both boot tests sent "root" with an empty password. nextbsd-overlays
f9dcd5b (#278) disabled root the way Darwin does -- its password field went
from empty to "*" -- so login rejects every password including the empty
one, and both stages failed with "Login incorrect".
Nothing was wrong with the image. The tests described an image that no
longer exists, and this branch is simply the first build since: overlays
landed that change on 2026-09-23 20:23 and this repo last built green on
2026-09-22 20:45.
That also explains the second failure. ROOT-NOATIME greps the transcript
for "on / (ufs, local, noatime", which `mount` prints -- and mount never
ran, because there was no shell. One cause, two red lines.
admin is the way in. nss_directory_services gives an account carrying
noPassword an EMPTY passwd field for a privileged caller
(dsdb_pack_passwd), and login is privileged, so an empty password is
[17 lines not shown]
rpi: boot normally, not verbose
The Pi boot partition shipped `FreeBSD: -v` in cmdline.txt, from when the
board was being brought up and its serial log was the only instrument
there was. That is no longer the right default, and leaving it was a trap
rather than a nicety.
The Pi rootfs carries nextbsd-overlays' loader.conf.d, so it already gets
boot_mutemsgs="YES" (#363). RB_VERBOSE is exactly what that mute exempts,
so a -v that arrived would turn the quiet console back off on this board
alone, while every other machine stayed quiet.
It does not arrive today: parse_fdt_bootargs() only parses when
fdt_get_chosen_bootargs() succeeds, and a tryboot with that line produced
a boot that was not verbose, so the firmware appears not to be writing
/chosen/bootargs at all (nextbsd-kernel#93). So the flag was inert -- and
would have started working silently, on every Pi image, the day that was
fixed. A latent regression rather than a current one.
[6 lines not shown]
build: seed /etc/sudoers as mode 0440 (#499)
sudo refuses a sudoers file that is not 0440 ("is mode 0644, should be
0440"), and git stores nextbsd-overlays' copy as 0644, so the image's
sudo from NextBSD-contrib could not load its policy. Same fix as
overlays' seed.sh and nextbsd-userland's assemble-image.sh.
Co-authored-by: Claude Fable 5.1 <noreply at anthropic.com>
build: seed /etc/sudoers as mode 0440
sudo refuses a sudoers file that is not 0440 ("is mode 0644, should be
0440"), and git stores nextbsd-overlays' copy as 0644, so the image's
sudo from NextBSD-contrib could not load its policy. Same fix as
overlays' seed.sh and nextbsd-userland's assemble-image.sh.
Co-Authored-By: Claude Fable 5.1 <noreply at anthropic.com>
E15 Stage A: noatime and #467 boot gates, build.sh fstab comments (#494)
* img-test: gate on noatime for / and match the whole mount line
Adds ROOT-NOATIME to the verdict: the serial transcript must show
' on / (ufs, local, noatime'. noatime moves from the fstab root line into
launchd's own remount (nextbsd-userland#185, E15 A2), and the overlay
stops shipping fstab (nextbsd-overlays#5, A1). This gate catches a
build where A1 ships before A2.
The ROOT-IS-UFS step matched "ufs" in the echoed device name
(/dev/ufs/...), so it said nothing about mount flags, and it sent
halt -p before the line finished. It now waits for the whole
'on / (ufs...)' line. The verdict checks login first, so a boot that
never reaches login still reports as a boot failure.
Refs nextbsd/nextbsd-userland#185
Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
[30 lines not shown]
build.sh: stop describing the overlay fstab as the source of the root entry
Root comes from the kernel's baked-in ROOTDEVNAME (#188), not from an
fstab line or loader.conf.d, and nextbsd-overlays#5 (E15 A1) stops
shipping /etc/fstab. Rewrite the three comments that said otherwise and
drop the doubled "the the".
Refs #472
Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
iso/img-test: fail on the #467 strings (FSTAB-ROOT-REMOUNT)
Once nextbsd-overlays#5 stops shipping /etc/fstab, launchctl's boot-time
mount -vat nonfs is skipped, so neither 'Cannot union mount root
filesystem' nor launchctl's 'fwexec(mount_tool' assert should appear.
Gate both harnesses on it so the root remount can't come back unnoticed.
The ISO verdict now checks pivot+login first, then the gate, so a boot
that never completes still reports as a boot failure.
Red until nextbsd-overlays#5 is in the image: merge together with it.
Refs #472, #467
Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
img-test: gate on noatime for / and match the whole mount line
Adds ROOT-NOATIME to the verdict: the serial transcript must show
' on / (ufs, local, noatime'. noatime moves from the fstab root line into
launchd's own remount (nextbsd-userland#185, E15 A2), and the overlay
stops shipping fstab (nextbsd-overlays#5, A1). This gate catches a
build where A1 ships before A2.
The ROOT-IS-UFS step matched "ufs" in the echoed device name
(/dev/ufs/...), so it said nothing about mount flags, and it sent
halt -p before the line finished. It now waits for the whole
'on / (ufs...)' line. The verdict checks login first, so a boot that
never reaches login still reports as a boot failure.
Refs nextbsd/nextbsd-userland#185
Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
iso/img-test: fail on the #467 strings (FSTAB-ROOT-REMOUNT)
Once nextbsd-overlays#5 stops shipping /etc/fstab, launchctl's boot-time
mount -vat nonfs is skipped, so neither 'Cannot union mount root
filesystem' nor launchctl's 'fwexec(mount_tool' assert should appear.
Gate both harnesses on it so the root remount can't come back unnoticed.
The ISO verdict now checks pivot+login first, then the gate, so a boot
that never completes still reports as a boot failure.
Red until nextbsd-overlays#5 is in the image: merge together with it.
Refs #472, #467
Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
build.sh: stop describing the overlay fstab as the source of the root entry
Root comes from the kernel's baked-in ROOTDEVNAME (#188), not from an
fstab line or loader.conf.d, and nextbsd-overlays#5 (E15 A1) stops
shipping /etc/fstab. Rewrite the three comments that said otherwise and
drop the doubled "the the".
Refs #472
Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>
img-test: gate on noatime for / and match the whole mount line
Adds ROOT-NOATIME to the verdict: the serial transcript must show
' on / (ufs, local, noatime'. noatime moves from the fstab root line into
launchd's own remount (nextbsd-userland#185, E15 A2), and the overlay
stops shipping fstab (nextbsd-overlays#5, A1). This gate catches a
build where A1 ships before A2.
The ROOT-IS-UFS step matched "ufs" in the echoed device name
(/dev/ufs/...), so it said nothing about mount flags, and it sent
halt -p before the line finished. It now waits for the whole
'on / (ufs...)' line. The verdict checks login first, so a boot that
never reaches login still reports as a boot failure.
Refs nextbsd/nextbsd-userland#185
Co-Authored-By: Claude Opus 5 (1M context) <noreply at anthropic.com>