Files
omarchycn/test/shell.d/launch-shell-test.sh
T
1e7bb66556 Recover a session lock stranded by a dead shell (#6692)
* Detect a compositor session lock through one helper

omarchy-restart-shell decided whether the session was locked by looking for
"LOCK" anywhere in the hyprctl monitors payload. That works, but not for the
reason the code reads like: Hyprland reports no lock state of its own, and the
string comes from solitaryBlockedBy, the list of reasons a monitor cannot hand
a client the whole screen. An active ext-session-lock is one of those reasons.

A substring match over the whole payload also answers yes to a workspace or a
monitor description that merely spells LOCK, and locking a desktop nobody asked
to lock is the worst way to be wrong. Match the reason list itself, and put it
behind a helper now that a second caller needs the same answer.

That second caller needs a third answer too, because the reason list is not
always readable. Hyprland stops at the first reason on a monitor with no
workspace yet — one just coming back — and returns before it ever looks at the
lock, so a missing LOCK there means nothing was asked rather than nothing was
found. Neither that nor an unreachable compositor is an unlocked session, and
locks strand precisely while outputs are coming and going, so both exit 2.
Callers that only branch on success are unaffected.

The test fixture claimed the string came from a workspace name, so it was
encoding the wrong model of the compositor. It now returns what Hyprland
actually returns.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Retake a session lock stranded by a dead shell

ext-session-lock keeps the session locked when its client goes away — that is
the point of the protocol, so a crashing lock screen cannot expose the desktop.
The cost is that a shell which dies while locked leaves the compositor locked
with nothing left to authenticate against: Hyprland's failsafe, which takes a
TTY or another machine to clear.

Nothing carried the lock across a restart. Quickshell relaunches itself after a
crash and omarchy-restart-shell can be run by hand, but both bring back a shell
holding no lock, so the failsafe stayed up. A fresh shell never holds a lock, so
a session already locked as the lock service starts can only be that orphan:
take it back and let the user type their way out.

Asking once is not enough. These deaths happen while outputs are going away,
and the replacement shell comes up inside that same window, where there is
nothing to read a lock off. So the question is asked until the answer means
something: on a short timer while the session settles, and again when a screen
comes back, since a display asleep for hours outlasts any timer worth running
and returns through a state the compositor cannot answer for either. Once an
answer does arrive the search ends, so the timer stops and later screen changes
cost nothing.

Three ways this could lock a desktop nobody asked to lock, all closed. A lock
this shell took itself is not an orphan, including one taken while the question
was in flight — omarchy-restart-shell re-locks a fresh shell, and the answer
cannot tell whose lock it found. Recovery runs once and clears the flag, so
nothing lingers to fire after an unlock. And PAM landing late reopens the
question rather than answering it: clearing the failsafe from a TTY is the
documented way out, so a yes from before there was anything to do about it may
be stale by the time it can be acted on.

The check has to live here rather than in the launcher. Quickshell's crash
handler re-execs in place, keeping the same pid, so a supervising process never
sees the restarts that recovery matters most for.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Relaunch the shell when it dies without a signal

Quickshell restarts itself after a crash, but only from its signal handlers:
SIGSEGV, SIGABRT, SIGFPE, SIGILL, SIGBUS, SIGTRAP. Qt does not always leave
that way. When the Wayland connection fails, QWaylandDisplay::checkWaylandError
calls _exit() directly, which raises no signal at all — so the crash handler
never runs, no report lands in ~/.cache/quickshell/crashes, and the desktop is
left with no bar and no explanation.

That is how #6684 ends: the lock path meets a screen with no valid Wayland
output, declines to create a lock surface for it, and the connection dies with
EINVAL. Supervise the launcher so those deaths come back.

A clean exit is deliberate — omarchy-restart-shell stops the shell over IPC and
starts its own replacement — and a signal to the supervisor means the session is
going away, so neither relaunches. Neither does a shell that outlived its
compositor, though that takes more than one unanswered query to conclude: the
shell dies while outputs are being reconfigured, which is also when a busy
compositor can miss one without being gone. A shell that cannot stay up gives
up after five tries in a minute rather than spinning.

Signals need care now that a launcher stands between the session and the shell.
Bash defers a trap until a foreground command returns, so the shell runs as a
job and the supervisor waits on it. Stopping the launcher used to stop the shell
with it, back when this script exec'd Quickshell, so the signal is passed on
rather than leaving a desktop nobody is watching. One arriving during the
backoff sleep only reaches the trap afterwards, so the flag is read again at the
top of the loop: a shutdown racing a crash would otherwise get one more
Quickshell on its way out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 11:29:35 +02:00

190 lines
6.2 KiB
Bash
Executable File

#!/bin/bash
set -euo pipefail
source "$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)/base-test.sh"
test_tmp=$(mktemp -d)
launch_pid=""
# A supervisor that fails to stop would hang the run instead of failing it.
cleanup() {
if [[ -n $launch_pid ]]; then
pkill -TERM -P "$launch_pid" 2>/dev/null || true
kill -KILL "$launch_pid" 2>/dev/null || true
wait "$launch_pid" 2>/dev/null || true
fi
rm -rf "$test_tmp"
}
trap cleanup EXIT
fake_bin="$test_tmp/bin"
shell_root="$test_tmp/root"
mkdir -p "$fake_bin" "$shell_root/shell"
# Each launch consumes the next status from OMARCHY_TEST_QS_STATUSES; "run"
# stands in for a healthy shell that keeps going until stopped.
cat >"$fake_bin/quickshell" <<'SH'
#!/bin/bash
printf '%s\n' "$*" >>"$OMARCHY_TEST_QS_LOG"
launches=$(wc -l <"$OMARCHY_TEST_QS_LOG")
status=$(awk -v n="$launches" 'NR == n { print; found = 1 } END { if (!found) print "0" }' <<<"$OMARCHY_TEST_QS_STATUSES")
if [[ $status == "run" ]]; then
trap 'touch "$OMARCHY_TEST_QS_TERMINATED"; exit 143' TERM
while true; do sleep 0.05; done
fi
exit "${status:-0}"
SH
cat >"$fake_bin/systemd-cat" <<'SH'
#!/bin/bash
while (( $# > 0 )); do
[[ $1 == "--" ]] && { shift; break; }
shift
done
exec "$@"
SH
cat >"$fake_bin/hyprctl" <<'SH'
#!/bin/bash
[[ ${OMARCHY_TEST_COMPOSITOR_GONE:-0} == 1 ]] && exit 4
# Refuse the first OMARCHY_TEST_HYPRCTL_MISSES queries, then answer.
if (( ${OMARCHY_TEST_HYPRCTL_MISSES:-0} > 0 )); then
misses=$(cat "$OMARCHY_TEST_HYPRCTL_MISS_COUNT" 2>/dev/null || printf '0')
if (( misses < OMARCHY_TEST_HYPRCTL_MISSES )); then
printf '%s\n' "$(( misses + 1 ))" >"$OMARCHY_TEST_HYPRCTL_MISS_COUNT"
exit 4
fi
fi
printf '[]\n'
SH
cat >"$fake_bin/logger" <<'SH'
#!/bin/bash
shift 2
printf '%s\n' "$*" >>"$OMARCHY_TEST_LOGGER_LOG"
SH
chmod +x "$fake_bin/quickshell" "$fake_bin/systemd-cat" "$fake_bin/hyprctl" "$fake_bin/logger"
qs_log="$test_tmp/quickshell.log"
logger_log="$test_tmp/logger.log"
qs_terminated="$test_tmp/quickshell-terminated"
hyprctl_misses="$test_tmp/hyprctl-misses"
launch_shell() {
: >"$qs_log"
: >"$logger_log"
PATH="$fake_bin:$PATH" \
OMARCHY_PATH="$shell_root" \
OMARCHY_TEST_QS_LOG="$qs_log" \
OMARCHY_TEST_QS_STATUSES="$1" \
OMARCHY_TEST_COMPOSITOR_GONE="${2:-0}" \
OMARCHY_TEST_LOGGER_LOG="$logger_log" \
OMARCHY_TEST_QS_TERMINATED="$qs_terminated" \
OMARCHY_TEST_HYPRCTL_MISSES="${3:-0}" \
OMARCHY_TEST_HYPRCTL_MISS_COUNT="$hyprctl_misses" \
timeout 30 "$ROOT/bin/omarchy-launch-shell"
}
launches() {
wc -l <"$qs_log" | tr -d ' '
}
launch_shell '0' || fail "a clean launch succeeds"
[[ $(launches) == 1 ]] || fail "a shell that exits cleanly is not relaunched" "$(<"$qs_log")"
grep -F -- "-n -p $shell_root/shell" "$qs_log" >/dev/null || fail "the shell launches from OMARCHY_PATH"
pass "a shell that exits cleanly is left alone"
# Qt leaves through _exit(), so Quickshell's crash handler never relaunches it.
launch_shell $'255\n0' || fail "a shell that died on a Wayland error is relaunched"
[[ $(launches) == 2 ]] || fail "the dead shell is relaunched exactly once" "$(<"$qs_log")"
grep -F 'exited with status 255' "$logger_log" >/dev/null || fail "the relaunch is recorded in the journal"
pass "a shell that dies without a signal is relaunched"
launch_shell $'255\n255\n255\n255\n255\n255\n255\n255' && fail "a shell that keeps dying is given up on"
[[ $(launches) == 6 ]] || fail "relaunches stop after the attempt budget" "$(<"$qs_log")"
grep -F 'Giving up' "$logger_log" >/dev/null || fail "giving up is recorded in the journal"
pass "a shell that keeps dying is not relaunched forever"
# The compositor takes the shell with it, and the session is already going.
launch_shell $'255\n0' 1 || fail "a shell outliving the compositor exits cleanly"
[[ $(launches) == 1 ]] || fail "the shell is not relaunched into a dead session" "$(<"$qs_log")"
pass "the shell is not relaunched once the compositor is gone"
# A compositor mid-modeset can miss a query without being gone.
rm -f "$hyprctl_misses"
launch_shell $'255\n0' 0 2 || fail "a shell survives a compositor that misses a query"
[[ $(launches) == 2 ]] || fail "a missed compositor query does not end supervision" "$(<"$qs_log")"
pass "a compositor too busy to answer is not mistaken for one that is gone"
# A signal mid-backoff only reaches the trap once the sleep is over.
: >"$qs_log"
: >"$logger_log"
PATH="$fake_bin:$PATH" \
OMARCHY_PATH="$shell_root" \
OMARCHY_TEST_QS_LOG="$qs_log" \
OMARCHY_TEST_QS_STATUSES=$'255\n0' \
OMARCHY_TEST_COMPOSITOR_GONE=0 \
OMARCHY_TEST_LOGGER_LOG="$logger_log" \
OMARCHY_TEST_QS_TERMINATED="$qs_terminated" \
"$ROOT/bin/omarchy-launch-shell" &
launch_pid=$!
for (( waited = 0; waited < 100; waited++ )); do
[[ $(launches) == 1 ]] && break
sleep 0.05
done
[[ $(launches) == 1 ]] || fail "the supervised shell launched before the signal" "$(<"$qs_log")"
kill -TERM "$launch_pid"
wait "$launch_pid" || fail "a signalled supervisor exits cleanly"
launch_pid=""
[[ $(launches) == 1 ]] || fail "the shell is not relaunched after the session asked to stop" "$(<"$qs_log")"
pass "a signal during backoff stops the supervisor before it relaunches"
# Stopping the launcher used to stop the shell, back when it exec'd Quickshell.
: >"$qs_log"
: >"$logger_log"
rm -f "$qs_terminated"
PATH="$fake_bin:$PATH" \
OMARCHY_PATH="$shell_root" \
OMARCHY_TEST_QS_LOG="$qs_log" \
OMARCHY_TEST_QS_STATUSES='run' \
OMARCHY_TEST_COMPOSITOR_GONE=0 \
OMARCHY_TEST_LOGGER_LOG="$logger_log" \
OMARCHY_TEST_QS_TERMINATED="$qs_terminated" \
"$ROOT/bin/omarchy-launch-shell" &
launch_pid=$!
for (( waited = 0; waited < 100; waited++ )); do
[[ $(launches) == 1 ]] && break
sleep 0.05
done
[[ $(launches) == 1 ]] || fail "the healthy shell launched before the signal" "$(<"$qs_log")"
kill -TERM "$launch_pid"
for (( waited = 0; waited < 100; waited++ )); do
kill -0 "$launch_pid" 2>/dev/null || break
sleep 0.05
done
kill -0 "$launch_pid" 2>/dev/null && fail "a signalled supervisor stops instead of waiting on a live shell"
wait "$launch_pid" 2>/dev/null || true
launch_pid=""
[[ -f $qs_terminated ]] || fail "the running shell is signalled when the supervisor is"
[[ $(launches) == 1 ]] || fail "the signalled shell is not relaunched" "$(<"$qs_log")"
pass "stopping the supervisor stops the shell it is watching"