Run the timers every 5 minutes, with overlap and failure guards

A push reaching the mirror should take minutes, not up to six hours. All four
units now fire every 5 minutes, staggered a minute apart. Three guards make
that cadence safe:

- Scheduled runs take the release lock NON-BLOCKING (try_release_lock) and
  skip the tick when a build is running. Blocking would stack one stalled
  process per tick behind a long build and stampede when it finished. Manual
  commands still wait, as an operator expects.
- check-versions takes the lock too, and now owns its git pull (--pull, passed
  by the unit) instead of an ExecStartPre: at this cadence an unlocked pull
  would swap PKGBUILDs out from under a running build.
- A failed release records .build-failed-<channel> and backs off
  exponentially (10m, 20m, 40m … capped at 6h) rather than rebuilding the same
  broken tree every 5 minutes. Any new commit clears the backoff, since a push
  is the most likely fix.

Idle ticks exit without output so the journal keeps showing the runs that
matter, and bin/repo timers reports backoff state — a paused channel is
otherwise indistinguishable from an idle one.
This commit is contained in:
Ryan Hughes
2026-08-27 01:10:33 -04:00
parent 9d86123264
commit 71e72e581b
10 changed files with 243 additions and 57 deletions
+37 -13
View File
@@ -728,14 +728,33 @@ The repository includes GitHub workflows and systemd services for automated rele
#### Systemd Services #### Systemd Services
1. **check-versions** (Every 6 hours at :30): Pulls latest from git, compares PKGBUILD versions to published versions, creates state files if builds are needed All four units run **every 5 minutes**, staggered by a minute each, so a push
2. **auto-release-edge** (Every 6 hours at +1:00): If state file exists, builds all edge packages that need updates reaches the mirror in minutes rather than hours:
3. **auto-release-rc** (Every 6 hours at +1:00): Retry backstop — rc builds are normally triggered immediately over SSH by the release orchestrator; the timer re-runs any build whose state file survived a failure. Builds from the `rc` branch worktree (`/root/omarchy-pkgs-rc`), publishing into the shared channel tree
4. **auto-release-stable** (Every 6 hours at +1:00): If state file exists, builds `release_ring=fast` packages for stable and replicates them to rc (runs in parallel with edge)
All channel-mutating runs share a host-wide release lock 1. **check-versions** (`*:0/5`): Pulls latest from git, compares PKGBUILD versions to published versions, creates state files if builds are needed
(`pkgs.omarchy.org/.release.lock`), so overlapping timers and manual runs 2. **auto-release-edge** (`*:1/5`): If a state file exists, builds all edge packages that need updates
serialize instead of interleaving. 3. **auto-release-rc** (`*:2/5`): Builds the rc channel from the `rc` branch worktree (`/root/omarchy-pkgs-rc`), publishing into the shared channel tree. The orchestrator also triggers this immediately over ssh when cutting an RC
4. **auto-release-stable** (`*:3/5`): If a state file exists, builds `release_ring=fast` packages for stable and replicates them to rc
That cadence is only safe because of three guards:
- **No overlap.** Every channel-mutating run takes a host-wide lock
(`pkgs.omarchy.org/.release.lock`). Scheduled runs take it
**non-blocking**: if a build is already going, the tick exits immediately
instead of queuing. Waiting would stack one stalled process per tick behind
a long build and stampede when it finished. Manual commands still wait, as
an operator expects. `check-versions` takes it too — its `git pull` would
otherwise swap PKGBUILDs out from under a running build.
- **Backoff on failure.** A failed release records the attempt in
`.build-failed-<channel>` and backs off exponentially — 10m, 20m, 40m, up to
a 6h ceiling — instead of rebuilding the same broken tree every 5 minutes.
**Any new commit clears the backoff immediately**, since a push is the most
likely fix. Clear it by hand with `rm /root/.state/.build-failed-<channel>`.
- **Quiet when idle.** With nothing queued a tick exits without output, so the
journal shows the runs that mattered rather than 288 no-ops a day.
`bin/repo timers` reports all of this: schedules, last results, what is
queued, what is failing and when it will retry, and whether the lock is held.
Check on all of it with `bin/repo timers` — schedule, each unit's last run and Check on all of it with `bin/repo timers` — schedule, each unit's last run and
whether it succeeded, what is queued, whether a release is running right now, whether it succeeded, what is queued, whether a release is running right now,
@@ -748,16 +767,21 @@ bin/repo timers --local # inspect this machine instead
``` ```
State files are stored in `/root/.state/`: State files are stored in `/root/.state/`:
- `.sync-needed-edge` - `.sync-needed-<channel>` — a build is queued for that channel
- `.sync-needed-rc` - `.build-failed-<channel>` — consecutive failure count, timestamp, and the
- `.sync-needed-stable` commit it failed on (drives the backoff; removing it forces a retry)
### Schedule (America/New_York) ### Schedule (America/New_York)
| Time | Action | | Minute of every hour | Action |
|------|--------| |------|--------|
| 00:30, 06:30, 12:30, 18:30 | check-versions (git pull + creates state files) | | :00, :05, :10, … | check-versions (git pull + creates state files) |
| 01:00, 07:00, 13:00, 19:00 | auto-release-edge + auto-release-stable (parallel) | | :01, :06, :11, … | auto-release-edge |
| :02, :07, :12, … | auto-release-rc |
| :03, :08, :13, … | auto-release-stable |
Each unit is a no-op unless its channel has queued work, another run holds the
lock, or the channel is in failure backoff.
### Installation ### Installation
+83 -11
View File
@@ -1,16 +1,33 @@
#!/bin/bash #!/bin/bash
# Process sync for a specific mirror if state file exists # Run the release workflow for a channel when work is queued.
# Usage: process-sync <mirror> # Usage: auto-release <edge|rc|stable>
# Example: process-sync edge #
# Safe to run on a tight schedule. Three guards make that true:
#
# 1. Nothing queued -> exit immediately (the common case).
# 2. A release already running -> skip this tick. The lock is taken
# non-blocking on purpose: waiting would pile up one stalled process per
# tick behind a long build, and they would all stampede when it finished.
# 3. The last attempt failed -> back off exponentially rather than rebuild
# the same broken tree every few minutes. A change to the repository
# (new commit) clears the backoff immediately, because that is the thing
# most likely to have fixed it.
set -e set -e
BUILD_ROOT=$(realpath "${BASH_SOURCE[0]%/*}/..") BUILD_ROOT=$(realpath "${BASH_SOURCE[0]%/*}/..")
source "$BUILD_ROOT/helpers/message-helpers.sh" source "$BUILD_ROOT/helpers/message-helpers.sh"
source "$BUILD_ROOT/helpers/paths.sh"
source "$BUILD_ROOT/helpers/lock-helpers.sh"
MIRROR="${1:-}" MIRROR="${1:-}"
STATE_DIR="${OMARCHY_STATE_DIR:-/root/.state}" STATE_DIR="${OMARCHY_STATE_DIR:-/root/.state}"
# Backoff schedule: 10m, 20m, 40m, 80m, 160m, 320m, then hourly-ish forever
# (capped at 6h, the cadence this system ran at before frequent timers).
BACKOFF_BASE_SECONDS="${OMARCHY_BACKOFF_BASE:-600}"
BACKOFF_MAX_SECONDS="${OMARCHY_BACKOFF_MAX:-21600}"
if [[ -z "$MIRROR" ]]; then if [[ -z "$MIRROR" ]]; then
print_error "Usage: $0 <mirror>" print_error "Usage: $0 <mirror>"
echo " mirror: edge, rc, or stable" echo " mirror: edge, rc, or stable"
@@ -23,27 +40,82 @@ if [[ "$MIRROR" != "edge" && "$MIRROR" != "rc" && "$MIRROR" != "stable" ]]; then
fi fi
STATE_FILE="$STATE_DIR/.sync-needed-$MIRROR" STATE_FILE="$STATE_DIR/.sync-needed-$MIRROR"
FAIL_FILE="$STATE_DIR/.build-failed-$MIRROR"
# Nothing queued: stay quiet. At a 5-minute cadence this is most invocations,
# and a header for each would bury the runs that matter in the journal.
if [[ ! -f "$STATE_FILE" ]]; then
exit 0
fi
print_header "Processing Sync for $MIRROR" print_header "Processing Sync for $MIRROR"
# Check if state file exists # The build inputs are this checkout's contents; its HEAD identifies them.
if [[ ! -f "$STATE_FILE" ]]; then current_fingerprint() {
print_info "No sync needed for $MIRROR (state file not found)" git -C "$BUILD_ROOT" rev-parse HEAD 2>/dev/null || echo "unknown"
}
backoff_seconds() {
local count="$1" delay="$BACKOFF_BASE_SECONDS"
while ((count > 1)); do
delay=$((delay * 2))
((delay >= BACKOFF_MAX_SECONDS)) && { delay=$BACKOFF_MAX_SECONDS; break; }
count=$((count - 1))
done
echo "$delay"
}
FAIL_COUNT=0
if [[ -f "$FAIL_FILE" ]]; then
# shellcheck disable=SC1090
source "$FAIL_FILE" 2>/dev/null || true
FAIL_COUNT="${FAILURE_COUNT:-0}"
failed_at="${FAILURE_AT:-0}"
failed_fingerprint="${FAILURE_FINGERPRINT:-}"
if [[ "$failed_fingerprint" != "$(current_fingerprint)" ]]; then
print_info "Repository changed since the last failure — clearing backoff and retrying"
rm -f "$FAIL_FILE"
FAIL_COUNT=0
else
delay=$(backoff_seconds "$FAIL_COUNT")
now=$(date +%s)
retry_at=$((failed_at + delay))
if ((now < retry_at)); then
print_warning "$MIRROR has failed $FAIL_COUNT time(s) on this tree — not retrying until $(date -d "@$retry_at" '+%H:%M:%S' 2>/dev/null || echo "+$((retry_at - now))s")"
echo " Push a fix (any new commit clears this), or: rm $FAIL_FILE"
exit 0
fi
print_info "Backoff elapsed — retrying $MIRROR (failure #$((FAIL_COUNT + 1)) if this fails)"
fi
fi
# Non-blocking: a build in progress means this tick has nothing to do.
if ! try_release_lock; then
holder=$(release_lock_holder)
print_info "A release is already running (${holder:-holder unknown}) — skipping this tick"
exit 0 exit 0
fi fi
print_info "State file found: $STATE_FILE" print_info "State file found: $STATE_FILE"
print_info "Starting release workflow for $MIRROR..." print_info "Starting release workflow for $MIRROR..."
# Run the release workflow
if "$BUILD_ROOT/bin/repo" release --mirror "$MIRROR" --skip-prod-check; then if "$BUILD_ROOT/bin/repo" release --mirror "$MIRROR" --skip-prod-check; then
print_success "Release completed successfully for $MIRROR" print_success "Release completed successfully for $MIRROR"
# Remove state file on success
rm -f "$STATE_FILE" rm -f "$STATE_FILE"
rm -f "$FAIL_FILE"
print_success "State file removed: $STATE_FILE" print_success "State file removed: $STATE_FILE"
else else
print_error "Release failed for $MIRROR" status=$?
FAIL_COUNT=$((FAIL_COUNT + 1))
cat >"$FAIL_FILE" <<EOF
FAILURE_COUNT=$FAIL_COUNT
FAILURE_AT=$(date +%s)
FAILURE_FINGERPRINT=$(current_fingerprint)
EOF
next=$(backoff_seconds "$FAIL_COUNT")
print_error "Release failed for $MIRROR (attempt $FAIL_COUNT)"
print_warning "State file retained for retry: $STATE_FILE" print_warning "State file retained for retry: $STATE_FILE"
exit 1 print_warning "Backing off $((next / 60))m before the next attempt; a new commit retries sooner"
exit "$status"
fi fi
+19
View File
@@ -8,14 +8,33 @@ BUILD_ROOT=$(realpath "${BASH_SOURCE[0]%/*}/..")
source "$BUILD_ROOT/helpers/message-helpers.sh" source "$BUILD_ROOT/helpers/message-helpers.sh"
source "$BUILD_ROOT/helpers/paths.sh" source "$BUILD_ROOT/helpers/paths.sh"
source "$BUILD_ROOT/helpers/package-metadata.sh" source "$BUILD_ROOT/helpers/package-metadata.sh"
source "$BUILD_ROOT/helpers/lock-helpers.sh"
STATE_DIR="${OMARCHY_STATE_DIR:-/root/.state}" STATE_DIR="${OMARCHY_STATE_DIR:-/root/.state}"
ARCH="${ARCH:-x86_64}" ARCH="${ARCH:-x86_64}"
PULL=false
[[ "${1:-}" == "--pull" ]] && PULL=true
mkdir -p "$STATE_DIR" mkdir -p "$STATE_DIR"
# Runs on a tight schedule, so it must never disturb a build in progress:
# comparing versions reads the PKGBUILDs a running build is reading, and
# --pull would swap them underneath it. Take the lock non-blocking and skip
# the tick when a release owns it — the queue is unchanged, so the next tick
# picks up exactly where this one left off.
if ! try_release_lock; then
exit 0
fi
print_header "Package Version Check" print_header "Package Version Check"
# The scheduled run pulls; a manual run leaves the operator's tree alone.
if [[ "$PULL" == true ]]; then
print_info "Updating repository..."
git -C "$BUILD_ROOT" pull --ff-only || print_warning "git pull failed — checking the current tree"
fi
get_repo_version() { get_repo_version() {
local pkg="$1" local pkg="$1"
local mirror="$2" local mirror="$2"
+35
View File
@@ -97,6 +97,41 @@ done
[[ "$queued" == false ]] && echo " (nothing queued — all channels up to date)" [[ "$queued" == false ]] && echo " (nothing queued — all channels up to date)"
echo "" echo ""
# A channel in backoff looks identical to an idle one from the outside, so
# say so plainly: it is queued but deliberately not being retried yet.
paused=false
for channel in edge rc stable; do
fail_file="$STATE_DIR/.build-failed-$channel"
[[ -f "$fail_file" ]] || continue
if [[ "$paused" == false ]]; then
print_error "Failing builds (backoff active)"
paused=true
fi
FAILURE_COUNT=0 FAILURE_AT=0 FAILURE_FINGERPRINT=""
# shellcheck disable=SC1090
source "$fail_file" 2>/dev/null || true
delay=600
for ((i = 1; i < FAILURE_COUNT; i++)); do
delay=$((delay * 2))
((delay >= 21600)) && { delay=21600; break; }
done
retry_at=$((FAILURE_AT + delay))
now=$(date +%s)
if ((now < retry_at)); then
when="retries at $(date -d "@$retry_at" '+%H:%M:%S' 2>/dev/null || echo "+$((retry_at - now))s")"
else
when="retries on the next tick"
fi
printf ' ✗ %-7s %s consecutive failure(s), %s\n' "$channel" "$FAILURE_COUNT" "$when"
printf ' last attempt %s on commit %s\n' \
"$(date -d "@$FAILURE_AT" '+%Y-%m-%d %H:%M:%S' 2>/dev/null || echo "$FAILURE_AT")" \
"${FAILURE_FINGERPRINT:0:12}"
done
if [[ "$paused" == true ]]; then
echo " Any new commit clears the backoff; or: rm $STATE_DIR/.build-failed-<channel>"
echo ""
fi
# --- release lock ------------------------------------------------------------ # --- release lock ------------------------------------------------------------
print_info "Release lock" print_info "Release lock"
+56 -17
View File
@@ -6,39 +6,78 @@
# #
# The lock lives beside the published tree (REPO_ROOT), not the checkout, so # The lock lives beside the published tree (REPO_ROOT), not the checkout, so
# the primary checkout and the rc branch worktree contend on the same file. # the primary checkout and the rc branch worktree contend on the same file.
# Reentrant across child scripts: acquire_release_lock exports # Reentrant across child scripts: acquiring exports OMARCHY_RELEASE_LOCK_HELD,
# OMARCHY_RELEASE_LOCK_HELD, and children skip acquisition when they see it # and children skip acquisition when they see it (the flock fd is inherited,
# (the flock fd is inherited, so the lock stays held for the whole tree). # so the lock stays held for the whole tree).
#
# Two acquisition modes:
# acquire_release_lock waits — for humans, who want the command to run
# try_release_lock fails immediately — for timers, which must never
# queue up behind a long build and stampede when it
# finishes
RELEASE_LOCK_FD=9 RELEASE_LOCK_FD=9
release_lock_file() {
echo "${REPO_ROOT:-$BUILD_ROOT/pkgs.omarchy.org}/.release.lock"
}
release_lock_holder() {
tail -1 "$(release_lock_file)" 2>/dev/null
}
# True when the recorded holder is a live process. A holder line left behind by
# a killed run describes nothing that is still running.
release_lock_is_held() {
local holder pid
holder=$(release_lock_holder) || return 1
[[ -n "$holder" ]] || return 1
pid=$(sed -n 's/^pid \([0-9]\+\).*/\1/p' <<<"$holder")
[[ -n "$pid" ]] && kill -0 "$pid" 2>/dev/null
}
_release_lock_open() {
local lock_file
lock_file=$(release_lock_file)
mkdir -p "$(dirname "$lock_file")"
eval "exec $RELEASE_LOCK_FD>>\"\$lock_file\""
}
_release_lock_record() {
local lock_file
lock_file=$(release_lock_file)
# Truncate first so a crashed holder's stale line does not linger.
: >"$lock_file"
echo "pid $$ ($0) since $(date '+%Y-%m-%d %H:%M:%S')" >>"$lock_file"
export OMARCHY_RELEASE_LOCK_HELD=1
}
acquire_release_lock() { acquire_release_lock() {
local timeout="${1:-3600}" local timeout="${1:-3600}"
if [[ -n "${OMARCHY_RELEASE_LOCK_HELD:-}" ]]; then [[ -n "${OMARCHY_RELEASE_LOCK_HELD:-}" ]] && return 0
return 0
fi
local lock_file="${REPO_ROOT:-$BUILD_ROOT/pkgs.omarchy.org}/.release.lock" _release_lock_open
mkdir -p "$(dirname "$lock_file")"
eval "exec $RELEASE_LOCK_FD>>\"\$lock_file\""
if ! flock -n "$RELEASE_LOCK_FD"; then if ! flock -n "$RELEASE_LOCK_FD"; then
local holder local holder
holder=$(cat "$lock_file" 2>/dev/null | tail -1) holder=$(release_lock_holder)
echo "Waiting for release lock (up to ${timeout}s)${holder:+ — held by: $holder}" >&2 echo "Waiting for release lock (up to ${timeout}s)${holder:+ — held by: $holder}" >&2
if ! flock -w "$timeout" "$RELEASE_LOCK_FD"; then if ! flock -w "$timeout" "$RELEASE_LOCK_FD"; then
echo "Could not acquire release lock within ${timeout}s: $lock_file" >&2 echo "Could not acquire release lock within ${timeout}s: $(release_lock_file)" >&2
echo "If no release is actually running, remove the file and retry." >&2 echo "If no release is actually running, remove the file and retry." >&2
return 1 return 1
fi fi
fi fi
# Record the holder for the "waiting for" message above. Truncate first so a _release_lock_record
# crashed holder's stale line does not linger once we own the lock. }
: >"$lock_file"
echo "pid $$ ($0) since $(date '+%Y-%m-%d %H:%M:%S')" >>"$lock_file"
export OMARCHY_RELEASE_LOCK_HELD=1 # Non-blocking. Returns 1 immediately when another run holds the lock, so
# scheduled work can skip this tick instead of piling up.
try_release_lock() {
[[ -n "${OMARCHY_RELEASE_LOCK_HELD:-}" ]] && return 0
_release_lock_open
flock -n "$RELEASE_LOCK_FD" || return 1
_release_lock_record
} }
+3 -3
View File
@@ -1,10 +1,10 @@
[Unit] [Unit]
Description=Auto-release edge mirror every 6 hours Description=Release queued edge builds every 5 minutes
[Timer] [Timer]
# Run every 6 hours at :00 (after version check at :30) OnCalendar=*:1/5
OnCalendar=*-*-* 01,07,13,19:00:00 America/New_York
Persistent=true Persistent=true
RandomizedDelaySec=20
[Install] [Install]
WantedBy=timers.target WantedBy=timers.target
+3 -5
View File
@@ -1,12 +1,10 @@
[Unit] [Unit]
Description=Retry pending rc releases every 6 hours Description=Release queued rc builds every 5 minutes
[Timer] [Timer]
# Backstop only: rc builds are normally triggered immediately over SSH by the OnCalendar=*:2/5
# release orchestrator (touch .sync-needed-rc + systemctl start). This timer
# retries builds whose state file survived a failure.
OnCalendar=*-*-* 01,07,13,19:00:00 America/New_York
Persistent=true Persistent=true
RandomizedDelaySec=20
[Install] [Install]
WantedBy=timers.target WantedBy=timers.target
+3 -3
View File
@@ -1,10 +1,10 @@
[Unit] [Unit]
Description=Auto-release stable mirror every 6 hours Description=Release queued stable builds every 5 minutes
[Timer] [Timer]
# Run every 6 hours at :00 (after version check at :30) OnCalendar=*:3/5
OnCalendar=*-*-* 01,07,13,19:00:00 America/New_York
Persistent=true Persistent=true
RandomizedDelaySec=20
[Install] [Install]
WantedBy=timers.target WantedBy=timers.target
+1 -2
View File
@@ -5,8 +5,7 @@ Wants=network-online.target
[Service] [Service]
Type=oneshot Type=oneshot
ExecStartPre=/usr/bin/git pull --ff-only ExecStart=/root/omarchy-pkgs/bin/check-versions --pull
ExecStart=/root/omarchy-pkgs/bin/check-versions
Environment=OMARCHY_STATE_DIR=/root/.state Environment=OMARCHY_STATE_DIR=/root/.state
WorkingDirectory=/root/omarchy-pkgs WorkingDirectory=/root/omarchy-pkgs
+3 -3
View File
@@ -1,10 +1,10 @@
[Unit] [Unit]
Description=Check package versions every 6 hours Description=Check package versions every 5 minutes
[Timer] [Timer]
# Run every 6 hours at :30 (after GitHub sync at :00) OnCalendar=*:0/5
OnCalendar=*-*-* 00,06,12,18:30:00 America/New_York
Persistent=true Persistent=true
RandomizedDelaySec=20
[Install] [Install]
WantedBy=timers.target WantedBy=timers.target