From 71e72e581b826e99eedd8540b59c8bb12c718425 Mon Sep 17 00:00:00 2001 From: Ryan Hughes Date: Wed, 26 Aug 2026 23:35:49 -0400 Subject: [PATCH] Run the timers every 5 minutes, with overlap and failure guards MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A push reaching the mirror should take minutes, not up to six hours. All four units now fire every 5 minutes, staggered a minute apart. Three guards make that cadence safe: - Scheduled runs take the release lock NON-BLOCKING (try_release_lock) and skip the tick when a build is running. Blocking would stack one stalled process per tick behind a long build and stampede when it finished. Manual commands still wait, as an operator expects. - check-versions takes the lock too, and now owns its git pull (--pull, passed by the unit) instead of an ExecStartPre: at this cadence an unlocked pull would swap PKGBUILDs out from under a running build. - A failed release records .build-failed- and backs off exponentially (10m, 20m, 40m … capped at 6h) rather than rebuilding the same broken tree every 5 minutes. Any new commit clears the backoff, since a push is the most likely fix. Idle ticks exit without output so the journal keeps showing the runs that matter, and bin/repo timers reports backoff state — a paused channel is otherwise indistinguishable from an idle one. --- README.md | 50 ++++++++---- bin/auto-release | 94 ++++++++++++++++++++--- bin/check-versions | 19 +++++ bin/timers | 35 +++++++++ helpers/lock-helpers.sh | 73 ++++++++++++++---- systemd/omarchy-auto-release-edge.timer | 6 +- systemd/omarchy-auto-release-rc.timer | 8 +- systemd/omarchy-auto-release-stable.timer | 6 +- systemd/omarchy-check-versions.service | 3 +- systemd/omarchy-check-versions.timer | 6 +- 10 files changed, 243 insertions(+), 57 deletions(-) diff --git a/README.md b/README.md index 1de5c12..887f7bf 100644 --- a/README.md +++ b/README.md @@ -728,14 +728,33 @@ The repository includes GitHub workflows and systemd services for automated rele #### Systemd Services -1. **check-versions** (Every 6 hours at :30): Pulls latest from git, compares PKGBUILD versions to published versions, creates state files if builds are needed -2. **auto-release-edge** (Every 6 hours at +1:00): If state file exists, builds all edge packages that need updates -3. **auto-release-rc** (Every 6 hours at +1:00): Retry backstop — rc builds are normally triggered immediately over SSH by the release orchestrator; the timer re-runs any build whose state file survived a failure. Builds from the `rc` branch worktree (`/root/omarchy-pkgs-rc`), publishing into the shared channel tree -4. **auto-release-stable** (Every 6 hours at +1:00): If state file exists, builds `release_ring=fast` packages for stable and replicates them to rc (runs in parallel with edge) +All four units run **every 5 minutes**, staggered by a minute each, so a push +reaches the mirror in minutes rather than hours: -All channel-mutating runs share a host-wide release lock -(`pkgs.omarchy.org/.release.lock`), so overlapping timers and manual runs -serialize instead of interleaving. +1. **check-versions** (`*:0/5`): Pulls latest from git, compares PKGBUILD versions to published versions, creates state files if builds are needed +2. **auto-release-edge** (`*:1/5`): If a state file exists, builds all edge packages that need updates +3. **auto-release-rc** (`*:2/5`): Builds the rc channel from the `rc` branch worktree (`/root/omarchy-pkgs-rc`), publishing into the shared channel tree. The orchestrator also triggers this immediately over ssh when cutting an RC +4. **auto-release-stable** (`*:3/5`): If a state file exists, builds `release_ring=fast` packages for stable and replicates them to rc + +That cadence is only safe because of three guards: + +- **No overlap.** Every channel-mutating run takes a host-wide lock + (`pkgs.omarchy.org/.release.lock`). Scheduled runs take it + **non-blocking**: if a build is already going, the tick exits immediately + instead of queuing. Waiting would stack one stalled process per tick behind + a long build and stampede when it finished. Manual commands still wait, as + an operator expects. `check-versions` takes it too — its `git pull` would + otherwise swap PKGBUILDs out from under a running build. +- **Backoff on failure.** A failed release records the attempt in + `.build-failed-` and backs off exponentially — 10m, 20m, 40m, up to + a 6h ceiling — instead of rebuilding the same broken tree every 5 minutes. + **Any new commit clears the backoff immediately**, since a push is the most + likely fix. Clear it by hand with `rm /root/.state/.build-failed-`. +- **Quiet when idle.** With nothing queued a tick exits without output, so the + journal shows the runs that mattered rather than 288 no-ops a day. + +`bin/repo timers` reports all of this: schedules, last results, what is +queued, what is failing and when it will retry, and whether the lock is held. Check on all of it with `bin/repo timers` — schedule, each unit's last run and whether it succeeded, what is queued, whether a release is running right now, @@ -748,16 +767,21 @@ bin/repo timers --local # inspect this machine instead ``` State files are stored in `/root/.state/`: -- `.sync-needed-edge` -- `.sync-needed-rc` -- `.sync-needed-stable` +- `.sync-needed-` — a build is queued for that channel +- `.build-failed-` — consecutive failure count, timestamp, and the + commit it failed on (drives the backoff; removing it forces a retry) ### Schedule (America/New_York) -| Time | Action | +| Minute of every hour | Action | |------|--------| -| 00:30, 06:30, 12:30, 18:30 | check-versions (git pull + creates state files) | -| 01:00, 07:00, 13:00, 19:00 | auto-release-edge + auto-release-stable (parallel) | +| :00, :05, :10, … | check-versions (git pull + creates state files) | +| :01, :06, :11, … | auto-release-edge | +| :02, :07, :12, … | auto-release-rc | +| :03, :08, :13, … | auto-release-stable | + +Each unit is a no-op unless its channel has queued work, another run holds the +lock, or the channel is in failure backoff. ### Installation diff --git a/bin/auto-release b/bin/auto-release index 584e841..4d6d064 100755 --- a/bin/auto-release +++ b/bin/auto-release @@ -1,16 +1,33 @@ #!/bin/bash -# Process sync for a specific mirror if state file exists -# Usage: process-sync -# Example: process-sync edge +# Run the release workflow for a channel when work is queued. +# Usage: auto-release +# +# Safe to run on a tight schedule. Three guards make that true: +# +# 1. Nothing queued -> exit immediately (the common case). +# 2. A release already running -> skip this tick. The lock is taken +# non-blocking on purpose: waiting would pile up one stalled process per +# tick behind a long build, and they would all stampede when it finished. +# 3. The last attempt failed -> back off exponentially rather than rebuild +# the same broken tree every few minutes. A change to the repository +# (new commit) clears the backoff immediately, because that is the thing +# most likely to have fixed it. set -e BUILD_ROOT=$(realpath "${BASH_SOURCE[0]%/*}/..") source "$BUILD_ROOT/helpers/message-helpers.sh" +source "$BUILD_ROOT/helpers/paths.sh" +source "$BUILD_ROOT/helpers/lock-helpers.sh" MIRROR="${1:-}" STATE_DIR="${OMARCHY_STATE_DIR:-/root/.state}" +# Backoff schedule: 10m, 20m, 40m, 80m, 160m, 320m, then hourly-ish forever +# (capped at 6h, the cadence this system ran at before frequent timers). +BACKOFF_BASE_SECONDS="${OMARCHY_BACKOFF_BASE:-600}" +BACKOFF_MAX_SECONDS="${OMARCHY_BACKOFF_MAX:-21600}" + if [[ -z "$MIRROR" ]]; then print_error "Usage: $0 " echo " mirror: edge, rc, or stable" @@ -23,27 +40,82 @@ if [[ "$MIRROR" != "edge" && "$MIRROR" != "rc" && "$MIRROR" != "stable" ]]; then fi STATE_FILE="$STATE_DIR/.sync-needed-$MIRROR" +FAIL_FILE="$STATE_DIR/.build-failed-$MIRROR" + +# Nothing queued: stay quiet. At a 5-minute cadence this is most invocations, +# and a header for each would bury the runs that matter in the journal. +if [[ ! -f "$STATE_FILE" ]]; then + exit 0 +fi print_header "Processing Sync for $MIRROR" -# Check if state file exists -if [[ ! -f "$STATE_FILE" ]]; then - print_info "No sync needed for $MIRROR (state file not found)" +# The build inputs are this checkout's contents; its HEAD identifies them. +current_fingerprint() { + git -C "$BUILD_ROOT" rev-parse HEAD 2>/dev/null || echo "unknown" +} + +backoff_seconds() { + local count="$1" delay="$BACKOFF_BASE_SECONDS" + while ((count > 1)); do + delay=$((delay * 2)) + ((delay >= BACKOFF_MAX_SECONDS)) && { delay=$BACKOFF_MAX_SECONDS; break; } + count=$((count - 1)) + done + echo "$delay" +} + +FAIL_COUNT=0 +if [[ -f "$FAIL_FILE" ]]; then + # shellcheck disable=SC1090 + source "$FAIL_FILE" 2>/dev/null || true + FAIL_COUNT="${FAILURE_COUNT:-0}" + failed_at="${FAILURE_AT:-0}" + failed_fingerprint="${FAILURE_FINGERPRINT:-}" + + if [[ "$failed_fingerprint" != "$(current_fingerprint)" ]]; then + print_info "Repository changed since the last failure — clearing backoff and retrying" + rm -f "$FAIL_FILE" + FAIL_COUNT=0 + else + delay=$(backoff_seconds "$FAIL_COUNT") + now=$(date +%s) + retry_at=$((failed_at + delay)) + if ((now < retry_at)); then + print_warning "$MIRROR has failed $FAIL_COUNT time(s) on this tree — not retrying until $(date -d "@$retry_at" '+%H:%M:%S' 2>/dev/null || echo "+$((retry_at - now))s")" + echo " Push a fix (any new commit clears this), or: rm $FAIL_FILE" + exit 0 + fi + print_info "Backoff elapsed — retrying $MIRROR (failure #$((FAIL_COUNT + 1)) if this fails)" + fi +fi + +# Non-blocking: a build in progress means this tick has nothing to do. +if ! try_release_lock; then + holder=$(release_lock_holder) + print_info "A release is already running (${holder:-holder unknown}) — skipping this tick" exit 0 fi print_info "State file found: $STATE_FILE" print_info "Starting release workflow for $MIRROR..." -# Run the release workflow if "$BUILD_ROOT/bin/repo" release --mirror "$MIRROR" --skip-prod-check; then print_success "Release completed successfully for $MIRROR" - - # Remove state file on success rm -f "$STATE_FILE" + rm -f "$FAIL_FILE" print_success "State file removed: $STATE_FILE" else - print_error "Release failed for $MIRROR" + status=$? + FAIL_COUNT=$((FAIL_COUNT + 1)) + cat >"$FAIL_FILE" </dev/null || true + delay=600 + for ((i = 1; i < FAILURE_COUNT; i++)); do + delay=$((delay * 2)) + ((delay >= 21600)) && { delay=21600; break; } + done + retry_at=$((FAILURE_AT + delay)) + now=$(date +%s) + if ((now < retry_at)); then + when="retries at $(date -d "@$retry_at" '+%H:%M:%S' 2>/dev/null || echo "+$((retry_at - now))s")" + else + when="retries on the next tick" + fi + printf ' ✗ %-7s %s consecutive failure(s), %s\n' "$channel" "$FAILURE_COUNT" "$when" + printf ' last attempt %s on commit %s\n' \ + "$(date -d "@$FAILURE_AT" '+%Y-%m-%d %H:%M:%S' 2>/dev/null || echo "$FAILURE_AT")" \ + "${FAILURE_FINGERPRINT:0:12}" +done +if [[ "$paused" == true ]]; then + echo " Any new commit clears the backoff; or: rm $STATE_DIR/.build-failed-" + echo "" +fi + # --- release lock ------------------------------------------------------------ print_info "Release lock" diff --git a/helpers/lock-helpers.sh b/helpers/lock-helpers.sh index 01a110e..6f6f9a5 100644 --- a/helpers/lock-helpers.sh +++ b/helpers/lock-helpers.sh @@ -6,39 +6,78 @@ # # The lock lives beside the published tree (REPO_ROOT), not the checkout, so # the primary checkout and the rc branch worktree contend on the same file. -# Reentrant across child scripts: acquire_release_lock exports -# OMARCHY_RELEASE_LOCK_HELD, and children skip acquisition when they see it -# (the flock fd is inherited, so the lock stays held for the whole tree). +# Reentrant across child scripts: acquiring exports OMARCHY_RELEASE_LOCK_HELD, +# and children skip acquisition when they see it (the flock fd is inherited, +# so the lock stays held for the whole tree). +# +# Two acquisition modes: +# acquire_release_lock waits — for humans, who want the command to run +# try_release_lock fails immediately — for timers, which must never +# queue up behind a long build and stampede when it +# finishes RELEASE_LOCK_FD=9 +release_lock_file() { + echo "${REPO_ROOT:-$BUILD_ROOT/pkgs.omarchy.org}/.release.lock" +} + +release_lock_holder() { + tail -1 "$(release_lock_file)" 2>/dev/null +} + +# True when the recorded holder is a live process. A holder line left behind by +# a killed run describes nothing that is still running. +release_lock_is_held() { + local holder pid + holder=$(release_lock_holder) || return 1 + [[ -n "$holder" ]] || return 1 + pid=$(sed -n 's/^pid \([0-9]\+\).*/\1/p' <<<"$holder") + [[ -n "$pid" ]] && kill -0 "$pid" 2>/dev/null +} + +_release_lock_open() { + local lock_file + lock_file=$(release_lock_file) + mkdir -p "$(dirname "$lock_file")" + eval "exec $RELEASE_LOCK_FD>>\"\$lock_file\"" +} + +_release_lock_record() { + local lock_file + lock_file=$(release_lock_file) + # Truncate first so a crashed holder's stale line does not linger. + : >"$lock_file" + echo "pid $$ ($0) since $(date '+%Y-%m-%d %H:%M:%S')" >>"$lock_file" + export OMARCHY_RELEASE_LOCK_HELD=1 +} + acquire_release_lock() { local timeout="${1:-3600}" - if [[ -n "${OMARCHY_RELEASE_LOCK_HELD:-}" ]]; then - return 0 - fi + [[ -n "${OMARCHY_RELEASE_LOCK_HELD:-}" ]] && return 0 - local lock_file="${REPO_ROOT:-$BUILD_ROOT/pkgs.omarchy.org}/.release.lock" - mkdir -p "$(dirname "$lock_file")" - - eval "exec $RELEASE_LOCK_FD>>\"\$lock_file\"" + _release_lock_open if ! flock -n "$RELEASE_LOCK_FD"; then local holder - holder=$(cat "$lock_file" 2>/dev/null | tail -1) + holder=$(release_lock_holder) echo "Waiting for release lock (up to ${timeout}s)${holder:+ — held by: $holder}" >&2 if ! flock -w "$timeout" "$RELEASE_LOCK_FD"; then - echo "Could not acquire release lock within ${timeout}s: $lock_file" >&2 + echo "Could not acquire release lock within ${timeout}s: $(release_lock_file)" >&2 echo "If no release is actually running, remove the file and retry." >&2 return 1 fi fi - # Record the holder for the "waiting for" message above. Truncate first so a - # crashed holder's stale line does not linger once we own the lock. - : >"$lock_file" - echo "pid $$ ($0) since $(date '+%Y-%m-%d %H:%M:%S')" >>"$lock_file" + _release_lock_record +} - export OMARCHY_RELEASE_LOCK_HELD=1 +# Non-blocking. Returns 1 immediately when another run holds the lock, so +# scheduled work can skip this tick instead of piling up. +try_release_lock() { + [[ -n "${OMARCHY_RELEASE_LOCK_HELD:-}" ]] && return 0 + _release_lock_open + flock -n "$RELEASE_LOCK_FD" || return 1 + _release_lock_record } diff --git a/systemd/omarchy-auto-release-edge.timer b/systemd/omarchy-auto-release-edge.timer index 7df3a26..0e97307 100644 --- a/systemd/omarchy-auto-release-edge.timer +++ b/systemd/omarchy-auto-release-edge.timer @@ -1,10 +1,10 @@ [Unit] -Description=Auto-release edge mirror every 6 hours +Description=Release queued edge builds every 5 minutes [Timer] -# Run every 6 hours at :00 (after version check at :30) -OnCalendar=*-*-* 01,07,13,19:00:00 America/New_York +OnCalendar=*:1/5 Persistent=true +RandomizedDelaySec=20 [Install] WantedBy=timers.target diff --git a/systemd/omarchy-auto-release-rc.timer b/systemd/omarchy-auto-release-rc.timer index 2108a38..ec16d3c 100644 --- a/systemd/omarchy-auto-release-rc.timer +++ b/systemd/omarchy-auto-release-rc.timer @@ -1,12 +1,10 @@ [Unit] -Description=Retry pending rc releases every 6 hours +Description=Release queued rc builds every 5 minutes [Timer] -# Backstop only: rc builds are normally triggered immediately over SSH by the -# release orchestrator (touch .sync-needed-rc + systemctl start). This timer -# retries builds whose state file survived a failure. -OnCalendar=*-*-* 01,07,13,19:00:00 America/New_York +OnCalendar=*:2/5 Persistent=true +RandomizedDelaySec=20 [Install] WantedBy=timers.target diff --git a/systemd/omarchy-auto-release-stable.timer b/systemd/omarchy-auto-release-stable.timer index 237726e..fde42b0 100644 --- a/systemd/omarchy-auto-release-stable.timer +++ b/systemd/omarchy-auto-release-stable.timer @@ -1,10 +1,10 @@ [Unit] -Description=Auto-release stable mirror every 6 hours +Description=Release queued stable builds every 5 minutes [Timer] -# Run every 6 hours at :00 (after version check at :30) -OnCalendar=*-*-* 01,07,13,19:00:00 America/New_York +OnCalendar=*:3/5 Persistent=true +RandomizedDelaySec=20 [Install] WantedBy=timers.target diff --git a/systemd/omarchy-check-versions.service b/systemd/omarchy-check-versions.service index dc4d228..94ec968 100644 --- a/systemd/omarchy-check-versions.service +++ b/systemd/omarchy-check-versions.service @@ -5,8 +5,7 @@ Wants=network-online.target [Service] Type=oneshot -ExecStartPre=/usr/bin/git pull --ff-only -ExecStart=/root/omarchy-pkgs/bin/check-versions +ExecStart=/root/omarchy-pkgs/bin/check-versions --pull Environment=OMARCHY_STATE_DIR=/root/.state WorkingDirectory=/root/omarchy-pkgs diff --git a/systemd/omarchy-check-versions.timer b/systemd/omarchy-check-versions.timer index 0c399eb..41cded0 100644 --- a/systemd/omarchy-check-versions.timer +++ b/systemd/omarchy-check-versions.timer @@ -1,10 +1,10 @@ [Unit] -Description=Check package versions every 6 hours +Description=Check package versions every 5 minutes [Timer] -# Run every 6 hours at :30 (after GitHub sync at :00) -OnCalendar=*-*-* 00,06,12,18:30:00 America/New_York +OnCalendar=*:0/5 Persistent=true +RandomizedDelaySec=20 [Install] WantedBy=timers.target