Run the timers every 5 minutes, with overlap and failure guards
A push reaching the mirror should take minutes, not up to six hours. All four units now fire every 5 minutes, staggered a minute apart. Three guards make that cadence safe: - Scheduled runs take the release lock NON-BLOCKING (try_release_lock) and skip the tick when a build is running. Blocking would stack one stalled process per tick behind a long build and stampede when it finished. Manual commands still wait, as an operator expects. - check-versions takes the lock too, and now owns its git pull (--pull, passed by the unit) instead of an ExecStartPre: at this cadence an unlocked pull would swap PKGBUILDs out from under a running build. - A failed release records .build-failed-<channel> and backs off exponentially (10m, 20m, 40m … capped at 6h) rather than rebuilding the same broken tree every 5 minutes. Any new commit clears the backoff, since a push is the most likely fix. Idle ticks exit without output so the journal keeps showing the runs that matter, and bin/repo timers reports backoff state — a paused channel is otherwise indistinguishable from an idle one.
This commit is contained in:
@@ -728,14 +728,33 @@ The repository includes GitHub workflows and systemd services for automated rele
|
||||
|
||||
#### Systemd Services
|
||||
|
||||
1. **check-versions** (Every 6 hours at :30): Pulls latest from git, compares PKGBUILD versions to published versions, creates state files if builds are needed
|
||||
2. **auto-release-edge** (Every 6 hours at +1:00): If state file exists, builds all edge packages that need updates
|
||||
3. **auto-release-rc** (Every 6 hours at +1:00): Retry backstop — rc builds are normally triggered immediately over SSH by the release orchestrator; the timer re-runs any build whose state file survived a failure. Builds from the `rc` branch worktree (`/root/omarchy-pkgs-rc`), publishing into the shared channel tree
|
||||
4. **auto-release-stable** (Every 6 hours at +1:00): If state file exists, builds `release_ring=fast` packages for stable and replicates them to rc (runs in parallel with edge)
|
||||
All four units run **every 5 minutes**, staggered by a minute each, so a push
|
||||
reaches the mirror in minutes rather than hours:
|
||||
|
||||
All channel-mutating runs share a host-wide release lock
|
||||
(`pkgs.omarchy.org/.release.lock`), so overlapping timers and manual runs
|
||||
serialize instead of interleaving.
|
||||
1. **check-versions** (`*:0/5`): Pulls latest from git, compares PKGBUILD versions to published versions, creates state files if builds are needed
|
||||
2. **auto-release-edge** (`*:1/5`): If a state file exists, builds all edge packages that need updates
|
||||
3. **auto-release-rc** (`*:2/5`): Builds the rc channel from the `rc` branch worktree (`/root/omarchy-pkgs-rc`), publishing into the shared channel tree. The orchestrator also triggers this immediately over ssh when cutting an RC
|
||||
4. **auto-release-stable** (`*:3/5`): If a state file exists, builds `release_ring=fast` packages for stable and replicates them to rc
|
||||
|
||||
That cadence is only safe because of three guards:
|
||||
|
||||
- **No overlap.** Every channel-mutating run takes a host-wide lock
|
||||
(`pkgs.omarchy.org/.release.lock`). Scheduled runs take it
|
||||
**non-blocking**: if a build is already going, the tick exits immediately
|
||||
instead of queuing. Waiting would stack one stalled process per tick behind
|
||||
a long build and stampede when it finished. Manual commands still wait, as
|
||||
an operator expects. `check-versions` takes it too — its `git pull` would
|
||||
otherwise swap PKGBUILDs out from under a running build.
|
||||
- **Backoff on failure.** A failed release records the attempt in
|
||||
`.build-failed-<channel>` and backs off exponentially — 10m, 20m, 40m, up to
|
||||
a 6h ceiling — instead of rebuilding the same broken tree every 5 minutes.
|
||||
**Any new commit clears the backoff immediately**, since a push is the most
|
||||
likely fix. Clear it by hand with `rm /root/.state/.build-failed-<channel>`.
|
||||
- **Quiet when idle.** With nothing queued a tick exits without output, so the
|
||||
journal shows the runs that mattered rather than 288 no-ops a day.
|
||||
|
||||
`bin/repo timers` reports all of this: schedules, last results, what is
|
||||
queued, what is failing and when it will retry, and whether the lock is held.
|
||||
|
||||
Check on all of it with `bin/repo timers` — schedule, each unit's last run and
|
||||
whether it succeeded, what is queued, whether a release is running right now,
|
||||
@@ -748,16 +767,21 @@ bin/repo timers --local # inspect this machine instead
|
||||
```
|
||||
|
||||
State files are stored in `/root/.state/`:
|
||||
- `.sync-needed-edge`
|
||||
- `.sync-needed-rc`
|
||||
- `.sync-needed-stable`
|
||||
- `.sync-needed-<channel>` — a build is queued for that channel
|
||||
- `.build-failed-<channel>` — consecutive failure count, timestamp, and the
|
||||
commit it failed on (drives the backoff; removing it forces a retry)
|
||||
|
||||
### Schedule (America/New_York)
|
||||
|
||||
| Time | Action |
|
||||
| Minute of every hour | Action |
|
||||
|------|--------|
|
||||
| 00:30, 06:30, 12:30, 18:30 | check-versions (git pull + creates state files) |
|
||||
| 01:00, 07:00, 13:00, 19:00 | auto-release-edge + auto-release-stable (parallel) |
|
||||
| :00, :05, :10, … | check-versions (git pull + creates state files) |
|
||||
| :01, :06, :11, … | auto-release-edge |
|
||||
| :02, :07, :12, … | auto-release-rc |
|
||||
| :03, :08, :13, … | auto-release-stable |
|
||||
|
||||
Each unit is a no-op unless its channel has queued work, another run holds the
|
||||
lock, or the channel is in failure backoff.
|
||||
|
||||
### Installation
|
||||
|
||||
|
||||
+83
-11
@@ -1,16 +1,33 @@
|
||||
#!/bin/bash
|
||||
# Process sync for a specific mirror if state file exists
|
||||
# Usage: process-sync <mirror>
|
||||
# Example: process-sync edge
|
||||
# Run the release workflow for a channel when work is queued.
|
||||
# Usage: auto-release <edge|rc|stable>
|
||||
#
|
||||
# Safe to run on a tight schedule. Three guards make that true:
|
||||
#
|
||||
# 1. Nothing queued -> exit immediately (the common case).
|
||||
# 2. A release already running -> skip this tick. The lock is taken
|
||||
# non-blocking on purpose: waiting would pile up one stalled process per
|
||||
# tick behind a long build, and they would all stampede when it finished.
|
||||
# 3. The last attempt failed -> back off exponentially rather than rebuild
|
||||
# the same broken tree every few minutes. A change to the repository
|
||||
# (new commit) clears the backoff immediately, because that is the thing
|
||||
# most likely to have fixed it.
|
||||
|
||||
set -e
|
||||
|
||||
BUILD_ROOT=$(realpath "${BASH_SOURCE[0]%/*}/..")
|
||||
source "$BUILD_ROOT/helpers/message-helpers.sh"
|
||||
source "$BUILD_ROOT/helpers/paths.sh"
|
||||
source "$BUILD_ROOT/helpers/lock-helpers.sh"
|
||||
|
||||
MIRROR="${1:-}"
|
||||
STATE_DIR="${OMARCHY_STATE_DIR:-/root/.state}"
|
||||
|
||||
# Backoff schedule: 10m, 20m, 40m, 80m, 160m, 320m, then hourly-ish forever
|
||||
# (capped at 6h, the cadence this system ran at before frequent timers).
|
||||
BACKOFF_BASE_SECONDS="${OMARCHY_BACKOFF_BASE:-600}"
|
||||
BACKOFF_MAX_SECONDS="${OMARCHY_BACKOFF_MAX:-21600}"
|
||||
|
||||
if [[ -z "$MIRROR" ]]; then
|
||||
print_error "Usage: $0 <mirror>"
|
||||
echo " mirror: edge, rc, or stable"
|
||||
@@ -23,27 +40,82 @@ if [[ "$MIRROR" != "edge" && "$MIRROR" != "rc" && "$MIRROR" != "stable" ]]; then
|
||||
fi
|
||||
|
||||
STATE_FILE="$STATE_DIR/.sync-needed-$MIRROR"
|
||||
FAIL_FILE="$STATE_DIR/.build-failed-$MIRROR"
|
||||
|
||||
# Nothing queued: stay quiet. At a 5-minute cadence this is most invocations,
|
||||
# and a header for each would bury the runs that matter in the journal.
|
||||
if [[ ! -f "$STATE_FILE" ]]; then
|
||||
exit 0
|
||||
fi
|
||||
|
||||
print_header "Processing Sync for $MIRROR"
|
||||
|
||||
# Check if state file exists
|
||||
if [[ ! -f "$STATE_FILE" ]]; then
|
||||
print_info "No sync needed for $MIRROR (state file not found)"
|
||||
# The build inputs are this checkout's contents; its HEAD identifies them.
|
||||
current_fingerprint() {
|
||||
git -C "$BUILD_ROOT" rev-parse HEAD 2>/dev/null || echo "unknown"
|
||||
}
|
||||
|
||||
backoff_seconds() {
|
||||
local count="$1" delay="$BACKOFF_BASE_SECONDS"
|
||||
while ((count > 1)); do
|
||||
delay=$((delay * 2))
|
||||
((delay >= BACKOFF_MAX_SECONDS)) && { delay=$BACKOFF_MAX_SECONDS; break; }
|
||||
count=$((count - 1))
|
||||
done
|
||||
echo "$delay"
|
||||
}
|
||||
|
||||
FAIL_COUNT=0
|
||||
if [[ -f "$FAIL_FILE" ]]; then
|
||||
# shellcheck disable=SC1090
|
||||
source "$FAIL_FILE" 2>/dev/null || true
|
||||
FAIL_COUNT="${FAILURE_COUNT:-0}"
|
||||
failed_at="${FAILURE_AT:-0}"
|
||||
failed_fingerprint="${FAILURE_FINGERPRINT:-}"
|
||||
|
||||
if [[ "$failed_fingerprint" != "$(current_fingerprint)" ]]; then
|
||||
print_info "Repository changed since the last failure — clearing backoff and retrying"
|
||||
rm -f "$FAIL_FILE"
|
||||
FAIL_COUNT=0
|
||||
else
|
||||
delay=$(backoff_seconds "$FAIL_COUNT")
|
||||
now=$(date +%s)
|
||||
retry_at=$((failed_at + delay))
|
||||
if ((now < retry_at)); then
|
||||
print_warning "$MIRROR has failed $FAIL_COUNT time(s) on this tree — not retrying until $(date -d "@$retry_at" '+%H:%M:%S' 2>/dev/null || echo "+$((retry_at - now))s")"
|
||||
echo " Push a fix (any new commit clears this), or: rm $FAIL_FILE"
|
||||
exit 0
|
||||
fi
|
||||
print_info "Backoff elapsed — retrying $MIRROR (failure #$((FAIL_COUNT + 1)) if this fails)"
|
||||
fi
|
||||
fi
|
||||
|
||||
# Non-blocking: a build in progress means this tick has nothing to do.
|
||||
if ! try_release_lock; then
|
||||
holder=$(release_lock_holder)
|
||||
print_info "A release is already running (${holder:-holder unknown}) — skipping this tick"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
print_info "State file found: $STATE_FILE"
|
||||
print_info "Starting release workflow for $MIRROR..."
|
||||
|
||||
# Run the release workflow
|
||||
if "$BUILD_ROOT/bin/repo" release --mirror "$MIRROR" --skip-prod-check; then
|
||||
print_success "Release completed successfully for $MIRROR"
|
||||
|
||||
# Remove state file on success
|
||||
rm -f "$STATE_FILE"
|
||||
rm -f "$FAIL_FILE"
|
||||
print_success "State file removed: $STATE_FILE"
|
||||
else
|
||||
print_error "Release failed for $MIRROR"
|
||||
status=$?
|
||||
FAIL_COUNT=$((FAIL_COUNT + 1))
|
||||
cat >"$FAIL_FILE" <<EOF
|
||||
FAILURE_COUNT=$FAIL_COUNT
|
||||
FAILURE_AT=$(date +%s)
|
||||
FAILURE_FINGERPRINT=$(current_fingerprint)
|
||||
EOF
|
||||
next=$(backoff_seconds "$FAIL_COUNT")
|
||||
print_error "Release failed for $MIRROR (attempt $FAIL_COUNT)"
|
||||
print_warning "State file retained for retry: $STATE_FILE"
|
||||
exit 1
|
||||
print_warning "Backing off $((next / 60))m before the next attempt; a new commit retries sooner"
|
||||
exit "$status"
|
||||
fi
|
||||
|
||||
@@ -8,14 +8,33 @@ BUILD_ROOT=$(realpath "${BASH_SOURCE[0]%/*}/..")
|
||||
source "$BUILD_ROOT/helpers/message-helpers.sh"
|
||||
source "$BUILD_ROOT/helpers/paths.sh"
|
||||
source "$BUILD_ROOT/helpers/package-metadata.sh"
|
||||
source "$BUILD_ROOT/helpers/lock-helpers.sh"
|
||||
|
||||
STATE_DIR="${OMARCHY_STATE_DIR:-/root/.state}"
|
||||
ARCH="${ARCH:-x86_64}"
|
||||
PULL=false
|
||||
|
||||
[[ "${1:-}" == "--pull" ]] && PULL=true
|
||||
|
||||
mkdir -p "$STATE_DIR"
|
||||
|
||||
# Runs on a tight schedule, so it must never disturb a build in progress:
|
||||
# comparing versions reads the PKGBUILDs a running build is reading, and
|
||||
# --pull would swap them underneath it. Take the lock non-blocking and skip
|
||||
# the tick when a release owns it — the queue is unchanged, so the next tick
|
||||
# picks up exactly where this one left off.
|
||||
if ! try_release_lock; then
|
||||
exit 0
|
||||
fi
|
||||
|
||||
print_header "Package Version Check"
|
||||
|
||||
# The scheduled run pulls; a manual run leaves the operator's tree alone.
|
||||
if [[ "$PULL" == true ]]; then
|
||||
print_info "Updating repository..."
|
||||
git -C "$BUILD_ROOT" pull --ff-only || print_warning "git pull failed — checking the current tree"
|
||||
fi
|
||||
|
||||
get_repo_version() {
|
||||
local pkg="$1"
|
||||
local mirror="$2"
|
||||
|
||||
+35
@@ -97,6 +97,41 @@ done
|
||||
[[ "$queued" == false ]] && echo " (nothing queued — all channels up to date)"
|
||||
echo ""
|
||||
|
||||
# A channel in backoff looks identical to an idle one from the outside, so
|
||||
# say so plainly: it is queued but deliberately not being retried yet.
|
||||
paused=false
|
||||
for channel in edge rc stable; do
|
||||
fail_file="$STATE_DIR/.build-failed-$channel"
|
||||
[[ -f "$fail_file" ]] || continue
|
||||
if [[ "$paused" == false ]]; then
|
||||
print_error "Failing builds (backoff active)"
|
||||
paused=true
|
||||
fi
|
||||
FAILURE_COUNT=0 FAILURE_AT=0 FAILURE_FINGERPRINT=""
|
||||
# shellcheck disable=SC1090
|
||||
source "$fail_file" 2>/dev/null || true
|
||||
delay=600
|
||||
for ((i = 1; i < FAILURE_COUNT; i++)); do
|
||||
delay=$((delay * 2))
|
||||
((delay >= 21600)) && { delay=21600; break; }
|
||||
done
|
||||
retry_at=$((FAILURE_AT + delay))
|
||||
now=$(date +%s)
|
||||
if ((now < retry_at)); then
|
||||
when="retries at $(date -d "@$retry_at" '+%H:%M:%S' 2>/dev/null || echo "+$((retry_at - now))s")"
|
||||
else
|
||||
when="retries on the next tick"
|
||||
fi
|
||||
printf ' ✗ %-7s %s consecutive failure(s), %s\n' "$channel" "$FAILURE_COUNT" "$when"
|
||||
printf ' last attempt %s on commit %s\n' \
|
||||
"$(date -d "@$FAILURE_AT" '+%Y-%m-%d %H:%M:%S' 2>/dev/null || echo "$FAILURE_AT")" \
|
||||
"${FAILURE_FINGERPRINT:0:12}"
|
||||
done
|
||||
if [[ "$paused" == true ]]; then
|
||||
echo " Any new commit clears the backoff; or: rm $STATE_DIR/.build-failed-<channel>"
|
||||
echo ""
|
||||
fi
|
||||
|
||||
# --- release lock ------------------------------------------------------------
|
||||
|
||||
print_info "Release lock"
|
||||
|
||||
+56
-17
@@ -6,39 +6,78 @@
|
||||
#
|
||||
# The lock lives beside the published tree (REPO_ROOT), not the checkout, so
|
||||
# the primary checkout and the rc branch worktree contend on the same file.
|
||||
# Reentrant across child scripts: acquire_release_lock exports
|
||||
# OMARCHY_RELEASE_LOCK_HELD, and children skip acquisition when they see it
|
||||
# (the flock fd is inherited, so the lock stays held for the whole tree).
|
||||
# Reentrant across child scripts: acquiring exports OMARCHY_RELEASE_LOCK_HELD,
|
||||
# and children skip acquisition when they see it (the flock fd is inherited,
|
||||
# so the lock stays held for the whole tree).
|
||||
#
|
||||
# Two acquisition modes:
|
||||
# acquire_release_lock waits — for humans, who want the command to run
|
||||
# try_release_lock fails immediately — for timers, which must never
|
||||
# queue up behind a long build and stampede when it
|
||||
# finishes
|
||||
|
||||
RELEASE_LOCK_FD=9
|
||||
|
||||
release_lock_file() {
|
||||
echo "${REPO_ROOT:-$BUILD_ROOT/pkgs.omarchy.org}/.release.lock"
|
||||
}
|
||||
|
||||
release_lock_holder() {
|
||||
tail -1 "$(release_lock_file)" 2>/dev/null
|
||||
}
|
||||
|
||||
# True when the recorded holder is a live process. A holder line left behind by
|
||||
# a killed run describes nothing that is still running.
|
||||
release_lock_is_held() {
|
||||
local holder pid
|
||||
holder=$(release_lock_holder) || return 1
|
||||
[[ -n "$holder" ]] || return 1
|
||||
pid=$(sed -n 's/^pid \([0-9]\+\).*/\1/p' <<<"$holder")
|
||||
[[ -n "$pid" ]] && kill -0 "$pid" 2>/dev/null
|
||||
}
|
||||
|
||||
_release_lock_open() {
|
||||
local lock_file
|
||||
lock_file=$(release_lock_file)
|
||||
mkdir -p "$(dirname "$lock_file")"
|
||||
eval "exec $RELEASE_LOCK_FD>>\"\$lock_file\""
|
||||
}
|
||||
|
||||
_release_lock_record() {
|
||||
local lock_file
|
||||
lock_file=$(release_lock_file)
|
||||
# Truncate first so a crashed holder's stale line does not linger.
|
||||
: >"$lock_file"
|
||||
echo "pid $$ ($0) since $(date '+%Y-%m-%d %H:%M:%S')" >>"$lock_file"
|
||||
export OMARCHY_RELEASE_LOCK_HELD=1
|
||||
}
|
||||
|
||||
acquire_release_lock() {
|
||||
local timeout="${1:-3600}"
|
||||
|
||||
if [[ -n "${OMARCHY_RELEASE_LOCK_HELD:-}" ]]; then
|
||||
return 0
|
||||
fi
|
||||
[[ -n "${OMARCHY_RELEASE_LOCK_HELD:-}" ]] && return 0
|
||||
|
||||
local lock_file="${REPO_ROOT:-$BUILD_ROOT/pkgs.omarchy.org}/.release.lock"
|
||||
mkdir -p "$(dirname "$lock_file")"
|
||||
|
||||
eval "exec $RELEASE_LOCK_FD>>\"\$lock_file\""
|
||||
_release_lock_open
|
||||
|
||||
if ! flock -n "$RELEASE_LOCK_FD"; then
|
||||
local holder
|
||||
holder=$(cat "$lock_file" 2>/dev/null | tail -1)
|
||||
holder=$(release_lock_holder)
|
||||
echo "Waiting for release lock (up to ${timeout}s)${holder:+ — held by: $holder}" >&2
|
||||
if ! flock -w "$timeout" "$RELEASE_LOCK_FD"; then
|
||||
echo "Could not acquire release lock within ${timeout}s: $lock_file" >&2
|
||||
echo "Could not acquire release lock within ${timeout}s: $(release_lock_file)" >&2
|
||||
echo "If no release is actually running, remove the file and retry." >&2
|
||||
return 1
|
||||
fi
|
||||
fi
|
||||
|
||||
# Record the holder for the "waiting for" message above. Truncate first so a
|
||||
# crashed holder's stale line does not linger once we own the lock.
|
||||
: >"$lock_file"
|
||||
echo "pid $$ ($0) since $(date '+%Y-%m-%d %H:%M:%S')" >>"$lock_file"
|
||||
_release_lock_record
|
||||
}
|
||||
|
||||
export OMARCHY_RELEASE_LOCK_HELD=1
|
||||
# Non-blocking. Returns 1 immediately when another run holds the lock, so
|
||||
# scheduled work can skip this tick instead of piling up.
|
||||
try_release_lock() {
|
||||
[[ -n "${OMARCHY_RELEASE_LOCK_HELD:-}" ]] && return 0
|
||||
_release_lock_open
|
||||
flock -n "$RELEASE_LOCK_FD" || return 1
|
||||
_release_lock_record
|
||||
}
|
||||
|
||||
@@ -1,10 +1,10 @@
|
||||
[Unit]
|
||||
Description=Auto-release edge mirror every 6 hours
|
||||
Description=Release queued edge builds every 5 minutes
|
||||
|
||||
[Timer]
|
||||
# Run every 6 hours at :00 (after version check at :30)
|
||||
OnCalendar=*-*-* 01,07,13,19:00:00 America/New_York
|
||||
OnCalendar=*:1/5
|
||||
Persistent=true
|
||||
RandomizedDelaySec=20
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
|
||||
@@ -1,12 +1,10 @@
|
||||
[Unit]
|
||||
Description=Retry pending rc releases every 6 hours
|
||||
Description=Release queued rc builds every 5 minutes
|
||||
|
||||
[Timer]
|
||||
# Backstop only: rc builds are normally triggered immediately over SSH by the
|
||||
# release orchestrator (touch .sync-needed-rc + systemctl start). This timer
|
||||
# retries builds whose state file survived a failure.
|
||||
OnCalendar=*-*-* 01,07,13,19:00:00 America/New_York
|
||||
OnCalendar=*:2/5
|
||||
Persistent=true
|
||||
RandomizedDelaySec=20
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
|
||||
@@ -1,10 +1,10 @@
|
||||
[Unit]
|
||||
Description=Auto-release stable mirror every 6 hours
|
||||
Description=Release queued stable builds every 5 minutes
|
||||
|
||||
[Timer]
|
||||
# Run every 6 hours at :00 (after version check at :30)
|
||||
OnCalendar=*-*-* 01,07,13,19:00:00 America/New_York
|
||||
OnCalendar=*:3/5
|
||||
Persistent=true
|
||||
RandomizedDelaySec=20
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
|
||||
@@ -5,8 +5,7 @@ Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
ExecStartPre=/usr/bin/git pull --ff-only
|
||||
ExecStart=/root/omarchy-pkgs/bin/check-versions
|
||||
ExecStart=/root/omarchy-pkgs/bin/check-versions --pull
|
||||
Environment=OMARCHY_STATE_DIR=/root/.state
|
||||
WorkingDirectory=/root/omarchy-pkgs
|
||||
|
||||
|
||||
@@ -1,10 +1,10 @@
|
||||
[Unit]
|
||||
Description=Check package versions every 6 hours
|
||||
Description=Check package versions every 5 minutes
|
||||
|
||||
[Timer]
|
||||
# Run every 6 hours at :30 (after GitHub sync at :00)
|
||||
OnCalendar=*-*-* 00,06,12,18:30:00 America/New_York
|
||||
OnCalendar=*:0/5
|
||||
Persistent=true
|
||||
RandomizedDelaySec=20
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
|
||||
Reference in New Issue
Block a user