Build on #306 without replacing its package or updater workaround. Exhaust the component feed before selecting a release and cover quarantine, version ordering, invalid input, and transport failures with offline fixtures.
Incremental advances collide with same-version artifacts that reached the
destination through another lineage — the stable -> rc bootstrap seed and
native rc builds are not byte-identical to edge's builds of the same
version. A collision means the package has not moved in the source since,
so keep the destination's published copy and continue; the hard failure
also aborted before the db rebuild and sync, leaving the advance half-done.
Also stop advancing pinned packages into rc: eligibility deferred to
package_builds_for_mirror, whose OMARCHY_RC_PINS gate is never set during
an advance, so edge's omarchy/omarchy-settings looked movable — and could
have overtaken an in-flight RC pin under a fresh filename.
makepkg always exports CARCH, so PKGBUILDs may branch on it at file
scope. Every place the tooling sourced a PKGBUILD did so without CARCH,
taking the wrong branch or aborting partway, and check-versions turned
the resulting empty pkgver/pkgrel into the version '-', which never
matches a published version and queues an endless rebuild.
package_pkgbuild_var reads one variable the way makepkg would see it,
for the architecture being checked, and reports whether the source
succeeded. check-versions, sync-rebuilds and omarchy-pkgs use it.
check-versions now warns and skips a package whose pkgver or pkgrel is
empty instead of comparing a partial version. A self-test covers a
PKGBUILD that branches on CARCH before assigning its version.
One list, PUBLISHED_ARCHES in helpers/paths.sh (default x86_64,
overridable with OMARCHY_ARCHES), now drives everything the repository
host schedules. check-versions compares PKGBUILDs against each
architecture's channel databases and writes one queue per channel and
architecture; auto-release works through the queues one architecture at
a time, each with its own backoff, so a failing build on one never
blocks the other; advance-channel --arch all re-runs an advance for every
published architecture and omarchy-release uses it for start and ship,
building the pinned pair once per architecture in its rc trigger; the
train observes channels through the reference (first) architecture
instead of a hard-coded x86_64. Queue and backoff files written under
the old per-channel names are treated as x86_64 until consumed.
Two things made an aarch64 builder image impossible to create: the
keyring bootstrap fetched omarchy-keyring from the target architecture's
own channel tree, which does not exist before that architecture has
published anything, and the QEMU probe only knew the x86_64-host,
aarch64-target case. The keyring (arch=any) now always comes from the
x86_64 tree, and the probe compares host and target architectures and
runs a container for the target platform.
clean-repo grouped versions with a regex that only knew any, x86_64 and
i686, so aarch64 packages would never have been pruned.
Schist is a layered image editor with PSD, Affinity and camera raw support,
developed by Infrawrench and packaged by its upstream author. The package
re-wraps the pacman-format payloads Schist's release workflow publishes for
x86_64 and aarch64, so the builder does no compiling, and both assets are
pinned by SHA-256.
Releases are tracked declaratively through the GitHub upstream provider,
which gains a "digests": true mode here: a vendor that publishes no checksum
manifest can have each asset's SHA-256 read from the digest GitHub's release
API reports, so the sync never downloads the artifacts. Exactly one of
"checksums" or "digests" must be set, and the provider enforces that itself
because scheduled runs reach it without the metadata validator.
Fresh releases wait 24 hours before the scheduled sync picks them up, as
mise-bin already does. vulkan-driver is an optional dependency rather than a
hard one: makepkg -s would otherwise satisfy the virtual package with
nvidia-utils in the build container, and Omarchy installs a Vulkan driver per
machine.
Co-authored-by: David Heinemeier Hansson <david@hey.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Three opt-in knobs for bin/build, each defaulting to today's behaviour:
OMARCHY_KEEP_BUILD_WORKSPACE=1 keeps build-output/$MIRROR/$ARCH instead
of wiping it, and build/build.sh now folds any packages already there
into omarchy-build.db even when no database exists yet, so packages
built by an earlier job (or a previous, interrupted run) resolve as
dependencies of what builds next.
OMARCHY_SKIP_BUILDER_IMAGE=1 uses the omarchy-pkg-builder image already
present instead of building it, so a workflow can build the image once
with an external BuildKit cache and fan out over package jobs that all
run the same bytes. A missing image is an error, not a silent rebuild.
OMARCHY_DEFER_RUNTIME_DEPS=true builds the omarchy/omarchy-settings pair
with --nodeps, installing only their makedepends and checkdepends
explicitly. The pair depends on each other and on packages a sharded
pipeline builds in other jobs, so they cannot resolve in isolation; the
assembled set is installed in one verified transaction downstream. The
request is refused for anything but exactly that pair, on the host
before Docker starts and again inside the container.
Also fix make_dir_writable: chown -R can succeed on part of the tree
and fail on files a previous container left behind as another uid, and
the old `|| chmod` fallback only ran when chown failed outright. Always
follow with chmod.
cmd_rc handed maybe_iso the bare train version, so omarchy-iso-release
named every candidate omarchy-X.Y.Z-rc.iso — rc2 uploaded over rc1's
URL. Pass the pinned rcN version instead; paired with omarchy-iso's
f9be60b, the artifact becomes omarchy-4.0.2rc2.iso. The ship path
keeps passing the final version.
pkgs.omarchy.org sits behind Cloudflare, and right after a sync the
plain db URL keeps serving the previous file from the edge cache
(observed: a cache HIT with age 900s+ returning the pre-release db).
That made status report the old version, wait_for_published poll a
frozen file for its full timeout, and would let RC auto-numbering
count from last release's versions. A unique query string per fetch
skips the cached entry.
package_builds_for_mirror gates pinned packages' rc builds on
OMARCHY_RC_PINS, but that check runs inside the container via
build.sh, and docker run only forwarded ARCH/MIRROR/PACKAGES —
so the rc trigger's pinned builds were always skipped as 'not
configured for direct rc builds'.
The rc trigger passes '--package omarchy omarchy-settings', and bin/build
already consumes every name up to the next --option, but bin/release only
took one — the second name fell through to 'Unknown option' and the
release exited before building anything. Parse greedily, like bin/build.
`gh pr list` orders by creation date, so `--limit 30` cut the candidate set by
when PRs were opened, not when they merged. The `sort_by(.mergedAt)` that
followed only reordered whatever survived that cut. A PR opened before the
window but merged inside it — exactly the kind a release branch still needs —
never reached the already-on-branch check at all. #7649 and #7709 were both
missing from v4-0-2's list for this reason.
The window is now bounded by the branch point instead of a count: pull a wide
page and keep what merged after the merge-base's commit date, since anything
merged into the dev branch before the release branch left it is already there
by ancestry. v4-0-2 went from 11 candidates to 31.
Widening it surfaced a second gap. A change re-applied by hand carries neither
a PR number nor a cherry-pick trailer, so already_on_branch could not see it
and offered it again (#6939, applied as 33d7363c). It now also compares the PR
title against the branch's subjects, with any trailing "(#N)" stripped.
The rc channel's Arch base can sit anywhere between stable's snapshot and
edge's, so a package built against stable's libraries is not necessarily
correct for rc. Copying stable's fast-ring artifacts into rc therefore shipped
possibly-mislinked packages to RC testers. Fast-ring packages now build
natively for all three channels, each in its own image against its own base
mirror, and the stable release's replication step is gone.
That required separating 'may be built here' from 'whose version wins'. The
release pair is now marked "pinned": its version is set per release on the rc
branch, so it builds for rc only from that branch's worktree
(OMARCHY_RC_PINS=1, set by omarchy-release rc) — master's shipped pins can
never overwrite an in-flight RC, even though check-versions now discovers rc
work like it does for edge and stable.
The fast-ring replication announced itself before the release that triggered
it, so chat read as though an rc job had run on its own — the reverse of what
happened. The publish report now fires first, and the replication says only
how many packages were kept in parity: the release report immediately above it
already lists them by name.
Reporting to BASECAMP_CHATBOT_URL is a working setup — the second variable
exists only for those who want release traffic in its own chat. setup now
states which chat receives reports instead of warning about the common case,
and the README frames the split as optional rather than expected.
Two failures from the first live run:
notify_basecamp declared 'local BASECAMP_CHATBOT_URL' and then called
release_chatbot_url, whose fallback reads that same global — bash locals are
visible to called functions, so the fallback saw the empty local and every
notification silently went nowhere for anyone with only the legacy variable
set. The local is now named 'url'.
check-versions compared PKGBUILD and published versions with !=, so a checkout
BEHIND the channel queued a rebuild of an older version every cycle: the
builder produced it and promotion refused it, because that exact filename is
already published with different bytes. It now skips (with a warning naming
the package) when an artifact for the PKGBUILD's version already exists in the
channel, whichever direction the versions differ.
The rc service fetched and reset /root/omarchy-pkgs-rc in ExecStartPre, but
that worktree does not exist until the first RC creates the rc branch — so on
a freshly set up host the unit failed every five minutes, forever, and showed
up as a failed unit in the timer report.
It now runs bin/auto-release-rc from the main checkout, which always exists:
nothing queued exits silently, no rc branch is a clean no-op, and a missing
worktree is created on demand before handing off to the worktree's own
auto-release.
Restores the no-change report now that it is clear it cannot fire on an idle
timer tick: releases only run when the version check queued work, so a run
that publishes nothing means check-versions and the builder disagree about
what is out of date. The report names the packages that were queued but never
built, so a recurring disagreement is diagnosable rather than invisible.
A release that published nothing is not news, and at a 5-minute cadence those
messages would bury the ones that matter — the log and bin/repo timers still
show the run happened.
Removing it would have left a start report with no follow-up, so the start
report now carries its own answer: check-versions writes the package names it
queued into the state file instead of touching an empty one, and the release
run reads them. 'A build is running' becomes 'your package is in this build',
which is the question the reports exist to answer.
Only failures were reported, so a push could reach the mirror with no way to
know short of querying the database by hand. Release runs now report:
- start: channel, arch, host, and the commit being built
- published: the packages and versions that went out, duration, channel URL
- no-changes: the run found nothing to build
- promoted: what advance moved between channels, including the rc bootstrap
and fast-ring replication
- failed: unchanged, plus the commit context the other reports carry
Release traffic goes to OMARCHY_RELEASE_CHATBOT_URL, falling back to
BASECAMP_CHATBOT_URL, so build reports stop drowning the repository chat the
sync workflows post to. bin/setup reports which destination is configured.
The published list is captured after the build step because promote moves the
files out of build-output, and is capped at 25 entries so a full rebuild does
not produce an unreadable wall of chat.
A push reaching the mirror should take minutes, not up to six hours. All four
units now fire every 5 minutes, staggered a minute apart. Three guards make
that cadence safe:
- Scheduled runs take the release lock NON-BLOCKING (try_release_lock) and
skip the tick when a build is running. Blocking would stack one stalled
process per tick behind a long build and stampede when it finished. Manual
commands still wait, as an operator expects.
- check-versions takes the lock too, and now owns its git pull (--pull, passed
by the unit) instead of an ExecStartPre: at this cadence an unlocked pull
would swap PKGBUILDs out from under a running build.
- A failed release records .build-failed-<channel> and backs off
exponentially (10m, 20m, 40m … capped at 6h) rather than rebuilding the same
broken tree every 5 minutes. Any new commit clears the backoff, since a push
is the most likely fix.
Idle ticks exit without output so the journal keeps showing the runs that
matter, and bin/repo timers reports backoff state — a paused channel is
otherwise indistinguishable from an idle one.
Wraps systemctl list-timers with the things you actually want when checking on
the build host: per-unit enabled state, last run and whether it succeeded,
which channels have builds queued (state files), whether the release lock is
held by a live process, and any failed units. Units are discovered from
systemd/*.timer so the report cannot drift from what setup installs.
It forwards over ssh like the other host commands, so the build box's timer
state is one command away from a workstation (--local to inspect this machine).
--check warned 'rc worktree would be created' whenever the directory was
absent, implying a plain setup run would create it — but with no rc branch in
existence setup skips it, so the warning described something that would not
happen and asked for action that was not possible. It now reports the branch
state: present, would-create (branch exists), or nothing-to-do (no branch
yet — the first RC cut creates the branch and the build trigger creates the
worktree on demand). Branch detection is read-only, as --check must be.
Eligibility used package_moves_to_channel, which excludes packages built
natively in the destination — right for ongoing edge -> rc advances (a native
build must not be raced under the same filename) but wrong for the bootstrap,
whose entire purpose is rc == stable. omarchy and omarchy-settings were
therefore left out, so a machine switched to the rc channel could not install
or update the release pair until the first RC was cut. The bootstrap now
requires only destination membership; nothing is built in rc yet, so there is
no native artifact to conflict with. The dev pair stays edge-only and
fast-ring replication is unchanged.
The builder stage never declared ARG MIRROR, so the keyring [omarchy] repo
pointed at the channel-less legacy pkgs.omarchy.org/$arch path — it works
only because a stale copy of the old layout still answers there, and it would
miss a keyring rotation. Each image now pulls omarchy-keyring from its own
channel (edge/rc/stable), matching the base mirror it already selects.
update-repo and remove-package switch to the edge x86_64 image: repo-add and
repo-remove compile nothing, and using the channel image would deadlock
bootstrap-rc — the rc image can only build once the rc channel it pulls the
keyring from exists remotely.
With a repository host configured (OMARCHY_REPO_HOST / .repo-host), release,
build, sign, promote, update, clean, advance, bootstrap-rc, remove, sync, and
migrate exec on the host over ssh — same code, run where the published tree
lives, after sourcing the host credentials and a --ff-only pull. --local
forces local execution. list/push/deploy/setup never forward. The host itself
has no .repo-host, so ssh'd-in manual use is unchanged. omarchy-release's
advance now rides the same forwarding (one code path), and --host exports
OMARCHY_REPO_HOST so child bin/repo calls follow it.
This closes the gap where bootstrap-rc ran against a workstation's stale
local tree despite .repo-host being set.
The published-db marker also exists on any workstation that once ran a full
local release, so --host / OMARCHY_REPO_HOST / .repo-host are now checked
before on_repo_host everywhere (triggers, advances, doctor). README documents
the detection, the caveat, and the .repo-host tie-breaker.
Release commands now work from anywhere: on_repo_host (the published database
living in this checkout) routes build triggers, advances, and promotion to
local execution; other machines go over ssh to the configured destination.
The host setting is any ssh destination — root@<ip>, root@<hostname>, or an
~/.ssh/config alias — resolved from --host, OMARCHY_REPO_HOST, then the
one-line .repo-host file. doctor reports which mode applies, and the README
documents the format and precedence. All host connections are plain ssh.
The work/mirror clones were HTTPS end to end, so pushes went through git's
credential-helper config — which breaks the moment a stale absolute gh path
is baked into it (as gh auth setup-git once did with /usr/bin/gh). Reads stay
anonymous HTTPS; pushes now use an SSH push URL (derived from the clone URL,
overridable with OMARCHY_UPSTREAM_PUSH_URL), set idempotently on every run so
existing cached clones self-repair.
Changes reach a release branch as backports — cherry-picks with new SHAs — so
the original quattro merge commit is never an ancestor of a patch branch and
the ancestry filter let already-applied PRs through. Candidates are now also
matched against the branch's own commit messages since it left quattro:
'backport of #N' / squash '(#N)' references and 'cherry picked from commit
<sha>' trailers (which pick -x itself writes). Explicitly named PRs/commits
that are already on the branch are skipped with a note instead of re-picked.
Verified against the live v4-0-2 branch: the five backported PRs it carries
filter out; un-backported ones are still offered.
- ship: no interactive override of the untested-commit guard; the tag targets
the pinned commit the artifacts were built from (never the branch head); a
tagged-but-incomplete train is found and resumed instead of vanishing from
open-train detection; a fully shipped train reports as such
- start: a failed edge→rc advance fails the command loudly (both start and
the advance are idempotent) instead of opening a train against stale rc
- rc trigger: bootstraps the server's rc worktree on first use, so a host set
up before the rc branch existed can run its first RC build
- advance-channel: fast-ring packages are excluded from edge→rc (the stable
build replicated by parity is authoritative for rc — same filename, other
bytes); differing destination bytes abort instead of warn; a package whose
signature copy was interrupted gets its .sig restored on resume
- the release lock now also covers direct promote/update/clean/remove/sync
invocations, not just release/advance/upload-prebuilt
A release train has three human moments, each one command: start (release
branch on basecamp/omarchy + notes staging PR + edge→rc advance for
minor/major), rc (pin both PKGBUILDs to the branch head as X.Y.ZrcN on the
pkgs rc branch, trigger the rc channel build, wait for publish, optional RC
ISO), and ship (final pins → promote rc→stable → tag → pins to master → GitHub
release from the staging PR body → final ISO → website bump, each step
skip-if-done so a crashed run resumes).
Bare omarchy-release is the shepherd: it derives the train state from observed
reality (remote branches, rc-branch pins, published channel dbs, tags — no
state files) and offers the correct next step. Versions are inferred from
branch names (v4-0-2 ⇒ 4.0.2rcN ⇒ v4.0.2). ship refuses to promote a commit
no RC was cut from. pick is a multi-select over merged quattro PRs,
cherry-picking merge commits. doctor pre-flights every credential and
connection. self-test wired into CI.
bin/omarchy-pkgs stays as the pin engine, driven with its db URL pointed at
the rc channel and pins committed to the standing rc branch (rebuilt as
master + pins per cut and force-pushed; the server rc worktree follows with
reset --hard).
The rc service builds from the rc branch worktree (/root/omarchy-pkgs-rc,
created by bin/setup) but publishes into the primary checkout's channel tree
via OMARCHY_REPO_ROOT. The timer is a retry backstop: rc builds are normally
triggered immediately over SSH by the release orchestrator.
bin/repo advance --from/--to moves packages forward through the pipeline
(edge → rc → stable), driven by the source channel's database rather than the
raw directory, copying packages AND their detached signatures (fixing the old
migrate bug that left promoted packages unverifiable), refusing to rewrite any
published filename, and requiring a .sig for everything it moves. stable → rc
is allowed only as --fast-ring parity replication or the one-time
--bootstrap seed (bin/repo bootstrap-rc). 'migrate' stays as a deprecated
alias for the transition.
helpers/lock-helpers.sh adds a host-wide flock shared by bin/release,
advance-channel, and upload-prebuilt (reentrant via OMARCHY_RELEASE_LOCK_HELD)
so timers and operators serialize instead of interleaving partial publishes.
bin/release gains a stable-only step 7: replicate fast-ring artifacts to rc so
rc and stable stay in parity between release trains (skipped until rc is
bootstrapped).
- helpers/paths.sh: validate_mirror/require_valid_mirror for the edge|rc|stable
set, and REPO_ROOT (OMARCHY_REPO_ROOT override) so a secondary checkout like
the rc branch worktree publishes into the same channel tree as the primary
- validate --mirror everywhere it previously accepted any string (sync-repo,
promote-build, update-repo, clean-repo, remove-package) and widen the
edge|stable checks in build, deploy, push-build, auto-release
- build/Dockerfile: rc builds compile against rc-mirror.omarchy.org
GNU date accepts relative expressions like '2 days ago', which would let a
buggy hook fabricate a release age; the backstop now insists on an ISO 8601
timestamp before date parses it. The provider header now states, rather than
contradicts, the code's behavior for a feed with no stable releases: that is
a loud failure by design, while quarantined releases report no update.
An empty min_release_age string now maps to unparseable rather than absent,
so "min_release_age": "" fails validation instead of silently running
with a zero-second quarantine. The end-to-end fixtures extend the
checked-in pkgver (.90/.91) so the test keeps working at any future mise
version. A Tests workflow runs bin/sync-upstream self-test and
bin/omarchy-pkgs self-test on every PR in the Arch container, making the
proof machine-checked instead of author-supplied. The README package
metadata field list documents upstream and min_release_age.
The self-test now runs sync_package over the checked-in mise-bin package --
its real metadata and PKGBUILD, the full selection/validation/backstop/
rewrite/read-back path -- with only the two network fetches replaced by
mise-shaped fixtures, asserting the final PKGBUILD holds the quarantine-
cleared version, pkgrel 1, and both architecture checksums.
Review fixes: the release-row builder uses "" fallbacks instead of empty
so a malformed row cannot shift columns past the per-field checks, and
provider discovery now keys on the presence of an upstream declaration
rather than a well-formed one, with sync_package failing loudly on a
declaration it cannot use -- a malformed manifest can no longer silently
drop a package out of scheduled synchronization.