15 KiB
Plan: reliable pushes to multiple grasp servers
Status: implemented (2026-09-07), see the change summary below.
A push to a repo announced with several grasp servers sometimes reaches only
one of them. The grasp server on relay.ngit.dev rejected the git push with
remote: ERR authorisation failed: No state events in purgatory, while
gitnostr.com accepted it. The push still reported success because the app
counts an operation as pushed when at least one grasp server accepted;
individual failures only hit the log as WARN grasp push failed: .... The
failed server silently stays behind until something pushes to it again.
Background: how a grasp push is authorized
A push to a GRASP server is a two-stage transaction, not a plain git push:
- The client publishes a NIP-34 kind-30618 state event listing the refs it
is about to push (
refs/heads/main -> <sha>etc.). - The client git-pushes the objects to
https://<host>/<owner>/<repo>.git.
The grasp server does not trust the git push on its own. State events it
accepts are held in an in-memory purgatory - accepted but not served,
"until git data arrives" - and the git push is authorized only against
purgatory state events, never against the database (the database is the
current state; purgatory holds the intended future state). When a push arrives
and purgatory holds no state event for that repo id, the server rejects with
No state events in purgatory (ngit-grasp git/authorization.rs), surfaced to
the client through the git protocol as the remote: ERR authorisation failed: ... line above. Parked events that are never matched by git data are discarded
after ~30 minutes.
The relay and the git server of a grasp entry share a host: wss://gitnostr.com
and wss://relay.ngit.dev are relay URLs, and the git URL is derived with
grasp_base_url (https://<host>/<owner>/<repo>.git). The grasp learns about
state events from its own relay. A relay accept for a state event parks it in
purgatory before the grasp sends its OK, and the nostr-sdk default ack policy
(AckPolicy::all) waits for that OK - so a confirmed stage means the state is
in place to authorize the push.
Root causes in the app
| # | Problem | Consequence |
|---|---|---|
| 1 | State event is published to the whole relay pool, then git is pushed to every grasp server - no per-server staging, no confirmation that the target grasp's own relay accepted it | a push can reach a grasp whose purgatory is still empty |
| 2 | No retry for that denial class | a self-resolving race permanently leaves one server behind |
| 3 | push_to_grasp_servers returns Ok when >= 1 server accepts; failures only log::warn! |
UI shows success; the failed grasp silently out of sync |
| 4 | When a grasp's relay is down, the client still fires a doomed git push | confusing git-level ERR instead of a clear "state not accepted by relay" |
Relevant code (crates/signed_state/src/backend.rs):
push_repo_from(L752-834): broadcast the state event to the whole pool, thenpush_to_grasp_servers.push_to_grasp_servers(L1536-1570): sequential git pushes,Okwhen at least one server accepted,warn!per failure.broadcast_event(L1336-1350): whole-poolclient.send_event, errors only when no relay accepted.send/publish_task(L1239-1292): sign, broadcast, store locally (the client saves accepted events), emitBackendEvent::Published.create_repository(L554-575) andpublish_local_repo(L692-708): directpush_to_grasp_serverscalls; retract announcement+state when zero servers accept.
The fix mirrors ngit's own client (its state_transaction.rs stages the state
event on the grasp relays, gates which servers are pushed, and fans the state
out to other relays only after a git server accepted) and adds retries, because
the app must also cope with a denial after a relay ack.
Target flow: staged push with retries
sequenceDiagram
participant App as signed app
participant R1 as gitnostr.com relay+grasp
participant R2 as relay.ngit.dev relay+grasp
participant O as other relays
Note over App: build state event S0 from local refs
App->>R1: stage S0 to grasp relay only (.to(url))
App->>R2: stage S0 to grasp relay only (.to(url))
R1-->>App: OK (parked in purgatory) -> eligible
R2-->>App: OK (parked in purgatory) -> eligible
App->>R1: git push
R1-->>App: authorized by S0 in purgatory
App->>R2: git push
R2-->>App: ERR ... No state events in purgatory (transient)
Note over App: retry: stage a fresh S1 to R2, wait ~1s, git push again
App->>R2: stage fresh S1 (new event id)
App->>R2: git push
R2-->>App: authorized
Note over App: >= 1 server ok -> fan out state to the pool
App->>O: broadcast state (only after a git server holds it)
Retries stage a fresh state event (new created_at -> new id) rather than
resending the same one: a grasp relay treats a same-id event as a duplicate and
will not re-run its policy, so a lost purgatory entry (e.g. a server restart
without its state file) cannot be re-parked by a resend. A fresh id re-runs
validation and re-parks. Superseded parked events expire server-side after 30
minutes, so the residue is bounded.
Changes
Implemented:
crates/signed_state/src/backend.rs:push_to_grasp_serversreplaced by the staged orchestrationpush_staged_to_graspsplusstage_event_on_relay,sign_state_event,is_transient_grasp_denial, and theGraspServerResult/PushOutcomeresult types.create_repository,publish_local_repoandpush_repo_from(repo push and checkout push) all push through it and fan the state out to the relay pool only after at least one git server accepted.crates/signed_state/src/repo.rs: push outcomes surface partial failures asRepoStore::last_push_warning.crates/workspace/src/views/repo_detail/mod.rs: a warning banner with a Republish action shows when a push did not reach every grasp server.- Unit tests for the denial classifier and the outcome reporting.
Decisions taken while implementing:
- State events are no longer broadcast before any git data exists. They are staged per grasp relay, and only fanned out to the relay pool after a git server accepted. A total push failure therefore leaves nothing served to retract except the announcement (create/publish flows retract that, as before). Staged-but-unmatched state events expire in the grasp's purgatory after ~30 minutes.
- Servers whose relay rejects the state event are skipped, not pushed (the eligibility gate): a doomed git push would only produce the same denial.
- Stale-ref races are retried and verified. The grasp server runs its own
background sync that aligns repository refs to parked state events as soon
as the git objects are present - including objects an earlier denied attempt
of this same push already uploaded.
git receive-packthen rejects the ref update against its stale advertisement withcannot lock ref ... is at ... but expected .../incorrect old value provided. Such rejections are retried against a fresh advertisement and probed for convergence: when the sync already aligned the refs to this push's target (git ls-remotematches,signed_git::remote_has_refs), the server is counted as accepted even though the push's compare-and-swap never returnedOk- because the data is already there. The sync can land a moment after the retry window, so the probe is what makes these pushes succeed without a manual republish. - Create/publish partial failures (some servers accepted, some not) stay
Okand are logged, matching the pre-existing contract; the repo-push paths additionally setlast_push_warningso the failure is visible and one-click republishable. - The push warning lives in a dedicated
RepoStore::last_push_warningso it never collides with the PR-flowlast_warningon the shared store.
Staged orchestration (crates/signed_state/src/backend.rs)
-
Per-relay publish helper (
stage_event_on_relay). Publishes only to one relay and verifies its acceptance, using the pinned nostr-sdk 0.45 API (send_event_tois deprecated at 0.45):// client.send_event(event).to([relay.clone()]).await // -> Ok only when the relay is in the output's success mapThe relay is added and connected first (reusing the
add_relaypattern); a failed connect is a staged failure with a clear reason. -
Transient-denial classifier (
is_transient_grasp_denial, pure, unit-tested). Two retry families:- Purgatory denials (ngit-grasp
git/authorization.rs):No state events in purgatory(the originally reported failure)no matching state event found in purgatoryin purgatory ... doesn't match pushnone from authorized publishersNo repository announcement found(new repos whose announcement is still propagating)
- Stale-ref races (git receive-pack compare-and-swap against the
advertised value, when the grasp's background sync moved the ref):
cannot lock refincorrect old value provided
Stale-ref race rejections are additionally probed for convergence with
git ls-remote; a server that already advertises the pushed refs counts as accepted (remote_has_refsinsigned_git).Everything else - HTTP auth rejection, not a maintainer, a genuine non-fast-forward divergence, network failure - is permanent for that attempt. Only the transient classes are retried.
- Purgatory denials (ngit-grasp
-
Orchestration (
push_staged_to_grasps, replacespush_to_grasp_servers):pub struct GraspServerResult { relay: RelayUrl, git_url: String, reason: Option<String> } pub struct PushOutcome { servers: Vec<GraspServerResult>, state_event: Option<Event> } async fn push_staged_to_grasps( client: &Client, signer: &UniversalSigner, repo_id: &str, refs: &[(String, String)], head: Option<&str>, path: &Path, owner: &str, servers: &[RelayUrl], push: fn(&Path, &str, &str, &str) -> Result<(), Error>, ) -> PushOutcomePer server, in
serversorder:- Stage: publish the state event to that relay (one retry for a connect blip). Not accepted => record the reason and skip the git push (the eligibility gate - no doomed push).
- Push: git push; on a transient denial, stage a fresh state event (see below), back off ~1s, and retry. Budget: three attempts total.
- Permanent git error => record it and stop for that server.
Retries stage a fresh event (new
created_at-> new id) rather than resending the same one: a grasp relay treats a same-id event as a duplicate and will not re-run its policy, so a lost purgatory entry (e.g. a server restart without its state file) cannot be re-parked by a resend. When a retry would fall in the same wall-clock second,sign_state_eventnudgescreated_atforward by one second so the event id differs. Superseded parked events expire server-side after 30 minutes. -
Reorder
push_repo_from: replace the pre-push whole-poolsend(state)with staging + push throughpush_staged_to_grasps, then, after at least one server accepted, fan the state out to the pool (broadcast_event) and emitBackendEvent::Publishedonce. Moving the fan-out after the first git success also closes today's leak where state is broadcast before any server holds the objects (ngit made the same change). -
Other callers updated (
create_repository,publish_local_repo): same orchestration. Their announcement broadcast stays first - grasps park brand-new announcements in announcement purgatory until the first git data, so there is no leak - and "zero servers accepted => retract announcement" is preserved. An empty repository (no refs) is announced without staging or pushing anything.
Outcome plumbing (crates/signed_state/src/repo.rs)
-
RepoStore::push_repository/push_checkout(L1075-1150) gainlast_push_warning: Option<String>populated from
PushOutcome::partial_warning()(a one-line summary naming the rejected servers and reasons).last_error(zero-success hard failure) and the PR-flowlast_warningare unchanged.
UI (crates/workspace/src/views/repo_detail/mod.rs)
-
When
store.last_push_warningis set,render_push_warning_bannershows a warning banner above the header:Pushed to 1 of 2 grasp servers: wss://relay.ngit.dev: remote: ERR authorisation failed: No state events in purgatory fatal: ... Republish to sync.
-
The banner's Republish action reuses
push_repository(an idempotent re-push of every announced server - the simplest catch-up for any drifted server); a dismiss control clears the warning. -
Zero-success failures continue through the existing error banner (
self.error/store.last_error, rendered inrender).
Tests
- Unit (backend tests module): the classifier (both retry families - the
reported purgatory stderr line and the reported
cannot lock refrace are transient; HTTP 403, not-a-maintainer and a non-fast-forward divergence are not), the outcome reporting (partial warnings are single-line and name the failing servers, an all-ok outcome has no warning) andflatten_whitespace. Allcargo test -p signed_statetests pass. - The staged orchestration itself is exercised against live grasps by manual validation; the pure decision helpers are unit-tested.
Manual validation
- Create and push a repo with both
gitnostr.comandrelay.ngit.dev; after each push,git ls-remote https://<host>/<npub>/<repo>.gitagainst both hosts to confirm convergence; repeat over several commits to shake out the race. Then repeat withrelay.ngit.devtemporarily unreachable to exercise the banner and the Retry action.
Behavior decisions
- At least one server accepted => the operation succeeds (existing contract), but partial failure is explicit and retryable (warning banner + Republish).
- Announcement publishing stays first; state fan-out moves after the first git success (fixes the state-before-data leak).
- No wire-protocol changes; error text stays user-readable in the banner.
Open questions (status)
- Auto-retry after a partial success: not implemented - retries happen inside the push (up to 3 attempts per server for transient denials), and a residual failure stays visible with a manual Republish. A delayed background re-push of failed servers could be added later.
- Post-push verification (
git ls-remoteper grasp after a successful push to confirm the refs are advertised): not implemented; could catch ack-but-not-promoted cases at the cost of one extra round-trip per server. - Scope: the create, publish-local and repo/checkout push paths were all converted in one change.