Ver Fonte

openspec/darkirc-delivery-status: Revise with durable FIFO for messages

x há 6 dias atrás
pai
commit
9d3db43c61

+ 145 - 94
openspec/changes/darkirc-delivery-status/design.md

@@ -22,15 +22,19 @@ Constraints that shape the design:
 - RLN is currently inactive on the network (empty blobs), but the
   enabled path must stay correct: recreations must reserve a fresh
   message slot or risk self-slashing.
+- This plan is not crypto sign-off. Stop for human review before any
+  implementation change touching RLN, crypto, circuits, or canonical
+  serialization; the security-agent verdict does not replace that review.
 
 ## Goals / Non-Goals
 
 **Goals:**
 
 - Outcome replies for `EventPut` with zero policy in the generic layer.
-- darkirc: durable outbound tracking, rebroadcast-first, bounded
-  recreate only on explicit evidence; survives restart; safe with RLN
-  on or off.
+- darkirc: a durable global FIFO, oldest logical message first, retained
+  until acknowledged; rebroadcast unchanged and recreate only after
+  expiry. Preserve order across restart and use the existing RLN-safe
+  send path when rate limiting is enabled.
 - Stable wire format (`u8` discriminants), no new panics on untrusted
   input.
 
@@ -38,7 +42,7 @@ Constraints that shape the design:
 
 - Pull-based possession verification (`EventReq` challenges) — dropped;
   status replies plus ancestry references are the only signals, treated
-  as hints whose worst-case lie is bounded by attempt caps.
+  as unverified hints, not proof of delivery to a recipient.
 - `bin/app` integration — out of scope entirely (no UI, subscriptions,
   or send-path changes there). darkirc is the reference implementation;
   the app consumes the same generic `receipt_pub` pipe in a follow-up
@@ -46,8 +50,13 @@ Constraints that shape the design:
 - `Privmsg` payload changes (`uid`/dedup field) — deferred (see
   Alternatives Considered). No content serialization changes in this
   change.
-- Status for `StaticPut`, IRC-surface receipt display, read-by-recipient
-  semantics, changes to strike/flood policing or the relay path.
+- Status for `StaticPut`, IRC-surface receipt/status display,
+  read-by-recipient semantics, changes to strike/flood policing or the
+  relay path. Enqueue rejection and blocked-send error notices to
+  connected clients remain in scope; a delivery-status UI does not.
+- Recipient display ordering and wire-level deduplication. FIFO governs
+  local submission only; propagation and delayed generations can still
+  arrive out of order or render twice.
 
 ## Decisions
 
@@ -98,7 +107,8 @@ versioned together.
 | `!is_synced()` skip | `Nack { NotSynced }` |
 | `main_tree` duplicate | `Has { inserted: false }` |
 | `timestamp < genesis_ts` | check retained slots first: present → `Has { inserted: false }`; absent → `Nack { TooOld }` |
-| `validate_new` / structural / RLN / parent-fetch failure | `Nack { Invalid }` |
+| established structural / proof / parent invalidity | `Nack { Invalid }` |
+| transient parent-fetch failure or internal processing failure | `Nack { Busy }` |
 | internal insert error after verification | `Nack { Busy }` |
 | `insert_verified_signal` success | `Has { inserted: true }` |
 
@@ -127,69 +137,100 @@ chat, taud would want its own rules, the app embeds evgr in-process and
 needs the events for UI, not policy). Keeping the generic layer stateless
 also keeps it out of the security-sensitive blast radius.
 
-### D4: darkirc outbound table
-
-New kvdb tree `darkirc_outbound`, key = event id, value (serial):
-`{ event_id, plaintext Privmsg fields, created_ts, state:
-Pending|Delivered|Failed, attempts: u16, last_broadcast_ts,
-superseded_by: Option<event_id> }`. Written in `publish_events` before
-`p2p.broadcast`. Plaintext-at-rest is inside the existing local-wallet
-trust boundary (same kvdb that holds RLN identity secrets).
-
-The `superseded_by` chain is the sender-side correlation handle: when a
-record is recreated, the new event id links back to the old, so clients
-can map statuses for any generation onto one logical message without
-any wire-level dedup field.
-
-### D5: Delivery monitor (darkirc)
-
-One task on `IrcServer` (started alongside the client loop):
-
-- Subscribes `receipt_pub`; per (event_id, channel-address) aggregates
-  replies; any `Has` (live or historical) closes the record as
-  Delivered.
-- Ancestry references count as positive evidence: when a new foreign
-  event is observed whose ancestry includes a tracked outbound event
-  (local walk via the existing `get_ancestors` machinery), the record
-  closes as Delivered. Zero wire cost; works among mixed-version peers.
-- Periodic sweep (30 s) over `Pending` records, gated by
+### D4: Durable global outbound FIFO
+
+Use the `darkirc_outbound` kvdb tree for ordered logical-message records,
+not an unordered set of independently retried event ids. Persist a
+stable queue position, plaintext Privmsg fields, enqueue time, optional
+active event/blob, supersession metadata, and retry/parked state.
+Concurrent enqueues must obtain one durable order; restart preserves it.
+Accepting a send means its queue record is durable, including while
+offline or initially unsynced. On enqueue/storage-limit failure, report
+rejection to the connected client instead of claiming the send is queued.
+
+All initial sends go through this queue. Only the head may materialize
+an event or expose it to the network, including via local DAG insertion
+and relay/sync. Later messages remain plaintext queue entries until every
+predecessor is acknowledged. The sender's queue order is not a guarantee
+of recipient display order.
+
+Keep the head's exact active event and blob independently of DAG
+retention for unchanged retries after pruning. Keep generation ids and
+old-to-new `superseded_by` links until the logical message is acknowledged.
+A replacement keeps the same queue position; it is not appended as a
+new logical message. Persist its event/blob, active-generation pointer,
+and supersession link as one crash-recoverable transition before network
+exposure. Recovery must never create two active generations or bypass
+an unacknowledged head.
+
+Bound queue storage, including active event/blob and generation metadata.
+At capacity, reject new enqueues explicitly; if a replacement cannot be
+retained, park the head rather than evicting unacknowledged state.
+Acknowledged entries are removable. There is no lifetime or recreation
+attempt cap: storage pressure can block progress, not justify dropping
+the head. Plaintext-at-rest adds sensitive message history within the
+local kvdb trust boundary. Never log or publish that stored plaintext;
+rollback does not erase it.
+
+### D5: ACK-driven head-only worker
+
+One task on `IrcServer` owns queue advancement and retry scheduling:
+
+- Subscribe to `receipt_pub`; filter to tracked generation ids and use
+  ephemeral connection identity, never peer address, for any per-channel
+  aggregation. Do not persist connection identity or log peer addresses.
+- An ACK is either `Has` (either inserted value) for any generation of
+  the head, or an accepted foreign event whose ancestry includes such
+  a generation. Ancestry is local computation, not an additional query.
+  Locally generated events alone are not foreign delivery evidence.
+- ACK processing durably dequeues the logical message exactly once,
+  cancels pending retries/recreation, and permits the next head to run.
+  Duplicate/late statuses cannot dequeue another message. Serialize
+  ACK, materialization, and supersession transitions so stale work
+  cannot resurrect an acknowledged head. Already-broadcast generations
+  cannot be recalled.
+- A periodic wakeup (30 s) services only the head, with sends gated by
   `is_synced() && connection_count >= K` (K = 2, both session directions
-  counted via the existing session APIs).
-- No evidence → rebroadcast the original `EventPut` unchanged (event and
-  blob refetched from the local DAG + `dag_blob_fetch`), backoff
-  doubling from 30 s, at most `R_MAX` (= 5) rounds per window, then
-  keep rebroadcasting at the slow rate.
-
-**Recreate only on explicit evidence (the complete trigger set):**
-
-1. *Window closed, no holder:* the rotation window for the record's slot
-   has closed (known locally from the rotation schedule, or indicated by
-   `TooOld` nacks) and no peer answered `Has` in the rebroadcast round →
-   recreate. This covers: offline across rotation (the canonical case),
-   sent into the void while "synced" with zero peers, and clock-skew
-   `TooOld` while we believed the window open.
-2. *Explicit rejection while open:* every observed reply is a nack and
-   at least one is `Invalid` → recreate (bounded by `A_MAX`; reasons
-   surface to the user).
-
-Never recreates on: silence (open window — this is what prevents
-mass-duplication during mixed-version rollout), `NotSynced`, `Busy`,
-or any `Has`. Silence after window close, with no status reply ever
-received from any status-capable peer, still recreates — local rotation
-knowledge is explicit evidence, and the alternative (gray forever,
-possibly lost) is worse than a rare duplicate for old-version holders.
-
-### D6: Recreate procedure
-
-Reuse the existing send path internals: stored plaintext →
-`try_encrypt` (fresh nonce — never reuse the old ciphertext) →
-`Event::new` (fresh tips/timestamp) → RLN branch:
-`reserve_rln_message_id` → `BudgetExhausted` parks the record until
-epoch rollover; `MissingIdentity` closes it as Failed → `create_signal`
-→ `insert_signal_with_blob` → broadcast → write the new record with
-`superseded_by` pointing at the old; `attempts` is shared across the
-chain and caps at `A_MAX` (= 2), then Failed + client notice (IRC error
-reply to the connected client).
+  counted via the existing session APIs). ACK processing is not gated
+  on connectivity or the retry timer.
+- For a materialized head without ACK, rebroadcast its exact active
+  `EventPut`. Backoff doubles from 30 s for `R_MAX` (= 5) rounds, then
+  continues at the capped slow rate until ACK. There is no total-round
+  limit. Negative replies must not bypass the retry schedule.
+- Recreate only after the active event's rotation window has closed
+  (local rotation knowledge or `TooOld`) and a rebroadcast round has
+  completed without ACK for any generation. This includes locally
+  known expiry with no replies ever received. A retained-slot `Has`
+  cancels recreation. Silence, `Invalid`, `NotSynced`, and `Busy` alone
+  never trigger recreation or advance the queue.
+
+The global FIFO intentionally accepts head-of-line blocking: an
+unacknowledged message blocks all later sends, even to other conversations.
+There is no automatic fail-and-skip policy.
+
+### D6: Head materialization, recreation, and parking
+
+Reuse the existing send-path internals only for an eligible head:
+stored plaintext → `try_encrypt` (fresh nonce for each new generation)
+→ `Event::new` (fresh tips/timestamp) → existing RLN reservation/proof
+path when enabled. Durably persist the prepared event/blob and any
+supersession transition before local insertion or broadcast makes it
+network-visible. Then use the existing insertion/broadcast path. Recover
+a committed generation by retrying that generation, not by regenerating
+it. Unchanged rebroadcast never re-encrypts or reserves a new RLN slot.
+
+Recreation stays at the head and uses a fresh nonce and, when enabled,
+a fresh RLN message slot. `BudgetExhausted` parks until epoch rollover;
+`MissingIdentity` parks until an identity is available. Processing and
+storage failures likewise preserve the head and prevent later sends
+from bypassing it. Enqueue/blocked-send error notices may inform a
+connected client but do not imply dequeue or terminal failure. Retry
+these conditions with bounded rate; never send an unproven replacement
+or turn an error into a queue-advancement signal.
+
+The enabled RLN reservation/recovery path requires human review before
+implementation. This design does not authorize slot reuse, crypto
+changes, or altered RLN semantics.
 
 ## Alternatives Considered
 
@@ -210,26 +251,32 @@ verification with extra steps; (iv) no rejection reasons.
 **`Privmsg.uid` dedup field.** Deferred. It protects against duplicate
 renders when a recreate fires while someone already holds the original.
 With D2's historical check and D5's explicit-evidence rule, that overlap
-shrinks to: mixed-version rollout windows (old peers can't answer
-`Has`), adversarial fake-nack griefing, and recreate-of-recreate — all
-rare and cosmetic. Asymmetry decides: adding a payload field later is an
+is reduced but not eliminated: mixed-version rollout windows (old peers
+can't answer `Has`), unreachable holders, lost replies, adversarial
+fake-nack griefing, and recreate-of-recreate can produce duplicates.
+Their frequency is not established. Asymmetry decides: adding a payload field later is an
 additive `Privmsg.version` bump; removing one after shipping is a wire
 break. Revisit on field evidence of annoying duplicates.
 
-**Hold-at-send (don't broadcast with zero connections).** Candidate
-follow-up: gate `publish_events` on connection count so doomed events
-aren't created at all. Not required for correctness (D5 handles them);
-deferred as a small polish item.
+**Independent retries and hold-at-send.** Independent record sweeps were
+rejected because later messages could overtake an unacknowledged send.
+Holding messages before event creation is now part of the FIFO design:
+only an eligible head is materialized. Per-conversation queues and
+automatic fail-and-skip are not part of this change.
 
 ## Risks / Trade-offs
 
 - [Fake statuses: a peer can lie `Has` (suppresses recreate → silent
-  loss) or spam `Nack` (forces recreates → budget burn + injected
-  duplicates)] → bounded: statuses only count for ids in our own
-  outbound table (blake3 ids are unguessable to non-recipients),
-  attempts capped at `A_MAX`, and a peer that has the event relays it
-  anyway. Total eclipse defeats this — accepted (an eclipsed node has
-  larger problems).
+  loss) or lie `TooOld` (forces recreates → budget burn + duplicates)]
+  → only tracked ids affect state; backoff, storage bounds, and enabled
+  RLN budget bound resource use/rate, not total lifetime attempts. Even
+  one malicious connected peer that knows an id can falsely acknowledge
+  it without forwarding. ACK means observed network-delivery evidence,
+  not verified delivery to a recipient. This limitation is accepted.
+- [Head-of-line blocking and indefinite retention] → intentional
+  ACK-only dequeue; no liveness guarantee without ACK. Cap storage and
+  reject new enqueues explicitly rather than silently evicting messages.
+  Storage exhaustion may also park recreation; no later send bypasses it.
 - [Reply amplification: one reply per relay edge] → ~70 bytes on the
   wire against events measured in hundreds of bytes to kilobytes, on
   RLN-rate-limited volume; the flood window is untouched (replies are
@@ -241,30 +288,34 @@ deferred as a small polish item.
   replies until a capability flag exists. First implementation task
   resolves this.
 - [Duplicate renders without uid: mixed-version transition, adversarial
-  nacks, reply loss] → accepted, rare, cosmetic; documented in the
+  nacks, reply loss] → accepted possibility, frequency unknown; documented in the
   proposal's non-goals; additive fix available later if needed.
-- [Clock skew: slightly-future local timestamps can earn spurious
-  `TooOld` nacks] → worst case is a harmless recreate.
-- [RLN interplay: recreate with a stale slot would self-slash] →
-  impossible by construction: recreations go through
-  `reserve_rln_message_id`, parking on exhaustion, never reusing a
-  reserved slot.
+- [Clock skew: disagreement about rotation boundaries can earn spurious
+  `TooOld` nacks] → recreation can duplicate a message and consume
+  budget; bounded in rate and storage, not in total attempts or harm.
+- [RLN interplay: recreation with a reused slot risks self-slashing] →
+  use `reserve_rln_message_id`, park on exhaustion, never reuse a
+  reserved slot; require human review and enabled-path tests, including
+  restart handling, before relying on this guarantee.
 - [Plaintext messages persisted in kvdb] → same trust boundary as the
   existing wallet/RLN secrets; tree is local-only.
 
 ## Migration Plan
 
 Additive wire message first (`src/event_graph`), verified against a
-mixed-version two-node test; then darkirc table+monitor. No content
+mixed-version two-node test; then darkirc queue+worker. No content
 serialization changes, so no coordination with `darkirc-mod` is
 required. Rollback: the message and records are inert for old code;
-reverting leaves a harmless `darkirc_outbound` tree. App integration is
+reverting leaves a `darkirc_outbound` tree containing sensitive
+plaintext and pending state, not a harmless cache. Handle retained data
+under the same local storage protections. App integration is
 a follow-up change that reuses the same `receipt_pub` pipe and copies
-the darkirc monitor logic.
+the darkirc queue policy.
 
 ## Open Questions
 
-- Exact values of K, `R_MAX`, `A_MAX`, sweep interval — tuning consts,
-  safe to adjust after rollout.
+- Exact values of K, `R_MAX`, wakeup interval, and queue storage limit —
+  tuning parameters; they must not change head-only service, ACK-only
+  dequeue, bounded resource use, or the absence of a lifetime retry cap.
 - Whether taud later reuses the same policy for task events — deferred,
   out of scope.

+ 56 - 30
openspec/changes/darkirc-delivery-status/proposal.md

@@ -29,46 +29,65 @@ network.
   `event_pub`) that republishes every inbound `EventPutStatus` unfiltered.
   The event layer is a dumb pipe; delivery policy lives in applications.
 - darkirc (`bin/darkirc`):
-  - Persistent outbound table (kvdb tree, written before broadcast):
-    event id, plaintext privmsg, state, attempts, `superseded_by` link.
-  - Receipt aggregation keyed by (event_id, peer channel); a foreign
-    event whose DAG ancestry includes our outbound event also counts as
-    positive delivery evidence (free secondary signal, no wire change).
-  - Delivery monitor — rebroadcast-first when reconnected and synced.
-    Silence alone never triggers recreation while the rotation window
-    is open; recreate only when the window is closed (locally or via
-    `TooOld` nacks) with no holder answering, or when every observed
-    reply is an explicit rejection. `NotSynced`/`Busy`/silence mean
-    retry, not evidence.
+  - Durable global FIFO of logical messages, preserving plaintext and
+    enqueue order across restart. Only the oldest unacknowledged message
+    creates/exposes an event when synced and sufficiently connected;
+    later messages cannot bypass it. Every generation is retained before
+    network exposure, including its unchanged event/blob independently
+    of DAG pruning and old-to-new `superseded_by` links.
+  - ACK = `Has` for any head generation, or an accepted foreign event
+    whose ancestry includes that generation. ACK durably dequeues the
+    logical message exactly once and permits the next send. Receipt
+    aggregation uses ephemeral peer-channel identity, not peer addresses.
+  - Head-only worker retries unchanged with capped backoff until ACK,
+    without a lifetime or attempt cap. Recreate only after expiry
+    (local rotation knowledge or `TooOld`) and a rebroadcast round
+    without ACK. Silence, `Invalid`, `NotSynced`, and `Busy` alone
+    neither recreate nor dequeue. Local expiry permits recreation even
+    when no replies have ever arrived.
+  - Bound storage, including retained generation metadata. Reject new
+    enqueues explicitly when full; park the head if recreation cannot
+    fit. Never evict unacknowledged messages to make progress.
 - Recreate = new event from stored plaintext (fresh parents/timestamp,
-  fresh saltbox nonce), new RLN slot when RLN is enabled
-  (`BudgetExhausted` parks the attempt until the next epoch), bounded
-  attempt count, `superseded_by` chain for local correlation.
+  fresh saltbox nonce), new RLN slot when RLN is enabled, same queue
+  position, and `superseded_by` chain for local correlation. Exhausted
+  budget, missing identity, and processing/storage failures park the
+  head rather than discard it or send an unproven replacement.
 
 Non-goals: read receipts stored in the event graph (pollutes the DAG,
-burns RLN budget, leaks linkability); IRC-surface receipt display;
+burns RLN budget, leaks linkability); IRC-surface receipt/status display
+(enqueue rejection and blocked-send error notices to connected clients
+remain in scope);
 `bin/app` integration (delivery-state UI, subscriptions, or any other
 app-side changes — darkirc is the reference implementation here and the
 app consumes the same generic pipe in a follow-up change); pull-based
 possession verification via `EventReq`; `StaticPut` status (nickserv
 already has a deferred-broadcast queue); recipient-identifying receipts
-(nodes are anonymous; this is delivery-to-network evidence only). A
+(nodes are anonymous; this is delivery-to-network evidence only);
+recipient display ordering (FIFO governs local submission, not network
+propagation); automatic fail-and-skip or per-conversation queues. A
 `Privmsg` dedup field (`uid`) was considered and deliberately deferred:
-with the historical-slot check and the silence-never-recreates rule,
-recreation almost never overlaps with "someone already rendered the
-original", and adding such a field later is an additive version bump
-while removing it after shipping would be a wire break. Known accepted
-cost: rare duplicate renders during mixed-version rollout windows and
-under adversarial fake-nack griefing.
+the historical-slot check and the no-recreation-on-silence rule while
+the window is open reduce, but do not eliminate, overlap with "someone
+already rendered the original". Adding such a field later is an
+additive version bump while removing it after shipping would be a wire
+break. Accepted cost: duplicate renders during mixed-version rollout,
+with unreachable holders or lost replies, and under adversarial
+fake-nack griefing; their frequency is not established. Statuses are
+unverified evidence: even one peer with knowledge of an event id can
+lie `Has` and suppress recovery without forwarding the event.
+The global FIFO intentionally accepts indefinite head-of-line blocking,
+including across conversations. Backoff and storage bounds constrain
+resource use, not total lifetime retries or successful delivery.
 
 ## Capabilities
 
 ### New Capabilities
 
 - `msg-delivery-status`: outcome reporting for `EventPut` broadcast,
-  delivery-state exposure to applications, and the sender-side
-  rebroadcast/recreate policy implemented by darkirc as the reference
-  consumer.
+  delivery-state exposure to applications, and a durable ACK-driven FIFO
+  with head-only rebroadcast/recreation implemented by darkirc as the
+  reference consumer.
 
 ### Modified Capabilities
 
@@ -80,13 +99,20 @@ under adversarial fake-nack griefing.
   status emission, receipt publisher. Protocol-handler surface: must stay
   panic-free on untrusted input, keep flood/strike policing unchanged.
 - `bin/darkirc` (client.rs, server.rs, lib.rs, crypto/rln.rs interplay) —
-  outbound table, monitor task, recreate path.
+  outbound FIFO, head-only worker, enqueue/error path, recreate path.
 - `bin/tau/taud` and `bin/app` — untouched consumers; they may adopt the
   same `receipt_pub` pipe in follow-up changes.
 - No `Privmsg` or other consensus/content serialization changes in this
   change.
-- Wire compatibility: peers without `EventPutStatus` support simply never
-  reply; senders treat silence as "retry later, never a verdict", so
-  correctness does not depend on reply availability.
+- Wire compatibility is conditional on the first task proving unknown
+  message tolerance. If old channels reject the new message type, stop
+  and revisit capability gating before implementation continues. Old
+  peers do not reply; silence alone is not a verdict, but local expiry
+  can justify recreation without replies and may cause duplicates.
+- Persisted plaintext is sensitive local data, including after rollback.
+- Stop for human review before implementation changes touching RLN,
+  crypto, circuits, or canonical serialization. This plan does not
+  authorize changes to those protected areas.
 - Per repo policy, `event_graph` protocol changes require
-  `@anon-security-review` before apply/archive.
+  `@anon-security-review` before marking ready to apply/archive; this
+  triage does not replace CI or human review.

+ 178 - 97
openspec/changes/darkirc-delivery-status/specs/msg-delivery-status/spec.md

@@ -3,11 +3,12 @@
 ## Purpose
 
 Gives senders of event-graph broadcast messages evidence about whether
-peers accepted them, and defines the sender-side rebroadcast/recreate
-policy that prevents messages written while offline from being silently
-lost to DAG rotation. Covers the reply wire format, delivery-state
-exposure to applications, and darkirc's outbound tracking, rebroadcast,
-and recreate flow as the reference implementation.
+peers accepted them, and defines a durable, oldest-first broadcast queue
+that retains messages until acknowledged rather than silently losing
+them to DAG rotation. Covers the reply wire format, delivery-state
+exposure to applications, and darkirc's head-only retry/recreate flow as
+the reference implementation. ACKs are unverified network-delivery
+evidence, not recipient receipts or a guarantee of recipient display order.
 
 ## ADDED Requirements
 
@@ -29,8 +30,8 @@ A node that receives an `EventPut` SHALL send the sender an
   initial DAG sync and skipped the event
 - `Nack { reason: Invalid }` when the event fails structural or
   proof validation
-- `Nack { reason: Busy }` when an internal condition prevented
-  processing
+- `Nack { reason: Busy }` when a transient parent-fetch failure or an
+  internal condition prevented processing, without established invalidity
 
 Replies apply to relayed events equally: the receiver cannot and does not
 distinguish originator from relay.
@@ -68,6 +69,12 @@ distinguish originator from relay.
   `EventPut`
 - **THEN** it replies `Nack { reason: NotSynced }`
 
+#### Scenario: Transient processing failure is not invalidity
+
+- **WHEN** parent retrieval fails transiently without establishing that
+  the event is invalid
+- **THEN** the receiver replies `Nack { reason: Busy }`, not `Invalid`
+
 #### Scenario: Malicious input keeps existing policing
 
 - **WHEN** a peer receives an `EventPut` that trips existing
@@ -119,125 +126,199 @@ or make retry decisions; that policy belongs to applications.
 - **THEN** the subscription still delivers it; consuming it is the
   application's choice
 
-### Requirement: darkirc persists outbound records before broadcast
+### Requirement: Durable oldest-first broadcast queue
+
+darkirc SHALL durably enqueue each accepted outgoing chat message in one
+global FIFO, preserving its plaintext rebuild fields and enqueue order
+across restart. Only the oldest unacknowledged logical message, the head,
+SHALL be eligible for event creation or network exposure. Later messages
+SHALL remain queued without publishing events, including through DAG
+relay/sync. The head SHALL be materialized and sent only when darkirc is
+DAG-synced and has the configured minimum number of peer connections.
+FIFO guarantees local send sequencing, not recipient display ordering.
+
+#### Scenario: Later messages cannot overtake the head
+
+- **WHEN** messages A and B are accepted in that order and A is not ACK'd
+- **THEN** B remains queued without an event being created or exposed to
+  peers, even while A is disconnected, retrying, or parked
+
+#### Scenario: Offline enqueue order survives restart
+
+- **WHEN** messages are accepted while offline or initially unsynced and
+  darkirc restarts before sending them
+- **THEN** their contents and queue order survive, and only the oldest
+  becomes eligible when sync/connectivity requirements are met
+
+### Requirement: Recoverable head generations before network exposure
+
+Before exposing a head generation to the network, darkirc SHALL durably
+retain its exact event and blob. Unchanged retries SHALL remain possible
+after local DAG pruning. Replacements SHALL keep the logical message's
+queue position and preserve old-to-new supersession links and generation
+ids until ACK. Committing a replacement SHALL be crash-recoverable, with
+only one active generation; restart SHALL NOT bypass the head or generate
+a second replacement for an already committed transition.
+
+#### Scenario: Pruning does not prevent unchanged rebroadcast
+
+- **WHEN** a head event is pruned from all local DAG slots before ACK and
+  darkirc restarts
+- **THEN** its exact event and blob remain available for rebroadcast
+
+#### Scenario: Replacement persistence fails
+
+- **WHEN** a replacement cannot be durably retained
+- **THEN** it is not exposed to peers and the existing head remains queued
+
+#### Scenario: Crash after replacement commit
+
+- **WHEN** darkirc restarts after committing a replacement but before
+  broadcasting it
+- **THEN** it resumes that generation at the same queue position, with
+  the old generation's supersession link intact
+
+### Requirement: ACK-only queue advancement
+
+A `Has` outcome with either inserted value for any generation of the head
+SHALL ACK its logical message. An accepted foreign event whose ancestry
+includes any such generation SHALL also ACK it, using local computation
+over received events without additional wire queries. Locally generated
+events alone SHALL NOT constitute foreign delivery evidence.
+
+ACK SHALL durably dequeue the logical message exactly once, cancel its
+pending retries/recreation, and permit service of the next head. ACK
+processing SHALL NOT depend on retry timers or the current connection
+count. Duplicate, unrelated, or late statuses SHALL NOT remove another
+message. Concurrent recreation work SHALL NOT resurrect an ACK'd entry.
+An ACK is unverified network evidence, not proof of forwarding or recipient
+receipt; already-broadcast generations cannot be recalled.
+
+#### Scenario: Either Has variant advances the queue
+
+- **WHEN** the head receives `Has { inserted: true }` or
+  `Has { inserted: false }` for one of its generations
+- **THEN** it is durably dequeued and the next message becomes eligible
+
+#### Scenario: Foreign descendant acknowledges the head
+
+- **WHEN** an accepted foreign event's ancestry contains a head generation
+- **THEN** the head is ACK'd without waiting for a status reply
+
+#### Scenario: Own events are not ACKs
 
-When darkirc publishes a chat message event, it SHALL first persist an
-outbound record containing the event id, the plaintext message fields
-needed to rebuild it, a state, an attempt counter, and a link to any
-superseding replacement event. Records SHALL survive restart. Records
-SHALL be closed (removable) once a positive outcome is observed or the
-attempt cap is reached.
+- **WHEN** only a locally generated event references the head generation
+- **THEN** that reference does not advance the queue
 
-#### Scenario: Restart does not lose pending sends
+#### Scenario: Old-generation ACK races with recreation
 
-- **WHEN** darkirc commits an outbound event locally, broadcasts it, and
-  the process restarts before any reply arrives
-- **THEN** after restart the outbound record is still present and the
-  delivery policy resumes evaluating it
+- **WHEN** an ACK for an older head generation arrives during recreation
+- **THEN** the logical message is dequeued exactly once and stale work
+  cannot restore it or start further retries
 
-### Requirement: Rebroadcast-first delivery policy
+#### Scenario: Duplicate ACK after restart
 
-For an outbound record with no positive outcome, darkirc SHALL wait
-until it is DAG-synced and has at least a configured minimum number of
-peer connections, then rebroadcast the original event unchanged. It
-SHALL retry with backoff, bounded by a configured maximum number of
-rounds. A single `Has` outcome (either `inserted` value) SHALL close the
-record as delivered.
+- **WHEN** an ACK'd head has been durably dequeued, darkirc restarts, and
+  another ACK arrives for that removed message
+- **THEN** the current head remains queued unless independently ACK'd
 
-#### Scenario: Reconnect triggers rebroadcast
+### Requirement: Retry the head until ACK
 
-- **WHEN** a pending record exists, the node is synced, and the minimum
-  connection count is reached
-- **THEN** the original event is rebroadcast unchanged
+For a materialized, unacknowledged head, darkirc SHALL rebroadcast the
+active event unchanged when sync/connectivity requirements are met. It
+SHALL use backoff for a configured number of rounds and continue at a
+capped slow rate thereafter. There SHALL be no total retry-round,
+recreation-attempt, or lifetime limit that discards an unacknowledged
+message or advances the queue. Negative replies SHALL NOT bypass backoff.
 
-#### Scenario: Prior propagation is detected, not recreated
+#### Scenario: Reconnect retries only the head
 
-- **WHEN** a rebroadcast reaches peers that already hold the event via
-  earlier propagation or in a retained older rotation slot
-- **THEN** their `Has { inserted: false }` replies close the record and
-  no recreation happens
+- **WHEN** queued messages exist and sync/connectivity requirements are
+  restored for a materialized head
+- **THEN** the head's active event is rebroadcast unchanged and later
+  messages are not sent
 
-### Requirement: Ancestry reference counts as delivery evidence
+#### Scenario: Retry limits never discard the head
 
-The delivery monitor SHALL treat a foreign event whose DAG ancestry
-includes one of its tracked outbound events as positive delivery
-evidence, closing the outbound record as delivered. This is a local
-computation over already-received events and adds no wire traffic.
+- **WHEN** repeated retry rounds and rotations pass without ACK
+- **THEN** the logical message remains at the head; retries and eligible
+  recreations continue at bounded rate while required resources exist
 
-#### Scenario: Foreign child closes the record
+### Requirement: Recreate the head only after expiry
 
-- **WHEN** a new event arrives whose ancestry (walked through parent
-  references) contains a pending outbound event
-- **THEN** the outbound record is closed as delivered without waiting
-  for any explicit status reply
+darkirc SHALL recreate the active head event only when its rotation
+window has closed, known locally or indicated by `TooOld`, and a
+rebroadcast round completes without ACK for any head generation. This
+SHALL apply even when no status replies have ever arrived. The replacement
+SHALL use fresh timestamp, parents, and encryption nonce, preserve queue
+position, and link the old generation to the new one locally.
 
-### Requirement: Recreate only on explicit evidence
+Silence, `Invalid`, `NotSynced`, and `Busy` alone SHALL NOT cause recreation
+or dequeue. Any ACK SHALL stop further recreation of that logical message.
 
-darkirc SHALL create a replacement event from the stored plaintext —
-fresh timestamp and DAG parents, fresh encryption nonce — only when
-either:
+#### Scenario: Local expiry with silent peers
 
-- the rotation window for the original event is closed (known locally
-  from the rotation schedule, or indicated by `TooOld` nacks) and no
-  peer has answered `Has` for it, or
-- every observed reply is a nack and at least one carries reason
-  `Invalid`
+- **WHEN** the head's active event is locally known to be expired and a
+  rebroadcast round completes without ACK, even with no replies ever
+  received, and materialization resources are available
+- **THEN** a replacement is durably retained and sent at the same queue
+  position; no later message is sent
 
-The following observations SHALL NOT trigger recreation while the
-rotation window is open: silence, `Nack { NotSynced }`, and
-`Nack { Busy }` mean retry later. Silence SHALL NOT trigger recreation
-even after the window closes when no status reply of any kind has ever
-been received from status-capable peers. A replacement links back via
-the supersession chain so the sender can correlate attempts locally.
+#### Scenario: Retained-slot holder answers before recreation
 
-Recreation SHALL consume a fresh rate-limit slot when rate limiting is
-enabled; if the epoch budget is exhausted the attempt SHALL be parked
-and retried after epoch rollover rather than dropped or sent unproven.
-Recreation attempts SHALL be capped; once the cap is reached the record
-SHALL be closed as failed and the failure surfaced to the client.
+- **WHEN** an expired head receives `Has { inserted: false }` from a peer
+  holding it in a retained slot
+- **THEN** it is ACK'd and dequeued without recreation
 
-#### Scenario: Rotation makes the original undeliverable
+#### Scenario: Rejection or silence during an open window
 
-- **WHEN** a pending record's rotation window has closed and no peer
-  answers `Has` after a rebroadcast round
-- **THEN** a replacement event with fresh parents/timestamp is published
-  and linked via the supersession chain
+- **WHEN** the active window is open and the head sees only silence,
+  `Invalid`, `NotSynced`, or `Busy`
+- **THEN** it retries unchanged with backoff, without recreation or dequeue
 
-#### Scenario: Holder answers before recreation
+### Requirement: Park rather than discard a blocked head
 
-- **WHEN** the rotation window has closed but a rebroadcast round
-  returns `Has { inserted: false }` from any peer holding the event in a
-  retained slot
-- **THEN** the record is closed as delivered and no replacement is
-  created
+Every new generation SHALL use a fresh rate-limit slot when rate limiting
+is enabled. Exhausted budget SHALL park the head until epoch rollover;
+missing identity SHALL park it until identity is available. Processing
+or storage failures SHALL preserve the head and prevent later sends from
+bypassing it. Retrying the same event SHALL NOT re-encrypt it or consume
+a new rate-limit slot. No unproven replacement SHALL be sent. Retryable
+errors SHALL be retried at bounded rate, not treated as terminal dequeue.
+Blocked-send error notices to connected clients SHALL NOT imply removal
+from the queue; receipt/status display remains out of scope.
 
-#### Scenario: Silence never recreates on its own
+#### Scenario: Budget exhaustion parks the whole queue
 
-- **WHEN** no status replies of any kind are observed (for example all
-  peers run versions without status support)
-- **THEN** the monitor keeps rebroadcasting at a bounded rate and does
-  not create replacements
+- **WHEN** the head needs a new generation but the rate-limit epoch budget
+  is spent
+- **THEN** the head waits for rollover without publishing an unproven
+  event, and later messages remain queued
 
-#### Scenario: Transient nacks mean retry
+#### Scenario: Missing identity does not lose the head
 
-- **WHEN** observed replies are `Nack { NotSynced }` or `Nack { Busy }`
-  while the rotation window is open
-- **THEN** the monitor retries later and does not recreate
+- **WHEN** an enabled rate-limit path cannot obtain an identity
+- **THEN** the head remains queued and can resume when identity becomes
+  available, without advancing to another message
 
-#### Scenario: Explicit rejection recreates
+### Requirement: Bounded queue storage with explicit backpressure
 
-- **WHEN** every observed reply across the retry rounds is a nack and at
-  least one carries reason `Invalid`
-- **THEN** a replacement event is published (bounded by the attempt cap)
+Queue storage SHALL have a configured bound covering pending messages,
+active event/blob data, and retained generation metadata. When a new
+enqueue cannot fit or cannot be persisted, darkirc SHALL explicitly reject
+it to the connected client rather than report acceptance. Storage pressure
+SHALL NOT evict unacknowledged entries or their required retry/correlation
+data. Recreation that cannot fit SHALL park the head. ACK'd entries MAY
+be removed to reclaim queue storage.
 
-#### Scenario: Budget exhaustion parks, not drops
+#### Scenario: Full queue rejects a new send
 
-- **WHEN** recreation is due but the rate-limit epoch budget is spent
-- **THEN** no unproven replacement is broadcast and the record waits for
-  the next epoch
+- **WHEN** a new message would exceed the queue storage bound
+- **THEN** its enqueue is explicitly rejected and existing queued messages
+  retain their contents and order
 
-#### Scenario: Attempt cap surfaces failure
+#### Scenario: Generation metadata fills the storage budget
 
-- **WHEN** the configured recreation attempt cap is reached without any
-  positive outcome
-- **THEN** the record is closed as failed and the client is informed
+- **WHEN** recreation would exceed the bound for retained generation data
+- **THEN** the head parks without dropping previous generation ids or
+  allowing later messages to bypass it

+ 92 - 50
openspec/changes/darkirc-delivery-status/tasks.md

@@ -1,28 +1,39 @@
 # Tasks: darkirc-delivery-status
 
-## 1. Wire compatibility spike (blocks everything)
+## 1. Compatibility and protected-area review prerequisites
 
 - [ ] 1.1 Verify unknown-message tolerance: using the event-graph test
   harness (`src/event_graph/test_helpers.rs`), connect two nodes where
-  only one registers an `EventPutStatus` dispatch, send one from the
-  other, and confirm the receiving channel stays up (no stop/strike/
+  only one registers an `EventPutStatus` dispatch, send a status from
+  that node to the unregistered receiver, and confirm the receiving
+  channel stays up (no stop/strike/
   panic). Record the outcome in the change notes; if unknown ids
   destabilize channels, stop and re-discuss gating before proceeding.
   Verify via a new test in `src/event_graph/tests.rs`.
+- [ ] 1.2 Before implementation changes touching RLN, crypto, circuits,
+  or canonical serialization, stop for human review of the affected
+  path, especially reservation/recovery for head materialization and
+  recreation. Record the review outcome and constraints. Do not treat
+  planning approval or an agent verdict as crypto sign-off. Flag any
+  required dependency, build-script, or proc-macro change for human
+  review rather than silently expanding scope.
 
 ## 2. Event-graph layer (`src/event_graph`)
 
 - [ ] 2.1 Define `EventPutStatus`/`EventPutResult`/`NackReason` in
   `proto.rs` with explicit `u8` discriminants, `impl_p2p_message!` with
   default metering, and register dispatch/subscription in
-  `ProtocolEventGraph::init`. Verify `make` builds and clippy is clean.
+  `ProtocolEventGraph::init`. Verify wire tags with round-trip and
+  explicit-byte tests without modifying canonical serialization code;
+  run `make` and `make clippy`.
 - [ ] 2.2 Emit statuses from `handle_event_put` at every outcome per the
   design table (NotSynced, Has{false} duplicate, historical-slot check
   before TooOld, Invalid, Busy, Has{true} inserted), unicast on the
   receiving channel; strike/flood paths unchanged and silent. Verify
   each emission point with a two-node test asserting the exact reply
   variant, including the retained-slot `Has{false}` case across a
-  rotation.
+  rotation. Distinguish established invalidity (`Invalid`) from
+  transient parent-fetch/processing failure (`Busy`) in tests.
 - [ ] 2.3 Add `receipt_pub: Publisher<(EventPutStatus, ChannelPtr)>` and
   `receipt_subscribe()` to `EventGraph`; republish every inbound status
   unfiltered from the per-channel handler. Verify with a test that
@@ -30,57 +41,88 @@
 - [ ] 2.4 Decode hardening: unknown `EventPutResult`/`NackReason`
   variant or malformed body is dropped with a warning, no panic, no
   strike, channel stays connected. Verify with a fuzz-style unit test
-  feeding truncated/random payloads through the decoder.
+  feeding truncated/random payloads through the decoder and a channel
+  test confirming the connection remains usable after malformed input.
 
-## 3. darkirc outbound tracking (`bin/darkirc`)
+## 3. Durable darkirc FIFO (`bin/darkirc`)
 
-- [ ] 3.1 Add the `darkirc_outbound` kvdb tree and serializable record
-  (event id, plaintext Privmsg fields, state, attempts, ts,
-  `superseded_by`), plus helpers to insert/close/update records. Verify
-  with round-trip serialization + tree unit tests in `server.rs`.
-- [ ] 3.2 Write the outbound record in `publish_events` before
-  `p2p.broadcast`; close as Delivered on any `Has`. Verify with a unit
-  test that a crash-restart (reopen kvdb) still finds the record
-  Pending.
+- [ ] 3.1 Add the `darkirc_outbound` ordered logical-message records and
+  durable enqueue/head/dequeue operations, with plaintext rebuild
+  fields, stable queue position, optional active event/blob, generation
+  links, and retry/parked state. Verify serialization, concurrent
+  enqueue ordering, and reopen/restart preservation with local tests.
+- [ ] 3.2 Route every accepted outgoing chat message through durable
+  enqueue, including offline and initial-sync sends. Report enqueue
+  failure explicitly; do not create events for later entries. Verify
+  that accepted messages survive restart, failed enqueues do not report
+  success, and neither direct broadcast nor DAG relay/sync exposes a
+  later queued message.
+- [ ] 3.3 Persist the head's exact event/blob independently of DAG
+  retention, and commit replacements with the active-generation pointer
+  and old-to-new `superseded_by` link before network exposure. Verify
+  fault-injected persistence failures, restart before/after commit,
+  unchanged rebroadcast after all local slots prune the event, and one
+  active generation at the original queue position.
+- [ ] 3.4 Enforce a storage bound covering queued plaintext, active
+  event/blob data, and generation metadata. Explicitly reject new
+  enqueues at capacity; park recreation if it cannot fit. Verify no
+  unacknowledged entry or required correlation data is evicted, and
+  ACK'd entries can be removed to reclaim queue space.
 
-## 4. darkirc delivery monitor (`bin/darkirc`)
+## 4. ACK-driven head-only worker (`bin/darkirc`)
 
-- [ ] 4.1 Receipt aggregation task: subscribe `receipt_pub`, filter to
-  tracked ids, dedupe per (event id, channel address), update records;
-  treat an ancestry reference (foreign event whose parents' closure
-  contains a tracked id) as positive delivery evidence. Verify with
-  unit tests driving synthetic statuses and a synthetic child event
-  through the aggregator.
-- [ ] 4.2 Sweep + rebroadcast: periodic sweep gated by
-  `is_synced() && connection_count >= K`, rebroadcasting the original
-  `EventPut` (event + blob from local DAG) with backoff up to `R_MAX`
-  rounds, then continuing at the slow rate. Verify with a multi-node
-  harness test that a peer which already has the event answers
-  `Has{inserted:false}` and the record closes without recreation, and
-  that pure silence (status-incapable peer) never produces a
-  replacement while the window is open.
-- [ ] 4.3 Recreate triggers per design: window-closed-with-no-holder
-  (local rotation knowledge or TooOld nacks) and explicit all-nack with
-  at least one `Invalid`; `NotSynced`/`Busy`/silence mean retry. Build
-  a fresh event from stored plaintext (fresh nonce, fresh
-  parents/timestamp), park on RLN `BudgetExhausted`, cap attempts at
-  `A_MAX`, surface failure to the client on cap. Verify with harness
-  tests: rotation-forced recreate writes a superseding record linked
-  via `superseded_by`; budget exhaustion parks; attempt cap reaches
-  Failed; `NotSynced`-only replies do not recreate.
-- [ ] 4.4 End-to-end darkirc scenario test: node A publishes while
-  disconnected from B, A reconnects after DAG rotation, B must end up
-  holding a recreate-generation event and A's record must show the
-  `superseded_by` link; a second scenario where B already holds the
-  original in a retained slot must close A's record as Delivered with
-  no recreate. Verify both assertions in the multi-node harness.
+- [ ] 4.1 Subscribe to receipts and accepted foreign events; filter to
+  tracked head generations and use ephemeral connection identity, not
+  channel addresses, for any aggregation. Verify ACK from both `Has`
+  variants and foreign ancestry, and no ACK from locally generated
+  events or unrelated ids. Do not persist or log peer addresses.
+- [ ] 4.2 Serialize ACK/dequeue with head materialization and recreation.
+  Verify exactly-once durable advancement, ACK processing while sends
+  are gated, duplicate/late replies after restart, and an old-generation
+  ACK racing with replacement work without resurrecting the entry or
+  removing the next head.
+- [ ] 4.3 Service only the head when synced and connection count reaches
+  K. Materialize its initial event only then; otherwise retry its exact
+  active event with doubling backoff up to `R_MAX`, then a capped slow
+  rate indefinitely. Verify with controlled-time tests that negative
+  replies cannot bypass backoff, no retry/lifetime limit discards the
+  head, and later entries never create or expose events before ACK.
+- [ ] 4.4 Recreate only after local expiry or `TooOld` and a completed
+  rebroadcast round without ACK. Keep queue position, use fresh
+  parents/timestamp/nonce and the reviewed existing RLN path, and retain
+  old-generation correlation. Verify silent expiry, retained-slot `Has`
+  cancellation, multiple rotations without an attempt cap, and no
+  recreation from silence/`Invalid`/`NotSynced`/`Busy` alone.
+- [ ] 4.5 Park the head on budget exhaustion, missing identity, and
+  processing/storage errors; retry at bounded rate without skipping it.
+  Verify enabled-RLN budget rollover and identity recovery under the
+  human-reviewed path, unchanged retries consuming no new slot, and no
+  unproven enabled-RLN replacement. Verify blocked-send notices do not
+  imply dequeue or terminal failure, and later messages remain queued.
 
-## 5. Hardening and review gate
+## 5. End-to-end FIFO scenarios
 
-- [ ] 5.1 Full workspace gates green: `make`, `make clippy`, `make test`
+- [ ] 5.1 Test offline enqueue of A then B, restart, and reconnect in a
+  multi-node harness with the configured connectivity threshold met.
+  Assert only A materializes/sends, B stays unexposed until A's ACK,
+  and B subsequently becomes the head. Repeat with A parked to confirm
+  cross-conversation head-of-line blocking rather than implicit skip.
+- [ ] 5.2 Test an already-materialized, unacknowledged head spanning
+  disconnection and rotation: first rebroadcast the original, then
+  recreate without advancing the queue if no holder ACKs. In a separate
+  scenario a peer retains the original and replies `Has{inserted:false}`:
+  assert dequeue without recreation. Include a late ACK for an older
+  generation after replacement and verify the next message advances
+  only once. Do not assume recipient display order or deduplication.
+
+## 6. Hardening and review gates
+
+- [ ] 6.1 Full workspace gates green: `make`, `make clippy`, `make test`
   (proofs + contracts built first per AGENTS.md); confirm no
   `unwrap`/`expect`/`panic!` on any new attacker-controlled decode path
-  and no peer addresses logged with receipt state.
-- [ ] 5.2 Invoke `@anon-security-review` on the full diff; treat FAIL as
+  and no secrets or peer addresses logged with queue/receipt state.
+- [ ] 6.2 Invoke `@anon-security-review` on the full implementation diff;
+  treat FAIL as
   blocking and address findings before marking the change ready to
-  apply. Verify the review verdict is recorded in the change notes.
+  apply/archive. Verify the review verdict is recorded in the change
+  notes; planning-only review does not satisfy this implementation gate.