R2D2-MERIDIAN/DIAGRAM.md
Joshua Belke 610b9cac2e
Some checks failed
helm chart / lint + unittest + render matrix (push) Has been cancelled
helm chart / install on kind (gated) (push) Has been cancelled
helm chart / publish chart to GHCR (push) Has been cancelled
Meridian Harness / Build (aarch64-unknown-linux-musl) (push) Has been cancelled
Meridian Harness / Build (x86_64-unknown-linux-musl) (push) Has been cancelled
Meridian Harness / Publish rolling release (push) Has been cancelled
Meridian Harness / Publish tagged release (push) Has been cancelled
CI / Detect Changed Paths (push) Has been cancelled
CI / Dead Token Reference Guard (push) Has been cancelled
Docker image / Build (linux/amd64) (push) Has been cancelled
Docker image / Build (linux/arm64) (push) Has been cancelled
Docker image / Build public push gateway (linux/amd64) (push) Has been cancelled
Docker image / Build public push gateway (linux/arm64) (push) Has been cancelled
CI / Rust Lint (push) Has been cancelled
CI / Unit Tests (push) Has been cancelled
CI / Isolated DB Gate (push) Has been cancelled
CI / Desktop Core (push) Has been cancelled
CI / Desktop Smoke E2E (1) (push) Has been cancelled
CI / Desktop Smoke E2E (2) (push) Has been cancelled
CI / Desktop Smoke E2E (3) (push) Has been cancelled
CI / Desktop Smoke E2E (4) (push) Has been cancelled
CI / Desktop (push) Has been cancelled
CI / Desktop E2E Relay (push) Has been cancelled
CI / Desktop E2E Integration (1/2) (push) Has been cancelled
CI / Desktop E2E Integration (2/2) (push) Has been cancelled
CI / Desktop E2E Integration (push) Has been cancelled
CI / Backend Integration (relay e2e) (push) Has been cancelled
CI / Relay E2E (push) Has been cancelled
CI / Web (push) Has been cancelled
CI / Admin Web (push) Has been cancelled
CI / Mobile (push) Has been cancelled
CI / Security (push) Has been cancelled
CI / Server Cross-Compile (push) Has been cancelled
CI / Server Cross-Compile-1 (push) Has been cancelled
CI / Windows Rust (x86_64-pc-windows-msvc) (push) Has been cancelled
CI / Desktop Build (macOS) (push) Has been cancelled
Docker image / Merge release multi-arch manifest (push) Has been cancelled
Docker image / Merge debug multi-arch manifest (push) Has been cancelled
Docker image / Publish public push gateway image (push) Has been cancelled
docs: withdraw the >=4 KB crossover — shared memory was on for every measurement
Both SHM keys default true in Zenoh 1.8, the eclipse-zenoh wheel IS built
with the shared-memory feature, and transport_optimization's threshold is
3,072 B. So every measured payload at 4 KB and above took a POSIX-SHM fast
path -- one the relay build cannot take, because shared-memory is not in
its zenoh feature list -- while the Redis side had no equivalent. The error
points toward Zenoh.

RELAY_BUS_SCALING.md's own bullet said "Nothing about shared memory. SHM
was not enabled." That was false, and nothing could have caught it: the
static contract checked transport.shared_memory.enabled and was blind to
the second switch beside it. This adds that check.

The "Router mode prices the Docker boundary" reading goes with it. SHM
works host-to-host and cannot cross into the Docker VM, so an unknown
share of the peer-vs-router divergence at >=64 KB is the SHM path dropping
out rather than the boundary appearing. Both readings are unlicensed until
re-measured under the shipped posture.

Below 4 KB stands, including the 256 B verdict that failed the >=5x gate:
256 B is well under the SHM threshold, and gossip's extra transports cost
the measured process work rather than saving it.

Worth stating plainly, because it is the second time: the 0A.2 result has
now been invalidated twice for two unrelated reasons -- unmatched publish
semantics, then transport posture -- and neither was visible in the
numbers. A bus measurement is not licensed by its spread. It is licensed by
its posture being pinned, read back, and stamped beside the result, which
is what af4b92eba now does.

Signed-off-by: Joshua Belke <joshua@innovationhub-act.org>
2026-08-21 00:54:18 -04:00

144 KiB
Raw Permalink Blame History

Status: active system diagram + maturity map · refreshed 2026-08-05 · current system at S0; local production hardening is ahead of live/provider qualification; the S1 consumer plane remains a proposal with no cursor migration or consumer crate. Framing: how the client, relay, and stack tiers actually fit together today; how that shape scores against the GOAT/EFDI principles; what it can be expected to carry compared with a central-broker deployment; and how it matures from one relay that does everything into a thin spine plus independently deployable consumers, without adding a second read path, a second policy store, or a client-visible protocol change. Related: ARCHITECTURE.md — the component reference this document builds on · VISION.md · AGENTS.md § Design Law · .settings/epics/README.md · .settings/features/README.md · .settings/features/feature-zenoh-transport.md · docs/git-on-object-storage.md · .settings/reference-docs/goat/00-DISTILLED.md

Meridian — how it fits together, and where it goes

ARCHITECTURE.md is the component reference: what each crate is, what it does, what it explicitly does not do. This document is the system view built on top of it — how the tiers compose, what each is allowed to assume about the others, and what happens to that composition under load.

Part Answers
I — The system as it runs today How client, relay, and stack work together right now
II — GOAT/EFDI conformance Which federated-data-access principles we meet, will meet, or deliberately deviate from
III — Throughput, and the case against a central broker What this can carry, how that compares with TBMQ, and why the broker is the wrong shape as we scale
IV — Maturation map The portfolio, the gaps, and the spine/bus/consumer road from S0

Two rules govern every number and every claim below, both from AGENTS.md, and they are why this document is longer than a slide:

  • No throughput figure without its profile — traffic class, auth profile, parsed or not, stored or not.
  • Nothing is resolved until deployed to a live relay and re-probed. Green CI is not evidence that an enforcement path runs in production.

Part I — The system as it runs today

The three tiers

flowchart TB
    classDef client fill:#4a4a4a,stroke:#2b2b2b,color:#fff
    classDef relay fill:#b03030,stroke:#6e1c1c,color:#fff
    classDef sub fill:#7d6608,stroke:#4d3f05,color:#fff
    classDef store fill:#1e7a45,stroke:#0f4527,color:#fff
    classDef opt fill:#1f6aa5,stroke:#123f63,color:#fff

    subgraph CL["CLIENT TIER — speaks NIP-01 JSON and a narrow HTTP surface, nothing else"]
        direction LR
        D["Desktop<br/>Tauri 2 + React 19"]:::client
        M["Mobile<br/>Flutter + Riverpod"]:::client
        W["Web<br/>repo browser + invite accept"]:::client
        A["Agents<br/>meridian-cli · ACP harness · dev-MCP"]:::client
        N["Any Nostr client<br/>gitworkshop · ngit · generic"]:::client
    end

    subgraph RL["RELAY TIER — meridian-relay, one Axum process, the ONLY enforcement point"]
        direction TB
        B0["STEP 0 · community bind<br/>resolve_host to TenantContext<br/>before AUTH, EVENT, REQ, REST, media, git"]:::relay
        B1["admission · NIP-42 challenge / NIP-98<br/>conn semaphore · handler semaphore 1024"]:::relay
        B2["EVENT pipeline<br/>kind gate · BIP-340 verify · drift · size<br/>identity match · scope · membership"]:::relay
        B3["REQ pipeline<br/>access check BEFORE registration<br/>historical query then EOSE"]:::relay
        B4["SubscriptionRegistry<br/>DashMap 3-tier fan-out index"]:::relay

        SUBSYS["SUBSYSTEMS — called only by the relay, never by each other<br/>meridian-core · -db · -auth · -pubsub · -search · -audit · -workflow · -media"]:::sub

        B0 --> B1 --> B2
        B1 --> B3
        B2 --> B4
        B3 --> B4
        B2 -.-> SUBSYS
        B3 -.-> SUBSYS
    end

    subgraph ST["STACK TIER — one system of record per fact"]
        direction LR
        PG[("Postgres 17<br/>THE EVENT LOG<br/>events · channels · members<br/>workflows · audit chain · FTS")]:::store
        DF[("Dragonfly<br/>STATE, never events<br/>pub/sub · presence TTL<br/>typing ZSET · caches")]:::store
        S3[("MinIO / S3<br/>BYTES, never events<br/>media blobs · git objects")]:::store
        OPT[("Opt-in profiles<br/>keycloak · prometheus · adminer")]:::opt
    end

    CL -->|"WSS · NIP-01 EVENT / REQ / CLOSE / AUTH"| RL
    CL -->|"HTTPS · /events /query /count /media /git /hooks /info"| RL
    RL <--> PG
    RL <--> DF
    RL <--> S3
    RL -.-> OPT

The composition rule that makes this legible: the relay orchestrates every subsystem by direct call, and the subsystems never call each other (ARCHITECTURE.md § Crate Dependency Hierarchy). meridian-workflow never calls meridian-pubsub; meridian-search never calls meridian-db. Every cross-subsystem interaction is a line of code in the relay, which is why one process can be reasoned about at all.

Tier contracts — what each owns, and what it must never do

Tier Owns Must never
Client Rendering, local drafts, local-only settings, key custody, optimistic UI Hold authority. A client is never asked to enforce membership, and never talks to anything but the relay — no client connects to Postgres, Dragonfly, S3, or a consumer
Relay Community binding, admission, signature verification, ordering, membership enforcement, fan-out, re-authorization of every candidate Delegate an access decision. It is the only enforcement point, which is what makes the no-third-shim law free rather than expensive
Postgres The event log, and every fact derived by transaction — channels, membership, workflow runs, the audit chain Be second. Nothing else may claim truth for a fact it stores (R2, R3)
Dragonfly Ephemeral state with a TTL — presence, typing, rate windows, replay set, caches — plus cross-pod delivery Become a durable log. It is a keyspace and a transport, never a system of record
Object store Bytes whose size makes them unfit for a row — media blobs, git packs Hold anything the relay must interpret to make a decision

Where the community boundary is drawn

Every tenant-visible path is fenced before a handler sees data, and the fence is derived from the request host, not from anything the client can assert.

flowchart LR
    classDef gate fill:#7d6608,stroke:#4d3f05,color:#fff
    classDef ok fill:#1e7a45,stroke:#0f4527,color:#fff
    classDef bad fill:#b03030,stroke:#6e1c1c,color:#fff

    H["Request arrives<br/>host: acme.meridian.example"] --> R{"resolve_host"}:::gate
    R -->|"known host"| T["TenantContext bound<br/>req.community = acme"]:::ok
    R -->|"unknown host"| X["fail closed · generic reject<br/>NEVER falls through to a default tenant"]:::bad

    T --> U["Every downstream path inherits it:<br/>AUTH · EVENT · REQ · REST · media · git · search · workflow · pub/sub"]:::ok
    U --> V["Client-supplied #h tags stay CHANNEL identifiers.<br/>They must resolve INSIDE the host-derived community.<br/>A NIP-98 stamp that disagrees with the host loses."]:::gate

This is the property that lets one deployment host many communities without a per-tenant relay, and it is why community_id is immutable on the row and part of every primary key.

Connection lifecycle

sequenceDiagram
    autonumber
    participant CL as Client
    participant RL as meridian-relay
    participant DF as Dragonfly
    participant PG as Postgres

    CL->>RL: WSS upgrade — Host header decides the community
    RL->>RL: Step 0 · resolve_host to TenantContext, or fail closed
    RL->>RL: Step 1 · conn_semaphore.try_acquire_owned, else reject before reading a byte
    RL-->>CL: Step 2 · AUTH challenge, random nonce
    CL->>RL: Step 3 · AUTH signed kind:22242 event
    RL->>RL: verify Schnorr, timestamp within 60s
    Note over RL: AuthState Pending becomes Authenticated(AuthContext).<br/>The pubkey is the identity for humans AND agents alike.

    rect rgba(30,122,69,0.15)
    Note over RL: Step 4 · three concurrent loops for the connection's life
    RL->>RL: recv_loop — parse frames, dispatch handlers
    RL->>CL: send_loop — drain mpsc, write frames, try_send with a 3-strike slow-client rule
    RL->>CL: heartbeat_loop — ping every 30s, 3 missed pongs disconnects
    end

    CL->>RL: REQ / EVENT / CLOSE
    RL<<->>PG: stored reads and writes
    RL<<->>DF: presence, typing, cross-pod publish

    CL-->>RL: disconnect
    RL->>RL: Step 5 · cancel token, await loops, drop subscriptions,<br/>deregister connection, release the semaphore permit

Write path — what an EVENT actually costs

Grounded in crates/meridian-relay/src/handlers/ingest.rs and handlers/event.rs, not in prose. The ordering matters: everything that can reject runs before anything that can persist.

sequenceDiagram
    autonumber
    participant CL as Client
    participant RL as Relay
    participant PG as Postgres
    participant DF as Dragonfly
    participant SUB as Local subscribers
    participant AU as Audit worker
    participant WF as Workflow engine

    CL->>RL: ["EVENT", signed event]

    rect rgba(176,48,48,0.12)
    Note over RL: ADMISSION — cheap rejects first, crypto next
    RL->>RL: reject KIND_AUTH, relay-only kinds, membership notifications
    RL->>RL: HTTP-forbidden kinds if this came via POST /events
    RL->>RL: BIP-340 verify + id hash — spawn_blocking, THE CPU CEILING
    RL->>RL: timestamp drift within 15 min
    RL->>RL: content size within 256 KiB
    RL->>RL: event.pubkey == authenticated identity, except gift wrap
    RL->>RL: required scope for kind
    RL->>RL: channel membership — 10s TTL cache
    end

    alt kind 20000-29999 — ephemeral
        RL->>DF: presence SET EX 90, or publish to the channel topic
        RL->>SUB: local fan-out
        Note over RL: never stored, never audited, never indexed
    else persistent
        rect rgba(30,122,69,0.15)
        Note over RL,PG: THE DURABILITY BOUNDARY
        RL->>PG: INSERT ... ON CONFLICT DO NOTHING<br/>search_tsv generated column fills here — FTS is the same row
        PG-->>RL: committed
        end

        par post-commit dispatch
            RL->>DF: PUBLISH — unconditional today (neither condition is built)
        and
            RL->>SUB: fan_out_scoped, filter_fanout_by_access,<br/>serialize once then a frame per subscription
        and
            RL->>AU: bounded mpsc(1000) — backpressures rather than dropping
        and
            RL->>WF: spawned trigger, excluding workflow-execution and command kinds
        end
        RL-->>CL: ["OK", id, true, ""]
    end

Four properties worth stating because they are easy to get wrong on a rewrite:

  • Search is not a step. Under Postgres FTS the searchable row is the persisted row — search_tsv is a generated column populated by the insert. The old out-of-band indexer and its search_index_tx queue are gone (handlers/event.rs:580). That removes a step and adds a write-path tax; see W3.
  • Audit backpressures on purpose. audit_tx.send().await on a bounded channel means a saturated audit DB slows ingest instead of silently losing chain entries. The advisory lock already serializes writes, so a full queue means the audit store is genuinely overloaded.
  • The bus is touched only to reach another pod. Local subscribers are served from the in-process registry. At one pod, the pub/sub hop costs nothing — this is the structural difference from a broker, and Part III turns on it.
  • Fan-out re-authorizes. filter_fanout_by_access runs over the matched recipient set before any frame is built, with additional owner-only gates for viewer-private kinds. Honest cost: this is per-event authorization work, and Part II scores it as such.

Read path — access is checked before the subscription exists

sequenceDiagram
    autonumber
    participant CL as Client
    participant RL as Relay
    participant PG as Postgres
    participant SE as meridian-search

    CL->>RL: ["REQ", sub_id, filters...]
    RL->>RL: parse filters, extract channel_id
    RL->>RL: scope check — else CLOSED "restricted: insufficient scope"
    RL->>RL: load accessible_channel_ids for this pubkey — 10s TTL cache

    alt not a member
        RL-->>CL: CLOSED "restricted: not a channel member"
        Note over CL: EXPLICIT DENIAL — never a silently empty result
    else authorized
        RL->>RL: register in SubscriptionRegistry only NOW<br/>no race window between registration and the check
        opt NIP-50 search filter
            RL->>SE: FTS over search_tsv, community_id always in the predicate
            SE-->>RL: CANDIDATES — permission filtering is the caller's job
            RL->>RL: re-authorize every hit
        end
        RL->>PG: historical query, 500 per filter hard cap
        PG-->>RL: stored events
        RL-->>CL: EVENT frames
        RL-->>CL: EOSE
        Note over CL: live events arrive on the fan-out path from here
    end

The p-gate is the other half of this: a REQ that omits kinds is rejected rather than served, so an open-ended query cannot walk a community's history.

Three data planes that deliberately avoid the event path

Not everything is an event, and the exceptions are principled rather than accidental. Each is bytes whose size or latency profile would poison the event pipeline.

Plane Transport Where truth lives Why it is not an event
Git objects Git smart HTTP under /git/{owner}/{repo}/* S3/MinIO — manifest pointer plus CAS, no authoritative per-repo filesystem state on the relay Packs are megabytes and immutable; the relay's remaining cost is hydrate/pack CPU, not storage
Media blobs Blossom PUT /media/upload, GET /media/{sha256} S3/MinIO A 50 MB upload through the event pipeline would occupy a handler slot and blow the 256 KiB content cap. The relay is still in the byte path, which is B5
Huddle audio Dedicated WS at /huddle/{channel}/audio, 8-byte header plus opaque Opus Nowhere — forwarded, never stored Real-time frames are worthless a second late. Per-peer channels are bounded and drop-on-full; the control channel never drops join/leave

All three still authenticate through the same identity and membership gates — they leave the event path, not the trust boundary.

More than one pod

flowchart TB
    classDef pod fill:#b03030,stroke:#6e1c1c,color:#fff
    classDef bus fill:#1f6aa5,stroke:#123f63,color:#fff
    classDef store fill:#1e7a45,stroke:#0f4527,color:#fff
    classDef note fill:#7d6608,stroke:#4d3f05,color:#fff

    C1["Clients on pod A"] --> PA["Relay pod A"]:::pod
    C2["Clients on pod B"] --> PB["Relay pod B"]:::pod

    PA --> PG[("Postgres — shared, the log")]:::store
    PB --> PG

    PA -->|"PUBLISH meridian:{community}:channel:{uuid}"| DF[("Dragonfly PUB/SUB")]:::bus
    DF -->|"only topics this pod RETAINED"| PB
    DF -->|"local echo skipped via local_event_ids"| PA

    NOTE["Interest scoping is the measured lever:<br/>retain_topic / release_topic drive dynamic SUBSCRIBE.<br/>64 communities, 1 subscribed: 64x cluster ingress reduction<br/>perf/RELAY_BUS_SCALING.md --mode redis, harness asserts >=95% of ideal"]:::note
    DF -.-> NOTE

Three facts hold this together and each is a real failure mode if broken:

  1. Local-echo dedup. The publishing pod marks the event id in local_event_ids before publishing, so the Redis round trip does not deliver it twice. If the publish fails, the mark is invalidated — otherwise a failed publish would suppress a later legitimate delivery.
  2. Interest, not firehose. A pod subscribes only to topics for which it has local subscribers. Without this, every pod pays every community's traffic — the harness has a mutant row that fails precisely when someone reintroduces the global firehose.
  3. Postgres is the arbiter, not the bus. Two pods writing the same addressable event resolve in the database, not in the delivery order. That is also where the LWW revocation defect lives.

What Part I does not claim

  • Rate limiting is not enforced. The RateLimiter trait exists; the only implementation is a test stub. RateLimitConfig's four tiers are a design target, not a control.
  • Every authenticated connection currently receives every scope. There is no participant class, so "guest", "partner", and "third-party agent" are not yet expressible (C2).
  • One undifferentiated delivery queue. Typing indicators, agent-observer frames, and chat share a lane. The eight-class QoS map is designed, not built.

Part II — GOAT/EFDI conformance

Why these principles bind here

The .settings/reference-docs/goat/ set distils 101 pages describing EFDI — a federated, unclassified data backbone built on a Zenoh fabric with SPIFFE workload identity, a Postgres registry, and grant-keyed payload encryption. It is relevant to Meridian for one reason, and it is not the security model:

Authorization has three possible binding times — publish time, subscribe time, and per-message time. A high-throughput bus must do zero authorization work at per-message time.

Everything distinctive in their design is that principle applied, and it is the same discipline this relay needs to carry a mixed human-and-machine workload. Two caveats travel with the whole comparison, taken from their own status ledger: their encryption core is proven but inert, their key service is built but not deployed, and their federation is designed. We are scoring against a well-reasoned design, not a validated deployment — and their scope note applies to us too: this pattern is built for the unclassified releasable environment and would be re-examined, not assumed to carry up.

flowchart LR
    classDef free fill:#1e7a45,stroke:#0f4527,color:#fff
    classDef cost fill:#b03030,stroke:#6e1c1c,color:#fff
    classDef gate fill:#7d6608,stroke:#4d3f05,color:#fff

    A["ADMISSION TIME<br/>NIP-42 once per connection<br/>community bound from the host"]:::free
    B["SUBSCRIBE TIME<br/>REQ access check before registration<br/>accessible_channel_ids, 10s TTL"]:::free
    C["PER-MESSAGE TIME<br/>filter_fanout_by_access per event<br/>membership_cache, 10s TTL"]:::cost

    A --> B --> C
    C --> V["The residual we own honestly.<br/>GOAT's target is ZERO here.<br/>We are cached, not clean."]:::gate

Status legend

Mark Means
✅ Comply In tree today, with evidence named
🔷 Comply differently Same guarantee and same binding time, different vocabulary — the deviation is stated
◐ Partial Met at some enforcement points and not others; the unmet half is named
🟡 Will comply Decided and gated, with a named issue; not built
🔶 Deviate Deliberate divergence, with the reason
⚠️ Non-compliant Known gap or latent defect, not yet fixed
⛔ Reject Refused, and GOAT refuses it too or the refusal is ours

The five deviations that actually matter

The register below compresses to five statements. If someone reads only one part of Part II, it should be this list — and the register is deliberately placed after it, because the answer matters more than the evidence for it.

Read the count honestly: thirty principles, five deviations, and exactly one of them is unblocked work today. Deviations 1 and 5 are deliberate at this rung, 3 is ratified as build-at-the-second-node, and 4 is cheap but waits on nothing in particular. Deviation 2 — Scope::all_known() — is the one item that is neither deliberate nor blocked, and it gates A2, B4, and every class of traffic that is not first-party. A thirty-row register is not a thirty-item work list, and reading it as one is the failure mode this paragraph exists to prevent.

# The deviation Why it exists What closes it
1 We authorize per event on the fan-out path (A1) Channel membership can change mid-session, and the cheap correct answer was a cached re-check Moving the decision fully to subscribe time with fenced invalidation — a design, not a patch
2 Every authenticated connection gets every scope (C2) Single-tenant first-party origins; nobody needed a partner box yet NIP-PC, resolved once at AUTH and cached for the session — zero per-event cost
3 Revocation resolves last-write-wins (D4) NIP-33 semantics, inherited from the protocol Epoch-monotonic absorbing revoke over an append-only log, before a second writer exists
4 A private channel is invisible rather than gated (B4) It was the simplest correct thing Existence + owner + labels readable by any authenticated member; contents governed
5 We are a central point of decision (E5) Deliberate. It is what makes one enforcement point true Nothing, at this rung. It is re-opened by a second writer, not by preference

A · Placement discipline — where decisions bind

# GOAT principle Status Why, or how we deviate
A1 Never do authorization work per message. Bind at publish, subscribe, or admission — never per message ◐ Partial We bind at admission (NIP-42 once, AuthContext for the session) and at subscribe (req.rs checks channel access before registration, so there is no leak window). We do not meet it at fan-out: filter_fanout_by_access evaluates per event over the matched recipient set, backed by a 10-second-TTL membership_cache. That is cached per-message authorization, not eliminated per-message authorization. Naming it beats claiming the principle
A2 The namespace you publish to is the audience. Pre-cut release lanes cost nothing at runtime 🔷 Comply differently Our audience is (community, channel) — the community from the request host, the channel from the h tag — resolved into accessible_channel_ids at REQ time. Same binding time, different vocabulary. We have no REL-* release lanes because we have no participant class to release to yet (A2 depends on C2), and shipping the vocabulary first would fail just check-kinds
A3 Anything the relay must act on lives in the header; the payload is opaque 🔶 Deviate today · 🟡 machine path Today the relay parses full NIP-01 JSON and re-serializes per fan-out. That is measured as the single largest per-event term: a fixed ~2–3 µs serde_json::to_string plus ~430 ns per recipient (fanout_cost.rs, M3 Max, release, single-thread). The opaque-frame envelope is specified (MIP-OF, meridian-tdq) and not built
A4 Batch is the unit of work through the whole pipeline, with a synchronous inner loop 🟡 Will comply Per-message channel hops today — 21 mpsc:: uses in connection.rs alone. meridian-0zd is the batching refactor; fanout_cost.rs is deliberately its pre-refactor baseline. The named trap: adopting a batching transport and then re-atomising the batch into per-message sends buys nothing
A5 Interest-based routing — forward only what a downstream subscriber declared interest in ✅ Comply retain_topic / release_topic drive dynamic Dragonfly SUBSCRIBE against meridian:{community}:channel:{uuid}. Measured 64× cluster ingress reduction at 64 communities with one subscribed (perf/relay_bus_scaling.py --mode redis); the harness fails below 95% of ideal, so the claim is load-bearing rather than decorative

B · Policy vocabulary

# GOAT principle Status Why, or how we deviate
B1 Labels: three closed lanes — fidelity, context, access — applied per stream, never per message 🟡 Will comply Zero labels today. Decided (meridian-06p): reuse NIP-32's L/l tag grammar but carry it inside the addressable NIP-29 kind 39000 record, not as separate kind 1985 events — a subscribe-time decision must cost one lookup, and a 1985 event costs two. The evidence behind per-stream: DoD IG measured 70% of 220 documents and GAO 84% of 111 with marking defects; per-object marking is the documented burden
B2 Grants are enumerable, dated, owner-issued rows carrying granter, justification, and expiry — never rule formulas ◐ Partial · 🟡 NIP-GR The negative half we already meet: there is no policy DSL in the relay and we refuse one (F3). The positive half is thin — channel_members gives us enumerable rows with joined_at / removed_at / invited_by, so membership is half-dated, but there is no expiry, no justification, no preemptive grant, and no request→approve loop. NIP-GR (meridian-ar6) is specified and parked
B3 Closed grant grammar — anything that does not parse cannot exist; CI rejects it 🟡 Will comply Depends on B2. The guard pattern is already proven here by just check-kinds
B4 Discovery ≠ access. Existence, owner, and labels are visible; contents are governed. "You cannot ask for data you do not know exists" 🔶 Deviate A private channel is currently invisible, not visible-but-gated. This is the item with the largest product effect and the smallest technical cost (meridian-l9i)
B5 Explicit denial, never a silently empty result, and every refusal names the deciding record ◐ Partial The denial half is real and is written law: restricted: not a channel member, restricted: insufficient scope, restricted: p-gated events require #p matching your pubkey (req.rs:90, :209, :231). The explainability half does not exist — no can / why / --as-of surface, and meridian-search returns a filtered candidate set with no record of what was dropped
B6 One-sided friction — restricting demands a written justification, sharing demands nothing — plus published restriction-drift telemetry 🟡 Will comply Not built (meridian-8ak); cheap once B1 and B2 land. The evidence: >60% of derivative classifiers meeting conflicting labels resolve upward, and the audited sharing damage came from erroneously restrictive markings
B7 No vocabulary without an enforcement point — aspirational vocabulary fails CI, not an audit months later ✅ Comply just check-kinds runs inside just check. This guard matters more here than in the source design: we carry a large draft NIP/MIP set under docs/nips/, so the gap between declared and enforced vocabulary is the standing risk

C · Identity and standing

# GOAT principle Status Why, or how we deviate
C1 Standing is registry-derived and effective-dated, never stamped into a credential 🔷 Comply in shape A Meridian pubkey carries no org, role, or tenancy claim whatsoever — standing is resolved server-side against the host-derived community, which is the load-bearing half and we get it for free. Incomplete on dating: [since, until) exists as joined_at / removed_at but is not queryable --as-of, so history is still reconstructed rather than read
C2 Two identity flavors boxed at onboarding — a partner is not a member with permissions removed ⚠️ Non-compliant Every NIP-42-authenticated connection currently receives Scope::all_known(). There is no guest, partner, contractor, or third-party-agent class. NIP-PC (meridian-u35) is parked. This is the single largest gap for any traffic that is not first-party, and it blocks A2
C3 Names never encode ownership ✅ Comply, and it is law Channels are UUIDs; bus topics are meridian:{community_uuid}:channel:{channel_uuid}. AGENTS.md § Design Law forbids org/team/channel paths in repo, topic, or namespace keys. Their reasoning is the one to keep: a producer key freezes into identity, ACL entry, topic prefix, and storage key at once, while ownership stays mutable — so an encoded name turns every reorg into a fleet event
C4 Virtual roles derived from registry rows; no role database ✅ Comply Role is an enum column on the membership row. It exists because the row exists and vanishes with it. No role store, no role explosion
C5 A change of command prompts a grant; it never confers one 🟡 Will comply No org hierarchy exists yet, so this is not live. Recorded now so nobody wires automatic inheritance later — automatic inheritance would open data silently and leave no record of a decision

D · Cryptographic enforcement

# GOAT principle Status Why, or how we deviate
D1 Grant-keyed envelope encryption — one AES-256-GCM DEK per (release-set, generation); O(1) per message, O(N) per rotation only 🟡 Will comply NIP-KB (meridian-9wy), reusing NIP-44 for the wrap and NIP-EE kind 443 KeyPackages as the recipient-key directory — the source design independently calls its own equivalent "the MLS KeyPackage pattern", so the ecosystem shipped it first. Today the only stream encryption we have is NIP-17 gift wrap, which is O(N) per message and correct only for DMs. Deliberate deviation: they mint keys in a separate daemon off the data path; our bundle is just an event riding existing fan-out, so there is still no key server in the read path
D2 Revocation is rotation, with a generation tag on every message and no flag day 🟡 Will comply, with an advantage Ciphertext stays self-describing, so recipients never switch in lockstep. Our hard cut is strictly better than theirs: they revoke a certificate and wait for propagation; we can cancel the authenticated WebSocket session immediately
D3 Crypto agility — suite identifier bound into the AAD, TBS-style signing, a hybrid-PQ path named but not built 🟡 Will comply Three things that are free on day one of D1 and expensive to retrofit. The KEM is the urgent PQ piece because harvest-now-decrypt-later targets confidentiality; a signature cannot be forged retroactively
D4 Revocation is monotonic and absorbing — never timestamp-wins ⚠️ Known latent defect The highest-value catch in the whole set, because it is not an idea we lack — it is a defect we would ship. NIP-33 addressable events resolve last-write-wins by created_at, and KIND_NIP29_GROUP_MEMBERS (39002) and KIND_NIP29_GROUP_ADMINS (39001) sit in that range. A revoke and a concurrent grant healing across a partition or across two relays resolve by later timestamp and resurrect the revoked grant. Ratified in AGENTS.md § Design Law as decided now, built at the second node (meridian-3wh) — dormant while one relay is the only writer, silent and catastrophic with two
D5 Two doors, one policy source — coarse-but-brokerless or fine-but-mediated, chosen per stream, never both 🔷 Structurally ahead, currently behind They cannot fork their router, so they are forced into encryption for machine reads and a static ACL that needs a restart. We wrote the relay, so we own both checkpoints and can choose per stream — but that choice only exists once D1 is built. Today we have only the mediated door, and saying otherwise would be a claim the binary cannot meet

E · Federation and governance

# GOAT principle Status Why, or how we deviate
E1 Networks federate rather than merge — O(N) shared standards, not O(N²) bespoke bridges 🟡 Will comply, hard-blocked Zero federation today. NIP-FP (meridian-w0f) for live interest-based peering, with NIP-77 negentropy reused for backfill rather than inventing sync. Blocked by written law on D4: a live peering NIP over LWW grant semantics is a security bug with extra steps
E2 Peering scopes releasability, not reach — per-namespace ingress/egress ACLs on each link 🟡 Will comply Folds into NIP-XP's XP-ACCEPT negotiation (meridian-2e2), which is the natural carrier — the link is already negotiating a profile, so the releasable prefix set belongs in the same handshake
E3 Recognition ≠ authorization. Compliance is admission; a badge is not a key 🟡 Will comply Adopted verbatim as language. Recognizing a peer relay's identity must never by itself grant access to its data
E4 Nine federation compliance criteria (C1–C9) as a dated, expiring evidence package — standards, not products; evidence, not assertion 🟡 Will comply NIP-11 is the natural carrier (meridian-e4e). C1–C7 are zero-trust hygiene we largely meet; C8 and C9 are exactly B1, D1, and E1, so the criteria document is mostly assembly — but it cannot be claimed before those exist
E5 No central point of decision. Policy is data; each enforcement point decides from a locally held, freshness-stamped copy 🔶 Deviate, deliberately, at this rung We are a central point of decision, and that is the design: one relay, one enforcement point. Their rule exists because their edge must survive disconnection from the network; our edge is a chat/agent client with nothing to do while disconnected from its own community. The deviation is what makes the no-third-shim law free for us rather than expensive. It must be revisited at S2/S3, when a second writer makes local caches into a correctness question
E6 The bus is a live-sharing source, never the enterprise sync engine. The resourced tier subscribes and lifts; it never becomes a routing peer ✅ Comply in shape Ratified below as Law 2. meridian-push-gateway and meridian-control-plane are already separate binaries that subscribe and never route, and git/media/audio already sit off the event path for the same instinct

F · Data shape and process law

# GOAT principle Status Why, or how we deviate
F1 Movement contracts declared by the publisher — persisted / volatile / conflated, with the class deciding congestion behavior ◐ Partial The ephemeral range 20000–29999 already skips storage, audit, and search by compile-time fence — that is volatile in all but name, and it is free correctness. What is missing is the declaration and the drop behavior: today every event shares one undifferentiated queue, so typing indicators can head-of-line block chat. The eight-class QoS map is designed (Zenoh Phase 3), conflated is decided (meridian-9j8), neither is built
F2 Add attributes, not mechanisms. A new capability is a new attribute plus a new event kind, never a new policy engine ✅ Comply, and it is law AGENTS.md § Design Law, and it is the authorization-side twin of this repo's existing "prefer Nostr events over new HTTP endpoints" rule
F3 Reject rule/formula policy expressions — un-auditable by enumeration; per-user/per-object review is NP-complete in general ⛔ Reject, same as they do No policy DSL in the relay. One clarification a reader will otherwise trip over: meridian-workflow embeds evalexpr with a 100 ms timeout, and meridian-acp uses it for subscription rules. Neither decides who may read anything. They are automation triggering and agent-side filtering; authorization never touches an expression evaluator, and must not
F4 No third shim — one grant store, every enforcement point reads it, caches are freshness-stamped projections ✅ Comply Structurally easy for us because there is exactly one enforcement point. The caveat that keeps it honest: membership_cache and accessible_channels_cache are moka caches with a 10-second TTL and invalidation closures — projections with a stated staleness bound, never writers. A mapping table or a sync job between two enforcement points would be a design violation, not an optimization
F5 Nothing is resolved until deployed to a live network and re-probed ✅ Comply, as law AGENTS.md § Design Law. Their two incidents behind it are both named there: a forgeable proxy header granting full admin including the audit stream, and two enforcement flags that were silent no-ops because the deployed image predated the enforcement code. Both are failure modes this repo can reproduce
F6 Every number carries its profile ✅ Comply, and stricter AGENTS.md § Published Throughput Ceilings requires traffic class, auth profile, parsed-or-not, and stored-or-not on every published figure, and forbids quoting the frame-path number for chat
F7 Scope honesty — the pattern is built for the unclassified releasable environment and would be re-examined, not assumed to carry up ✅ Adopt verbatim Inherited honestly. Any pitch of this relay into a higher-sensitivity setting inherits the same caveat

What we take, and what we decline

flowchart TB
    classDef take fill:#1e7a45,stroke:#0f4527,color:#fff
    classDef no fill:#b03030,stroke:#6e1c1c,color:#fff
    classDef law fill:#7d6608,stroke:#4d3f05,color:#fff

    subgraph T["TAKE"]
        direction LR
        T1["Placement discipline<br/>bind at admission or subscribe,<br/>never per message"]:::take
        T2["Labels as closed lanes<br/>per stream, never per message"]:::take
        T3["Grants as enumerable dated rows<br/>with justification and expiry"]:::take
        T4["Envelope encryption to a release-set<br/>O(1) per message"]:::take
        T5["Fail-closed absorbing revoke"]:::take
        T6["Movement contracts + interest routing"]:::take
    end

    subgraph D["DECLINE"]
        direction LR
        D1["Rule/formula policy DSL<br/>un-auditable by enumeration"]:::no
        D2["Row/column masking<br/>we carry events, not tables"]:::no
        D3["Network segmentation as<br/>content access control"]:::no
        D4["Gift wrap on high-rate streams<br/>O(N) per message"]:::no
        D5["A frozen standing claim<br/>inside a credential"]:::no
    end

    T --> LAW["Both lists are already written into<br/>AGENTS.md Design Law and this document"]:::law
    D --> LAW

One thing we decline that they cannot: their two-doors problem is not ours. They cannot fork their router, so machine reads are forced through encryption and a restart-bound static ACL. We wrote the relay. That is a real advantage and it is worth naming — but it is only cashable once D1 exists.


Part III — Throughput, and the case against a central broker

Who this is for, and two rules

Who. This part is decision support for one question: should a workload that is currently on — or being proposed for — a central MQTT broker be carried by the relay instead? It is not a claim that anyone should migrate today, and at current volume nobody has to: human message volume at 10M users is ~3.3k events/s, which one verification pod covers. The section exists so that the argument is on the shelf before somebody has to make the call under schedule pressure, and so the answer is not assumed in either direction.

If no such decision is in flight, the two rows that still pay for themselves are what a broker is genuinely better at and the edge/backbone seam. The rest is dormant until it is needed.

And two rules:

  1. Every figure carries its profile. A bare number misleads in both directions here, because one relay has two ceilings that differ by roughly two orders of magnitude and share a process and almost nothing else.
  2. No head-to-head benchmark exists. Nothing below is a measurement of TBMQ, and nothing below is a measurement of Meridian against TBMQ on one workload. The architectural argument is decidable from shape; the numeric comparison is not, and is marked accordingly throughout.

What a message costs on each path

flowchart TB
    classDef broker fill:#b03030,stroke:#6e1c1c,color:#fff
    classDef mer fill:#1e7a45,stroke:#0f4527,color:#fff
    classDef gate fill:#7d6608,stroke:#4d3f05,color:#fff

    subgraph TB1["TBMQ — central broker, Kafka-backed"]
        direction TB
        K1["MQTT packet parse"]:::broker
        K2["topic ACL check<br/>PER MESSAGE, at the broker"]:::broker
        K3["Kafka produce<br/>MANDATORY HOP — network + durable append<br/>even when pub and sub are the same node"]:::broker
        K4["Kafka consume by the owning node"]:::broker
        K5["topic-tree match over wildcards"]:::broker
        K6["per-subscriber deliver<br/>QoS 1/2 state in Redis / Kafka"]:::broker
        K1 --> K2 --> K3 --> K4 --> K5 --> K6
    end

    subgraph MR["Meridian — relay, Postgres-backed"]
        direction TB
        M1["WS frame + NIP-01 JSON parse"]:::mer
        M2["auth bound at CONNECT<br/>membership cached 10s"]:::mer
        M3["BIP-340 verify<br/>THE COST TBMQ DOES NOT PAY<br/>37.7 us/event at batch 64, MEASURED"]:::gate
        M4["Postgres INSERT<br/>the log, the query store, and the FTS index<br/>ONE durable write, not a second log"]:::mer
        M5["in-process fan-out to LOCAL subscribers<br/>NO BROKER HOP"]:::mer
        M6["Dragonfly PUBLISH only to reach OTHER pods,<br/>only on retained topics — 64x measured scoping"]:::mer
        M1 --> M2 --> M3 --> M4 --> M5 --> M6
    end

The two paths differ in three structural places, and every argument below is one of these three.

TBMQ Meridian
The hop Kafka is on the path for every message, including same-node publisher→subscriber The bus is touched only to reach another pod, and only for topics a pod retained. At one pod it costs nothing
The log Kafka is a durable log alongside PostgreSQL and Redis — truth for "what was published" lives apart from truth for "what exists" The events table is the log, the query store, and the FTS index. One durable write, one answer
The proof The broker vouches for the publisher. Once a message leaves the broker, nothing carries proof of who wrote it Every event is a signed object — id = SHA-256 of the canonical serialization, sig = BIP-340 over it. Authorship survives the relay, survives storage, and is verifiable by a third party years later

That third row is the honest trade: we pay a signature the broker does not, and we get a fact the broker cannot produce. It is what the audit chain, federation, and offline verification all rest on. It is also, at ~37.7 µs, the single largest per-event cost on the persistent path.

Meridian's ceilings, with profiles attached

Profile Figure Basis
Opaque pre-decoded frames · trusted P0 link · payload never parsed · not stored ~600k msg/s Estimate, not measured. The path does not exist in tree; the pre-/post-dedup question behind it is open at ~10×
Signed NIP-01 events · per 16-core pod · BIP-340 bound ~250–320k events/s Measured per-core, extrapolated. just bench: verify_event runs ~30k/s on one core, observed 21.6–33.1k over 3 runs (M3 Max, release, single-thread), so 16 cores tops out near ~480k/s for verification alone (350–530k across that spread)
Stored and indexed chat events · one Postgres ~10–30k/s Estimate. Persistence latency is measured — p95 74.6 ms direct, 77.1 ms nested, synchronous_commit, concurrency 1 (reply-path benchmark). The common reply path is now one data-modifying CTE, not the five or six statements the estimate derives from
Fan-out, per event ~4.1 µs at N=1 → ~56 µs at N=64 Measured. fanout_cost.rs — a fixed ~2–3 µs JSON serialize plus ~430 ns per recipient. Past roughly 40 recipients, fan-out exceeds signature verification
Bus — Dragonfly today 64× cluster ingress reduction from interest scoping Measured. perf/relay_bus_scaling.py --mode redis, 64 communities, 1 subscribed, pods 1/2/4
Bus — Zenoh at 256 B 1.15–1.51× Redis. Target withdrawn Measured, and it failed the ≥5× gate. 1.51× vs Dragonfly-in-Docker, 1.15× [0.98–1.34] vs same-footing host Redis, 0.71× with express; the ≥4 KB crossover is WITHDRAWN — shared memory was on (3,072 B threshold), giving every larger payload a fast path the relay build cannot take. Below 4 KB stands. Opaque bytes, nothing signed/parsed/stored, Python binding — perf/RELAY_BUS_SCALING.md § Phase 0A.2
Symmetric MAC alternative to per-event signatures ~120× (HMAC-SHA256) · ~400× (keyed BLAKE3) than BIP-340 Measured at batch 1, bands ~110–170× and ~360–550×. Compare at a matched batch size — Schnorr is flat across the 1/8/64 sweep, the MAC groups are not. At batch 64: ~200× and ~420× (HMAC-SHA256 0.19 µs, keyed BLAKE3 0.089 µs vs 37.7 µs Schnorr)

The arithmetic that changes the intuition: at 10M users / 1M concurrent, human message volume is only ~3.3k events/s, and even with agents at 5× human volume it is ~20k/s — one verification pod covers it. Fan-out at an average 50 recipients is ~1M frames/s. At vast scale this system is fan-out bound, not crypto bound, which is the opposite of the per-pod ceiling story and is the number that should drive the comparison.

TBMQ's numbers

TBMQ publishes cluster performance results in the range of millions of concurrent connections and order-10⁶ messages per second, for MQTT QoS 0/1 with small payloads on a multi-node cluster with Kafka partitioned behind it.

Those figures are not reproduced here and must not be quoted as though they were. The same rule we apply to ourselves applies to them: take the vendor's own published profile — node count, instance types, QoS level, payload size, persistent versus clean sessions, fan-in versus fan-out ratio — and carry it with the number. A TBMQ throughput figure compared against a Meridian throughput figure measured on different hardware, at a different QoS, with a different durability contract, is not a comparison. It is two unrelated numbers next to each other.

What is decidable without a benchmark is the shape of the work each system does per message, and that is the rest of this section.

Why a central broker is the wrong shape as we scale

Argument 1 is the only irreducible one — a capability a broker cannot acquire without becoming something else. Arguments 2 through 8 are cost, operational, and shape differences, and a sufficiently determined broker deployment can narrow several of them: ACL decisions can be cached, Kafka can be made the only store, topic prefixes can be disciplined. Nothing closes argument 1 short of per-message asymmetric authorship proofs, which is the thing MQTT does not have and Nostr is.

# Argument The mechanism
1 Message authenticity dies at the broker An MQTT message is authenticated by its connection. Downstream of the broker there is no proof of authorship, so a tamper-evident audit chain, third-party verification, and cross-operator federation all have to be rebuilt above the transport. A signed event carries its own proof to every consumer, forever. This is the capability the ~37.7 µs buys
2 Kafka is a second durable log, and we already have one This is refusal R2, written before TBMQ entered the conversation. The events table is already ordered, partitioned, and queryable, and it carries the FTS index and the audit chain. A broker buys retention we have and costs an operational tier plus truth ambiguity — "what happened" answerable from two places is the beginning of every reconciliation job nobody wants to own
3 The broker hop is mandatory; ours is conditional Every TBMQ message crosses Kafka even when publisher and subscriber sit on the same node. Meridian serves local subscribers from an in-process DashMap registry and touches the bus only for cross-pod reach — and then only on topics a pod retained, measured at 64× ingress reduction. As pods scale, the broker's cost grows with total traffic; ours grows with cross-pod traffic, which tenant sharding (A3) actively reduces. Load-bearing caveat: this advantage is a function of connection affinity we do not implement. With naive round-robin load balancing, a channel's members spread across pods and most fan-out becomes cross-pod — the hop we call conditional is then paid nearly always. The measurement that decides it is the cross-pod fan-out ratio under realistic LB, and it does not exist
4 Authorization binds per message at the broker MQTT topic ACLs are evaluated against the topic string on every PUBLISH and SUBSCRIBE. That is exactly the per-message authorization the GOAT placement principle forbids. We bind at connect and at subscribe. Stated honestly: we are better placed, not clean — filter_fanout_by_access is our own residual (A1)
5 QoS is a delivery dial, not a congestion dial MQTT QoS 0/1/2 says how hard to try to deliver. It says nothing about which class to shed under load. A 100 Hz telemetry stream and a chat message are not the same data shape and should not share a movement contract. Our ephemeral range already skips storage, audit, and search by compile-time fence; the eight-class map and the conflated contract are designed. Honest caveat: neither is built, so today we too have one undifferentiated queue
6 Tenancy is a topic-prefix convention MQTT has no tenant boundary in the protocol, so multi-tenancy becomes prefixes plus ACLs — the exact "names encode ownership" failure C3 refuses, where a rename or reorganization becomes a fleet-wide ACL edit. Meridian's community is host-derived, bound before any handler sees data, immutable on the row, and part of every primary key
7 One scaling axis, not five TBMQ scales by adding broker nodes and Kafka partitions — axis A1 plus a durable log. It offers nothing for A2 (peel unbounded derived state into consumers), A3 (shard whole tenants so every kind of a tenant stays co-located), A4 beyond QoS, or A5 (trust-domain federation). Those four axes are what the rest of this document is built on
8 Bridging is O(N²) MQTT bridges are point-to-point configuration. N interconnected networks need N² hand-maintained bridges — the tangle that interest-based transitive routing with releasability-scoped ingress/egress ACLs exists to replace at O(N). Blocked here on D4, and we say so rather than implying it is available

What TBMQ is genuinely better at

A comparison that only flatters one side is not evidence. Five things a broker gives you that this relay does not, and the last three are not close:

Capability Why we do not have it
Broker-held offline queues and persistent sessions MQTT QoS 1/2 with a non-clean session hands a device that has been offline for a day exactly the messages it missed, from broker-held state. A Meridian client reconnects and re-runs a REQ over stored events. That works for a chat client and does not work for a constrained device that cannot page history. This is the single strongest reason to keep a broker somewhere in the topology
Shared subscriptions $share/group/topic load-balances a topic across a consumer group at the broker. We have no equivalent: every matching subscription receives every event. Competing consumers must coordinate themselves
Retained messages and Last Will & Testament Both are broker primitives. Presence with a 90-second TTL does part of LWT's job; NIP-33 addressable events resemble retained messages without being them
Device SDK ecosystem Every embedded platform ships an MQTT client. Almost none ships a Nostr client. A constrained device reaching a Meridian relay needs a bridge or a purpose-built client, and that is a real integration cost
QoS 2 exactly-once We are at-most-once on the bus and idempotent at the store (ON CONFLICT DO NOTHING), which composes to at-least-once-with-dedup end to end. Useful, but not the same contract, and it should not be described as if it were

The shape that actually follows: broker at the edge, relay as the backbone

The conclusion is not "rip out TBMQ." It is the same boundary discipline the GOAT set applies to its own tactical/enterprise split: the resourced tier subscribes and lifts; it never becomes a routing peer.

flowchart LR
    classDef dev fill:#4a4a4a,stroke:#2b2b2b,color:#fff
    classDef edge fill:#1f6aa5,stroke:#123f63,color:#fff
    classDef feed fill:#7d6608,stroke:#4d3f05,color:#fff
    classDef spine fill:#b03030,stroke:#6e1c1c,color:#fff
    classDef bad fill:#8a1c1c,stroke:#500f0f,color:#fff

    DEV["Constrained devices<br/>MQTT clients, intermittent links,<br/>need broker-held offline queues"]:::dev
    BRK["TBMQ at the EDGE<br/>keeps what it is good at:<br/>persistent sessions · QoS 1/2 ·<br/>retained · LWT · shared subs"]:::edge
    FD["FEEDER — the boundary<br/>decode · dedup · compute the routing header ·<br/>sign or MAC · batch"]:::feed
    SP["MERIDIAN SPINE<br/>admission · verification · ordering ·<br/>tenancy fence · audit · fan-out"]:::spine

    DEV --> BRK --> FD --> SP

    X["NEVER: the broker as a routing peer.<br/>A broker outage would become a spine routing failure,<br/>and a compromised broker could transit between communities."]:::bad
    SP -.->|refused| X

Three properties make this the right seam rather than a compromise:

  • The feeder is where the expensive, content-aware work belongs. Decode, deduplicate, and compute the routing header upstream — because a relay that never reads the payload categorically cannot deduplicate by content, and that refusal to read is the enabling property for the opaque-frame ceiling, not a limitation to work around.
  • The broker keeps the device contract it is good at, and the backbone keeps the contract a broker cannot offer: signed authorship, one tenancy fence, a tamper-evident chain, and an event log that is also the query store.
  • The boundary is one-directional. The spine ingests from the edge. The edge never routes for the spine.

The measurement that would settle the numeric half

Nothing above licenses a performance claim. What would:

. ./bin/activate-hermit

# Ours, today — these run and are the basis of every measured figure above
cargo bench -p meridian-core  --bench event_cost      # BIP-340 ceiling, per core
cargo bench -p meridian-relay --bench fanout_cost     # fan-out, N = 1 / 8 / 64
./perf/relay_bus_scaling.py --mode redis              # interest-scoping reduction

# Missing, and required before ANY head-to-head statement
#  1. one workload definition — payload size, fan-out ratio, QoS/durability
#     contract, persistent vs clean sessions, tenant count
#  2. identical hardware for both systems
#  3. Meridian's stored-and-indexed ceiling MEASURED, replacing the ~10-30k/s estimate
#  4. TBMQ run by someone who tunes TBMQ, not by us

Point 4 is not politeness. A benchmark of an unfamiliar system tuned by its competitor is worth nothing, and publishing one would cost more credibility than the number could buy.

Council record — Architecture & Scale

Convened per AGENTS.md § User Preferences. Decision under review: Meridian should be positioned as a replacement for a TBMQ central-broker deployment on the grounds that it removes a mandatory broker hop, a second durable log, and per-message topic-ACL authorization — and this document should publish a throughput comparison supporting that.

Seat Position Strongest objection or named risk What would change it
Systems Architect (chair) Approve with conditions, narrowed "Replacement" invites someone to promise broker-held offline queueing, which we do not have and cannot add without a second durable queue — an axis violation A workload profile dominated by intermittently-connected devices; then the broker stays at the edge permanently and the document must say so
Record Keeper Approve Kafka under TBMQ is precisely R2, a second durable log. If MQTT is ever bridged in, the bridge must be a feeder and must not retain authoritative state A fact Postgres genuinely cannot hold — device offline queues may be exactly that, in which case it is a leaf consumer's store, never a second truth
Capacity Engineer Approve with conditions We have measured verification and fan-out. We have not measured the stored path, the opaque-frame path, or TBMQ at all. Any head-to-head msg/s claim would be an estimate published as a measurement A committed harness running both systems under one workload definition on identical hardware
Protocol Steward (borrowed) Approve Claiming "zero per-message authorization" would be false — filter_fanout_by_access runs per event per recipient. Publish the honest version Moving the access decision fully to subscribe time with fenced invalidation
Provenance Auditor (borrowed) Approve with conditions None of this is deployed or re-probed. The section must read as analysis, not as a migration plan or a benchmark result A live side-by-side probe on one workload

Call — approve with conditions. Publish the comparison as an architecture and cost-model comparison, never as a benchmark result. Every Meridian figure carries its profile and its measured-or-estimated status; every TBMQ figure is labelled vendor-published and unverified here; no head-to-head throughput claim is made. Frame the position as Meridian replaces the broker's bus, policy, and audit role, with a broker legitimate and often correct at the device edge, feeding the spine through a feeder and never as a routing peer. State the per-message authorization residual honestly. Record R2 as the load-bearing refusal.

Overruled: the Capacity Engineer's instinct to publish nothing until a head-to-head exists. The architectural argument — mandatory hop, second log, binding time, traffic classes, scaling axes, federation cost — is decidable from shape and is what a reader needs. Only the numeric claim waits.

Reversal evidence — any one of these reopens the call:

  1. A workload profile dominated by intermittently-connected devices needing broker-held offline queues. Then the broker is not being replaced; it is being kept, and this section is reframed as an integration document.
  2. A committed side-by-side harness showing Meridian's stored path materially below TBMQ's on one workload and one set of hardware.
  3. Postgres stored-event throughput measuring materially below the ~10–30k/s estimate — which would make argument 2 an argument against us until the write path is fixed.

LLM Council — Parts I–III under review

A second review with a different remit. The Architecture & Scale record above gated one decision — the broker call — against a chartered seat list. This one reviews the document itself: whether Parts I–III are true, useful, and honest about what they do not know. It is convened on the advisory roster rather than a charter, so it supersedes nothing; the chartered call stands and this sits beside it.

Seat Position Strongest falsifiable objection What would change it
Contrarian Approve with conditions Part III has no reader who can act on it. It argues against adopting something at a scale we have not reached, for a workload nobody has named. The measured arithmetic in this very document says one pod covers human volume at 10M users. An essay dressed as decision support is worse than no document, because it invites someone to cite it as though a decision were made A named decision in flight — someone proposing MQTT or TBMQ for a real workload. Absent that, Part III must say plainly that it is dormant until needed
First Principles Thinker Approve with conditions Strip the naming and only one of the eight broker arguments is irreducible: a signed event carries authorship proof to every downstream consumer and a broker message does not. The other seven are cost and shape differences a determined broker deployment can partly close — cache the ACL decisions, make Kafka the only store, discipline the topic prefixes. The document leads with the hop and the log and buries the only argument that cannot be answered A broker that ships per-message asymmetric authorship proofs. Then even argument 1 collapses and the comparison is purely operational — which would be worth knowing
Network Engineer & Architect Approve with conditions The "conditional hop" claim is conditional on connection affinity we do not implement. In-process fan-out only beats a broker hop when a channel's members land on the same pod. Under naive round-robin load balancing they do not — members spread, most fan-out goes cross-pod, and the hop we called conditional is paid nearly always. The 64× interest-scoping number is real and measured, but it bounds which topics a pod receives, not how often the bus is on the path A measurement of cross-pod fan-out ratio under realistic load balancing, or community-affinity routing at the edge — which is A3-adjacent and not built
Sr. Full Stack Developer Approve with conditions The write and read paths are traced to real code, which is the half that will stay true — but they are traced to line-level facts (req.rs:90, :209, handlers/event.rs:580) that the next refactor moves. Nothing fails when this document stops being true. check-kinds guards vocabulary because vocabulary drift is silent; pipeline-description drift is silent in exactly the same way, and a confidently wrong diagram costs more than no diagram A doc-drift guard, or citing functions and invariants rather than line numbers. Standing risk either way — worth naming in the verification section rather than pretending it is solved
Expansionist Approve, with a risk to escalate Part II is the more valuable artifact and it is undersold. C8/C9 of the federation compliance criteria are B1/D1/E1 — the register is already most of a conformance package a third party could audit us against, which turns the relay from a product into a pattern others conform to. The risk is the mirror image: a public conformance register with ⚠️ rows is a public defect list, and D4 (timestamp-resurrectable revocation) is a security-relevant gap now written down in prose designed to be read Whether this document is intended to leave the repository. That is a human call, not ours — escalate before any external publication
Outsider Approve with conditions Nobody reads nineteen hundred lines. Everywhere in this document the summary comes after the evidence — thirty rows of register before the five deviations that are the actual answer, eight broker arguments before the one that matters. A reader who does not already know the system has no entry point and will conclude the answer is "it's complicated" A reader who is not us reaching the five deviations without being told where they are
Executor Approve with conditions Parts I–III ship nothing. Of thirty principles and five deviations, exactly one is unblocked work: Scope::all_known() (C2), which gates A2, B4, and every non-first-party traffic class. Deviations 1 and 5 are deliberate, 3 is ratified as build-at-the-second-node, 4 waits on nothing but is cheap. If the register does not say this out loud it reads as thirty items of work and will be planned as such A second unblocked defect surfacing on a re-read. Then the work list is two, and it should say two

Chairman — approve with conditions. The document is true where it claims to be true and marked where it is not, which is the bar. Three of the seven objections are corrections rather than opinions and were applied before this record was written:

# Condition From Status
1 Name Part III's reader, and say plainly that it is dormant until a real decision exists Contrarian applied — § Who this is for
2 Lead the broker arguments with the irreducible one; state that the other seven are partly answerable First Principles applied — authorship is now argument 1
3 Qualify the conditional-hop claim with the connection-affinity dependency and name the missing measurement Network Engineer applied — load-bearing caveat in argument 3
4 Put the five-deviation summary ahead of the thirty-row register Outsider applied — § The five deviations now precedes group A
5 State that the register's actionable work list is one item, not thirty Executor applied — in the same section
6 Add a doc-drift guard, or stop citing line numbers Sr. Full Stack filed as meridian-9ao — a CI change rather than a doc edit, so it gets its own slice instead of a promise in a table
7 Disclosure review before this document leaves the repository Expansionist escalated to the human owner — the council does not decide outward-facing publication

Overruled: the Contrarian's implied stronger claim that Part III should not exist until a decision is in flight. Overruled because the cost asymmetry runs the other way: the argument is cheap to write now and expensive to assemble under schedule pressure, and the failure mode it prevents — adopting a broker by default because nobody had written down what it costs — is exactly the kind of silent, expensive-to-retrofit decision this repo's Design Law exists to catch. The condition, not the conclusion, is what survives: say it is dormant.

Reversal evidence — any one of these reopens this review:

  1. The cross-pod fan-out ratio measures high under realistic load balancing. Then broker argument 3 is materially weaker than written and must be rewritten rather than footnoted — the Network Engineer's objection becomes the finding.
  2. A second unblocked defect surfaces in the register. The "one item of work" claim is the most load-bearing sentence in Part II and the easiest to falsify by re-reading.
  3. The document is approved for external publication. Then Part II's ⚠️ rows need a disclosure pass before it leaves the repo, and D4's phrasing in particular is written for engineers who already have the source.

Part IV — Maturation map — now, next, spine, bus, consumers

The relay is the single source of truth. That sentence is the architecture (ARCHITECTURE.md §1), and it is also the thing that will not survive vast adoption unchanged — not because the relay is slow, but because truth and bloom are different jobs and today one process does both.

Part I described that one process as it stands. This part is the map from there to where it goes: the portfolio below is the cross-project status view, and the spine/bus/consumer design after it is the long-range architecture. It adds no client-facing NIP, no new policy engine, and no new durability system — the same three refusals Part III makes against a central broker, applied inward.


How to read this map

Maturity claims use four deliberately different evidence levels:

Level Means Does not mean
Implemented locally Code and focused/local gates exist in this worktree the immutable release artifact ran in CI or production
Qualified The exact digest passed its provider, upgrade, failure, recovery, and live probes every product surface has graduated
Graduated / supported The preview flag or evaluation warning is gone and the support contract is published the next architecture rung has started
Proposed / parked The design and trigger are recorded a crate, migration, deployment, or customer commitment exists

.settings/epics/ records intent and proof gates, Beads records live task state, and the code plus executed checks record implementation. When they disagree, the least mature reproducible state wins. In particular, a green local just ci is not a deployed proof, and a closed epic is not complete while it has an open acceptance child.

Executive map — where we are and where we are going

flowchart LR
    classDef now fill:#4a4a4a,stroke:#2b2b2b,color:#fff
    classDef gate fill:#b03030,stroke:#6e1c1c,color:#fff
    classDef next fill:#7d6608,stroke:#4d3f05,color:#fff
    classDef target fill:#1e7a45,stroke:#0f4527,color:#fff
    classDef horizon fill:#1f6aa5,stroke:#123f63,color:#fff

    N0["NOW · S0<br/>one Redis-backed relay owns admission, storage,<br/>search, audit, workflows, Git, media, and audio<br/>local hardening is strong; live qualification is incomplete"]:::now

    G1["RELEASE PROOF<br/>exact immutable digests · N/N-1 candidate run<br/>subscriber + alert + CNI drills · provider restore<br/>managed Redis HA · stored-event capacity<br/>production OIDC + provisioned 101 / unknown 404"]:::gate

    G2["PRODUCT GRADUATION<br/>close Projects P0 + structural/e2e gaps<br/>workflows dry-run + emitted history + versioning<br/>mobile preview gates · Forum/Pulse parity<br/>enumerable permissions + user blocking"]:::next

    G3["PORTABLE DISTRIBUTION<br/>profile parity contract · Dokploy single-node<br/>external S3/R2 · recovery · Swarm HA<br/>Cloudflare origin protection and live probes"]:::next

    T1["S1 · SPINE + CONSUMERS<br/>measure first · ingest_seq + commit watermark<br/>search is the first consumer · contract extracted<br/>audit/archive follow only after measured wins"]:::target

    H1["S2/S3 · HORIZON<br/>community sharding · blind consumers + DEK generations<br/>epoch-monotonic revocation before a second writer<br/>bus re-verification · federated networks"]:::horizon

    N0 --> G1
    G1 --> G2
    G1 --> G3
    G2 --> T1
    G3 --> T1
    T1 --> H1

This is a dependency map, not a promise that every box is sequential. Product graduation and portable distribution can run in parallel after the release proof is trustworthy. The consumer plane is intentionally later: its first gate is a measurement, not architectural enthusiasm.

Maturity portfolio — current evidence

Lane Where we are now Smallest next gate Where this is going Evidence
Relay production Redis-backed local hardening is functionally complete; subscriber fencing, N/N-1 harness, recovery contracts, alerts, NetworkPolicy, and capacity contracts exist at different proof levels. It is not production-qualified. Run the final published digest through subscriber-fault, N/N-1, alert, production-CNI, provider restore, managed Redis HA, and stored/authenticated/indexed capacity gates. A repeatable, provider-qualified Redis launch whose readiness, recovery, and ceiling claims are attached to evidence. relay hardening epic; meridian-4z3
Hosted community registration Production candidate. OIDC trust, ownership-safe detach, projection repair, bounded admission, browser E2E, supply-chain workflow, alerts, and local restore rehearsal are implemented. Deploy the exact digest, exercise production OIDC, observe enabled alerts, then re-probe provisioned 101 and unknown-host 404. Controlled/allowlisted GA with relay-authoritative ownership and an operable control plane. registration feature; meridian-zps.9
Portable deployment Compose and Helm exist; the portable profile contract, Dokploy adapter, external S3/R2 qualification, Swarm profile, and Cloudflare edge modes are scoped but not implemented. meridian-dlh.1: freeze the profile matrix and parity gate before adding a provider adapter. Honest evaluation, single-VPS, and HA profiles using the unchanged published relay image, each with exact recovery and failure evidence. portable profiles epic; meridian-dlh
Preview system Stage metadata, stage UI, shared resolution, side-effect registry, feature-contract CI, and desktop panel E2E have landed. Five features remain in the manifest; mobile has no shared flag reader. Add the Flutter reader/gates for Pulse and Forum; then close each feature's open high gaps. Graduation is deletion from preview-features.json, backed by protocol, platform, e2e, failure-path, agent, and docs proof. preview graduation; preview-features.json
Workflows Candidate, closest to graduation. Runtime, CLI, approvals, secrets, and e2e exist; execution history kinds still have no producer. Dry-run before save (W2) and emit the 46001–46012 run/history events (W8); then definition versioning and suspended-run bounds. A recoverable, observable automation surface that leaves preview without hidden execution state. workflow spec; meridian-ezr
Projects / NIP-34 Beta with broad protocol/desktop coverage and major QA/security fixes delivered. A new P0 says the desktop cannot fetch a repo the relay actually hosts when its announcement advertises an external clone URL. The QA epic is closed while that child remains open. Fix meridian-ssx.6, reopen/reconcile the parent task state, then split near-ceiling files and add repo-browser/failure e2e. A graduated project/forge surface with one documented REMAPPING/meridian-desktop/web Git contract and reliable external/relay-hosted interoperability. Projects spec; meridian-ssx.6; meridian-hmd
Pulse / Forum / agent profiles Pulse and Forum are beta on REMAPPING/meridian-desktop/mobile but mobile is ungated; Pulse still polls and vocabulary drifts; Forum lacks mobile voting and abuse semantics. Agent-managed profiles are alpha with no open high gap, but the outcome and off-semantics remain unclear. Close mobile gating first, then Pulse realtime/parity and Forum vote/abuse gaps. Keep agent-profile scope expansion behind a real user need. Truthful cross-platform previews with explicit failure/abuse semantics; agent autonomy that is visible and reversible. Pulse; Forum; profiles
Security and collaboration parity Granular permissions and user blocking are documented P1 gaps; kinds 39003 and 10000 exist without end-to-end readers. Link unfurl, nicknames, masquerade, slowmode, keybinds, and agent distribution remain later slices. Implement enumerable capabilities first, then self-serve blocking; both unlock or de-risk downstream parity work. One dated, enumerable grant model at every enforcement point and a user-visible remedy against abusive members. parity index; meridian-parity.1; meridian-parity.2
Client surface Settings search, shortcut search, audio input selection, and Identity/Git readiness are delivered. Plan/usage, member-visible tool grants, top-level MCP connections, extension distribution, and shared preview artifacts remain. Make agent tool authorization enumerable before adding marketplace/distribution polish. A coherent settings/control surface where identity, tools, limits, integrations, and previews are visible at the scope that owns them. client roadmap
Transport and performance Redis/Dragonfly is the launch transport. Fan-out and signature cost are measured; same-topology Redis-vs-Zenoh, opaque-frame, last-hop QoS, placement, and serialization work remain open. Zenoh is a committed direction as of 2026-08-06 — zenohd routers are the bus end-state — but it is still deferred out of the launch, with no dependency in tree. Qualify Redis HA/capacity for launch; separately complete Phase 0 measurements before authorizing a transport cutover. Phase 0 now sizes the move and licenses the public claim rather than deciding whether to proceed. QoS-classed, interest-routed transport, sequenced by measured evidence against the current path; no non-Redis launch claim by implication, and no bus-layer figure quoted as a relay number. Zenoh spec; meridian-ctg; meridian-vt1
Spine / consumer plane S0. No events.ingest_seq, commit watermark, insert-cost bench, consumer contract, or search consumer exists. Search maintains search_tsv on every insert and adds a NIP-34 trigger/vector path; audit can backpressure behind one bounded worker. Measure search/write amplification and audit queue/lock cost. If material, land the cursor/watermark invariant before the first consumer. S1: bounded spine plus independently deployable leaf consumers that return candidates and publish lag. This document; schema/schema.sql; migrations/0027_nip34_search.sql
Federation / cross-domain access Design law is decided; multi-master enforcement is deliberately not built. NIP-77 backfill, owner grants, participant classes, and monotonic grant merge are parked without a driver. Before the second ownership writer or relay federation link: epoch-monotonic, fail-closed revocation and cross-domain re-verification. S2/S3: shardable communities, blind consumers, and federation without credential-stamped standing or timestamp-based grant resurrection. AGENTS.md § Design Law; meridian-3wh; parked meridian-r9d, meridian-ar6, meridian-u35

Highest-leverage gap register

This register is the cross-project work list. The later consumer-plane register is narrower and remains authoritative for S0 → S1 architecture work.

Priority Gap Why it is next Exit evidence
P0 Local release evidence is being mistaken for production evidence. Relay and registration have strong local proof but the exact deployed digest has not passed every live/provider gate. Every support, security, and capacity claim depends on the binary that actually runs. Published digest; final N/N-1 and subscriber-fault artifacts; provider restore; alert fire→resolve; CNI canary; Redis HA/capacity; production OIDC plus 101/404.
P0 Projects has a newly open fetch failure after its QA epic closed. The repository is present and authenticated, but the main desktop action is unusable; closed-parent/open-child state hides it from a casual maturity read. meridian-ssx.6 accepted, parent state reconciled, relay-hosted/external-clone preference documented and e2e-proved.
P1 Workflows cannot be safely rehearsed and their event history is empty. It is the nearest preview to graduation and the missing dry-run/history pair is user-visible. W2 and W8 closed; versioning and suspended-run bounds proven before graduation.
P1 Mobile ships preview surfaces without the preview contract. The manifest now tells the truth, but users cannot enable/disable Pulse or Forum consistently across platforms. Shared Flutter manifest reader, Riverpod gates, mobile tests, and truthful default behavior.
P1 Permissions and blocking are vocabulary without a complete enforcement/read path. They are the largest current user-safety and dependency-unlocking gaps. Enumerable capability enforcement plus UI/CLI readers; NIP-51 block/mute behavior end to end; just check-kinds green.
P1 Portable deployment support is still a collection of plans. Distribution must not outrun recovery, origin protection, or provider qualification. Profile parity gate, live Dokploy single-node proof, external-state/R2 recovery, then HA/edge drills.
P1 later No correct consumer cursor exists. Starting a consumer without commit-safe ordering can silently miss events; building it before measuring the write-path tax creates inventory. Phase 0 numbers, then slow-commit acceptance test for ingest_seq + watermark.
Trigger-bound Revocation is timestamp-resurrectable at a second writer. Dormant with one writer, catastrophic and silent with two. Epoch-monotonic absorbing revocation lands before the topology adds a second writer or federation peer.

Preview graduation snapshot

Feature Stage Open high gaps Graduation-critical next slice
Workflows candidate W2 dry-run; W8 missing execution-event producers Rehearsable definitions plus event-backed run history
Projects beta P2 five files at the size ceiling; plus P0 ssx.6 live fetch failure Correct the fetch path, split the module, cover repo browse/failure states
Pulse beta PU1 mobile ungated; PU2 polling; PU3 platform vocabulary drift Shared flag gate, realtime subscription, one tab contract
Forum beta F1 no mobile voting; F2 mobile ungated; F5 vote abuse/retraction undefined Cross-platform voting plus explicit abuse semantics
Agent-managed profiles alpha none Visible outcome and documented off-semantics; per-agent scope remains deliberate low priority

No preview is ready to delete from the manifest today. Workflows is closest; Projects has the deepest surface but a live P0 and structural debt; Pulse and Forum cannot graduate while mobile bypasses the preview contract.


The thesis in five sentences

  1. The spine admits, verifies, orders, fences, and fans out. Its state is bounded by the hot set, not by history.
  2. Everything that grows without bound moves to a consumer — an independently deployable process that subscribes by interest, owns its own store, and can be docker run on hardware the spine does not own.
  3. Consumers never authorize. They label, store, and return candidates; the spine re-authorizes every byte before a client sees it. One grant store, one enforcement point — the no-third-shim law survives intact.
  4. The events table is already the log; the bus is only a low-latency tap on it. Durability for consumers is a durable cursor plus backfill, not a broker — which is why Kafka and JetStream are not in this design. The events table being the log is the part that is true today; the cursor is not built (meridian-5riu), so the refusal currently rests on the log, not on the recovery mechanism.
  5. NIP-01 does not move. A client does REQ over WebSocket and cannot tell whether a consumer exists, which is what makes the whole plane deployable without a flag day.

Decision

Adopt a three-plane architecture — spine / bus / consumers — with the consumer contract written down before the second consumer is built.

Plane Owns Scales by State
Spine (meridian-relay) admission, BIP-340 verify, ordering, tenancy fence, subscription fan-out, the hot event set adding pods bounded, shardable by community
Bus (the EventBus seam — zenoh plan Phase 1, not yet in tree; Redis pub/sub today, Zenoh next) delivery between spine pods and consumers, QoS classes, interest routing adding routers none — transport only
Consumers search, archive, feeds, forge, media, observability, apps adding containers, anywhere unbounded, owned per consumer

Kept as-is, deliberately:

Kept Why
Postgres as the ordered event store It is the log. A consumer plane that needs a second durable log has failed.
Redis/Dragonfly for state Presence TTL, rate limits, replay set, fenced-generation arbiter. A pub/sub fabric is not a keyspace.
NIP-01 JSON over WebSocket at the client edge Non-negotiable, and the reason this whole plan needs no client NIP.
The relay as the only enforcement point Consumers returning candidates is the existing meridian-search contract, generalized — not a new idea.

Rejected alternatives:

Rejected Why
Kafka / JetStream as the consumer log A second durable log next to a partitioned Postgres that is already ordered. Buys retention we have, costs an operational tier and a truth ambiguity.
Consumers as Zenoh routing peers A consumer outage becomes a spine routing problem, and a compromised consumer can transit between communities. Consumers are leaves. See Law 2.
Consumers holding their own ACL copy Direct violation of the no-third-shim law. It is the seam along which the human and agent paths drift apart.
A consumer read path for data the spine already stores Already examined and rejected once for NIP-CW: "a second read path is a correctness liability for zero measured gain." See Law 1.
Per-message policy evaluation at the consumer The GOAT finding: rule/formula policy cannot answer "who can read this?" by enumeration. Grants stay dated rows.
Sharding relays by NIP/kind (a grouped "super relay") Reads and invariants are cross-kind, gift wrap encrypts the routing key, NIP-33 needs one arbiter per address, and kind volume is skewed. A kind is a column in every shard, never a shard. Full reasoning: § Scaling axes.

Goals

  1. Nothing unbounded on the write path. Ingest cost per event must not grow with corpus size, repo count, upload volume, or member count.
  2. Consumers are independently deployable and independently failable. A consumer down means a view is stale, never that a message is lost or rejected.
  3. Remote and community-hosted consumers are safe by construction — a consumer that cannot read what it stores is safe to run on hardware you do not control.
  4. Zero client-visible protocol change. No new kind is required to have the plane; kinds are added only where a consumer publishes genuinely new facts.
  5. One grant store, one enforcement point, one audit story across spine and consumers.
  6. Integrations are consumers. No app framework, no plugin runtime, no second auth path — an app is an npub with a dated grant and a declared interest.
  7. Every number in the Scale section is profile-attached per the Published Throughput Ceilings rule.

Non-Goals

  1. Consumers at the client edge. Desktop, mobile, and web speak NIP-01 over WebSocket. A client never connects to a consumer.
  2. Replacing Postgres or Redis. See Decision.
  3. Deciding the Zenoh cutover. Phase 0A has now run and 256 B failed the ≥5× gate (1.15–1.51×); the cutover stays gated on Phase 0B — the native Rust adapter measured on a quiet host — per feature-zenoh-transport.md, and is not re-litigated here. This document specifies what rides the bus, not which bus.
  4. Federation across trust domains. Named as the S3 rung on the ladder, deliberately not designed here — it is blocked on the re-verify branch and the monotonic-revocation work already ratified in Design Law.
  5. A distributed transaction between spine and consumer. Consumers are eventually consistent, with a published lag. See Law 4.

Consumer-plane current state

This section narrows from the portfolio to the S0 → S1 architecture. It is grounded in the tree — every row is reproducible with a grep.

Fact Evidence
The relay orchestrates every subsystem by direct call; subsystems never call each other ARCHITECTURE.md § Crate Dependency Hierarchy — "Cross-subsystem coordination happens only through the relay"
Audit and workflow already leave the direct database path — audit via a bounded mpsc(1000) worker, workflow via a spawned task. Audit enqueue can still backpressure. Search cannot leave the insert: the old search_index_tx indexer is gone and vector maintenance is fused into the write state.rs:712 audit channel + worker; handlers/event.rs:654-680 awaited enqueue; handlers/event.rs:583 — the old search mpsc is gone
The search vector is a generated column on the hot write table, with a GIN index on it schema/schema.sql:211 — search_tsv TSVECTOR GENERATED ALWAYS AS (...) STORED; schema.sql:267 — CREATE INDEX idx_events_search_tsv ON events USING GIN (search_tsv)
NIP-34 search adds a second vector-maintenance path for work-item kinds: a per-row trigger on new writes, followed by an operator-run bounded backfill and partial GIN index migrations/0027_nip34_search.sql; scripts/backfill-nip34-search.sh
meridian-search already returns unauthorized candidates and makes the relay re-authorize ARCHITECTURE.md §6 — "permission filtering is caller's responsibility"
The audit chain serializes on a per-community advisory lock, one round trip per event crates/meridian-audit/src/service.rs:59 — SELECT pg_advisory_lock(hashtextextended($1, 0))
Git already has no authoritative relay-local state — object store is the truth docs/git-on-object-storage.md § Scope and Non-Goals — "no authoritative per-repo filesystem state"; manifest pointer + CAS in § System Model
Two out-of-relay services already exist and work crates/meridian-push-gateway and crates/meridian-control-plane, each a separate [[bin]]
There is no monotonic server-assigned sequence on events schema/schema.sql — PK is (community_id, created_at, id); created_at is client-supplied wall clock
received_at looks like a cursor but is not one received_at TIMESTAMPTZ NOT NULL DEFAULT NOW() — NOW() is transaction start time, so a transaction that starts earlier and commits later lands a lower value after a consumer has already passed it

That last row is the single most important finding in this document. It is a silent-missed-event bug waiting for the first consumer that tries to page through history by timestamp, and it is invisible under single-writer load.

The pattern already exists twice

This is not a new architecture. It is the generalization of two things the repo already ships:

flowchart LR
    classDef done fill:#1e7a45,stroke:#0f4527,color:#fff
    classDef new fill:#1f6aa5,stroke:#123f63,color:#fff
    classDef law fill:#7d6608,stroke:#4d3f05,color:#fff

    G["git-on-object-storage<br/>relay holds NO authoritative state<br/>S3 manifest + CAS is the truth"]:::done
    P["push-gateway<br/>separate binary, own store,<br/>leases, off the hot path"]:::done
    S["meridian-search<br/>returns CANDIDATES,<br/>relay re-authorizes"]:::done

    C["The consumer contract<br/>= these three, written down<br/>and made repeatable"]:::new

    G --> C
    P --> C
    S --> C

    C --> L["Two new design laws<br/>candidates-not-authorization<br/>leaves-not-peers"]:::law

What blooms

The register of unbounded state. "Growth driver" is what makes it unbounded; "today" is where it lives now; "verdict" is whether it belongs on the spine.

# Bloom Growth driver Today Verdict
B1 Search index corpus size × retention search_tsv generated column + GIN on events, plus the NIP-34 nip34_search_tsv trigger/backfill/index path — vector maintenance is on the write path ➡️ consumer. This one has a cost that can be measured today.
B2 Event history users × time Postgres, monthly range partitions ➡️ split: hot partitions on spine, cold to archive consumer
B3 Audit chain 1:1 with events per-community pg_advisory_lock per event, all communities behind one drain worker ➡️ consumer + batch-per-epoch entries
B4 Git objects repos × commits S3/MinIO, manifest CAS ✅ already offloaded — hydrate/pack/CI is the remaining relay CPU
B5 Media blobs uploads Blossom on S3/MinIO ⚠️ stored off-relay, but the relay is still in the byte path
B6 Home / activity feed users × subscriptions fan-out-on-read, assembled at query time ➡️ consumer (materialized, cursor-driven)
B7 Agent observer stream (KIND_AGENT_OBSERVER_FRAME 24200) agents × turns highest-rate kind in the system, ephemeral, shares one bus lane with chat ➡️ consumer on a Drop lane
B8 Workflow runs + traces automation volume Postgres + executor spawned in-relay ➡️ consumer (the runner leaves the relay process)
B9 Huddle audio concurrent calls × bitrate in-relay Opus forwarding ➡️ consumer (media sink), RealTime/Drop
B10 Apps & integrations ecosystem size does not exist ➡️ consumer — this is the whole app model
B11 Push delivery devices × notifications meridian-push-gateway ✅ already a consumer in all but name
B12 Reputation / web-of-trust contributions × time × projects designed, not built ➡️ consumer — never let this land on the spine

Read the verdict column as the work list. B1 is the only one with a cost you can measure today, which is why it is Phase 2 and everything else is later.


Consumer-plane gap register

# Gap Evidence Severity
1 No consumer cursor primitive. Nothing on events is monotonic and server-assigned, so no process can page history without silently skipping commits. schema.sql PK (community_id, created_at, id); received_at defaults to NOW() = txn start high — blocks everything
2 Search indexing taxes writes: generated search_tsv plus GIN maintenance on the hot table; NIP-34 work items also run maintain_nip34_search_tsv() and acquire their partial GIN path after backfill. The aggregate tax is still unmeasured. schema.sql:211, :267; migrations/0027_nip34_search.sql; scripts/backfill-nip34-search.sh high
3 Audit takes a per-community advisory lock per event, and every community drains through a single worker behind a bounded mpsc(1000). Off the direct DB path only until the queue fills — then audit_tx.send().await backpressures the event pipeline itself. Bounded, not solved. audit/src/service.rs:59; state.rs:712; handlers/event.rs:654-680 high
4 No consumer contract exists, so the third out-of-relay service will be designed from scratch a third time. push-gateway and control-plane share no seam medium
5 Consumers would hold cross-community plaintext on shared infrastructure with no encryption boundary. no DEK/generation machinery in tree high (latent) — blocks remote hosting
6 No app/integration surface. Every integration today is either a full agent or a relay patch. no app crate, no install grant kind medium
7 The relay is in the media byte path even though bytes live in S3. meridian-media medium
8 Fan-out is per-recipient with no large-channel strategy: one message to a 10k-member channel is 10k frames. subscription.rs three-tier fan-out; event.rs conn_manager.send_to high at vast scale
9 VISION.md publishes ~600K events/day (~7/sec avg) as the scale target, contradicting AGENTS.md § Published Throughput Ceilings. VISION.md § Scale medium — doc defect, cheap fix

Design

The three planes

flowchart TB
    classDef edge fill:#4a4a4a,stroke:#2b2b2b,color:#fff
    classDef spine fill:#b03030,stroke:#6e1c1c,color:#fff
    classDef bus fill:#1f6aa5,stroke:#123f63,color:#fff
    classDef cons fill:#1e7a45,stroke:#0f4527,color:#fff
    classDef store fill:#7d6608,stroke:#4d3f05,color:#fff

    subgraph E["EDGE — NIP-01 JSON over WebSocket, unchanged forever"]
        direction LR
        D["Desktop<br/>Tauri + React"]:::edge
        M["Mobile<br/>Flutter"]:::edge
        W["Web<br/>repo browser"]:::edge
        A["Agents<br/>CLI · ACP · MCP"]:::edge
        N["Any Nostr client<br/>gitworkshop · ngit"]:::edge
    end

    subgraph S["SPINE — meridian-relay · bounded state · horizontally scalable"]
        direction LR
        AU["community bind<br/>+ NIP-42 / NIP-98"]:::spine
        VF["BIP-340 verify<br/>THE CPU CEILING"]:::spine
        OR["order + assign<br/>ingest_seq"]:::spine
        FO["subscription fan-out<br/>+ re-authorize"]:::spine
    end

    subgraph B["BUS — EventBus trait · Redis today · Zenoh next · QoS classed"]
        direction LR
        BB["interest routing<br/>community / channel / kind"]:::bus
    end

    subgraph C["CONSUMERS — independently deployable · docker run · anywhere"]
        direction LR
        SC["search"]:::cons
        AR["archive"]:::cons
        FD["feed"]:::cons
        FG["forge / CI"]:::cons
        OB["observer"]:::cons
        AP["apps &<br/>integrations"]:::cons
        PG["push<br/>SHIPS TODAY"]:::cons
    end

    subgraph ST["STORES — owned per consumer, never shared"]
        direction LR
        PGSQL[("Postgres<br/>HOT SET ONLY")]:::store
        S3[("Object store<br/>packs · blobs · cold")]:::store
        RD[("Dragonfly<br/>STATE not events")]:::store
        CS[("consumer-owned<br/>whatever fits")]:::store
    end

    E -->|"REQ / EVENT"| S
    S <--> PGSQL
    S <--> RD
    S -->|"live tap · at-most-once"| B
    B --> C
    C -.->|"durable cursor backfill · Background lane"| S
    C --> CS
    C --> S3
    C -->|"results as ordinary signed events"| S

Three arrows carry the whole design:

  • S --> B is non-blocking. The client's OK is already sent.
  • C -.-> S is the backfill, and it is why no broker is needed for durability — once it exists. It does not yet (meridian-5riu).
  • C --> S is the only way a consumer's output reaches a client: as an ordinary signed event through the ordinary ingest path.

The maturation ladder

stateDiagram-v2
    direction LR

    [*] --> S0

    S0: S0 · Monolith relay
    S0: one process does admission, storage, index, audit, git, media, audio
    S0: TODAY · correct · the right shape at this size

    S1: S1 · Spine + consumers
    S1: unbounded state leaves the write path
    S1: consumers are leaves, return candidates, hold a cursor
    S1: ONE deployment, N containers

    S2: S2 · Multi-community shared spine
    S2: community is the shard key and is already immutable
    S2: blind consumers hold ciphertext, trusted consumers hold plaintext
    S2: ONE operator, thousands of tenants

    S3: S3 · Federated networks
    S3: networks federate rather than merge — O of N, not O of N squared
    S3: router peering with per-namespace ingress/egress ACLs
    S3: N operators

    S0 --> S1: ingest_seq + consumer contract
    S1 --> S2: DEK generations + community sharding
    S2 --> S3: monotonic revocation + bus re-verify branch

    S3 --> [*]

    note right of S1
        Each rung is reversible by
        config until the next one lands.
        No rung requires a client change.
    end note

    note right of S3
        BLOCKED by ratified Design Law:
        revocation must be epoch-monotonic
        BEFORE any second writer exists.
    end note

The ladder's discipline: you may not climb a rung until the rung below has a measured number attached to it. S0 → S1 is gated on Gap #2 having a benchmark, not on this document being persuasive.

The five design laws

These are the cheap-now, expensive-later constraints — the same class as the Design Law block already in the root contract, and proposed as additions to it.

flowchart TD
    classDef law fill:#7d6608,stroke:#4d3f05,color:#fff
    classDef bad fill:#b03030,stroke:#6e1c1c,color:#fff
    classDef ok fill:#1e7a45,stroke:#0f4527,color:#fff

    L1["LAW 1 — A consumer may own a read path<br/>ONLY for data the spine does not store."]:::law
    L2["LAW 2 — Consumers are leaves, never routing peers."]:::law
    L3["LAW 3 — Consumers never authorize.<br/>They return candidates; the spine decides."]:::law
    L4["LAW 4 — Consumer staleness is a published number,<br/>not an invisible property."]:::law
    L5["LAW 5 — A consumer outside the trust domain<br/>must not be able to read what it stores."]:::law

    L1 --> F1["Prevents: two answers to one REQ,<br/>and the NIP-CW mistake repeated"]:::bad
    L2 --> F2["Prevents: a consumer outage becoming a<br/>spine routing failure, and cross-community transit"]:::bad
    L3 --> F3["Prevents: the third shim — a second policy<br/>store that drifts from the first"]:::bad
    L4 --> F4["Prevents: 'search is broken' when search is<br/>merely 900ms behind"]:::bad
    L5 --> F5["Prevents: 'run a consumer on a member's box'<br/>silently becoming a tenant breach"]:::bad

    F1 --> OK["The plane stays deployable<br/>without a flag day"]:::ok
    F2 --> OK
    F3 --> OK
    F4 --> OK
    F5 --> OK

Law 3 is not new — it is meridian-search's existing contract, promoted from an implementation detail to a law. That is deliberate: it means the multi-tenant TLA+ isolation model gains one assumption (consumers emit candidates only) rather than needing a new model per consumer.

The consumer contract

classDiagram
    direction TB

    class Consumer {
        <<trait>>
        +interest() KeyExpr
        +cursor() Cursor
        +on_batch(events) Result
        +lag_seconds() f64
        +health() Health
    }

    class Projector {
        <<trait>>
        maintains derived state
        +emit() Vec~SignedEvent~
    }

    class Queryable {
        <<trait>>
        answers what the spine does not store
        +resolve(filter) Vec~Candidate~
    }

    class Cursor {
        +community_id: Uuid
        +ingest_seq: i64
        +watermark: i64
        +persist()
        +gap_detected() bool
    }

    class Candidate {
        NEVER authorized by the consumer
        +event_id: EventId
        +community_id: Uuid
        +channel_id: Option~Uuid~
        +payload: Bytes
    }

    class TrustedConsumer {
        runs INSIDE the trust domain
        holds PLAINTEXT
        NIP-XP profile P0 or P1
        search · feed · workflow
    }

    class BlindConsumer {
        runs ANYWHERE
        holds CIPHERTEXT ONLY
        NIP-XP profile P2
        archive · media · observer replay
    }

    Consumer <|.. Projector
    Consumer <|.. Queryable
    Consumer *-- Cursor
    Queryable ..> Candidate : returns
    Consumer <|-- TrustedConsumer
    Consumer <|-- BlindConsumer

    class Spine {
        THE ONLY ENFORCEMENT POINT
        +re_authorize(Candidate) Option~Event~
        +assign_ingest_seq() i64
        +publish_watermark()
    }

    Spine ..> Candidate : re-authorizes every one
    Spine ..> Consumer : publishes watermark to

Two shapes, and a consumer is usually only one of them:

  • A Projector reads the stream and produces new facts — a feed row, a CI result, a reputation score. Its output re-enters through normal ingest as a signed event. It has no read path at all.
  • A Queryable answers questions about data the spine does not hold. Its output is Candidates the spine re-authorizes. Per Law 1, it may not exist for data the spine already stores.

Write path — where the durability boundary sits

TARGET STATE, NOT AS-BUILT. Everything below this line describes the S1 spine. ingest_seq, the commit watermark and the consumer cursor do not exist: grep -rn ingest_seq over migrations/ and crates/**/*.rs returns zero matches, and the S0 row of the readiness table above says so too. Today's recovery is a client-driven REQ backfill capped at 2000 (crates/meridian-relay/src/handlers/req.rs:25) with no gap detection — so "consumer down? no message is lost" is the property this design would buy, not one the system has. Do not cite the durable cursor as the reason no broker is needed until the column exists (meridian-5riu).

sequenceDiagram
    autonumber
    participant CL as Client
    participant SP as Spine pod
    participant PG as Postgres
    participant BU as Bus
    participant CO as Consumer
    participant CS as Consumer store

    CL->>SP: EVENT — NIP-01 JSON over WS
    SP->>SP: community bind · auth · BIP-340 verify · membership

    rect rgba(176,48,48,0.15)
    Note over SP,PG: THE DURABILITY BOUNDARY
    SP->>PG: INSERT + assign ingest_seq monotonic per community
    PG-->>SP: committed
    SP-->>CL: OK true — the client is DONE here
    end

    par live fan-out, unchanged
        SP->>CL: EVENT frame to matching subscriptions
    and non-blocking tap
        SP-->>BU: publish · postcard + attachment · QoS by kind
        BU-->>CO: delivered if the consumer is up
        CO->>CS: apply
        CO->>CO: advance cursor to ingest_seq
    end

    Note over CO: consumer down? no message is lost.<br/>the cursor simply stops advancing.

    CO->>SP: on restart — backfill from cursor, Background lane
    SP->>PG: SELECT WHERE ingest_seq > cursor AND ingest_seq <= watermark
    PG-->>CO: ordered batch
    CO->>CS: apply · cursor catches up

    opt consumer produces a fact
        CO->>SP: EVENT — ordinary signed event, ordinary ingest
        SP->>CL: fan-out — client cannot tell a consumer produced it
    end

Read step 5–7 as the contract: the OK never waits on a consumer, and step 12 is why that is safe. The consumer plane cannot lose a message because it is not the system of record for one.

The one hard correctness detail

A monotonic ingest_seq is not enough on its own. With multiple spine pods writing concurrently, sequence 105 can commit before 104, so a consumer reading WHERE ingest_seq > cursor at that instant sees 105, advances past 104, and 104 is silently lost forever.

flowchart LR
    classDef bad fill:#b03030,stroke:#6e1c1c,color:#fff
    classDef ok fill:#1e7a45,stroke:#0f4527,color:#fff
    classDef gate fill:#7d6608,stroke:#4d3f05,color:#fff

    A["Pod A takes seq 104<br/>slow transaction"]:::bad
    B["Pod B takes seq 105<br/>commits FIRST"]:::bad
    R["Consumer reads<br/>seq greater than 103"]:::gate
    L["Sees 105 · advances cursor to 105<br/>104 LOST SILENTLY"]:::bad

    A --> R
    B --> R
    R --> L

    W["FIX — read only up to the WATERMARK<br/>watermark = highest seq S where every claim<br/>at or below S has committed or aborted"]:::ok
    R2["Consumer reads<br/>seq in cursor..watermark"]:::ok
    OK["104 and 105 both delivered, in order"]:::ok

    W --> R2 --> OK

One subtlety a future implementer must not blur: the watermark is a sequence fence, but Postgres snapshots fence transactions — and a transaction id says nothing about which ingest_seq values that transaction claimed. The publisher bridges the two by draining:

  1. Sample S = the sequence's last claimed value and snap = pg_current_snapshot() together.
  2. Hold S until pg_snapshot_xmin(pg_current_snapshot()) has advanced past every xid in-progress in snap — at that point every transaction that could have claimed a seq ≤ S has committed or aborted.
  3. Publish S as the watermark.

Two consequences, both cheap to state and expensive to rediscover:

  • Aborted transactions leave permanent gaps below the watermark. A consumer reading (cursor, watermark] gets committed rows only; a missing seq in that range will never arrive. Treat a gap as already resolved — waiting for it to fill is a livelock, and gap_detected() means "backfill the range," never "hold the cursor."
  • One global sequence is enough. Global monotonicity implies per-community monotonicity, so a single BIGINT DEFAULT nextval(...) on the partitioned parent serves every community (Postgres 17, which this repo runs, also permits identity columns on partitioned tables). Per-community sequences are inventory.

Build the watermark publisher; document Postgres logical decoding as the named upgrade (exact commit order, at the cost of a replication slot per consumer) — the same "build flat; document the key tree" discipline the GOAT analysis uses.

Read path — the client never learns a consumer exists

sequenceDiagram
    autonumber
    participant CL as Client
    participant SP as Spine pod
    participant PG as Postgres — hot set
    participant BU as Bus
    participant QC as Queryable consumer

    CL->>SP: REQ · filters · sub_id
    SP->>SP: channel access check BEFORE registering — unchanged
    SP->>SP: classify the filter

    alt filter is inside the hot set — the common case
        SP->>PG: query hot partitions
        PG-->>SP: events
        Note over SP: LAW 1 — no consumer is consulted.<br/>The spine already holds the answer.
    else filter needs data the spine does not store<br/>cold history · full-text · archive
        SP-->>BU: get on the queryable key expression
        BU-->>QC: resolve
        QC-->>SP: CANDIDATES — unauthorized, labelled
        SP->>SP: LAW 3 — re-authorize every candidate<br/>community fence · membership · p-gate · owner gate
    end

    SP->>CL: EVENT frames
    SP->>CL: EOSE
    Note over CL: byte-identical to today.<br/>No new message type, no new kind, no NIP change.

Cursor lifecycle

stateDiagram-v2
    direction LR

    [*] --> Cold

    Cold: Cold start
    Cold: no cursor · full backfill from seq 0

    Live: Live
    Live: consuming the bus tap
    Live: cursor advances per batch
    Live: lag published every tick

    Backfill: Backfill
    Backfill: reading Postgres on the Background lane
    Backfill: bounded batch · never preempts chat

    Degraded: Degraded
    Degraded: lag over SLO · surfaced in the UI
    Degraded: LAW 4 — visible, never silent

    Cold --> Backfill
    Backfill --> Live: caught up to watermark
    Live --> Backfill: gap in ingest_seq detected
    Live --> Backfill: bus reconnect
    Live --> Degraded: lag over SLO
    Degraded --> Backfill: operator or auto recovery
    Backfill --> Degraded: cannot keep up
    Degraded --> [*]: consumer retired · cursor kept

    note right of Live
        The bus is a LATENCY optimization.
        Correctness lives entirely in
        cursor + watermark + Postgres.
        Delete the bus and the plane is
        slower, never wrong.
    end note

That note is the load-bearing property. It is also the argument that settles the broker question: the bus can be Redis, Zenoh, or nothing at all, and the consumer plane stays correct. Which is exactly why this document does not re-decide the Zenoh question.

Trust zones — why a consumer can run on hardware you do not own

The user-facing ask is "launch new docker instances or services remotely." That is only safe with an answer to: what does that box get to read?

flowchart TD
    classDef trusted fill:#1e7a45,stroke:#0f4527,color:#fff
    classDef blind fill:#1f6aa5,stroke:#123f63,color:#fff
    classDef gate fill:#7d6608,stroke:#4d3f05,color:#fff
    classDef bad fill:#b03030,stroke:#6e1c1c,color:#fff

    E["Event leaving the spine toward a consumer"] --> Q{"Does the consumer need<br/>to READ the content?"}:::gate

    Q -->|"yes — search, feed, workflow"| T{"Is it inside<br/>the trust domain?"}:::gate
    Q -->|"no — archive, media, observer replay, packs"| BL["BLIND CONSUMER<br/>encrypt to community DEK<br/>generation-tagged"]:::blind

    T -->|yes| TC["TRUSTED CONSUMER<br/>plaintext · NIP-XP P0 or P1<br/>same operator, same cluster"]:::trusted
    T -->|no| X["NOT ALLOWED.<br/>A plaintext consumer outside the<br/>trust domain is a tenant breach<br/>with extra steps."]:::bad

    BL --> ANY["Runs ANYWHERE.<br/>Member hardware · another cloud ·<br/>a partner's datacentre.<br/>NIP-XP P2, ciphertext only."]:::blind

    TC --> LAW3["Both re-enter through the spine.<br/>LAW 3 — neither authorizes anything."]:::gate
    ANY --> LAW3

The encryption machinery is not invented here — it is GOAT §8: one AES-256-GCM DEK per (release-set, generation) — the community is the release set here — wrapped per recipient, generation tag on every message, revocation is rotation. Cost is indifferent to audience size, and the key service is off the read path so its outage blocks new rotations, never reads. The profile letters are the NIP-XP ladder (feature-zenoh-transport.md § NIP-XP): P0 none (mTLS link identity, same trust domain), P1 mac (HMAC on a handshake key, ~50–200 ns), P2 ed25519 (link-authenticated across operators).

This is what turns Meridian Mesh from "share your GPU" into "share your GPU and your disk" — a member can host the archive consumer for their own community without the community trusting them with its plaintext.

Apps and integrations — zero new mechanisms

Per the Design Law "Add attributes, not mechanisms," an app is not a plugin runtime. It is a consumer with four attributes.

sequenceDiagram
    autonumber
    participant OP as Community owner
    participant SP as Spine
    participant GR as Grant store — the ONE store
    participant AH as App container — a consumer
    participant CL as Client

    OP->>SP: install app — signed event
    SP->>GR: write dated grant row<br/>subject=app npub · since..until · epoch=n
    Note over GR: same store the human and agent<br/>paths already read. NO THIRD SHIM.

    AH->>SP: connect · NIP-98 · declare interest key-expression
    SP->>GR: resolve grant · check epoch · check kind allowlist
    SP-->>AH: admitted as a LEAF at NIP-XP P1 or P2 — never a peer

    loop steady state
        SP-->>AH: bus tap, scoped to the declared interest only
        AH->>SP: output as ordinary signed events · kinds in its allowlist
        SP->>CL: fan-out — indistinguishable from any other member
    end

    OP->>SP: uninstall
    SP->>GR: end-date the grant AND bump epoch to n+1
    Note over GR: monotonic + absorbing.<br/>A concurrent grant at epoch <= n<br/>can NEVER resurrect it.
    SP--xAH: interest revoked at the next fence tick

The entire app model:

Attribute Is Enforced at
Identity an npub — same as a human, same as an agent NIP-42 / NIP-98, unchanged
Installation a dated grant row with an epoch the one grant store
Interest a key-expression under meridian/v1/{community}/** bus ACL + spine fence
Output signed events in a declared kind allowlist ingest, unchanged
Hosting a container the operator or the vendor runs leaf link, NIP-XP P1/P2
Revocation end-date + epoch bump monotonic, absorbing, fail-closed

No app framework. No plugin API. No sandbox runtime. No second auth path. The blast radius of a malicious app is exactly interest ∩ kind allowlist ∩ grant window — three things an operator can read off a screen.

Bus key expressions — one extension

The Zenoh design already proposes the key space. Consumers need exactly one more chunk, and it exists so a consumer's interest is a routing fact rather than a filter applied after the bytes arrive.

flowchart LR
    classDef ns fill:#2b2b2b,stroke:#111,color:#fff
    classDef leaf fill:#1f6aa5,stroke:#123f63,color:#fff
    classDef new fill:#1e7a45,stroke:#0f4527,color:#fff

    R["meridian / v1"]:::ns --> C["community_uuid"]:::ns
    C --> CH["ch / channel_uuid / k / kind"]:::leaf
    C --> G["global / k / kind"]:::leaf
    C --> AU["audio / room_uuid / peer_index"]:::leaf
    C --> CT["ctl / conn · cache · fence"]:::leaf
    C --> CN["cons / consumer_id"]:::new

    CN --> Q["q / queryable_name<br/>the read path"]:::new
    CN --> W["wm<br/>watermark ticks"]:::new
    CN --> L["lag<br/>published staleness · LAW 4"]:::new

cons/** is the only addition, and it is the mechanical expression of Law 4: a consumer's lag is a topic, not a log line.


Scaling axes — how we scale, and how we refuse to

Every scaling move is a choice of axis. Four are sanctioned, one is deferred, and the register below names the moves this architecture refuses — each with its reason and its review-time tell, because the next person to propose one deserves the reasoning, not a bare no. The one-line rule:

Shard by who owns the data and by what job runs — never by what the data is called. A kind is a column in every shard, never a shard.

The sanctioned axes

# Axis You add What stays singular Status
A1 Replica — identical spine pods behind one URL a pod all state (Postgres, Dragonfly); the per-community serialization points ✅ today — cross-pod fan-out is wired
A2 Function — truth on the spine, derived views in consumers a consumer container admission, ordering, re-authorization this document, Phases 1–6
A3 Tenant — shard whole communities a community shard each community lives entirely on one shard: every kind co-located, so filters, thread counters, membership, and the audit chain never cross a boundary S2 — community_id is already immutable
A4 Traffic class — persistent vs ephemeral a QoS lane, an edge pod, a native Zenoh edge for machines and agents the persistent path's durability boundary; the single enforcement point both edges terminate on ephemeral range 20000–29999 already skips storage/audit/search; NIP-SF (meridian-owl) then the NIP-XP machine edge are next. Nostr clients stay on NIP-01/WebSocket permanently
A5 Trust domain — networks federate rather than merge a peered network each network's own truth; per-namespace egress ACLs S3 — deferred, blocked on epoch-monotonic revocation

An axis choice is also a failure-domain choice, and the failure story is the test: A1 loses a pod (nothing happens), A2 loses a view (staleness, published), A3 loses one tenant (the others keep running), A4 loses droppable traffic (by design), A5 loses a peer (the local network stays whole). If a proposed split's failure story is not one of those five sentences, it is on the wrong axis.

The invariants every axis must preserve

All five are already law elsewhere in this document or the root contract; a scaling move that needs to break one is not a scaling move — it is a redesign and gets its own document.

  1. The edge cannot tell (Goal 4) — NIP-01 unchanged; no client learns the internal shape.
  2. One system of record per fact (Decision table; Law 1) — everything else is a cursor-bearing, freshness-stamped projection.
  3. Stale, never wrong (Goal 2; Law 4) — a scaled-out part may lag, with the lag published; it may never lose or reject.
  4. No new authorizer (Law 3; the no-third-shim law) — scale-out never adds a place where policy is decided.
  5. Bounded write path, measured rungs (Goal 1; the ladder discipline; AGENTS.md § Published Throughput Ceilings) — no axis is exercised ahead of a number, and no number ships without its profile.

The refusal register

# Refusal Why it fails here The tell in review
R1 Shard by kind/NIP — "NIP-50 on relay B," grouped as one logical super relay Reads are cross-kind by construction (kinds: [9,7,45001] in one REQ; reply counters materialize on thread roots; reactions join messages; the audit chain covers every kind 1:1). Gift wrap (1059) encrypts the inner kind, so the router cannot see its own key. NIP-33 needs one arbiter per address or LWW resurrects. Kind volume is skewed (24200 dominates) — one hot shard, N idle ones. And no measured wall W1–W6 is relieved. a deployment map, topic plan, or ACL keyed by kind
R2 A second durable log (Kafka, JetStream) the events table is already the log; a second one buys retention we have and costs truth ambiguity plus an operational tier a broker with a retention setting
R3 A second read path for data the spine stores the NIP-CW verdict: "a correctness liability for zero measured gain" (Law 1) a consumer answering filters Postgres already serves
R4 Consumers promoted to routing peers a consumer outage becomes a spine routing failure, and a compromised consumer can transit between communities (Law 2; GOAT §13) a consumer in router peering config
R5 Grant copies in the scaled-out part the third shim: a second policy store drifts from the first, and a concurrent writer resurrects revoked grants until epoch revocation exists an ACL cache with no freshness stamp; a sync job between two stores
R6 Client-side multiplexing as the scale story (the outbox model) it is how open Nostr scales, and it works there precisely because there is no membership, tenancy, ordering, or audit contract — that contract is this product. Capability boxes live behind the spine, invisible at the edge a client holding a per-capability relay list
R7 Parallelizing what serializes by meaning the audit chain, the NIP-33 arbiter, and per-community ordering are serial because their semantics are serial; the sanctioned move is A3 — shard tenants so each serial thing stays whole "shard the chain" / "split the arbiter" proposals
R8 Scale-out ahead of a measurement machinery for load that does not exist is inventory — the phase plan's own gate a new tier whose PR carries no Phase 0 number
R9 Zenoh storages as an event log (zenoh-backend-*) the transport arriving does not make it a store: no transactions, no FTS, no partitioning, no audit chain. This is R2 wearing the spine's clothes, and it is the predictable next proposal once Zenoh is in tree a zenoh-backend- dependency, or a query answered from the bus instead of Postgres
R10 A second edge that authorizes the machine/agent Zenoh edge is a traffic class (A4) terminating on the one enforcement point. An edge that decides access itself is the third shim, and the human and agent paths drift apart along that seam admission or membership logic living in the Zenoh ingress path
flowchart TD
    classDef ax fill:#1e7a45,stroke:#0f4527,color:#fff
    classDef no fill:#b03030,stroke:#6e1c1c,color:#fff
    classDef gate fill:#7d6608,stroke:#4d3f05,color:#fff

    Q{"What is saturated?<br/>A MEASURED number, not a feeling"}:::gate
    Q -->|"nothing measured yet"| R8["STOP — measure first<br/>R8: machinery for absent load<br/>is inventory"]:::no
    Q -->|"stateless CPU<br/>verify · WS terminate"| A1["A1 REPLICA<br/>add a pod"]:::ax
    Q -->|"unbounded derived state<br/>index · archive · feed"| A2["A2 FUNCTION<br/>peel a consumer"]:::ax
    Q -->|"tenant locality<br/>inserts · chain · fan-out"| A3["A3 TENANT<br/>shard whole communities"]:::ax
    Q -->|"loss-tolerant volume<br/>presence · typing · observer"| A4["A4 CLASS<br/>ephemeral lane · NIP-SF"]:::ax
    Q -->|"another operator"| A5["A5 TRUST DOMAIN<br/>federate — BLOCKED on<br/>epoch revocation"]:::gate

    K["'put kind X on its own relay'"]:::no --> R1["R1 — a kind is a column in<br/>every shard, never a shard.<br/>Re-ask which axis you actually meant."]:::no

Scale — where the walls actually are

Per the Published Throughput Ceilings rule, every number carries its profile. Estimates are labelled as such.

# Wall Where it sits Horizontally scalable? This design's answer
W1 BIP-340 verify ~250–320k signed events/s per 16-core pod ✅ stateless — add pods untouched here. The levers skip verification rather than speed it up: NIP-SF session frames (ephemeral kinds drop to a ~150 ns MAC) and NIP-XP P0/P1 profile selection on qualifying links. The batch-verify branch (zenoh Phase 6.1) collapsed on its own feasibility gate — no stable batch BIP-340 API exists (AGENTS.md § Published Throughput Ceilings).
W2 Postgres event persistence throughput still ~10–30k/s per instance (estimate — the zenoh doc's cost budget; no throughput measurement exists). Persistence latency is measured: docs/benchmarks/reply-path-2026-08-01.md — direct p95 74.6 ms, nested p95 77.1 ms at 10 ms configured downstream delay, persistent_chat_synchronous_commit, concurrency 1 ⚠️ shard by community community is already the immutable shard key. Two levers landed: the tuned baseline (meridian-njv, runbook) and the round-trip collapse (meridian-5ox) — the common reply path is now one data-modifying CTE, not the five/six statements this row's estimate was derived from. COPY batching → ~200–500k/s (estimate) remains next.
W3 tsvector + GIN on the write path generated search_tsv taxes every insert; NIP-34 work items also run a trigger-maintained vector path; index cost grows with corpus size ❌ not scalable — it is on W2 B1 → search consumer, only if Phase 0 measures material cost. The first candidate win, still unquantified.
W4 Audit advisory lock one lock round trip per event, per community — one shared drain worker (state.rs:712) ❌ serialized by construction batch one chain entry per epoch, hash-covering the batch. Tamper-evidence survives; the per-event lock does not.
W5 Fan-out amplification measured (meridian-relay/benches/fanout_cost.rs — M3 Max, release, single-thread; WS send excluded): ~4.1 µs at N=1 → ~56 µs at N=64, a fixed ~2–3 µs JSON serialize plus ~430 ns per recipient — one message to 10k members is 10k frames ⚠️ partially the honest open problem. Large channels need fan-out-on-read or a broadcast class. Not solved here.
W6 WebSocket connection count ~50–100k conns/pod (estimate) ✅ add edge pods an edge tier terminating WS separately from the verify tier.
W7 Consumer lag per consumer, published ✅ add consumer replicas Law 4 makes it visible; partition by community to parallelize.
flowchart LR
    classDef wall fill:#b03030,stroke:#6e1c1c,color:#fff
    classDef fixed fill:#1e7a45,stroke:#0f4527,color:#fff
    classDef open fill:#7d6608,stroke:#4d3f05,color:#fff
    classDef plain fill:#4a4a4a,stroke:#2b2b2b,color:#fff

    A["W6 · WS terminate<br/>add edge pods"]:::plain
    B["W1 · BIP-340 verify<br/>250-320k/s per pod<br/>stateless, add pods"]:::plain
    C["W2 · Postgres persistence<br/>10-30k/s est · latency measured<br/>one CTE, not 5-6 statements"]:::wall
    D["W3 · tsvector + GIN<br/>ON the insert"]:::wall
    E["W4 · audit lock<br/>serialized per community"]:::wall
    F["W5 · fan-out · MEASURED<br/>~2-3us serialize + ~430ns x N"]:::open

    A --> B --> C --> D --> E --> F

    FIX1["THIS DOC · B1<br/>search to a consumer"]:::fixed -.removes.-> D
    FIX2["THIS DOC · B3<br/>batch audit per epoch"]:::fixed -.removes.-> E
    FIX3["Zenoh Phase 6<br/>COPY ingest"]:::fixed -.raises.-> C
    FIX4["UNSOLVED<br/>large-channel strategy"]:::open -.->F

Read it honestly: this document removes W3 and W4, and does nothing for W1, W2, or W5. W1 and W2 already have owners. W5 is the wall that vast adoption actually hits first for a social workload, and naming it unsolved is more useful than pretending the consumer plane addresses it.

The arithmetic for "vast"

Sizing at 10M users / 1M concurrent, all figures estimates pending measurement:

Quantity Estimate Implication
Concurrent WebSockets 1M ~10–20 edge pods at 50–100k conns each
Signed events/s at 1 msg/user/5min ~3.3k/s W1 is not the wall — one pod covers it
Same, with agents at 5× human volume ~20k/s still ~1 verify pod; W2 is the wall at one Postgres
Fan-out frames/s at avg 50 recipients ~1M frames/s W5 dominates everything above
Fan-out frames/s with one 10k-member announce channel at 1 msg/s +10k/s from a single channel why large channels need their own class

The conclusion this arithmetic forces: at vast scale Meridian is fan-out bound, not crypto bound. That is the opposite of the current per-pod ceiling story, and it should be said out loud in VISION.md rather than discovered by an operator. fanout_cost.rs now grounds it in a measurement: past roughly 40 recipients the per-event fan-out cost exceeds the 34–37.7 µs signature verify, and the largest single term is re-serializing the event to JSON — exactly the cost a broadcast class or opaque-frame envelope exists to remove.


How this grows and spreads

Three independent vectors. They compound, and only the third needs anything not already designed.

flowchart TB
    classDef v1 fill:#1e7a45,stroke:#0f4527,color:#fff
    classDef v2 fill:#1f6aa5,stroke:#123f63,color:#fff
    classDef v3 fill:#a5641f,stroke:#633c12,color:#fff
    classDef out fill:#7d6608,stroke:#4d3f05,color:#fff

    subgraph A["VECTOR 1 — DEEPEN · one community grows"]
        A1["more surfaces · more agents · more repos"]:::v1
        A2["growth = CONSUMER COUNT, not relay count"]:::v1
        A1 --> A2
    end

    subgraph B["VECTOR 2 — WIDEN · one deployment, many communities"]
        B1["onboard = a DB write + a DNS route"]:::v2
        B2["shared spine · shared BLIND consumers ·<br/>per-community TRUSTED consumers"]:::v2
        B1 --> B2
    end

    subgraph C["VECTOR 3 — FEDERATE · many deployments interoperate"]
        C1["networks federate rather than merge"]:::v3
        C2["O of N shared standards,<br/>NOT O of N squared bespoke bridges"]:::v3
        C1 --> C2
    end

    A2 --> D["THE FLYWHEEL"]:::out
    B2 --> D
    C2 --> D

    D --> E["NIP-01 compliance = distribution<br/>any Nostr client reaches you on day one"]:::out
    D --> F["consumer contract = ecosystem<br/>ship an integration as a docker image,<br/>no relay fork, no permission"]:::out
    D --> G["sovereign default = trust<br/>the exit is real, so the entry is cheap"]:::out

The adoption argument in one line: the relay is the platform, consumers are the app store, and NIP-01 is the distribution channel. Nobody has to adopt all three — the sovereign single-relay operator gets Vector 1 alone and never touches the other two. The vectors are the adoption view of § Scaling axes: deepen rides A2, widen is A3, federate is A5.

What each vector demands that does not exist yet:

Vector Needs Status
Deepen ingest_seq + watermark, consumer contract this document, Phase 1
Widen community sharding, DEK generations, blind consumers designed here, built at S2
Federate monotonic revocation epochs, bus re-verify branch, per-namespace egress ACLs already ratified in Design Law as "decided now, built at the second node"

Consumer-plane maturation plan

Strictly sequenced. Each phase is independently valuable and independently revertible. This is the S0 → S1 plan, not the whole-project release plan above, and no phase starts before the previous one has a number.

Phase What Gate to start New client surface
0 — Measure Benchmark the generated-vector/GIN insert tax, including the NIP-34 trigger path (Gap #2), and the audit lock/queue depth under load (Gap #3). Both against just test infrastructure. — ready; benchmark does not exist none
1 — The cursor primitive events.ingest_seq from one global sequence (global monotonicity covers every community — see the watermark section), plus the sequence-sample + snapshot-drain watermark publisher. Migration + publisher + a test that proves a slow-commit event is never skipped. Phase 0 is material and the first consumer is authorized none
2 — First consumer: search Move search_tsv off the events table into a consumer with its own store. It already returns candidates, so Law 3 costs nothing. Re-run Phase 0's benchmark. Phase 1 none
3 — The contract, written down Extract the seam from Phase 2 into Consumer / Projector / Queryable / Cursor. Retro-fit push-gateway onto it as the proof it generalizes. Phase 2 measured none
4 — Audit + archive Batch audit entries per epoch (W4). Cold partitions to the archive consumer as a Queryable. Phase 3 none
5 — Blind consumers DEK per (community, generation), generation tag on the wire, key service off the read path. Unlocks remote and member-hosted consumers. Phase 3 + a second community exists none
6 — Apps Install grant kind, interest declaration, kind allowlist, epoch revocation. Phase 5 — an app is a consumer and most apps are blind one new kind
7 — Large channels The W5 problem. Fan-out-on-read or a broadcast class. Scoped only after a real channel hurts. evidence, not a roadmap possibly a kind
flowchart TD
    classDef ready fill:#1e7a45,stroke:#0f4527,color:#fff
    classDef gate fill:#7d6608,stroke:#4d3f05,color:#fff
    classDef stop fill:#b03030,stroke:#6e1c1c,color:#fff
    classDef pend fill:#3d3d3d,stroke:#222,color:#fff

    P0["Phase 0 — Measure<br/>tsvector tax + audit queue<br/>READY"]:::ready
    G{"Is the write-path tax<br/>actually material?"}:::gate
    STOP["Stop. Say so in the doc.<br/>A consumer plane for a load<br/>that does not exist is inventory."]:::stop

    P1["Phase 1 — ingest_seq + watermark<br/>THE PREREQUISITE"]:::pend
    P2["Phase 2 — search consumer<br/>the only measured win"]:::pend
    P3["Phase 3 — the contract"]:::pend
    P4["Phase 4 — audit + archive"]:::pend
    P5["Phase 5 — blind consumers<br/>unlocks remote hosting"]:::pend
    P6["Phase 6 — apps"]:::pend
    P7["Phase 7 — large channels · W5<br/>INDEPENDENT of this chain<br/>needs a real 10k-member channel first"]:::pend

    P0 --> G
    G -->|no| STOP
    G -->|yes| P1
    P1 --> P2 --> P3
    P3 --> P4
    P3 --> P5
    P5 --> P6

Once the first durable consumer is authorized, Phase 1 is the one that cannot be deferred. Everything after it is reorderable; retro-fitting ordering onto a running multi-writer consumer system is not. Until Phase 0 measures a material tax and S1 work is authorized, keeping the migration unbuilt is the truthful state rather than a blocker to the Redis launch.


What this means for the VISION documents

The vision docs are strong on what it feels like and thin on what it costs at scale. Six concrete maturations, in priority order:

# Doc Change Why
1 VISION.md § Scale Replace ~600K events/day (~7/sec avg) and the flat metric table with the profile-attached ceilings from AGENTS.md, plus the fan-out-bound finding. A doc claiming 7 events/second cannot justify any of this, and the contradiction with AGENTS.md is already live. Gap #9.
2 VISION.md § Architecture Name the three planes. Today it says "Rust backend, TypeScript clients" — true and uninformative. The spine/consumer split is the architectural claim; it should appear where the architecture is claimed.
3 VISION.md — new § Apps & integrations. Six lines: an app is an npub with a dated grant, a declared interest, and a kind allowlist. It is the largest missing surface and the strongest adoption argument, and it costs no new mechanism.
4 VISION_SOVEREIGN.md § What You Give Up Add the honest cost: derived views are eventually consistent, and their lag is shown to you. Law 4 is a user-visible promise, and this doc's whole credibility rests on stating costs plainly.
5 VISION_MESH.md Extend from compute commons to capacity commons — a member can host a blind consumer, not only a GPU. Blind consumers make it safe, and it is a strictly larger version of the same story with the same trust gate.
6 VISION_PROJECTS.md § CI and Workflows State that the forge's compute is a consumer, not relay work. It is already true for storage (git-on-object-storage); the doc should stop implying the relay builds anything.

Not proposed: rewriting the vision docs' voice. They are good. The gap is quantitative honesty, not prose.


LLM Council

Voice Position
Contrarian The relay is not overloaded. VISION.md says 7 events/second. Building a consumer plane for a load nobody has measured is inventory, and every consumer is a new failure mode, a new runbook, and a new "why is search stale" ticket. Two things here are real today and both are one-line greps: the tsvector+GIN tax on every insert, and the per-event audit advisory lock. Do those. The relay's actual superpower is that it is one process you can reason about — spend that slowly.
First Principles Strip the naming and the ask is three things: (a) get unbounded state off the write path, (b) make derived views independently deployable, (c) do not create a second read path. Only (a) is urgent, and it is solvable without a consumer plane at all — move the tsvector out of the generated column. The genuinely important observation in this document is that the events table is already the log, so no broker is needed for durability. Say that loudly and repeatedly, or someone will bring Kafka.
Sr. Full Stack Developer ingest_seq + watermark is the whole thing. It is a migration, a column, a publisher, and one nasty test — and it is the only piece that cannot be retro-fitted onto a running system. The received_at-is-not-a-cursor finding is the real bug in this document; it would have shipped, and it would have been diagnosed as "search randomly misses messages" six months later. Also: do not invent a second seam. EventBus from Phase 1 of the Zenoh plan is the seam.
Network Engineer & Architect Four transports if this lands as written — WebSocket edge, Redis bus, iroh mesh, Zenoh — plus consumer links. Four identity systems, four metric surfaces, four on-call runbooks. Either consumers ride the same session as the pod bus with key-expression separation, or this is unmanageable. And Law 2 is not stylistic: GOAT §13 already ratified that the enterprise tier "subscribes to and ingests from" the fabric and "never becomes a routing peer." A consumer as a Zenoh peer means a consumer's routing table is a spine problem and a compromised consumer can transit between communities. Leaves. Always.
Expansionist The consumer contract is the product. The moment a third party can docker run meridian-consumer-jira against any compliant relay, this competes with GitHub Apps on the axis GitHub cannot move on — the operator owns the runtime. Publish the contract as a NIP so consumers work against any relay, not just this one. Ship the contract before the second consumer, not after.
Executor Seven phases is four quarters, and only three of them are load-bearing: measure, ingest_seq, search consumer. Everything after Phase 3 is the same shape repeated. Do not start with git or media — they are already offloaded, so the work is visible and the win is zero. And put a hard stop on Phase 0: if the tsvector tax measures under a few percent, publish that number and close the epic.
Outsider Nobody adopts a chat platform because of a consumer plane. They adopt because myproject.com works and their agent can talk to it. The user-visible value here is exactly two sentences — "search never slows down chat" and "your integration does not need our permission" — and if those two sentences are not in VISION.md, no one reads past the first diagram.
UI/UX Expert Eventual consistency is a UI contract, not a backend detail. VISION_ACTIVITY.md already ratified the rule: "the absence of an event is itself information… never an empty void." A consumer 900ms behind must render as 900ms behind. Law 4 is the backend half of a principle this project already committed to on the front end — make the lag a topic (cons/{id}/lag) so the client can actually show it.
Chairman Approve the contract; gate the build behind release truth and measurement. The immediate sequence is the exact-digest/provider proof, the open Projects P0, and the nearest preview graduation gaps. Preserve the five consumer laws now because they are cheap constraints, but do not confuse them with authorization to build S1. Measure the complete search write path first. If it is material and a first consumer is authorized, land ingest_seq + watermark before that consumer and make search the only initial extraction; received_at remains a trap, but an unused cursor migration is inventory. Fix VISION.md's scale table because a document claiming seven events per second cannot also justify a consumer plane. Defer federation exactly as Design Law already defers it: decided now, built at the second node.

Verification

Commands that actually run in this repo:

. ./bin/activate-hermit

# Phase 0 — the two measurements that gate everything
cargo bench -p meridian-db --bench insert_cost     # Phase 0 DELIVERABLE — bench does not exist yet; include generated FTS, NIP-34 trigger, and index variants
just test                                                # audit lock queue depth under integration load
./perf/relay_bus_scaling.py --mode redis                 # existing baseline

# Phase 1 — the acceptance test that matters
cargo test -p meridian-db ingest_seq_watermark_never_skips_a_slow_commit

# Phase 2 — the win must be re-measured, not assumed
cargo bench -p meridian-db --bench insert_cost      # compare against Phase 0
cargo test -p meridian-search

# Every phase
just ci
just test        # integration — this touches db, pubsub, and relay
just check-kinds # no vocabulary without an enforcement point

Green CI closes no phase. Per AGENTS.md § Design Law, a phase is done only when its migration or consumer is deployed against a live relay and re-probed — a deployed image predating the enforcement code makes any flag a silent no-op.

Metrics this plane must add (names follow the relay's existing meridian_* convention — subscription.rs:94, state.rs:1274):

  • meridian_consumer_lag_seconds{consumer,community} — Law 4, and the only metric a client is allowed to see indirectly
  • meridian_consumer_cursor_seq{consumer,community} vs meridian_ingest_watermark{community} — the gap is the backlog
  • meridian_consumer_backfill_total{consumer,reason} — reason distinguishes a cold start from a gap, and a rising gap rate is the bus failing quietly
  • meridian_candidates_rejected_total{consumer} — Law 3's enforcement point. A consumer whose candidates are never rejected is a consumer that is authorizing. This metric is how the law is audited rather than assumed.

Open questions

  1. Does Gap #2 measure? The generated tsvector/GIN tax and NIP-34 trigger path are architecturally visible and numerically unmeasured. Phase 0 either substantiates S1 extraction work or parks it. Nothing in the consumer-plane phase chain should be built before that number exists.
  2. Is the watermark publisher enough, or is logical decoding required from the start? The sequence-sample + snapshot-drain publisher is correct but polls every tick. Logical decoding gives exact commit order for free at the cost of a replication slot per consumer. Measure before choosing.
  3. W5 — large-channel fan-out. The honest wall at vast scale, unaddressed here. Fan-out-on-read for large channels, a broadcast channel class, or something else? Needs a real channel to hurt first.
  4. Does the app model need a kind, or is a grant row enough? Phase 6 assumes one new kind for install. Per just check-kinds, it does not get merged until something reads it.
  5. Do blind consumers break search forever? Encrypted archives cannot be full-text searched by the consumer that stores them. Either cold search stays trusted-plane, or history search degrades to metadata beyond the hot window. This is a real product tradeoff, not just an engineering one.
  6. Does the consumer contract belong in a NIP? The Expansionist case is that ecosystem value requires portability across relays. The counter is the repo's own rule against vocabulary without an enforcement point. Revisit after Phase 3.