R2D2-MERIDIAN/ARCHITECTURE.md
Joshua Belke 610b9cac2e
Some checks failed
helm chart / lint + unittest + render matrix (push) Has been cancelled
helm chart / install on kind (gated) (push) Has been cancelled
helm chart / publish chart to GHCR (push) Has been cancelled
Meridian Harness / Build (aarch64-unknown-linux-musl) (push) Has been cancelled
Meridian Harness / Build (x86_64-unknown-linux-musl) (push) Has been cancelled
Meridian Harness / Publish rolling release (push) Has been cancelled
Meridian Harness / Publish tagged release (push) Has been cancelled
CI / Detect Changed Paths (push) Has been cancelled
CI / Dead Token Reference Guard (push) Has been cancelled
Docker image / Build (linux/amd64) (push) Has been cancelled
Docker image / Build (linux/arm64) (push) Has been cancelled
Docker image / Build public push gateway (linux/amd64) (push) Has been cancelled
Docker image / Build public push gateway (linux/arm64) (push) Has been cancelled
CI / Rust Lint (push) Has been cancelled
CI / Unit Tests (push) Has been cancelled
CI / Isolated DB Gate (push) Has been cancelled
CI / Desktop Core (push) Has been cancelled
CI / Desktop Smoke E2E (1) (push) Has been cancelled
CI / Desktop Smoke E2E (2) (push) Has been cancelled
CI / Desktop Smoke E2E (3) (push) Has been cancelled
CI / Desktop Smoke E2E (4) (push) Has been cancelled
CI / Desktop (push) Has been cancelled
CI / Desktop E2E Relay (push) Has been cancelled
CI / Desktop E2E Integration (1/2) (push) Has been cancelled
CI / Desktop E2E Integration (2/2) (push) Has been cancelled
CI / Desktop E2E Integration (push) Has been cancelled
CI / Backend Integration (relay e2e) (push) Has been cancelled
CI / Relay E2E (push) Has been cancelled
CI / Web (push) Has been cancelled
CI / Admin Web (push) Has been cancelled
CI / Mobile (push) Has been cancelled
CI / Security (push) Has been cancelled
CI / Server Cross-Compile (push) Has been cancelled
CI / Server Cross-Compile-1 (push) Has been cancelled
CI / Windows Rust (x86_64-pc-windows-msvc) (push) Has been cancelled
CI / Desktop Build (macOS) (push) Has been cancelled
Docker image / Merge release multi-arch manifest (push) Has been cancelled
Docker image / Merge debug multi-arch manifest (push) Has been cancelled
Docker image / Publish public push gateway image (push) Has been cancelled
docs: withdraw the >=4 KB crossover — shared memory was on for every measurement
Both SHM keys default true in Zenoh 1.8, the eclipse-zenoh wheel IS built
with the shared-memory feature, and transport_optimization's threshold is
3,072 B. So every measured payload at 4 KB and above took a POSIX-SHM fast
path -- one the relay build cannot take, because shared-memory is not in
its zenoh feature list -- while the Redis side had no equivalent. The error
points toward Zenoh.

RELAY_BUS_SCALING.md's own bullet said "Nothing about shared memory. SHM
was not enabled." That was false, and nothing could have caught it: the
static contract checked transport.shared_memory.enabled and was blind to
the second switch beside it. This adds that check.

The "Router mode prices the Docker boundary" reading goes with it. SHM
works host-to-host and cannot cross into the Docker VM, so an unknown
share of the peer-vs-router divergence at >=64 KB is the SHM path dropping
out rather than the boundary appearing. Both readings are unlicensed until
re-measured under the shipped posture.

Below 4 KB stands, including the 256 B verdict that failed the >=5x gate:
256 B is well under the SHM threshold, and gossip's extra transports cost
the measured process work rather than saving it.

Worth stating plainly, because it is the second time: the 0A.2 result has
now been invalidated twice for two unrelated reasons -- unmatched publish
semantics, then transport posture -- and neither was visible in the
numbers. A bus measurement is not licensed by its spread. It is licensed by
its posture being pinned, read back, and stamped beside the result, which
is what af4b92eba now does.

Signed-off-by: Joshua Belke <joshua@innovationhub-act.org>
2026-08-21 00:54:18 -04:00

66 KiB
Raw Permalink Blame History

Meridian Architecture

1. Executive Summary

One binary. Two data planes. Meridian is a self-hosted relay carrying two kinds of traffic that share a process and almost nothing else: signed, stored, indexed events for humans and agents, and opaque, ephemeral, header-routed frames for machines.

This document describes the human and agent plane, which is built on the Nostr protocol (NIP-01 wire format), where AI agents and humans are first-class equals. Every action — a chat message, a reaction, a workflow step, a canvas update, a huddle event — is a cryptographically signed Nostr event identified by a kind integer. Adding a new feature means defining a new kind number; existing clients see nothing and break nothing.

The two planes have ceilings roughly two orders of magnitude apart, and the gap is architectural rather than incidental. Never quote a throughput number without its profile attached — see DIAGRAM.md and the published ceilings table in AGENTS.md.

The relay is the single source of truth. All reads and writes flow through it. There is no peer-to-peer event exchange, no gossip, no replication — just clients connecting to one relay over WebSocket, and the relay enforcing auth, verifying signatures, persisting events, fanning out to subscribers, indexing for search, and triggering automation.

A Meridian community is the tenant-visible workspace selected by the request host. The self-hosted default remains one host, one relay process, one implicit community. Multi-community deployments move that semantic boundary one level up: req.community = resolve_host(connection.host) is established before AUTH, EVENT, REQ, REST, media, git, search, workflow, or pub/sub handling. Unknown hosts fail closed, and NIP-98/API-token stamps must agree with the host-derived community rather than overriding it.

Meridian is a Rust monorepo of 29 crates, licensed MIT and built by STELLAR. It lives at git.office.ilab.zone/RAID/R2D2-MERIDIAN.

This document is the component reference — what each crate is, what it does, what it explicitly does not do. For the system view built on top of it — how the client, relay, and stack tiers compose, the GOAT/EFDI conformance register, published throughput ceilings with their profiles, and the spine/bus/consumer maturation road — see DIAGRAM.md.


System Architecture

┌─────────────────────────────────────────────────────────────────────┐
│                           CLIENTS                                    │
│                                                                      │
│  Human (Nostr app, web, mobile)    Agent (CLI tools via meridian-cli)    │
│           │                                    │                     │
│           └──────────── WebSocket ─────────────┘                    │
└─────────────────────────────────────────────────────────────────────┘
                               │
                               ▼
┌─────────────────────────────────────────────────────────────────────┐
│                         meridian-relay (Axum)                          │
│                                                                      │
│  ┌──────────┐  ┌──────────┐  ┌──────────┐  ┌─────────────────────┐ │
│  │ NIP-42   │  │  EVENT   │  │   REQ    │  │  HTTP bridge       │ │
│  │  auth    │  │ pipeline │  │ handler  │  │ /events            │ │
│  └──────────┘  └──────────┘  └──────────┘  │ /query             │ │
│                                             │ /count             │ │
│  ┌──────────────────────────────────────┐   │ /hooks/{id}        │ │
│  │       SubscriptionRegistry           │   │ /media/*           │ │
│  │  DashMap: (channel_id, kind) → conns │   │ /git/*             │ │
│  └──────────────────────────────────────┘   │ /info, NIP-05      │ │
│                                             └─────────────────────┘ │
└─────┬───────────────┬───────────────────────┬───────────────────────┘
      │               │                       │
    TRUTH        STATE + BUS                BYTES
      │               │                       │
┌─────▼──────┐  ┌─────▼──────┐        ┌───────▼────────┐
│  Postgres  │  │ Dragonfly  │        │  MinIO / S3    │  [in tree]
│  events    │  │ presence   │        │  media blobs   │  relay → S3
│  channels  │  │ SET EX     │        │  git CAS       │  direct
│  tokens    │  │ typing     │        └───────┬────────┘
│  workflows │  │ ZADD       │                │ replication
│  audit·FTS │  │ PUBLISH    │        ┌───────▼────────┐
└────────────┘  └─────┬──────┘        │  IPFS          │  [planned]
   [in tree]          │               │  CID replicas  │
                      │               └────────────────┘
              ════════▼════════
               EventBus seam          ┌────────────────┐
               (Phase 1 · 8rn)        │  ReductStore   │  [opt-in ·
                      │               │  REST · FIFO   │   no call
              ┌───────▼────────┐      │  recordings ·  │   sites]
              │ Zenoh peer     │      │  telemetry     │
              │ IN-PROCESS in  │      └────────────────┘
              │ every relay pod│       artifact axis — parallel to
              │ (Phase 2·bem)  │       S3, NOT in the media/git path
              └───────┬────────┘
                      │  pod <-> pod DIRECT — no broker hop.
                      │  This is the destination, not a waypoint.
                      ┆
              ┌───────┴────────┐
              │ zenohd router  │      [Phase 4 · 0ps — OPTIONAL]
              │ scale-out and  │      added only where a full peer
              │ cross-region   │      mesh stops being viable
              └────────────────┘

     Legend — [in tree] runs today · [opt-in] shipped but unwired
     · [committed] decided, not yet built · [planned] not yet specified.
     Nothing below the EventBus seam exists in the workspace today:
     there is no `zenoh` dependency in any Cargo.toml.

     Dragonfly speaks the Redis wire protocol (reports
     redis_version:7.4), so REDIS_URL, the `redis` crate,
     deadpool-redis, Lua Script calls and pub/sub are all
     unchanged. The compose service is `dragonfly`, the
     container `meridian-dragonfly`, and a `redis` network
     alias keeps in-network redis://redis:6379 resolving.

     Dragonfly keeps STATE after the bus moves. Presence TTLs,
     rate-limit INCR+EXPIRE, the NIP-98 replay set and the fenced
     arbiter all need a keyspace with atomics; Zenoh is a pub/sub
     fabric, not a keyspace. Only PUBLISH/PSUBSCRIBE migrates.

     Fan-out: sub_registry.fan_out() → conn_manager.send_to()
     (in-process for local events; bus round-trip for
     events from other relay instances)

     PUBLISH occurs for channel-scoped events.
     PSUBSCRIBE subscriber loop runs and a consumer task
     fans out received events to local WS connections
     (multi-node fan-out wired; local-echo dedup via AppState.local_event_ids).

     ┌──────────────┐
     │  Postgres    │  ← meridian-search (FTS over the search_tsv
     │ (full-text   │     generated column + GIN index)
     │   search)    │
     └──────────────┘

The bus seam — why Zenoh sits below EventBus, not beside Dragonfly

The client edge is NIP-01 JSON over WebSocket and does not change; this is a transport swap behind the relay, on the A1 replica axis, and the edge cannot tell the difference afterward. The sequence is fixed by .settings/features/feature-zenoh-transport.md: an EventBus trait seam first (meridian-8rn), a ZenohEventBus in peer mode behind an env gate with dual-publish shadow mode (meridian-bem), QoS classes, then zenohd routers for scale-out and cross-region (meridian-0ps).

The Zenoh peer runs inside the relay process — that is the whole point. Today every cross-pod event takes two hops (pod A → Dragonfly → pod B). In peer mode pods connect directly and the broker hop disappears, which is a latency win independent of any throughput figure. zenohd is not the destination and is not required for the win: it is added only where a full peer mesh stops being viable, because peer links grow with the square of the pod count. A small or single-region deployment may never run a router at all.

Two constraints travel with the router phase and are not optional:

  • zenohd goes under a new opt-in bus compose profile, default off. Core services must stay profile-free — just _ensure-services blocks on core health and would hang 120s against a disabled service.
  • A router that spans trust domains makes cross-node re-verification mandatory. Today cross-node events are accepted without re-running verify_event, which is safe only because every writer is inside one trust domain. Federating a router without that branch is an event-injection primitive: anything reaching the router is delivered to clients as authentic.

Zenoh as the spine — surface by surface

Zenoh is the transport spine for everything except the Nostr client edge. That exception is not a compromise; it is what lets the spine exist at all, and it keeps Nostr interoperability — the project's stated core value — intact.

Surface Zenoh's role Scaling axis Gate before it lands
Inter-pod bus replaces Dragonfly PUBLISH/PSUBSCRIBE, peer mode in-process A1 replica meridian-ctg → 8rn → bem
Machine / IoT / agent ingress native Zenoh edge, NIP-XP profile — not NIP-01 A4 traffic class NIP-XP + NIP-SF (meridian-owl)
Relay ↔ relay mesh candidate consolidation with the iroh QUIC mesh A1 replica Phase 5 — must be decided, not deferred
Cross-org federation zenohd routers peering across trust domains A5 trust domain blocked on epoch-monotonic revocation
Nostr client edge none — NIP-01 JSON over WebSocket stays — permanently out of scope (Non-Goal #1)

Two edges, one relay. Desktop, mobile, web and any third-party Nostr client speak NIP-01 over WebSocket; machines, feeders and server-to-server peers speak Zenoh natively. Both terminate on the same enforcement point and the same event log. A second edge is a traffic class, not a second authorizer — the moment it decides access on its own it has become one, which the no-third-shim law refuses.

Three properties keep this honest and are checkable in review:

  • Zenoh adds no event kind and no client-facing NIP. NIP-SF adds a message type (SFRAME) in the existing ephemeral range, not a kind. A kind without a reader is not progress.
  • Zenoh is never a store. zenoh-backend-* storages exist and will be proposed; they are not a transactional store with FTS, partitioning and an audit chain. Postgres remains the event log — one system of record per fact.
  • Dragonfly is not displaced, only narrowed. Presence TTLs, rate-limit INCR+EXPIRE, the NIP-98 replay set and the fenced arbiter need a keyspace with atomics. Zenoh is a pub/sub fabric, not a keyspace.

Throughput claims — what is measured and what is not

Published Zenoh benchmarks (~4–5M msg/s at small payloads, peer and brokered) are measured against Kafka, MQTT and CycloneDDS — not against Dragonfly, which is multi-threaded and Redis-wire-compatible and therefore a far closer baseline than the ~35K msg/s MQTT figure those charts compare to. Meridian's own gate was ≥5× over the current Redis path on one fixed topology at 256 B (meridian-ctg). That comparison has now been run, and 256 B failed it: 1.51× against Dragonfly, 1.15× against a same-footing host-native Redis, 0.71× with express. The ≥4 KB crossover is withdrawn: shared memory was enabled throughout (both SHM keys default true in Zenoh 1.8 and the eclipse-zenoh wheel is built with the feature), and transport_optimization's threshold is 3,072 B — so every point above 4 KB took a POSIX-SHM fast path the relay build cannot take, and the error points toward Zenoh. Below 4 KB stands (~2× at 2 KB). Full profile in perf/RELAY_BUS_SCALING.md § Phase 0A.2; it is a directional Python-binding probe, not the Rust adapter's capacity, and it licenses no ≥5× claim at any payload size.

Which path you are on decides whether any of it matters. There are two, and they have different ceilings for different reasons:

  • Signed persistent events are bounded by BIP-340 verification at ~30k events/s per core (observed 21.6–33.1k over 3 runs; just bench), and Postgres-bound below that. A faster bus does not move this number. Quoting a bus-layer figure here quotes the pipe, not the plumbing.

  • Machine, agent and realtime traffic — huddle audio, agent-observer frames (24200), typing (20002), presence (20001), IoT feeder frames — is high-rate, drop-tolerant, never federated and never stored. This is where the order of magnitude actually is, and it is measured: replacing BIP-340 with a symmetric MAC on this class runs ~120× faster with HMAC-SHA256 (3.67 M/s at batch 1) and ~400× faster with keyed BLAKE3 (11.9 M/s at batch 1), per crates/meridian-core/benches/event_cost.rs (M3 Max, release, one core; just bench). The ratios are stated to two significant figures on purpose: they divide by the BIP-340 number, whose observed spread is 21.6–33.1k/s, so the honest bands are ~110–170× and ~360–550×. They were previously quoted as exactly 125× and 405× against a single 29.4 K/s endpoint (meridian-9fqe).

    The batch size is part of the figure, not a footnote. The bench sweeps 1/8/64 and the MAC groups are not flat across it — HMAC-SHA256 runs 3.67 M/s at batch 1 and 5.27 M/s at batch 64, so quoting a MAC throughput without its batch size is the same class of error as quoting a throughput without its traffic class. At batch 64 the matched ratios are ~200× and ~420× against Schnorr's 26.5 K/s. Per-event costs quoted elsewhere in this repo — 0.19 µs HMAC and 0.089 µs keyed BLAKE3 — are batch-64 figures, and TASKS.md § 3.1 labels them as such.

Those two facts set the division of labour for the transport work, and the plan is deliberately two-headed. Zenoh is what makes the fast class separable — eight QoS lanes instead of one undifferentiated queue, so typing indicators stop head-of-line blocking chat and agent frames get their own drop policy. NIP-SF/NIP-XP are what remove its ceiling, by making signature cost a function of a message's destiny rather than a global setting. Neither half delivers the machine-path number alone.


Crate Dependency Hierarchy

meridian-core    (zero I/O — types, verification, filter matching, kind registry)
    │
    ├── meridian-db          (Postgres: events, channels, tokens, workflows, audit)
    ├── meridian-auth        (NIP-42, NIP-98, API tokens, scopes, rate limiting)
    ├── meridian-pubsub      (Dragonfly pub/sub, presence, typing indicators)
    ├── meridian-search      (Postgres FTS: query, delete)
    ├── meridian-audit       (hash-chain tamper-evident log)
    ├── meridian-media       (Blossom/S3 media storage)
    └── meridian-workflow    (YAML-as-code automation engine)
         │
         └── meridian-relay       (ties everything together — the server)

Agent surface
  meridian-acp            (harness — bridges relay @mentions → AI agents via ACP/JSON-RPC)
  meridian-agent          (minimal ACP-compliant agent — non-streaming, tool-calls-as-output)
  meridian-dev-mcp        (developer MCP server — shell + file-edit tools)
  meridian-persona        (agent persona packs)
  meridian-harness        (all-in-one bundle of ACP, agent, and dev MCP)

Clients + interop
  meridian-pair-relay     (ephemeral sidecar relay for NIP-AB device pairing)
  meridian-pairing-cli    (CLI for NIP-AB device pairing interop testing)
  meridian-ws-client      (shared NIP-42 WebSocket client — connect, auth, publish)
  git-sign-nostr          (sign git objects with a Nostr key)
  git-credential-nostr    (git credential helper for Nostr-authed push/fetch)

Distribution + operations
  meridian-relay-mesh     (pod-to-pod mesh within one deployment; not federation)
  meridian-control-plane  (control plane service)
  meridian-push-gateway   (mobile push fan-out)
  meridian-conformance    (TLA+ trace schema + replay checker; no production call sites yet)

Tooling + shared
  meridian-sdk            (typed Nostr event builders — used by meridian-acp and meridian-cli)
  meridian-cli            (agent-first CLI)
  meridian-admin          (operator CLI: relay membership + key generation)
  meridian-test-client    (integration test harness + manual CLI)

The workspace is 29 crates — 30 members, counting examples/countdown-bot, which is the number just check-architecture-map reports. The tree above shows the relay's dependency spine; the grouped lists below it are peers that do not sit under meridian-relay.

Key architectural principle: The relay is the single source of truth. meridian-relay orchestrates all subsystems by calling them directly — it imports meridian-db, meridian-auth, meridian-pubsub, meridian-search, meridian-audit, and meridian-workflow. However, those subsystems are isolated from each other: meridian-workflow never calls meridian-pubsub, meridian-search never calls meridian-db, etc. Cross-subsystem coordination happens only through the relay. In multi-community mode, the relay also owns propagation of TenantContext; service crates should receive community-scoped inputs rather than independently deriving tenancy from client-controlled event tags.


2. The Protocol

Meridian uses Nostr NIP-01 on the wire. Every action is a JSON event with six fields:

{
  "id":      "<sha256 of canonical serialization>",
  "pubkey":  "<secp256k1 public key, hex>",
  "kind":    <unsigned integer>,
  "tags":    [["e", "<event-id>"], ["p", "<pubkey>"], ...],
  "content": "<JSON payload or plain text>",
  "sig":     "<Schnorr signature over id>"
}

The kind integer is the only dispatch switch. The relay routes, stores, and fans out events based on kind. Clients filter subscriptions by kind. New feature = new kind number = zero breaking changes to existing clients.

Kind Ranges

Range Meaning
0–9999 Standard Nostr kinds (NIP-01 through NIP-XX)
10000–19999 Replaceable events (NIP-16)
20000–29999 Ephemeral events — not stored, not audited
30000–39999 Parameterized replaceable events
40000–49999 Meridian custom kinds

Meridian Custom Kinds (selected)

Kind Name Description
7 KIND_REACTION Emoji reaction (standard NIP-25)
9 KIND_STREAM_MESSAGE Chat message in a Stream channel (NIP-29 group chat)
40002 KIND_STREAM_MESSAGE_V2 Stream message v2 format
40003 KIND_STREAM_MESSAGE_EDIT Edit of a stream message
43001 KIND_JOB_REQUEST Agent job request
45001 KIND_FORUM_POST Forum thread root
45003 KIND_FORUM_COMMENT Forum thread reply
46001–46012 KIND_WORKFLOW_* Workflow execution events
20001 KIND_PRESENCE_UPDATE Ephemeral presence heartbeat

meridian-core defines all 81 kinds as pub const KIND_*: u32 and exports ALL_KINDS: &[u32]. Kinds are u32 (NIP-01 specifies unsigned integer; u32 covers the full range). Meridian uses both standard Nostr kinds (e.g., kind 7 for reactions) and custom ranges (40000+).

Note: KIND_AUTH (22242) is pub const KIND_AUTH: u32 in meridian-core/src/kind.rs and imported by meridian-relay/src/handlers/event.rs. KIND_CANVAS (40100) is likewise pub const KIND_CANVAS: u32 in meridian-core/src/kind.rs.

Wire Protocol (NIP-01 messages)

Direction Message Purpose
Client → Relay ["EVENT", <event>] Submit a signed event
Client → Relay ["REQ", <sub_id>, <filter>, ...] Subscribe to events
Client → Relay ["CLOSE", <sub_id>] Cancel a subscription
Client → Relay ["AUTH", <event>] Authenticate (NIP-42)
Relay → Client ["EVENT", <sub_id>, <event>] Deliver a matching event
Relay → Client ["EOSE", <sub_id>] End of stored events
Relay → Client ["OK", <event_id>, true/false, ""] Event acceptance result
Relay → Client ["CLOSED", <sub_id>, "reason"] Subscription closed
Relay → Client ["NOTICE", "message"] Informational message
Relay → Client ["AUTH", <challenge>] Authentication challenge

Max frame size: 524,288 bytes (default, MERIDIAN_MAX_FRAME_BYTES). Max event content size: 262,144 bytes. Max subscriptions per connection: 1024. Max historical results per filter: 500.


3. Connection Lifecycle

Every WebSocket connection follows this exact sequence:

Step 0: Community Binding

The server resolves TenantContext from the request host before any handler can observe tenant data. The URL/domain is authoritative for the community, matching today's "the relay URL is the workspace" behavior. In single-community mode the configured host maps to the default community. In multi-community mode, an unknown or unmapped host rejects generically and never falls through to a default tenant. Client-supplied #h tags are still channel identifiers; they must resolve to a channel inside the host-derived community.

Step 1: Semaphore Acquire

state.conn_semaphore.try_acquire_owned() — if the relay is at connection capacity, the connection is rejected immediately before any data is read. The permit is held for the entire connection lifetime and dropped on cleanup.

Step 2: NIP-42 Challenge

The relay immediately sends ["AUTH", "<challenge>"]. The challenge is a random string. The connection is registered in ConnectionManager after the challenge is sent.

Step 3: Authentication

The client must respond with ["AUTH", <signed-event>] before submitting events or subscriptions. Authentication paths:

Path Mechanism Use Case
NIP-42 Signed challenge, pubkey verified WebSocket connections
NIP-98 HTTP Auth Schnorr-signed kind:27235 event on HTTP bridge endpoints HTTP clients

On success, ConnectionState.auth_state transitions from Pending → Authenticated(AuthContext). On failure → Failed. Unauthenticated EVENT/REQ messages are rejected with ["CLOSED", ...] or ["OK", ..., false, "auth-required: ..."].

Step 4: Active Loops

Three concurrent tasks run for the lifetime of the connection:

  • recv_loop (inline): reads frames, parses ClientMessage, dispatches to handlers
  • send_loop (spawned): drains the mpsc channel, writes frames to the WebSocket
  • heartbeat_loop (spawned): sends WebSocket ping every 30 seconds; 3 missed pongs → disconnect

A CancellationToken coordinates shutdown across all three loops.

Slow clients: ConnectionState::send() uses try_send — if the send buffer is full, a grace counter increments. After SLOW_CLIENT_GRACE_LIMIT (3) consecutive full-buffer events, the connection is cancelled. A successful send resets the counter.

Step 5: Cleanup

On disconnect (any cause):

  1. cancel.cancel() — signals all loops
  2. Await send_loop and heartbeat_loop tasks
  3. sub_registry.remove_connection(conn_id) — removes all subscriptions from the DashMap indexes
  4. conn_manager.deregister(conn_id) — removes from the send-channel map
  5. drop(permit) — releases the connection semaphore slot

4. Event Pipeline

When the relay receives ["EVENT", <event>], the handler in handlers/event.rs runs this pipeline in order:

1. AUTH CHECK        — AuthState::Authenticated? MessagesWrite scope?
2. PUBKEY MATCH      — event.pubkey == auth_context.pubkey?
3. KIND_AUTH REJECT  — kind == 22242 (AUTH events never stored)
4. EPHEMERAL ROUTE   — kind 20000–29999 → ephemeral sub-pipeline (see below)
5. VERIFY            — spawn_blocking(verify_event) — Schnorr sig + ID hash
6. MEMBERSHIP        — channel_id in event tags? → check_channel_membership
7. DB INSERT         — db.insert_event (ON CONFLICT DO NOTHING — idempotent)
                       the generated `search_tsv` column is populated by this
                       write — there is no separate search-index step
8. REDIS PUBLISH     — pubsub.publish_event (if channel-scoped)
9. FAN-OUT           — sub_registry.fan_out_scoped → filter_fanout_by_access
                       → conn_manager.send_to
10. AUDIT LOG        — audit_tx.send (bounded mpsc, capacity 1000)
11. WORKFLOW TRIGGER — wf.on_event (spawned async, excludes kinds 46001–46012)

Step 11 is fire-and-forget (spawned as an independent async task). Step 10 is not: audit_tx.send().await on a bounded channel deliberately backpressures ingest when the audit store is saturated, rather than silently dropping hash-chain entries. Under Postgres FTS the searchable row is the persisted row, so the old out-of-band indexer and its search_index_tx queue no longer exist. A workflow-trigger failure does not fail the event submission. The client receives ["OK", <id>, true, ""] at the end of the pipeline, not immediately after DB insert.

Step 9 (fan-out) explicitly excludes global subscriptions (no channel_id constraint) from channel-scoped events — global subscriptions do NOT receive events from private channels, regardless of filter match. This is a deliberate security boundary: only subscriptions scoped to an accessible channel_id receive those events.

Workflow loop prevention: workflow execution kinds (46001–46012), relay-signed messages with meridian:workflow tag, and KIND_GIFT_WRAP are excluded from triggering workflows. All other stored events (including kind 9 stream messages) trigger workflow evaluation.

Ephemeral Sub-Pipeline (kinds 20000–29999)

Ephemeral events bypass DB storage, audit, and search. Two sub-paths:

Presence events (kind 20001):

1. VERIFY            — spawn_blocking(verify_event)
2. REDIS PRESENCE    — set_presence() or clear_presence() based on content
3. LOCAL FAN-OUT     — sub_registry.fan_out → conn_manager.send_to (no Redis PUBLISH)

Presence events skip membership checks and use local-only fan-out. Multi-node presence fan-out would require Redis pub/sub (documented as future work).

Other ephemeral events (e.g., typing indicators):

1. VERIFY            — spawn_blocking(verify_event)
2. MEMBERSHIP        — check_channel_membership (if channel-scoped)
3. MARK LOCAL        — state.mark_local_event (dedup before Redis round-trip)
4. REDIS PUBLISH     — pubsub.publish_event (no DB write)
5. LOCAL FAN-OUT     — sub_registry.fan_out → conn_manager.send_to

Ephemeral events are never stored in Postgres and never appear in REQ historical queries.

Handler Semaphore

Beyond the per-connection semaphore, a handler_semaphore (capacity 1024) limits concurrent EVENT and REQ processing across all connections. CLOSE is not rate-limited.


5. Subscription System

SubscriptionRegistry

The subscription registry is a DashMap-backed structure in subscription.rs:

pub struct SubscriptionRegistry {
    subs: DashMap<ConnId, HashMap<SubId, SubEntry>>,
    channel_kind_index: DashMap<IndexKey, Vec<(ConnId, SubId)>>,
    channel_wildcard_index: DashMap<Uuid, Vec<(ConnId, SubId)>>,
}

pub struct IndexKey {
    pub channel_id: Uuid,
    pub kind: Kind,
}

Three-Tier Fan-Out

When an event arrives, fan_out consults three indexes in order:

Tier Index Key Use Case
1 channel_kind_index (channel_id, kind) Subs with explicit channel + kind filter — O(1) lookup
2 channel_wildcard_index channel_id Subs with channel but no kinds constraint
3 subs (linear scan) — Global subs (no channel_id) — fallback scan

Global subs (tier 3) are checked for non-channel-scoped events only. Channel-scoped events are delivered exclusively to subscriptions that carry a matching channel_id — global subscriptions are explicitly excluded from channel fan-out as a security boundary.

NIP-01 Edge Cases

  • kinds: [] (explicit empty array) means "match nothing" — NOT a wildcard. Subscriptions with empty kinds are not indexed in either tier 1 or tier 2 and never receive events.
  • kinds absent (no field) means "match all kinds" — indexed in tier 2 (channel wildcard) or tier 3 (global).

REQ Handler Access Control

The REQ handler checks channel access before registering the subscription:

1. Parse filters, extract channel_id
2. Load accessible_channel_ids for this connection's pubkey
3. If channel_id not in accessible_channels → send CLOSED "restricted: not a channel member"
4. Only then: sub_registry.register(conn_id, sub_id, filters, channel_id)

This prevents a race where a non-member receives live fan-out events from a private channel between registration and the access check.

Historical Query (EOSE)

After registering, the REQ handler queries Postgres for stored events matching the filters (up to 500 per filter, hard cap). These are sent as ["EVENT", sub_id, event] frames before ["EOSE", sub_id]. New events arriving after EOSE are delivered via the fan-out path.


6. Crate Reference

meridian-core — Shared Types and Verification

Zero I/O. The foundation every other crate builds on. Explicitly prohibits tokio, sqlx, redis, and axum in its Cargo.toml.

Key types:

pub struct StoredEvent {
    pub event: nostr::Event,
    pub received_at: DateTime<Utc>,
    pub channel_id: Option<Uuid>,
    verified: bool,          // private — use is_verified()
}

pub const ALL_KINDS: &[u32]  // 80 entries (KIND_AUTH excluded — never stored)

Key functions:

Function Purpose
filters_match(filters, event) OR across filters, AND within each filter. Includes NIP-01 prefix matching on event IDs.
verify_event(event) Schnorr signature + SHA-256 ID check. CPU-bound — callers use spawn_blocking.
is_private_ip(ip) SSRF protection: IPv4 unspecified/loopback/private/link-local/CGNAT/benchmarking/broadcast + IPv6 loopback/ULA/link-local/multicast/documentation + IPv4-mapped IPv6.

Does NOT: store events, make network calls, spawn tasks, or depend on any async runtime.


meridian-auth — Authentication and Authorization

Handles authentication paths, scope enforcement, and token operations.

Auth paths:

Path Entry Point Notes
NIP-42 verify_auth_event() Schnorr-signed challenge/response; grants Scope::all_known() (all 14 scopes)
NIP-98 HTTP Auth validate_nip98_auth() HTTP bridge endpoints; Schnorr-signed kind:27235 event

Key types:

pub struct AuthContext { pub pubkey: PublicKey, pub scopes: Vec<Scope>, pub auth_method: AuthMethod }
pub enum AuthMethod { Nip42, Nip98 }
pub enum Scope { MessagesRead, MessagesWrite, ChannelsRead, ChannelsWrite,
                 AdminChannels, UsersRead, UsersWrite, AdminUsers,
                 JobsRead, JobsWrite, SubscriptionsRead, SubscriptionsWrite,
                 FilesRead, FilesWrite, Unknown(String) }
pub trait ChannelAccessChecker: Send + Sync { ... }
pub trait RateLimiter: Send + Sync { ... }

Security details:

  • NIP-98 auth: Schnorr-signed kind:27235 events with URL + method tags.
  • NIP-42 timestamp tolerance: ±60 seconds.
  • Dev-only key derivation: SHA-256("meridian-test-key:{username}") — gated behind #[cfg(any(test, feature = "dev"))]. The dev feature must not be enabled in production relay deployments.

Does NOT: implement RateLimiter beyond a test stub (AlwaysAllowRateLimiter, gated behind #[cfg(any(test, feature = "test-utils"))]). No Redis-backed rate limiter exists anywhere in the codebase — rate limiting is not currently enforced. RateLimitConfig defines 4 tiers (human, agent-standard, agent-elevated, agent-platform) as a design target.


meridian-db — Postgres Event Store

All database access. Uses sqlx::query() (runtime, not compile-time macros) — no .sqlx/ offline cache required.

Key operations:

Module Responsibility
event.rs insert_event (ON CONFLICT DO NOTHING), query_events (QueryBuilder), get_event_by_id
channel.rs Channel CRUD, membership management, role enforcement (transactional)
feed.rs query_mentions (INNER JOIN event_mentions), query_needs_action, query_activity
workflow.rs Full workflow/run/approval CRUD; SHA-256 hashed approval tokens
partition.rs Monthly range partitioning for events and delivery_log tables
dm.rs DM channel management
reaction.rs Reaction storage and retrieval
thread.rs Thread/reply tracking
user.rs User profile storage
error.rs Database error types

Channel types: Stream, Forum, Dm, Workflow
Member roles: Owner, Admin, Member, Guest, Bot
Workflow statuses: Active, Disabled, Archived
Run statuses: Pending, Running, WaitingApproval, Completed, Failed, Cancelled

Key behaviors:

  • ON CONFLICT DO NOTHING for event dedup — returns (StoredEvent, was_inserted: bool).
  • Rejects KIND_AUTH (22242) and ephemeral (20000–29999) with distinct error variants.
  • Transactional role enforcement in add_member/remove_member/create_channel — TOCTOU-safe.
  • Soft-delete for channel members: remove_member sets removed_at; re-adding reverses it.
  • Feed hard cap: FEED_MAX_LIMIT = 100 rows regardless of caller-requested limit.
  • query_mentions uses INNER JOIN event_mentions — normalized table with composite index on (pubkey_hex, created_at).
  • Approval tokens: create_approval receives the raw token and hashes it internally with SHA-256.
  • DDL injection protection in partition manager: allowlist of table names + strict suffix/date validators.

Does NOT: cache queries, implement connection pooling logic (delegated to sqlx), or make network calls outside Postgres.


meridian-pubsub — Bus Pub/Sub, Presence, Typing

Manages pub/sub fan-out, presence tracking, and typing indicators over Dragonfly. In multi-community mode all tenant-visible keys are prefixed or otherwise partitioned by community (meridian:{community}:...) so channel fan-out, presence, typing, and cache invalidation cannot cross hosts.

The crate is written against the Redis wire protocol and that is deliberate: Dragonfly reports redis_version:7.4, so REDIS_URL, the redis crate, deadpool-redis, Lua Script calls and pub/sub all work unchanged. Do not reintroduce a redis service — the compose service is dragonfly and the container is meridian-dragonfly, with a redis network alias so in-network redis://redis:6379 still resolves.

Architecture:

Publisher  → pool connection   → PUBLISH meridian:channel:{uuid}
Subscriber → dedicated PubSub  → PSUBSCRIBE meridian:channel:*
                                  → broadcast::channel(4096)

The subscriber uses a dedicated redis::aio::PubSub connection — not from the pool. This is intentional: pool connections cannot hold PSUBSCRIBE state.

Current state: The subscriber loop is spawned in meridian-relay/src/main.rs and populates the broadcast channel. A consumer task subscribes via pubsub.subscribe_local(), calls sub_registry.fan_out() on each received event, and delivers matches to local WebSocket connections via conn_manager.send_to(). Multi-node fan-out is now wired end-to-end. Local-echo deduplication is implemented via AppState.local_event_ids — events published by the local relay instance are tracked and skipped when received via the Redis round-trip.

Fan-out is lossy under lag, and the compensation is per-lane. The broadcast channel between the subscriber and each consumer is bounded (4096 per lane), so a consumer that falls behind gets RecvError::Lagged(n) and those n messages are gone — Redis pub/sub cannot replay them. Lag is not terminal: the consumer resumes, because a lag burst used to latch the lane off for the lifetime of the process and needed a pod restart to clear.

Resuming is only safe because the three lanes do not carry equivalent consequences, and the consumer reconciles each one before it resumes:

  • Events — no reconciliation. A client backfills a missed event with REQ.
  • CacheInvalidation — flush the authorization caches. A missed drop leaves a stale projection, and flushing is the conservative direction.
  • ConnectionControl — flush the authorization caches and restart every local socket. This lane carries ban and disconnect commands, and an established WebSocket never re-authenticates, so a dropped command would otherwise leave a banned member connected indefinitely. Failing the pod closed used to guarantee that drain as a side effect; the reconciliation hook now does it deterministically instead of depending on a 250 ms health poller sampling a microsecond-wide window.

Lag never moves a lane's phase. Phase is owned solely by the Redis subscriber tasks — one writer — so a lagging consumer cannot publish a Ready that masks a genuinely dropped subscription. The loss is recorded as fact instead: the per-lane *_lag_total counter, last_lag_unix, and meridian_pubsub_subscriber_skipped_messages, which is monotonic and never cleared. A lane at Ready with a non-zero skipped count is a pod serving traffic that has permanent gaps in its history — read the gauge, not the phase. Both fields are in the /ready diagnostic. mark_application_fault is reserved for the two genuinely unrecoverable outcomes: a closed channel and a panicked consumer.

Reconnection: exponential backoff 1s → 30s (backoff_secs * 2). Backoff resets to 1s only after a clean stream end, not on each reconnect attempt.

Presence: SET meridian:presence:{pubkey_hex} {status} EX 90 — 90-second TTL (3× the 30-second heartbeat interval). Single missed heartbeat does not cause presence flap.

Typing indicators:

ZADD meridian:typing:{channel_id} {now_unix} {pubkey_hex}
ZREMRANGEBYSCORE meridian:typing:{channel_id} -inf {now - 5.0}
EXPIRE meridian:typing:{channel_id} 60

5-second activity window. 60-second key TTL prevents orphaned empty sets.

Does NOT: implement the rate limiter. Does NOT store events. PubSubManager is not Clone — callers use Arc<PubSubManager>.


meridian-search — Postgres FTS Integration

Full-text search via Postgres FTS. Events are searchable through the events.search_tsv generated tsvector column (populated on insert, indexed by a GIN index) — there is no separate search service or out-of-band indexer. Privacy-sensitive kinds are excluded at the storage level (the search_tsv CASE WHEN kind IN (...) yields NULL, which never matches @@). In multi-community mode every query filter includes community_id, so the shared events table is infrastructure, not a cross-community result space; the relay re-authorizes every candidate hit before returning it.

Key behaviors:

  • SearchService::new(pool) wraps a PgPool; search(&SearchQuery) runs a parameterized FTS query against the events.search_tsv GIN index and returns SearchResult (candidate SearchHits).
  • ChannelScope makes the channel constraint explicit (Any / ChannelLessOnly / Channels / ChannelsOrChannelLess), closing the ambiguity the old Option<Vec<Uuid>> + bool matrix could not express.
  • Every query carries community_id; the FTS predicate is BitmapAnd-ed with the community-leading btree filters so a query never crosses tenants.
  • Permission filtering is caller's responsibility — meridian-search returns candidate hits; the relay re-authorizes each one (channel membership, #p, owner gates) before delivering it.

Does NOT: enforce channel membership or access control. Does NOT write events (indexing is the search_tsv generated column on the events insert).


meridian-audit — Hash-Chain Audit Log

Tamper-evident append-only log with SHA-256 hash chaining.

Hash chain: each entry stores prev_hash (hash of the previous entry). In multi-community mode audit heads/chains are per-community; operator metrics may aggregate, but tenant-readable audit verification walks one community chain. verify_chain() walks entries and recomputes hashes to detect tampering. Genesis entry uses GENESIS_HASH (64 zeros).

Hash covers: seq (big-endian bytes), timestamp (RFC3339), event_id, event_kind (big-endian), actor_pubkey, action string, channel_id (16 bytes or 16 zero bytes if None), canonical metadata JSON (BTreeMap for deterministic key ordering), prev_hash.

Single-writer guarantee: pg_advisory_lock before each transaction. Lock released in all branches including panic (catch_unwind).

10 audit actions: EventCreated, EventDeleted, ChannelCreated, ChannelUpdated, ChannelDeleted, MemberAdded, MemberRemoved, AuthSuccess, AuthFailure, RateLimitExceeded.

Does NOT: log KIND_AUTH (22242) events — returns AuditError::AuthEventForbidden immediately. Does NOT log ephemeral events (they never reach the audit pipeline).


meridian-workflow — YAML-as-Code Automation Engine

Parses, validates, and executes channel-scoped workflow definitions. In multi-community mode workflow definitions, runs, approvals, webhook routes, and schedules inherit the host-derived community and evaluate triggers only against events in that community.

Workflow definition structure:

name: "Incident Triage"
trigger:
  on: message_posted
  filter: "str_contains(trigger_text, 'P1')"
steps:
  - id: notify
    action: send_message
    text: "P1 incident detected: {{trigger.text}}"
  - id: page
    if: "str_contains(trigger_text, 'production')"
    action: request_approval
    from: "{{trigger.author}}"
    message: "Page on-call?"

Note: Both TriggerDef and ActionDef use serde internally-tagged enums. Triggers use on: as the tag field; actions use action: as the tag field. Fields are flattened into the parent struct, not nested.

4 trigger types: message_posted, reaction_added, schedule, webhook

7 action types:

Action Description
send_message Post to the workflow's channel (or override channel)
send_dm Direct message to a user (pubkey hex or {{trigger.author}})
set_channel_topic Update channel topic
add_reaction React to the trigger message
call_webhook HTTP POST to external URL (SSRF-protected, redirects disabled, 1 MiB response cap)
request_approval Suspend execution; fields: from, message, timeout (default 24h)
delay Pause execution (max 300 seconds)

Template variables: {{trigger.text}}, {{trigger.author}}, {{steps.ID.output.FIELD}}. Single-pass resolution (not recursive). Unknown variables left as literal text.

Condition evaluation: evalexpr with HashMapContext. Dot notation converted to underscores (trigger.text → trigger_text). Custom functions registered: str_contains, str_starts_with, str_ends_with, str_len. 100ms timeout prevents adversarial expressions from blocking.

Concurrency: Arc<Semaphore> with 100 permits. try_acquire() — returns CapacityExceeded immediately rather than queuing.

Approval gates: request_approval action returns StepResult::Suspended with a generated UUID token, but the engine does not yet persist the token or resume execution — runs that hit an approval gate are marked as failed (🚧 WF-08). execute_from_step() exists for future resumption support.

Cron scheduler: loop ticks every 60 seconds, evaluates cron expressions with window-based matching, and creates workflow runs for matched triggers. Fully implemented.

Does NOT: recursively resolve templates (single-pass only). Does NOT queue workflow runs when at capacity — returns CapacityExceeded immediately.


Huddle Audio — WebSocket Opus Relay

Real-time voice lives inside meridian-relay (src/audio/), not a separate crate. A WebSocket endpoint (wss://.../huddle/{channel_id}/audio) authenticates each participant with a NIP-42 challenge, checks channel membership, admits them to an in-memory room, and forwards opaque Opus frames between peers. No external SFU.

Frame protocol (v2): 8-byte big-endian header (sequence u16, 48 kHz timestamp u32, level dBov i8, flags u8) followed by an opaque Opus payload. Invalid level_dbov values are clamped rather than dropped — losing a metric beats losing audio.

Room state: an admission guard synchronizes joins against the room's ended flag; soft cap 25 peers (hard cap 255 via u8 peer index). Per-peer audio uses a bounded channel (drop-on-full); the control channel is separate and never drops join/leave.

Lifecycle events: the relay emits Nostr events for participant joined / left and huddle ended; the desktop client emits huddle started and guidelines. When the last peer leaves, the room ends and the channel archives atomically.

Not yet built: recording and per-track publishing (the corresponding kinds are reserved, no producer exists).


meridian-relay — The Server

Axum WebSocket server. Ties all other crates together. The only crate that imports and orchestrates all subsystems.

AppState (Arc-wrapped, shared across all connections — key fields shown, not exhaustive):

pub struct AppState {
    pub db: Db,
    pub audit: Arc<AuditService>,
    pub pubsub: Arc<PubSubManager>,
    pub auth: Arc<AuthService>,
    pub search: Arc<SearchService>,
    pub sub_registry: Arc<SubscriptionRegistry>,
    pub conn_manager: Arc<ConnectionManager>,
    pub workflow_engine: Arc<WorkflowEngine>,
    pub conn_semaphore: Arc<Semaphore>,       // connection limit
    pub handler_semaphore: Arc<Semaphore>,    // 1024 concurrent handlers
    pub relay_keypair: nostr::Keys,           // relay identity
    pub local_event_ids: moka::sync::Cache,   // local-echo dedup
    // + config, redis_pool, membership_cache, media_storage, shutdown state
}

ConnectionState (per-connection):

pub struct ConnectionState {
    pub auth_state: RwLock<AuthState>,
    pub subscriptions: Mutex<HashMap<String, Vec<Filter>>>,
    // + send_tx, cancel token
}
pub enum AuthState { Pending { challenge: String }, Authenticated(AuthContext), Failed }

HTTP endpoints:

Method Path Handler
GET / WebSocket upgrade or NIP-11 relay info
GET /info NIP-11 relay info
GET /.well-known/nostr.json NIP-05 identity
GET /health Health check
GET /_liveness Liveness probe
GET /_readiness Readiness probe
POST /events Submit a signed Nostr event over HTTP (same ingest path as WebSocket EVENT)
POST /query Query Nostr events over HTTP with NIP-01 filters
POST /count Count Nostr events over HTTP with NIP-45 filters
POST /hooks/{id} Workflow webhook trigger (secret-authenticated)
PUT /media/upload Upload media blob (Blossom, 50 MB limit)
GET/HEAD /media/{sha256_ext} Retrieve/probe media blob
GET /git/{owner}/{repo}/info/refs Git smart HTTP advertisement
POST /git/{owner}/{repo}/git-upload-pack Git smart HTTP fetch
POST /git/{owner}/{repo}/git-receive-pack Git smart HTTP push
POST /internal/git/policy Internal git hook policy check

Constants:

Constant Value Purpose
max_frame_bytes (MERIDIAN_MAX_FRAME_BYTES) 524,288 default Max WebSocket frame size
MAX_EVENT_CONTENT_BYTES 262,144 Max event content size — rejected with OK false, advertised as NIP-11 max_content_length
MAX_SUBSCRIPTIONS 1024 Per-connection subscription limit
MAX_HISTORICAL_LIMIT 500 Per-filter historical query cap
handler_semaphore capacity 1024 Concurrent EVENT/REQ handlers

Does NOT: implement business logic — delegates to the appropriate crate for every operation.


meridian-acp — Agent Communication Protocol Harness

Standalone binary that bridges Meridian relay events to AI agents via the Agent Communication Protocol (ACP).

Architecture:

Meridian Relay ──WS──→ meridian-acp ──stdio (ACP/JSON-RPC)──→ Agent (goose/codex/claude)

meridian-acp spawns AI agent subprocesses (1–32, default 1), connects to the relay via WebSocket with NIP-42 auth, discovers channels via REST API, and queues @mention events per channel. At most one prompt is in-flight per channel. Queued events are batched into a single prompt sent via session/prompt over ACP.

Key modules:

Module LOC Responsibility
relay.rs 3,143 WebSocket + REST relay connection, NIP-42 auth
queue.rs 2,565 Per-channel event queue, batching, dedup
main.rs 2,457 Event loop, pool orchestration, heartbeat
pool.rs 2,253 N-agent pool, claim/return lifecycle
config.rs 1,903 CLI/env/TOML configuration
acp.rs 1,785 ACP client, stdio JSON-RPC, timeouts
filter.rs 814 Subscription rules, evalexpr filtering

Key behaviors:

  • Pool of 1–32 agent subprocesses with claim/return lifecycle.
  • Per-channel queuing: at most one prompt in-flight per channel; subsequent @mentions queue until the agent responds.
  • Crash recovery: agent subprocess crashes are detected and the agent is respawned.
  • Depends on meridian-core (kind constants) and meridian-sdk (relay/REST utilities).

Does NOT: persist state.


meridian-admin — Operator CLI

Subcommands:

Subcommand Purpose
add-member Add a pubkey to the relay membership list (--pubkey, --role); accepts npub or hex; publishes kind:13534 roster
remove-member Remove a pubkey from the relay membership list (--pubkey, optional --role guard); publishes kind:13534 roster
list-members List all relay members
generate-key Generate a new Nostr keypair (for bootstrapping)
reconcile-channels Emit kind:39000/39002 discovery events for channels missing them (idempotent)

The meridian-admin binary is shipped in the relay Docker image (/usr/local/bin/meridian-admin) and is the recommended way to manage relay membership in production. Use ./run.sh add-member, ./run.sh remove-member, and ./run.sh list-members in Docker Compose deployments.


meridian-test-client — Integration Test Harness

MeridianTestClient wraps a WebSocket connection with a VecDeque<RelayMessage> buffer for message interleaving. Methods: connect, connect_unauthenticated, authenticate, send_event, send_text_message, subscribe, close_subscription, recv_event, collect_until_eose, disconnect.

Test coverage:

File Tests Scope
tests/e2e_relay.rs 27 WebSocket protocol (auth, subscriptions, filters, limits, NIP-11)
tests/e2e_media.rs 7 Media upload/download (Blossom)
tests/e2e_media_extended.rs 18 Extended media scenarios
tests/e2e_nostr_interop.rs 15 Nostr interoperability: NIP-50 search, NIP-10 threads, NIP-17 gift wraps, DM discovery

All e2e tests are #[ignore] — require a running relay. Total: 134 e2e tests.

src/main.rs is a manual testing CLI (meridian-test-cli) with --send, --subscribe, --channel, --url, --kind flags.

Defines parse_relay_message, OkResponse, RelayMessage directly in src/lib.rs.


7. Security Model

Every security-sensitive operation uses an explicit, verified pattern. No implicit trust.

Authentication

Concern Mechanism
NIP-42 timestamp ±60 second tolerance — prevents replay attacks
AUTH events Never stored in Postgres, never logged in audit chain
NIP-98 HTTP Auth Schnorr-signed kind:27235 events — URL and method verification

Input Validation

Concern Mechanism
Schnorr signatures verify_event() in meridian-core — every event verified before storage
Event ID SHA-256 of canonical serialization verified independently of signature
Frame size max_frame_bytes (512 KiB default, MERIDIAN_MAX_FRAME_BYTES) — oversized frames rejected by the WebSocket parser, connection closed
Event content size MAX_EVENT_CONTENT_BYTES = 262,144 — larger events rejected with OK false on WS and HTTP 400 on POST /events
Search event IDs 64-char hex validation before URL construction — prevents path injection
Workflow step IDs Alphanumeric + underscore only — prevents evalexpr variable injection
Partition names Allowlist of table names + strict suffix/date validators — prevents DDL injection

SSRF Protection

is_private_ip() in meridian-core covers:

  • IPv4: unspecified (0.0.0.0/8), loopback (127.0.0.0/8), private (10/8, 172.16/12, 192.168/16), link-local (169.254/16), CGNAT (100.64/10), benchmarking (198.18/15), broadcast (255.255.255.255)
  • IPv6: loopback (::1), ULA (fc00::/7), link-local (fe80::/10), multicast (ff00::/8), documentation (2001:db8::/32)
  • IPv4-mapped IPv6 (::ffff:0:0/96) — recursively checks the embedded IPv4 address

Applied in: meridian-workflow (CallWebhook action), meridian-core (shared utility).

Audit Integrity

  • Hash chain: each entry's SHA-256 covers all fields including prev_hash — tampering any entry breaks all subsequent hashes
  • Canonical JSON: BTreeMap for deterministic key ordering — hash is reproducible
  • Single-writer lock: pg_advisory_lock — prevents concurrent writes from breaking the chain
  • Panic-safe: catch_unwind ensures lock release even on panic

Access Control

  • Channel membership is the only gate — enforced by the relay at every operation
  • REQ handler checks access before subscription registration — no race window for private channel leaks
  • TOCTOU-safe membership operations: all check-then-modify sequences run inside Postgres transactions
  • Approval tokens: UUID (CSPRNG), stored as SHA-256 hash, single-use enforced with AND status = 'pending' in UPDATE

Webhook Security

  • Workflow webhooks: constant-time XOR comparison of stored UUID secret (not HMAC — compares the secret directly, not a body MAC)
  • Outbound webhooks (CallWebhook): SSRF protection + redirects disabled + 1 MiB response cap

8. Infrastructure

Docker Compose provides the full local development stack. All services include health checks and resource limits.

Services

Core services carry no profile and always start: postgres, dragonfly, minio*, and the *-db-init one-shots. Everything else is opt-in via COMPOSE_PROFILES. Do not put a core service behind a profile — just _ensure-services blocks on postgres/dragonfly health and would hang for 120s against a disabled service.

Service Image Port Profile Purpose
Postgres postgres:17-alpine 5432 core Primary event store — events, channels, tokens, workflows, audit; full-text search (search_tsv GIN)
Dragonfly dragonflydb/dragonfly:v1.39.0 6379 core Pub/sub fan-out, presence (SET EX), typing (sorted sets). Redis-wire-compatible drop-in; container meridian-dragonfly, redis network alias
MinIO minio/minio 9000 (API), 9001 (console) core S3-compatible object storage (media + Git/CAS)
Adminer adminer 8082 tools DB web UI (dev only)
Keycloak quay.io/keycloak/keycloak:26.0 8180 auth Local OAuth
Prometheus prom/prometheus 9090 observability Metrics collection
ReductStore reduct/store:v1.20.11 8383 artifacts Artifact/blob REST tier on local disk
ReductStore reduct/store:v1.20.11 8383 artifacts-s3 Same tier, MinIO behind it — requires a commercial licence, see below
zenohd not yet added — bus Zenoh router for bus scale-out and cross-region. Committed, not built — Phase 4 (meridian-0ps); profile must be opt-in and default off
IPFS not yet selected — none CID-addressed replication behind S3. Planned, not specified — see the bytes-tier note below

Artifact Tier (ReductStore) — opt-in

A time-indexed blob store with an HTTP REST API, for artifact bytes whose access shape is "give me this range of this stream": huddle recordings, large attachments, telemetry. Records are addressed by bucket + entry + timestamp and batched into blocks, so range reads and appends beat per-object S3 GET/PUT for that shape, and the REST interface removes the S3 SDK from the caller.

It is not a system of record. Event truth is Postgres; no reader resolves chat history from ReductStore. It sits on the A2 consumer axis — downstream of the spine, and invisible at the NIP-01 edge. This matters because a second store that also claims to hold truth is refused by name (DIAGRAM.md § Scaling axes, R2/R3, Invariant 2).

Two mutually exclusive profiles — same host port, two backings of one tier:

  • artifacts — free/OSS, bytes on the meridian-reduct-data volume. Verified working: bucket create, timestamped write, and read-back all return 200 against reduct/store:v1.20.11, and survive a container restart.
  • artifacts-s3 — MinIO behind it, with local disk demoted to a hot cache (RS_REMOTE_CACHE_PATH; RS_DATA_PATH is ignored once a remote backend is configured).

The artifacts-s3 profile cannot run on any public image today. The S3 remote backend is a ReductStore Pro commercial feature, and verification on 2026-08-05 found that reduct/store:v1.20.11 does not parse RS_REMOTE_* at all — the variables never appear in its startup config dump. It boots Apache-2.0 OSS mode on RS_DATA_PATH = /data and reports healthy, so a misconfiguration presents as a working MinIO tier rather than as a failure.

That silent-degradation mode is why the profile is gated by a blocking reduct-license-check one-shot: without a non-empty licence key it exits 1 with an explanatory message and depends_on: service_completed_successfully stops reductstore-s3 from ever starting. A named volume is also mounted at /data so that an OSS fallback, if one ever occurs, cannot lose bytes on docker compose down. Non-empty meridian-reduct-s3-fallback is itself the signal that the MinIO tier is not actually engaged.

Replication Tier (IPFS) — planned, not specified

The intended shape is S3 stays the resolution path; IPFS is asynchronous replication behind it, with CIDs recorded as content fingerprints on the events that reference a blob. That ordering is what keeps the tier legitimate: one system of record for bytes, with a downstream replica, rather than two places a blob can be authoritative.

Binding constraints, settled 2026-08-06 (meridian-sgp7):

  • Private IPFS only — no public DHT announcement, no public gateway fallback. A public DHT has no tenant boundary, and Meridian's boundary is host-derived and enforced at the relay.
  • Encrypted CAR packages, randomized authenticated encryption.
  • CIDs are computed over ciphertext, never plaintext. A ciphertext CID does not identify plaintext and supports cryptographic expiration, which is what stops the tier becoming an un-revocable public index.
  • MinIO remains the managed object origin. IPFS is replication; it is never a resolution path.

Still open, and sgp7 does not close until it is tested. The honest claim is controlled distribution and cryptographic revocation — not absolute takedown: a ciphertext CID cannot claw back plaintext or keys already retained by an authorized recipient. Closure requires a pin/provider inventory, time-bounded pin leases, removal from every enrolled provider, destruction of wrapped data-encryption keys, revocation of catalogue and manifest references, and two live tests — that the CID is no longer retrievable through controlled infrastructure, and that retained ciphertext cannot be decrypted after key expiration.

Revocation ordering is unchanged and still gates this tier. Cross-trust-domain replication is federation-shaped (axis A5), deferred behind monotonic revocation. That mechanism is now specified in MIP-DD — and note that session_epoch and policy_generation are not it; the absorbing counter is revocation_floor, and it needs council approval before MIP-DD may reach IMPLEMENTABLE.

Postgres Schema (key tables)

Table Purpose
events All stored Nostr events; monthly range-partitioned by PARTITION BY RANGE on created_at; multi-community mode keys every tenant-visible event by community_id
channels Channel records (type, visibility, canvas, topic); community_id is immutable after creation in multi-community mode
channel_members Membership with roles; soft-delete via removed_at
workflows Workflow definitions (YAML stored as canonical JSON); scoped by community in multi-community mode
workflow_runs Execution records with trigger context and trace
workflow_approvals Approval gates (token stored as SHA-256 hash)
audit_log Hash-chain audit entries; per-community chain/head in multi-community mode
delivery_log Delivery tracking (partitioned; Rust module pending)

Bus Key Patterns (Dragonfly)

Pattern Type TTL Purpose
meridian:channel:{uuid} Pub/Sub channel — Event fan-out (single-community form; shared multi-community Redis must use meridian:{community}:channel:{uuid} or equivalent)
meridian:presence:{pubkey_hex} String 90s Online/away status (single-community form; shared multi-community Redis must scope by community)
meridian:typing:{channel_uuid} Sorted Set 60s Active typers (5s window; shared multi-community Redis must scope by community)

Full-Text Search (Postgres FTS)

Search runs over the events.search_tsv generated tsvector column on the events table (no separate collection or service). The column is populated on insert — to_tsvector('simple', content) — and excludes privacy-sensitive kinds via CASE WHEN kind IN (1059, 30300, 30622) THEN NULL, so those rows are storage-level unsearchable (a NULL tsvector never matches @@). A GIN index (idx_events_search_tsv) backs the @@ probe; in multi-community mode the community-leading btree filters BitmapAnd with the GIN probe so every query is fenced to its community_id.


9. Known Limitations

These are verified gaps in the current implementation — not design aspirations.

# Limitation Detail
1 No sqlx offline query cache Uses sqlx::query() (runtime) not sqlx::query!() (compile-time). No .sqlx/ directory. Queries are not validated at compile time.
2 No rate limiting implementation RateLimiter trait exists in meridian-auth. Only implementation is AlwaysAllowRateLimiter (test stub, gated behind #[cfg(any(test, feature = "test-utils"))]). RateLimitConfig defines 4 tiers (human, agent-standard, agent-elevated, agent-platform) but none are enforced.
3 No dedicated typing REST endpoint Typing indicators (kind 20002) are delivered via both local fan-out and Redis pub/sub (cross-node). There is no REST endpoint to query current typers — /api/presence returns online/away status only, not typing state.
4 Huddle recording/tracks not built Voice, room lifecycle, and join/leave/end events are wired (see Huddle Audio above). Recording and per-track publishing have reserved kinds but no producer yet.
5 Approval gates not wired end-to-end The executor returns StepResult::Suspended and the relay has grant/deny API endpoints with DB CRUD, but the engine intercepts before creating WaitingApproval rows — runs that hit an approval gate are marked as Failed (🚧 WF-08).
6 Workflow actions partially stubbed The send_dm and set_channel_topic workflow actions are in the schema but return NotImplemented — a run that reaches one fails at execution (🚧 WF-07).
7 ReductStore artifact tier has no call sites The artifacts profile is verified working end-to-end, but no crate or client reads or writes it yet — it is an evaluation tier, not a wired dependency. It is opt-in precisely so it does not become inventory; a self-hosted Convex tier was removed from this stack for exactly that reason. Either wire a consumer or remove it.
8 artifacts-s3 cannot run on a public image The MinIO-backed backend is a ReductStore Pro commercial feature, and reduct/store:v1.20.11 ignores RS_REMOTE_* entirely while reporting healthy. Blocked closed by the reduct-license-check gate. Needs both a licence key and a licensed image before it can be exercised.