Both SHM keys default true in Zenoh 1.8, the eclipse-zenoh wheel IS built
with the shared-memory feature, and transport_optimization's threshold is
3,072 B. So every measured payload at 4 KB and above took a POSIX-SHM fast
path -- one the relay build cannot take, because shared-memory is not in
its zenoh feature list -- while the Redis side had no equivalent. The error
points toward Zenoh.
RELAY_BUS_SCALING.md's own bullet said "Nothing about shared memory. SHM
was not enabled." That was false, and nothing could have caught it: the
static contract checked transport.shared_memory.enabled and was blind to
the second switch beside it. This adds that check.
The "Router mode prices the Docker boundary" reading goes with it. SHM
works host-to-host and cannot cross into the Docker VM, so an unknown
share of the peer-vs-router divergence at >=64 KB is the SHM path dropping
out rather than the boundary appearing. Both readings are unlicensed until
re-measured under the shipped posture.
Below 4 KB stands, including the 256 B verdict that failed the >=5x gate:
256 B is well under the SHM threshold, and gossip's extra transports cost
the measured process work rather than saving it.
Worth stating plainly, because it is the second time: the 0A.2 result has
now been invalidated twice for two unrelated reasons -- unmatched publish
semantics, then transport posture -- and neither was visible in the
numbers. A bus measurement is not licensed by its spread. It is licensed by
its posture being pinned, read back, and stamped beside the result, which
is what af4b92eba now does.
Signed-off-by: Joshua Belke <joshua@innovationhub-act.org>
62 KiB
TASKS.md — Decentralized Relay Program
Program: Replace the centralized MQTT/TBMQ bus, policy, and audit tier with a decentralized Meridian relay backbone. Duration: 10 weeks · 5 × 2-week sprints Deliverable: Frozen MIP specifications + reference implementation + conformance suite, handed off to the implementing team. Status: Draft for review · not yet ratified · no work authorized
Related: DIAGRAM.md § Part III (the ratified architectural case) · AGENTS.md § Design Law · AGENTS.md § Published Throughput Ceilings ·
.settings/features/feature-zenoh-transport.md·.settings/reference-docs/goat/03-STREAMING-INGEST.md
0. Document control
| Field | Value |
|---|---|
| JIRA project key | MRDN |
| Issue hierarchy | Initiative → Epic → Story → Sub-task |
| Estimation | Story points, modified Fibonacci (1, 2, 3, 5, 8, 13) |
| Sprint length | 2 weeks |
| Priority scale | P0 Blocker · P1 Critical · P2 Major · P3 Minor |
| Source of truth for task state | Beads (bd show <id>) — JIRA mirrors it, never the reverse |
| Traceability | Every story carries its Bead ID. Where none exists, the story's first sub-task is file the Bead |
Assumptions requiring confirmation before Sprint 1 planning. Each changes sizing, not direction:
- Team: 6 engineers — 2 Rust data-plane, 1 Rust protocol/spec, 1 client/SDK, 1 frontend, 1 QA/performance. Velocity assumed 55 points/sprint with Sprint 1 discounted to 40 for ramp. Total budget ≈ 260 points.
- Hardware: one dedicated 16-core load-generation host and one 16-core system-under-test host, ≥25 GbE between them. § 3 shows why the NIC is a first-class dependency and not an afterthought.
- The 1.337M figure is TBMQ-class vendor-published, not measured on our
stack.
MRDN-101establishes our own baseline before anything commits to it.
1. Program summary — what is and is not being replaced
The scope of "replacement" is narrower than the phrase implies, and the narrowing is already ratified in DIAGRAM.md § Part III by the Architecture & Scale council. Restating it here so the implementing team does not acquire an obligation the binary cannot meet:
Meridian replaces the broker's bus, policy, and audit role. A broker remains legitimate — and is often correct — at the device edge, feeding the spine through a feeder, and never as a routing peer.
| Replaced by Meridian | Kept at the edge (TBMQ or equivalent) |
|---|---|
| The mandatory Kafka hop on every message | Broker-held offline queues / persistent sessions |
| Per-message topic-ACL authorization | MQTT QoS 1/2 delivery state for constrained devices |
| A second durable log alongside the query store | Retained messages and Last Will & Testament |
| Topic-prefix multi-tenancy | Shared subscriptions ($share/group/topic) |
| Unauthenticated-past-the-broker messages | The embedded MQTT client SDK ecosystem |
That right-hand column is not a gap list to close in 10 weeks. It is a standing
architectural boundary. The single strongest reason to keep a broker in the
topology is offline queueing for intermittently-connected devices: a Meridian
client reconnects and re-runs a REQ over stored events, which works for a chat
client and does not work for a constrained device that cannot page history.
Topology decision — hub-and-leaf forwarding
Edge relays terminate device connections locally and forward upstream to one
authoritative relay per community. Single-writer is preserved, which keeps
the epoch-monotonic revocation defect (D4 / meridian-3wh) dormant rather
than live.
flowchart LR
classDef dev fill:#4a4a4a,stroke:#2b2b2b,color:#fff
classDef edge fill:#1f6aa5,stroke:#123f63,color:#fff
classDef feed fill:#7d6608,stroke:#4d3f05,color:#fff
classDef spine fill:#b03030,stroke:#6e1c1c,color:#fff
classDef bad fill:#8a1c1c,stroke:#500f0f,color:#fff
DEV["Constrained devices<br/>MQTT clients · intermittent links"]:::dev
BRK["EDGE — MQTT bridge + leaf relay<br/>persistent sessions · QoS 1/2 ·<br/>retained · LWT · shared subs ·<br/>store-and-forward across outages"]:::edge
FD["FEEDER boundary<br/>decode · dedup · compute routing header ·<br/>MAC or sign · batch"]:::feed
SP["SPINE — authoritative relay<br/>admission · verification · ordering ·<br/>tenancy fence · audit · fan-out"]:::spine
DEV --> BRK --> FD --> SP
X["REFUSED: leaf as routing peer.<br/>A leaf outage would become a spine routing failure,<br/>and a compromised leaf could transit between communities."]:::bad
SP -.->|never| X
Three properties make this the correct seam rather than a compromise:
- The feeder owns the expensive, content-aware work. A relay that never reads the payload categorically cannot deduplicate by content — and that refusal to read is the enabling property for the opaque-frame ceiling, not a limitation to work around.
- The boundary is one-directional. The spine ingests from the edge; the edge
never routes for the spine. This is
E6in the conformance register and Law 2 in the design laws. - Nothing here needs federation. Multi-writer peering stays blocked behind
D4, exactly as Design Law already blocks it.
Phasing decision — greenfield core, then bridge
| Weeks | Track |
|---|---|
| 1–6 | Greenfield opaque-frame core. Native MIP frames end to end. Hits the throughput number on a path with no MQTT semantics in it |
| 6–10 | MQTT compatibility bridge. Ingress adapter at the edge, so devices that cannot be reflashed migrate without firmware work |
The order matters: building the bridge first would let MQTT semantics — per-message topic ACLs, broker-held delivery state, topic-string tenancy — leak into the core, which is precisely what this program exists to remove.
2. Ratified constraints — the non-negotiables
These are not preferences. Each is written law in this repo, each prevents a silent failure rather than a loud one, and a story that violates one is rejected at review regardless of its benchmark.
| # | Constraint | Source | What it forbids in this program |
|---|---|---|---|
| C-1 | Every published throughput figure carries its profile — traffic class, auth profile, parsed-or-not, stored-or-not | AGENTS.md § Published Throughput Ceilings | Quoting a frame-path number for chat, or any bare msg/s in README/NIP-11/marketing |
| C-2 | No vocabulary without an enforcement point | just check-kinds, conformance B7 |
Merging a kind, tag, or MIP name that nothing reads |
| C-3 | Names never encode ownership | AGENTS.md § Design Law, C3 |
MQTT-style org/site/device topic prefixes anywhere in key expressions, storage keys, or ACLs |
| C-4 | Add attributes, not mechanisms | Design Law, F2 |
A policy DSL, a rules engine, or a plugin runtime in the bridge |
| C-5 | No third shim — one grant store, every enforcement point reads it | Design Law, F4 |
The MQTT bridge holding its own ACL copy or a sync job between stores |
| C-6 | Explicit denial, never a silent empty result | Design Law, B5 |
A bridge that drops unauthorized publishes without a named deciding record |
| C-7 | Revocation is monotonic and absorbing, never timestamp-wins | Design Law, D4 |
Any second writer or peering link before meridian-3wh lands |
| C-8 | Nothing is resolved until deployed to a live relay and re-probed | Design Law, F5 |
Closing any story on green CI alone |
| C-9 | Scale by replica / function / tenant / traffic class — never by vocabulary | DIAGRAM.md § Scaling axes, R1 |
Sharding relays by kind, MIP, or payload type |
| C-10 | Measure before optimizing; machinery for absent load is inventory | R8, meridian-ctg |
Any optimization story merging without its Sprint 1 baseline number |
The refusals this program will be asked to reverse
Name them now so they are declined by citation rather than re-argued:
- "Put Kafka behind the relay for durability." Refusal
R2. Theeventstable is already the ordered, partitioned, queryable log and carries the FTS index and audit chain. A broker buys retention we have and costs an operational tier plus truth ambiguity. - "Shard telemetry kinds onto their own relays." Refusal
R1. A kind is a column in every shard, never a shard. Use A3 (tenant) or A4 (traffic class). - "Let the edge relay peer with the spine." Refusal
R4/ Law 2. - "Cache grants at the bridge for speed." Refusal
R5/C-5.
3. Throughput model — capability examples with profiles
Rule for this entire section, enforced at review: no figure appears without
[MEASURED], [ESTIMATE], or [TARGET] and its profile. Mixing the three is
how a demo becomes a support obligation.
3.1 Per-event cost — the measured floor
Profile for every measured row: Apple M3 Max, release build, single-threaded, one core. Per-core figures — do not quote as relay throughput.
| Stage | Cost | Status | Source |
|---|---|---|---|
| BIP-340 Schnorr verify (no batch API exists) | 34–37.7 µs | [MEASURED] |
crates/meridian-core/benches/event_cost.rs |
| Keyed BLAKE3 MAC | 0.089 µs (11.2 M/s @ batch 64) | [MEASURED] |
same bench |
| HMAC-SHA256 MAC | 0.19 µs (5.27 M/s @ batch 64) | [MEASURED] |
same bench |
| Fan-out routing only, N=1 | 0.70 µs (channel-scoped) / 0.54 µs (global) | [MEASURED] |
crates/meridian-relay/benches/fanout_cost.rs |
| Fan-out + JSON frame construction, N=1 | 3.43 µs | [MEASURED] |
same bench |
| Fan-out + JSON frame construction, N=64 | 30.6 µs | [MEASURED] |
same bench |
| Fitted fan-out model | fixed ~2–3 µs serde_json::to_string + ~430 ns per additional recipient |
[MEASURED] |
fitted from the above |
| Subscription register/remove round trip | ~1.3 µs | [MEASURED] |
same bench |
| Async pipeline overhead, opaque path, unbatched | ~1–3 µs/msg | [ESTIMATE] |
04-PIPELINE-OVERHEAD.md — calibration, not data |
| Async pipeline overhead, batched at 64 | ~100–150 ns/msg | [TARGET] |
meridian-0zd acceptance criterion |
| Postgres single-row insert | ~30–100 µs | [ESTIMATE] |
zenoh cost budget |
| Interest-scoping cluster ingress reduction | 64× | [MEASURED] |
perf/relay_bus_scaling.py --mode redis, 64 communities, 1 subscribed |
The single most important row is the first one. A signature costs two orders of magnitude what a MAC costs — ~200× (HMAC-SHA256) and ~420× (keyed BLAKE3) at batch 64, the batch size the rows above are quoted at and the one this program's frame path actually runs. That gap is the entire argument for an opaque, pre-authenticated frame envelope, and it is why the throughput target is only reachable on a path that carries no per-message signature.
Per C-1, the batch size is part of the figure. schnorr_verify_nostr is flat
across the 1/8/64 sweep; the MAC groups are not, so at batch 1 the same ratios
are ~125× and ~405× — and against Schnorr's 21.6–33.1 K/s run-to-run spread the
honest bands there are ~110–170× and ~360–550× (ARCHITECTURE.md). This
paragraph previously read 179–382×, which divided MAC throughput at batch 64
by Schnorr throughput at batch 1 and so matched neither profile.
Two findings constrain how this ceiling can be raised, both already closed:
- Batch signature verification is unavailable and would not be enough. No
stable batch BIP-340 API exists at any level of the pinned tree
(
meridian-pt1). Measured on ed25519, where batching does exist, batch-64 amortizes to only ~2.9× and is a net loss below about N=4. - The way past the ceiling is to skip verification, never to accelerate it.
3.2 What 1.337M msg/s costs on each path
Target profile, per the answer that scoped this program: ingress, QoS 0 equivalent, small payload, never stored.
Ingest CPU required to sustain 1,337,000 msg/s on one host:
| Path | Per-msg cost | Cores required | Verdict |
|---|---|---|---|
| Signed NIP-01 events | 37.7 µs [MEASURED] |
50.4 cores | Not viable. ~3.2 × 16-core pods for verification alone, before routing, framing, or egress |
| Opaque frame + HMAC-SHA256 | 0.19 µs [MEASURED] |
0.25 cores | viable |
| Opaque frame + keyed BLAKE3 | 0.089 µs [MEASURED] |
0.12 cores | viable — specify BLAKE3 |
| Opaque frame, P0 trusted link, no MAC | 0 | 0 cores | viable where the link is trust-domain-internal |
| Pipeline overhead, unbatched | ~2 µs [ESTIMATE] |
2.67 cores | the dominant cost today |
| Pipeline overhead, batched at 64 | ~150 ns [TARGET] |
0.20 cores | the meridian-0zd win |
Conclusion: 1.337M msg/s ingress is architecturally reachable on the opaque-frame path and architecturally unreachable on the signed path. This is the number that must appear in the handoff, and it must appear with both halves.
3.3 The binding constraint is egress, not ingest
\text{egress frames/s} = \text{ingress} \times \sum_{s \in \text{subscribers}} \text{keyspace fraction}(s)
At 1.337M msg/s ingress, with a 92-byte frame (44 B header + 48 B payload):
| Fan-out ratio | Egress frames/s | CPU @ 150 ns [TARGET] |
Egress bandwidth | NIC required |
|---|---|---|---|---|
| 1× | 1.34 M/s | 0.20 cores | 0.98 Gbps | 10 GbE |
| 8× | 10.7 M/s | 1.6 cores | 7.9 Gbps | 10 GbE |
| 32× | 42.8 M/s | 6.4 cores | 31.4 Gbps | 40 GbE |
| 64× | 85.6 M/s | 12.8 cores | 62.9 Gbps | 100 GbE |
| 128× | 171 M/s | 25.7 cores — exceeds one 16-core pod | 125.8 Gbps — exceeds 100 GbE | — |
The same table on the NIP-01 JSON path, using the measured fan-out curve:
| Fan-out ratio | Per-event fan-out cost | CPU required |
|---|---|---|
| 1× | 4.1 µs [MEASURED] |
5.5 cores |
| 64× | 56 µs [MEASURED] |
74.9 cores |
Read the two tables together. Interest scoping is the only lever on the egress term, and it is already measured at 64× cluster ingress reduction. Anything that raises the average subscriber's keyspace fraction — a wildcard subscription, a debugging firehose left on, a poorly-scoped bridge topic map — multiplies the entire system cost. Egress discipline is a design property, not a tuning knob.
3.4 Batching arithmetic
At 64 messages per frame, 1.337M msg/s becomes 20,900 pipeline traversals/s instead of 1,337,000 — a 64× reduction in syscalls, channel hops, cross-core wakeups, and allocations.
Header overhead falls from ~35% unbatched to ~6% batched, a ~25% bandwidth saving, and one TLS record per batch replaces one per message.
The Zenoh trap, to be named explicitly in every review of
MRDN-3xx: a batching transport whose win is destroyed byfor msg in batch { tx.send(msg).await }immediately above it. The batch must survive all the way to fan-out or the transport batching buys nothing.
3.5 Published capability examples
These are the figures the handoff may publish. Each is a pair — the number and the obligation it does not carry.
| # | Profile | Figure | Status |
|---|---|---|---|
| T1 | Opaque pre-decoded frames · P0 trusted link · payload never parsed · not stored · batched at 64 | ~600k msg/s per pod | [ESTIMATE] — path does not exist in tree; pre-/post-dedup question open at ~10× |
| T2 | Same as T1, post-meridian-0zd batching refactor |
1.0–1.5M msg/s per pod | [TARGET] — this program's goal; MRDN-810 is the gate |
| T3 | Signed NIP-01 events · per 16-core pod · BIP-340 bound | ~250–320k events/s | [MEASURED] per-core, extrapolated |
| T4 | Stored and indexed chat events · one Postgres | ~10–30k/s | [ESTIMATE] — latency measured at p95 74.6 ms direct |
| T5 | Bus — Dragonfly, interest-scoped | 64× cluster ingress reduction | [MEASURED] |
| T6 | Bus — Zenoh peer mode, 256 B, one pod pair | 1.15–1.51× Redis | [MEASURED] — meridian-ctg ran and 256 B failed the ≥5× gate; the ≥2M msg/s target is withdrawn. The ≥4 KB crossover is also withdrawn — shared memory was on above its 3,072 B threshold, on a path the relay build cannot take. Below 4 KB stands |
T1 through T6 span roughly two orders of magnitude and share a process and almost nothing else. Quoting T2 for a workload that looks like T4 is the failure mode this table exists to prevent.
4. Technology stack — frozen for this program
Versions are the pinned workspace values, verified in-tree. Changes go through a council per AGENTS.md § User Preferences, not through a PR.
4.1 Relay and data plane — Rust
| Concern | Choice | Version | Note |
|---|---|---|---|
| Language | Rust | 1.88.0, edition 2021 | rust-toolchain.toml; Hermit-pinned |
| Async runtime | tokio |
1.x, multi-thread | — |
| HTTP + WebSocket | axum |
0.8 (ws, macros) |
existing edge |
| Binary framing | postcard |
1.x | the MIP-OF wire encoder — already a workspace dep |
| Event store | sqlx + Postgres |
0.9 / PG 17 | the log, the query store, and the FTS index — one durable write |
| Bus (launch) | redis crate → Dragonfly |
1.0 | Redis wire protocol; do not reintroduce a redis service |
| Bus (candidate) | Zenoh | — | gated on meridian-ctg; not authorized by this program |
| Inter-relay mesh | iroh |
1.0.0-rc.0 | existing; leaf transport candidate for MIP-LF |
| Frame MAC | blake3 (keyed) |
1.8 — promote to workspace dep | Currently a meridian-core dev-dependency (benches only). Measured 0.089 µs vs HMAC's 0.19 µs — specify BLAKE3, and promote it to a production dependency in MRDN-312 |
| Existing MAC | hmac + sha2 |
0.13 / 0.11 | fallback / interop profile |
| KDF | nostr::util::hkdf via hkdf32 |
nostr 0.44 |
already exists at meridian-core/src/pairing/crypto.rs:34 — MIP-SF adds no new crypto dependency |
| Concurrency | dashmap, moka |
6 / 0.12 | subscription registry, freshness-stamped caches |
| Observability | tracing + opentelemetry + metrics-exporter-prometheus |
0.1 / 0.32 / 0.18 | meridian_* metric naming convention |
| Errors | thiserror / anyhow |
2 / 1 | no new unwrap()/expect() in production paths |
| Benchmarks | criterion |
— | existing event_cost.rs, fanout_cost.rs |
| Property tests | proptest |
1 | mandatory for MIP-OF header codec |
New crates created by this program:
| Crate | Purpose |
|---|---|
meridian-frame |
MIP-OF envelope: header codec, batch framing, movement contracts. No I/O, no async — pure, fuzzable, no_std-friendly where practical |
meridian-leaf |
Hub-and-leaf edge relay: local termination, store-and-forward, upstream forwarding |
meridian-mqtt |
MQTT 3.1.1 / 5.0 ingress adapter and feeder. Depends on meridian-frame, never the reverse |
meridian-loadgen |
Reproducible load generator and capacity harness |
Open decision — MQTT codec (MRDN-601). Evaluate mqttbytes (rumqtt
workspace, pure codec) against mqtt-protocol, on: MQTT 5.0 property coverage,
no_std/allocation behavior, deny.toml license compatibility, and maintenance.
Do not vendor a full broker — we need the codec, not the delivery engine.
4.2 Frontend — recommendation against Next.js
The brief named "Rust/Next.js/Vite". There is no Next.js anywhere in this repository, and I recommend against introducing it. The verified in-tree standard is uniform across all three web surfaces:
| Surface | Stack (verified) |
|---|---|
REMAPPING/meridian-desktop/ |
React 19.1 · Vite 8 · TypeScript 6 · Tailwind 4.3 · TanStack Query 5.90 · Tauri 2.11 |
REMAPPING/meridian-web/ |
React 19.1 · Vite 8 · TypeScript 6 · Tailwind 4.3 · TanStack Query 5.90 |
REMAPPING/meridian-admin-web/ |
React 19.1 · Vite 8 · TypeScript 6 |
Reasons to hold the line, in order of weight:
- The operator console is a live-telemetry SPA behind auth. Next.js's differentiators — SSR, RSC, file-system routing, edge middleware — buy nothing for a page whose content arrives over a WebSocket. The cost is a second build system, a second lint/format config, a second dependency tree, and a second set of CI gates.
REMAPPING/meridian-admin-web/already exists and already speaks the relay admin API. The throughput console is an extension of it, not a new application.- Biome, the px-text CI guard, and the shared token scale are configured per-workspace. A Next.js app either re-implements them or silently escapes them.
Where Next.js would be justified, and the one place I propose it: a
public, statically-generated MIP registry and specification portal — content
that genuinely benefits from SSG, SEO, and markdown-first authoring, and that has
no relay connection at all. That is MRDN-730, explicitly optional, and it must
not import from REMAPPING/meridian-admin-web/ or REMAPPING/meridian-web/.
| Surface | Stack | Story |
|---|---|---|
| Operator console — live throughput, lag, class shed, device registry | Extend REMAPPING/meridian-admin-web/ — React 19 + Vite 8 + TS 6 |
MRDN-720 |
| MIP registry / spec portal (public, static) | Next.js (SSG only, no server runtime) or Astro | MRDN-730 — optional |
| Device simulator UI | CLI first; no GUI in scope | MRDN-711 |
If the other team has a standing Next.js platform requirement I have not been told about, this is the decision to reverse first — say so and
MRDN-720re-scopes to a standalone app. Nothing downstream of it changes.
4.3 Deployment
Unchanged from the repo standard: Helm charts in deploy/charts/, Compose in
deploy/compose/, images to registry.r2d2.office.ilab.zone/meridian-*,
release flow per RELEASING.md. The leaf relay and MQTT bridge
each get a chart and a Compose profile; neither becomes a core (unprofiled)
Compose service, because just _ensure-services blocks on core health.
5. The MIP system — Meridian Implementation Possibilities
5.1 The deconfliction rule
Upstream Nostr NIPs are numeric (NIP-01, NIP-42). Repo-local proposals
are two uppercase letters (MIP-OF, MIP-SF). The namespaces cannot
collide because digits and letters are disjoint — that is the deconfliction,
and it needs no registry negotiation with upstream.
The repo currently has this half-done and inconsistent: 17 files under
docs/nips/ use NIP-XX letter codes, while beads already use MIP- for 14
newer proposals (MIP-OF, MIP-KN, MIP-SM, MIP-DM, MIP-PS, MIP-HU,
MIP-FR, MIP-CN, MIP-AL, MIP-MN, MIP-IB, MIP-GR, MIP-FU, MIP-TK).
MRDN-201 completes the migration: every repo-local proposal becomes
MIP-XX and moves to docs/mips/, with docs/nips/ retained only for
profiles of genuinely upstream NIPs.
Why the letter code and not MIP-0001: a numeric scheme invites the reader to
assume ordering and completeness, and it creates a merge conflict on every
concurrent proposal. A two-letter mnemonic is stable, greppable, and conflicts
loudly rather than silently.
5.2 Lifecycle
flowchart LR
classDef d fill:#4a4a4a,stroke:#2b2b2b,color:#fff
classDef r fill:#7d6608,stroke:#4d3f05,color:#fff
classDef i fill:#1f6aa5,stroke:#123f63,color:#fff
classDef e fill:#1e7a45,stroke:#0f4527,color:#fff
classDef x fill:#b03030,stroke:#6e1c1c,color:#fff
D["DRAFT<br/>problem + wire shape<br/>no code"]:::d
R["REVIEW<br/>council seated<br/>reversal evidence named"]:::r
I["IMPLEMENTABLE<br/>frozen wire format<br/>golden vectors published"]:::i
E["ENFORCED<br/>reader exists<br/>just check-kinds green"]:::e
W["WITHDRAWN<br/>reason recorded"]:::x
D --> R --> I --> E
R --> W
I --> W
ENFORCED is the only state that permits NIP-11 advertisement. This is
constraint C-2, and the failure it prevents is documented: NIP-17 was
advertised for months while kind:10050 was rejected by the ingest allowlist, so
conformant third-party clients feature-detected support and then failed silently.
A negotiation field that promises frames the relay will reject reproduces exactly
that bug on the hot path.
5.3 The MIP suite for this program
Eight specifications. Two exist as beads already; six are new.
| MIP | Title | Layer | Bead | Sprint | Status |
|---|---|---|---|---|---|
| MIP-RG | Proposal registry, lifecycle, and numbering | Process | new | 1 | Draft |
| MIP-OF | Opaque frame envelope — header, batch framing, movement contracts | Wire | meridian-tdq |
1–2 | Draft |
| MIP-XP | Negotiated exchange profiles — per-link auth floors (P0/P1/P2) | Link | meridian-2e2 |
2 | Draft |
| MIP-SF | Session frames — removing BIP-340 from the ephemeral hot path | Wire | new (from zenoh spec) | 2 | Draft |
| MIP-QC | QoS classes and congestion contracts — 8 lanes, class-aware shed | Scheduling | meridian-ygr, meridian-6kx, meridian-9j8 |
3 | Draft |
| MIP-LF | Leaf forwarding — hub-and-leaf topology, store-and-forward, no peering | Topology | new | 3–4 | Draft |
| MIP-MQ | MQTT interoperability — topic mapping, QoS mapping, session semantics | Interop | new | 4 | Draft |
| MIP-KN | Kinematic state payload schema | Payload | meridian-ss1 |
4 | Draft |
| MIP-TL | Scalar telemetry / sensor reading payload schema | Payload | new | 4 | Draft |
| MIP-AS | AIS feed profile — the first real workload on the machine surface | Profile | meridian-hgjj |
1–4 | Draft, written |
MIP-AS is the profile that makes the rest of the suite demand-driven.
Written at docs/mips/MIP-AS.md, it binds MIP-OF, MIP-KN
and MIP-XP to a live AIS feed, and it is the reason this program stops being
inventory: every spec above now has a named consumer. It adds no mechanism, and
it returns one resolution to MIP-OF — payload_type and kind are the same
header field (§ 5.1), because the 44-byte budget has room for one and the
relay's only two uses of either are the QoS lane and consumer discrimination.
Layering is strict and one-directional. A payload MIP may never require the relay to read it — that is the enabling property, not a limitation:
MIP-RG (process — governs the rest)
│
MIP-XP (link auth) ──► MIP-OF (envelope) ──► MIP-QC (scheduling)
│ │
├─► MIP-SF (frame auth) │
├─► MIP-KN (payload) │
└─► MIP-TL (payload) │
MIP-LF (topology)
│
MIP-MQ (interop)
5.4 MIP-OF at a glance — the load-bearing spec
The governing rule, inherited verbatim from 03-STREAMING-INGEST.md:
The relay reads a fixed-size routing header and forwards opaque bytes. It never interprets the payload.
\textbf{Anything the relay must act on MUST live in the header.}
That single rule allocates the whole system:
| The relay must… | Needs | Where it lives |
|---|---|---|
| Fence the tenant | community | header |
| Pick a QoS lane | kind | header |
| Route by region | geohash cell | header — the feeder already decoded the position |
| Conflate superseded state | conflation key | header |
| Detect gaps | sequence | header |
| Expire stale state | timestamp / TTL | header |
| Deduplicate across receivers | content comparison | upstream — the relay has forsworn reading |
Header budget ~44 bytes. Dedup moves to the feeder because it requires reading content; conflation stays at the relay only because its key is a header field. That is the whole trade in one line.
Three movement contracts, declared by the publisher (F1 in the conformance
register):
| Contract | Congestion behavior | Example |
|---|---|---|
persisted |
block; never drop | chat, commands, audit-relevant events |
volatile |
drop oldest under pressure | typing, presence, observer frames |
conflated |
replace superseded state by conflation key | position reports, sensor state |
conflated is the one MQTT has no equivalent for, and it is the correct answer
for telemetry: under congestion a stale position should be replaced, not queued
or dropped by luck.
5.5 Payload examples
Several payload types ride one envelope. The relay distinguishes them only by a
payload_type header field it never dereferences.
| Type | MIP | Shape | Batching axis |
|---|---|---|---|
kinematic |
MIP-KN | position, velocity, heading, altitude, mandatory non-defaulting provenance: self-reported | observed |
temporal (N samples, 1 platform) or spatial (N platforms, 1 sample) |
telemetry |
MIP-TL | typed scalar readings — unit, value, quality, sensor id | temporal |
mqtt-passthrough |
MIP-MQ | original MQTT payload bytes, untouched; topic preserved in an attribute | spatial |
opaque |
MIP-OF | uninterpreted bytes; consumer-defined | either |
The provenance field is mandatory and non-defaulting for a specific reason: a
drone signs its own telemetry, an ADS-B target signs nothing and the feeder
attests. That distinction must not be droppable by omission — a defaulted
provenance field would silently promote observed data to self-reported.
6. Epic register
| Epic | Key | Title | Priority | Points | Sprints |
|---|---|---|---|---|---|
| E1 | MRDN-100 |
Measurement baseline and capacity model | P0 | 34 | 1 |
| E2 | MRDN-200 |
MIP specification suite | P0 | 47 | 1–4 |
| E3 | MRDN-300 |
Opaque frame data plane | P0 | 63 | 2–5 |
| E4 | MRDN-400 |
Traffic class, QoS, and congestion | P1 | 34 | 3–4 |
| E5 | MRDN-500 |
Hub-and-leaf topology and store-and-forward | P1 | 42 | 3–5 |
| E6 | MRDN-600 |
MQTT compatibility bridge | P1 | 47 | 3–5 |
| E7 | MRDN-700 |
Client SDK and operator console | P2 | 34 | 2–5 |
| E8 | MRDN-800 |
Conformance, load, and handoff | P0 | 42 | 1–5 |
| Total | 343 |
343 points against a 260-point budget. This is deliberate and must be resolved at Sprint 1 planning rather than discovered in Sprint 4. § 11 names the cut line: E7 drops to CLI-only, E6 descopes MQTT 5.0 to 3.1.1, and the
MIP-TLpayload schema defers. That leaves 258 points. Do not resolve it by cutting E1 or E8 — those are the gates that make every other number true.
Dependency graph
flowchart TB
classDef gate fill:#b03030,stroke:#6e1c1c,color:#fff
classDef core fill:#1f6aa5,stroke:#123f63,color:#fff
classDef ext fill:#1e7a45,stroke:#0f4527,color:#fff
classDef opt fill:#4a4a4a,stroke:#2b2b2b,color:#fff
E1["E1 · Measurement baseline<br/>THE GATE — nothing optimizes<br/>before this has a number"]:::gate
E2["E2 · MIP specs<br/>frozen wire format"]:::core
E3["E3 · Opaque frame data plane"]:::core
E4["E4 · QoS + congestion"]:::core
E5["E5 · Hub-and-leaf"]:::ext
E6["E6 · MQTT bridge"]:::ext
E7["E7 · SDK + console"]:::opt
E8["E8 · Conformance + handoff<br/>runs continuously"]:::gate
E1 --> E2 --> E3
E3 --> E4
E3 --> E5 --> E6
E3 --> E7
E1 --> E8
E3 --> E8
E4 --> E8
E6 --> E8
7. Sprint plan — 10 weeks
| Sprint | Weeks | Theme | Exit gate |
|---|---|---|---|
| S1 | 1–2 | Measure and specify | Baseline numbers published; MIP-OF and MIP-XP at IMPLEMENTABLE; go/no-go on the whole program |
| S2 | 3–4 | Frame path exists | meridian-frame codec fuzzed and golden-vectored; relay parses and routes a frame end to end |
| S3 | 5–6 | Batch and class | Batch survives to fan-out; 8 QoS lanes wired; class-aware shed proven under a slow consumer |
| S4 | 7–8 | Edge and interop | Leaf relay forwards upstream with store-and-forward; MQTT 3.1.1 bridge publishes into the spine |
| S5 | 9–10 | Prove and hand off | Sustained-rate capacity run with attached profile; conformance suite green; handoff package frozen |
Sprint 1 gate — the one that can stop the program
MRDN-110 is a go/no-go, not a checkpoint. Per C-10 and refusal R8,
if the Sprint 1 measurements show the opaque-frame path is dominated by
something unlisted — a TLS setting, an accidental debug! in the fan-out loop, a
clone() on a hot struct — then the batching and QoS work in S3 is optimizing
the wrong term and must be re-planned before it starts.
The named outcomes:
| Measurement | If it lands | If it does not |
|---|---|---|
| Per-message frame-path cost within ~2× of the 1–3 µs estimate | Proceed as planned | Re-scope E3 around the actual dominant term |
| Batching projects ≥10× on the pipeline traversal count | Proceed with MRDN-310 |
Drop meridian-0zd; find the real cost |
Zenoh ≥5× over Dragonfly on one fixed topology (meridian-ctg) |
Zenoh becomes a candidate | Stay on Dragonfly; defer the implementation chain — this is the existing bead's own gate |
8. Backlog
Format: Key · Summary · Type · Priority · Points · Sprint · Bead · Depends-on.
Acceptance criteria are Given/When/Then where behavior is testable and a named
artifact where the deliverable is a document or a measurement.
E1 · MRDN-100 — Measurement baseline and capacity model · P0 · 34 pts
The gate. Every optimization story downstream cites a number produced here.
| Key | Summary | Type | Pri | Pts | Sprint | Bead |
|---|---|---|---|---|---|---|
MRDN-101 |
Establish our own baseline before adopting the 1.337M target | Story | P0 | 5 | 1 | new |
MRDN-102 |
Frame-path criterion bench: header parse + key match + fan-out at N=1/8/64 | Story | P0 | 8 | 1 | meridian-lq7 |
MRDN-103 |
Flamegraph of a sustained synthetic 600k msg/s run | Story | P0 | 5 | 1 | meridian-lq7 |
MRDN-104 |
Counters: allocations/s, cross-thread wakeups/s, syscalls/s | Story | P0 | 5 | 1 | meridian-lq7 |
MRDN-105 |
Zenoh Phase 0 — same-topology Dragonfly vs Zenoh at 256 B | Story | P1 | 8 | 1 | meridian-ctg |
MRDN-110 |
Capacity model + go/no-go decision record | Story | P0 | 3 | 1 | new |
MRDN-101 — Establish our own baseline before adopting the 1.337M target.
Given the 1.337M msg/s figure is vendor-published for a TBMQ-class broker and has not been reproduced here, When the capacity model is published, Then it states the target's origin, the profile it was published under (node count, instance types, QoS level, payload size, clean vs persistent sessions, fan-in vs fan-out ratio), and our own measured starting point on identical hardware — and it makes no head-to-head claim.
Per the ratified council call in DIAGRAM.md § Part III: publish the comparison as an architecture and cost-model comparison, never as a benchmark result. A benchmark of an unfamiliar system tuned by its competitor is worth nothing, and publishing one would cost more credibility than the number could buy.
MRDN-110 — Capacity model + go/no-go. Deliverable is
docs/benchmarks/frame-path-baseline-<date>.md containing: the measured
per-message cost breakdown replacing the estimate table; the projected ceiling
with its profile; the egress arithmetic from § 3.3 with measured per-recipient
cost substituted for the target; and an explicit PROCEED / RE-PLAN verdict
signed by the tech lead.
E2 · MRDN-200 — MIP specification suite · P0 · 47 pts
| Key | Summary | Type | Pri | Pts | Sprint | Bead | Depends |
|---|---|---|---|---|---|---|---|
MRDN-201 |
MIP-RG: registry, lifecycle, numbering; migrate docs/nips/ → docs/mips/ |
Story | P0 | 5 | 1 | new | — |
MRDN-210 |
MIP-OF: opaque frame envelope | Story | P0 | 13 | 1–2 | meridian-tdq |
MRDN-110 |
MRDN-211 |
MIP-OF golden test vectors (published artifact) | Story | P0 | 5 | 2 | new | MRDN-210 |
MRDN-220 |
MIP-XP: exchange profiles P0/P1/P2 + XP-ACCEPT negotiation | Story | P0 | 8 | 2 | meridian-2e2 |
MRDN-110 |
MRDN-230 |
MIP-SF: session frames — ephemeral-only fence | Story | P1 | 8 | 2 | new | MRDN-220 |
MRDN-240 |
MIP-QC: 8 QoS lanes + three movement contracts | Story | P1 | 5 | 3 | meridian-9j8 |
MRDN-210 |
MRDN-250 |
MIP-LF: leaf forwarding topology | Story | P1 | 5 | 3 | new | MRDN-210 |
MRDN-260 |
MIP-MQ: MQTT topic/QoS/session mapping | Story | P1 | 5 | 4 | new | MRDN-250 |
MRDN-270 |
MIP-KN: kinematic payload schema | Story | P2 | 3 | 4 | meridian-ss1 |
MRDN-210 |
MRDN-271 |
MIP-TL: scalar telemetry payload schema | Story | P2 | 3 | 4 | new | MRDN-210 |
MRDN-210 — MIP-OF. The load-bearing spec. Sub-tasks:
| Sub-task | Deliverable |
|---|---|
MRDN-210.1 |
Fixed ~44 B header field layout, byte-exact, with alignment and endianness stated |
MRDN-210.2 |
Batch framing: count, per-message delta encoding, max batch size |
MRDN-210.3 |
The opaque-payload contract, written as a negative requirement the relay must not violate |
MRDN-210.4 |
Three movement contracts + conflation key semantics |
MRDN-210.5 |
payload_type registry and its extension rule |
MRDN-210.6 |
Explicit statement of which costs are measured and which remain estimated |
Given MIP-OF reaches
IMPLEMENTABLE, When a reviewer reads it without access to this repository, Then they can implement a byte-compatible encoder and decoder, and every field is justified by a routing decision the relay must make.
Circular-dependency note, inherited from meridian-tdq: the bead says
"write MIP-OF after meridian-lq7 measures", and meridian-lq7 measures the
frame path that MIP-OF specifies. Neither can go first under a literal reading.
Resolution: write MIP-OF against the already-landed fanout_cost.rs and
event_cost.rs measurements, and state explicitly in the spec that header-parse
and key-match cost is the one term still estimated. The frame-specific bench then
lands against the implementation. MRDN-102 is re-scoped accordingly.
MRDN-230 — MIP-SF fence. The entire safety argument rests on one rule, so
it is enforced at compile time rather than by review:
// per eligible kind, in kind.rs
const _: () = assert!(is_ephemeral(k));
Given a session frame arrives carrying a persistent kind, When the relay processes it, Then the frame is rejected — persistent kinds always require a client signature — and a counter increments with the deciding reason.
The catastrophic version of this idea is extending session frames to persistent kinds: that converts "the relay can lie about who is typing" into "the relay can forge chat history." The compile-time assertion exists so that extension cannot land by accident.
E3 · MRDN-300 — Opaque frame data plane · P0 · 63 pts
| Key | Summary | Type | Pri | Pts | Sprint | Bead | Depends |
|---|---|---|---|---|---|---|---|
MRDN-301 |
Create meridian-frame crate — pure, no I/O, no async |
Story | P0 | 5 | 2 | new | MRDN-210 |
MRDN-302 |
Header codec + proptest round-trip + fuzz target |
Story | P0 | 8 | 2 | new | MRDN-301 |
MRDN-303 |
Batch framing encoder/decoder | Story | P0 | 5 | 2 | new | MRDN-302 |
MRDN-304 |
Frame ingress on the relay — parse header, never touch payload | Story | P0 | 8 | 3 | meridian-8ft |
MRDN-303 |
MRDN-305 |
Key-expression matcher for frame routing (geohash prefix nesting) | Story | P0 | 8 | 3 | meridian-8ft |
MRDN-304 |
MRDN-306 |
Zero-copy fan-out — one buffer, N refcounted sends | Story | P0 | 8 | 3 | new | MRDN-305 |
MRDN-310 |
Batch as the unit of work through the whole pipeline | Story | P0 | 13 | 3–4 | meridian-0zd |
MRDN-306 |
MRDN-311 |
MIP-XP P0/P1 link authentication on the frame ingress | Story | P1 | 5 | 4 | meridian-2e2 |
MRDN-220 |
MRDN-312 |
Keyed BLAKE3 frame MAC (MIP-SF steady state) | Story | P1 | 3 | 4 | new | MRDN-230 |
MRDN-306 — zero-copy fan-out. This is where the measured win comes from.
The fitted model shows a fixed ~2–3 µs serde_json::to_string dominating
per-event fan-out cost, and it is already serialized once and shared across
recipients via fanout_frame_cache — so batching does not remove it. Only
avoiding JSON entirely does.
Given a frame with N subscribers, When the relay fans it out, Then the payload bytes are never copied and never re-serialized — per-recipient cost is a refcount bump and a ring push — verified by an allocation counter showing O(1) allocations per frame regardless of N.
MRDN-310 — batching. Acceptance is a measured delta against the
pre-refactor baseline already committed in fanout_cost.rs:
Given the Sprint 1 baseline, When the batching refactor lands, Then measured per-message cost on the frame path drops from ~1–3 µs toward the ~100–150 ns residual (header parse + key match + refcount + ring push), verified by re-running the criterion bench and the flamegraph.
Two regression guards that must not break (landed under meridian-22y,
mutation-verified to fail on the exact mistakes batching invites):
- One frame per subscription, not per connection — batching invites per-connection coalescing, which is wrong.
- A full-buffer recipient is counted as backpressured without starving the rest — batching invites abort-on-first-failure.
Shape of the refactor, stated so review can check it: async at the batch boundary, synchronous loop inside, async at the egress boundary. Parsing a fixed header, matching a key, bumping a refcount and pushing to a ring is pure CPU work over a buffer and needs no async at all.
E4 · MRDN-400 — Traffic class, QoS, and congestion · P1 · 34 pts
| Key | Summary | Type | Pri | Pts | Sprint | Bead | Depends |
|---|---|---|---|---|---|---|---|
MRDN-401 |
kind_to_qos() in kind.rs, exhaustive, compile-time asserted |
Story | P1 | 5 | 3 | meridian-ygr |
MRDN-240 |
MRDN-402 |
Wire QoS class into frame publish path | Story | P1 | 5 | 4 | meridian-ygr |
MRDN-401 |
MRDN-403 |
Per-priority queue sizing + wait_before_close |
Story | P1 | 5 | 4 | meridian-ygr |
MRDN-402 |
MRDN-410 |
Class-aware last-hop shed — replace broadcast(4096) Lagged |
Story | P1 | 8 | 4 | meridian-6kx |
MRDN-402 |
MRDN-420 |
conflated movement contract — replace superseded state by key |
Story | P1 | 8 | 4 | meridian-9j8 |
MRDN-402 |
MRDN-430 |
Congestion load test across all three movement contracts | Story | P1 | 3 | 4 | new | MRDN-420 |
MRDN-410 is the story that makes the QoS work real. Three
broadcast::channel(4096) in meridian-pubsub/src/lib.rs shed via
RecvError::Lagged, which drops arbitrarily rather than by class. Wiring
eight priority lanes on the bus and then discarding the classification at the
socket buys scheduling quality on the wire and throws it away at the last hop.
Given a slow consumer subscribed to both a
volatileand apersistedclass, When the connection buffer saturates, Thenvolatile/Droptraffic sheds before anypersisted/Blocktraffic, and the shed is counted per class rather than reported as a generic lag.
MRDN-401 establishes a merge rule mirroring the existing "do not merge a
kind constant without a reader": do not merge a kind constant without a QoS
class. Enforced by an exhaustive match plus per-range compile-time assertions.
E5 · MRDN-500 — Hub-and-leaf topology and store-and-forward · P1 · 42 pts
| Key | Summary | Type | Pri | Pts | Sprint | Bead | Depends |
|---|---|---|---|---|---|---|---|
MRDN-501 |
Create meridian-leaf crate |
Story | P1 | 5 | 3 | new | MRDN-250 |
MRDN-502 |
Leaf ingress: terminate device connections locally | Story | P1 | 8 | 4 | new | MRDN-501 |
MRDN-503 |
Upstream forwarding over MIP-XP P1 link | Story | P1 | 8 | 4 | new | MRDN-311 |
MRDN-510 |
Store-and-forward across an upstream outage | Story | P1 | 13 | 4–5 | new | MRDN-503 |
MRDN-520 |
Leaf-as-peer refusal test (negative test for Law 2) | Story | P1 | 5 | 5 | new | MRDN-503 |
MRDN-530 |
Leaf Helm chart + Compose profile (non-core) | Story | P2 | 3 | 5 | new | MRDN-510 |
MRDN-510 — store-and-forward. The one place this program touches
durability, and the place a second durable log will be proposed. It must not be.
Given a leaf relay with devices publishing at rate R, When the upstream spine is unreachable for duration D, Then the leaf buffers to a bounded local store, publishes its backlog depth as a metric, applies the movement contract on overflow (
conflatedreplaces,volatiledrops oldest,persistedapplies backpressure to the device), and drains in order on reconnect without duplicating events the spine already holds.
The refusal to state at review: the leaf's buffer is a bounded forwarding
queue, never a system of record. If someone proposes making it durable and
authoritative, that is refusal R2 (a second durable log) and it is declined by
citation. A leaf that holds truth is a second writer, and a second writer
activates the dormant D4 revocation defect.
MRDN-520 is a negative test, and it is cheap insurance:
Given a leaf relay, When it attempts to register as a routing peer or transit traffic between two communities, Then the spine refuses the registration with an explicit named reason, and the refusal is covered by a test that fails if someone later "fixes" it.
E6 · MRDN-600 — MQTT compatibility bridge · P1 · 47 pts
| Key | Summary | Type | Pri | Pts | Sprint | Bead | Depends |
|---|---|---|---|---|---|---|---|
MRDN-601 |
Decide the MQTT codec crate (mqttbytes vs mqtt-protocol) |
Spike | P1 | 3 | 3 | new | — |
MRDN-602 |
Create meridian-mqtt crate; CONNECT/CONNACK, MQTT 3.1.1 |
Story | P1 | 8 | 4 | new | MRDN-601 |
MRDN-603 |
PUBLISH → MIP-OF frame, QoS 0 | Story | P1 | 8 | 4 | new | MRDN-602 |
MRDN-610 |
Topic filter → key expression mapping | Story | P1 | 8 | 4 | new | MRDN-603 |
MRDN-611 |
SUBSCRIBE/UNSUBSCRIBE → interest declaration | Story | P1 | 5 | 5 | new | MRDN-610 |
MRDN-620 |
QoS 1 (PUBACK) at the bridge, mapped to persisted |
Story | P1 | 8 | 5 | new | MRDN-603 |
MRDN-630 |
Retained messages + Last Will & Testament at the leaf | Story | P2 | 5 | 5 | new | MRDN-620 |
MRDN-640 |
MQTT 5.0 properties (descope candidate) | Story | P3 | 5 | 5 | new | MRDN-602 |
MRDN-610 is the story where constraint C-3 is won or lost. MQTT topics
are hierarchical strings that encode ownership by convention
(acme/site7/floor2/sensor9). Meridian key expressions must not.
Given an MQTT topic filter, When the bridge maps it to a key expression, Then the mapping is a table lookup against a registry, never a string transformation — the topic string is carried as an opaque attribute for the device's benefit and is never part of the routing key, the storage key, or any ACL.
The failure this prevents is named in C3: a producer key that freezes into
identity, ACL entry, topic prefix, and storage key all at once, while ownership
stays mutable — so a reorganization becomes a fleet-wide ACL edit.
Explicitly out of scope for the bridge, per C-5:
- Shared subscriptions (
$share/). We have no equivalent; every matching subscription receives every event. Document it as a known gap, do not emulate it, and do not let a partial emulation ship as if complete. - QoS 2 exactly-once. We are at-most-once on the bus and idempotent at the store, which composes to at-least-once-with-dedup. That is not the same contract and must not be described as if it were.
- A bridge-local ACL cache. Refusal
R5.
MRDN-620 note. QoS 1 PUBACK is answerable at the bridge only after the
spine has durably committed, or the acknowledgement is a lie. That couples the
bridge's ack latency to the spine's persistence latency — measured at p95
74.6 ms on the current reply path. Devices with tight ack timeouts will need
either the volatile contract at QoS 0 or a spine-side latency story. Surface
this at Sprint 4 planning; do not discover it in Sprint 5.
E7 · MRDN-700 — Client SDK and operator console · P2 · 34 pts
| Key | Summary | Type | Pri | Pts | Sprint | Bead | Depends |
|---|---|---|---|---|---|---|---|
MRDN-701 |
Rust client crate: connect, negotiate MIP-XP, publish frames | Story | P2 | 8 | 3 | new | MRDN-303 |
MRDN-702 |
Subscribe by key expression, receive batches | Story | P2 | 5 | 4 | new | MRDN-701 |
MRDN-711 |
meridian-loadgen — reproducible device simulator |
Story | P1 | 8 | 3 | new | MRDN-701 |
MRDN-720 |
Operator console — extend REMAPPING/meridian-admin-web/ (React 19 + Vite 8) |
Story | P2 | 8 | 5 | new | MRDN-810 |
MRDN-730 |
MIP registry portal (Next.js SSG) — optional | Story | P3 | 5 | 5 | new | MRDN-201 |
MRDN-711 is priority P1 despite sitting in a P2 epic: E8 cannot produce a
capacity number without it. Ship it before the console.
MRDN-720 renders, at minimum: ingress rate by class, egress rate and fan-out
ratio, per-class shed counters, leaf backlog depth, and bridge session count.
Per Law 4, consumer/leaf lag is a published number, not an invisible property.
E8 · MRDN-800 — Conformance, load, and handoff · P0 · 42 pts
| Key | Summary | Type | Pri | Pts | Sprint | Bead | Depends |
|---|---|---|---|---|---|---|---|
MRDN-801 |
Extend meridian-conformance with a MIP-OF suite |
Story | P0 | 8 | 2 | new | MRDN-211 |
MRDN-802 |
Payload golden vectors — every payload_type |
Story | P0 | 5 | 4 | new | MRDN-270 |
MRDN-803 |
Client conformance suite | Story | P1 | 5 | 4 | new | MRDN-702 |
MRDN-810 |
Sustained-rate capacity run with attached profile | Story | P0 | 13 | 5 | new | MRDN-310, MRDN-711 |
MRDN-820 |
Live-relay deployment and re-probe (C-8) |
Story | P0 | 5 | 5 | new | MRDN-810 |
MRDN-830 |
Handoff package freeze | Story | P0 | 6 | 5 | new | MRDN-820 |
MRDN-810 is the story the entire program is judged on.
Given the load generator and a system under test on the hardware from assumption 2, When a sustained run executes for ≥30 minutes at the highest rate sustainable with zero unintended drops, Then the report records: elapsed delivered messages/s (never configured or attempted counts), p50/p95/p99 latency, CPU per core, allocation and syscall rates, drops by class, the exact hardware and topology, payload size, fan-out ratio, auth profile, and durability contract — and labels the result
[MEASURED]with all of the above attached.
Three anti-patterns that fail this story at review:
- Configured event counts labeled as throughput. This is called out
explicitly in
meridian-ctg's acceptance criteria and it is the most common benchmark lie. - A number without its fan-out ratio. Per § 3.3, egress is the binding constraint; an ingress figure at an unstated fan-out is unfalsifiable.
- Piping a gate into
tail/head/grep— that reports the filter's exit status, so a failed run looks clean. Capture${PIPESTATUS[0]}or redirect to a file.
MRDN-820 exists because of C-8 and two named failure modes: a deployed
image predating the enforcement code makes a feature flag a silent no-op, and a
trusted proxy header is forgeable unless something strips it at the edge.
Green CI closes nothing.
9. Testing expectations
Three surfaces, three different contracts. Each has a gate that must be green before the corresponding epic closes.
9.1 Relay
| Layer | Requirement | Gate |
|---|---|---|
| Unit | meridian-frame codec: proptest round-trip over all header field ranges |
cargo test -p meridian-frame |
| Fuzz | Header decoder must not panic, allocate unboundedly, or read OOB on arbitrary input. ≥1 h clean run per release | cargo fuzz run frame_header |
| Compile-time | const _: () = assert!(is_ephemeral(k)) per MIP-SF-eligible kind; exhaustive kind_to_qos() |
build failure |
| Vocabulary | Every kind, tag, and MIP name has a reader | just check-kinds |
| Integration | Frame ingress → route → fan-out against live Postgres + Dragonfly | just test |
| Regression | The two meridian-22y guards: one frame per subscription; backpressured recipient does not starve the rest |
cargo test -p meridian-relay pubsub_fanout |
| Negative | Persistent kind on a session frame is rejected; leaf-as-peer is refused; unauthorized publish yields an explicit named denial (C-6) |
just test |
| Benchmark | No merge that regresses fanout_cost.rs or event_cost.rs beyond noise |
cargo bench in CI, compared |
| Load | Sustained ≥30 min, zero unintended drops, drops counted by class | MRDN-810 |
| Live | Deployed digest re-probed against a running relay | MRDN-820 |
Two traps specific to this repo, both documented and both expensive:
- A cold
fanout_costrun is ~13 minutes because cargo builds the whole dev-dependency tree. Budget for it in CI; do not let someone "fix" the timeout by dropping the bench. - Clippy passing does not mean fmt passes, and a green
cargo testsays nothing about the lint gate — a test build compiles through an unused import that-D warningsrejects. Runjust ci.
9.2 Client
| Layer | Requirement | Gate |
|---|---|---|
| Conformance | Encode/decode every published golden vector byte-exactly | MRDN-803 suite |
| Negotiation | Correct fallback when the relay advertises no MIP-XP profile the client supports — must fail closed, never silently downgrade | conformance |
| Feature detection | A client that feature-detects an advertised capability must not then fail silently. This is the NIP-17 bug — advertised for months while kind:10050 was rejected |
conformance |
| Reconnect | Sequence gap detection across a reconnect; gap is reported, never silently skipped | integration |
| Backpressure | Client honors the movement contract; a conflated stream replaces rather than queues |
integration |
| Frontend | Biome lint + format; pnpm check:px-text (rem-only text sizing) |
just ci |
9.3 Payload
Payload correctness is the feeder's responsibility, not the relay's — the relay has forsworn reading. The tests therefore live at the boundary.
| Requirement | Rationale | Gate |
|---|---|---|
Golden vectors per payload_type, versioned and published |
Third parties implement against bytes, not prose | MRDN-802 |
provenance is mandatory and non-defaulting in MIP-KN |
A defaulted value silently promotes observed data to self-reported | decoder rejects on absence |
| A relay integration test asserts the payload bytes are never dereferenced | The opaque contract is a negative requirement and needs a negative test | MRDN-304 |
| Round-trip through the MQTT bridge preserves payload bytes exactly | mqtt-passthrough must be byte-transparent |
MRDN-603 |
Unknown payload_type routes correctly and is delivered uninterpreted |
Forward compatibility is the point of the envelope | conformance |
| Batch decode is order-preserving and gap-detecting | Sequence lives in the header for exactly this reason | proptest |
10. Definition of Ready / Done
Definition of Ready
- Acceptance criteria written as Given/When/Then, or as a named artifact
- A Bead exists and is linked
- Dependencies are linked and not themselves blocked
- Any performance claim cites a Sprint 1 baseline number
- Any new vocabulary names its enforcement point (
C-2) - Estimated by the team, not by the author alone
Definition of Done
just cigreen — fmt and clippy and tests and buildsjust check-kindsgreenjust testgreen ifmeridian-relay,meridian-db, ormeridian-authchanged- Benchmarks re-run; no unexplained regression
- Committed with
git commit -s(DCO) - STELLAR pass: nearest owning
AGENTS.mdupdated, indexes refreshed - Any published figure carries its profile and its measured/estimate/target label (
C-1) - Bead closed with the evidence, not with an assertion
- For P0 stories: deployed to a live relay and re-probed (
C-8)
11. Risk register
| # | Risk | Likelihood | Impact | Mitigation | Owner |
|---|---|---|---|---|---|
| R-1 | 343 points against a 260-point budget | Certain | High | Cut line agreed at Sprint 1 planning: E7 → CLI-only (−13), E6 → 3.1.1 only (−5), MIP-TL defers (−3), MIP-KN defers to spec-only (−3). Never cut E1 or E8 | Tech lead |
| R-2 | Sprint 1 measurement invalidates the E3 plan | Medium | High | MRDN-110 is an explicit go/no-go with named re-plan outcomes, not a checkpoint |
Tech lead |
| R-3 | The 1.337M target is quoted without its profile and becomes a support obligation | High | Severe | C-1 enforced at review; § 3.5 publishes figures in pairs; MRDN-101 documents the target's origin |
Everyone |
| R-4 | Egress bandwidth, not CPU, is the real wall — and the NIC is discovered late | Medium | High | § 3.3 makes the NIC a Sprint 1 dependency (assumption 2), not a Sprint 5 surprise | Perf eng |
| R-5 | Batching win destroyed by re-atomising the batch above fan-out (the Zenoh trap) | Medium | High | Named explicitly in every MRDN-3xx review; MRDN-310 acceptance measures end-to-end, not per-stage |
Data plane |
| R-6 | MQTT QoS 1 ack latency coupled to spine persistence (p95 74.6 ms) exceeds device timeouts | Medium | Medium | Surfaced at Sprint 4 planning per MRDN-620; fallback is QoS 0 + volatile |
Bridge eng |
| R-7 | Someone proposes making the leaf buffer durable and authoritative | Medium | Severe | Refusal R2 cited by name; MRDN-520 negative test; a second writer activates the dormant D4 defect |
Architect |
| R-8 | MQTT topic strings leak into routing keys or ACLs | Medium | High | C-3; MRDN-610 mandates registry lookup over string transformation |
Bridge eng |
| R-9 | Shared subscriptions / QoS 2 partially emulated and shipped as complete | Low | High | Explicitly out of scope in E6; documented as known gaps in the handoff | Bridge eng |
| R-10 | MIP advertised in NIP-11 before its reader exists (the NIP-17 bug, on the hot path) | Medium | High | ENFORCED is the only state permitting advertisement; existing nip11.rs tests already fail on this |
Protocol |
| R-11 | Head-to-head benchmark published against an untuned TBMQ | Low | Severe | Ratified council call forbids it; MRDN-101 states the rule. A benchmark of an unfamiliar system tuned by its competitor is worth nothing |
Tech lead |
12. Handoff package — MRDN-830
The frozen artifact set the implementing team receives:
| # | Artifact | Location |
|---|---|---|
| 1 | MIP specifications at IMPLEMENTABLE or ENFORCED |
docs/mips/MIP-{RG,OF,XP,SF,QC,LF,MQ,KN,TL}.md |
| 2 | Golden test vectors, versioned | crates/meridian-conformance/vectors/ |
| 3 | Reference implementation | crates/meridian-{frame,leaf,mqtt} |
| 4 | Conformance suite + runner | crates/meridian-conformance |
| 5 | Load harness and its reproduction procedure | crates/meridian-loadgen + runbook |
| 6 | Capacity report — every figure with its profile and label | docs/benchmarks/frame-path-capacity-<date>.md |
| 7 | Deployment charts and Compose profiles | deploy/charts/meridian-{leaf,mqtt} |
| 8 | Operator runbook — congestion, backlog, class shed, bridge sessions | docs/runbooks/ |
| 9 | Known-gaps document — shared subscriptions, QoS 2, offline queueing, federation | docs/mips/MIP-MQ.md § Non-Goals |
| 10 | Decision record — every council call and its reversal evidence | this file § 2 + Bead history |
Artifact 9 is not optional and not a weakness. A handoff that lists only capabilities invites the receiving team to discover the gaps in production. The five things a broker does better than this relay are named in DIAGRAM.md § Part III, and the last three are not close.
13. Open questions
Each blocks a specific story, not the program:
Is the target rate pre- or post-dedup?CLOSED — pre-dedup. The deployment AIS feed runs ~10× receiver redundancy, so 400k raw frames/s resolve to ~40k unique vessel states/s. Deduplicating at the feeder is therefore nearly free and removes ~90% of everything downstream, and it is now a normative feeder requirement with two mandatory counters (MIP-AS § 4.1–4.2).MRDN-110sizing unblocked. One consequence must not be buried: at 40k states/s ingest is cheap on every path — even signed BIP-340 costs only ~1.5 cores — so the case for the opaque frame path here rests on zero-copy fan-out, key-expression interest scoping, and the stored-and-indexed ceiling, not on ingest CPU. Anyone arguing the reverse will be corrected by the first reviewer who does the arithmetic. (Was inherited frommeridian-8ft.)- What is the real fan-out ratio of the production workload? § 3.3 shows
the answer moves the CPU requirement by 100× and the NIC requirement past
100 GbE. Blocks
MRDN-810's test profile. - What device population cannot be reflashed? Sets whether E6 is a migration aid or a permanent tier. If permanent, per the ratified reversal evidence, "the broker is not being replaced; it is being kept, and this document is reframed as an integration plan."
- Does the operator require a Next.js platform? Blocks
MRDN-720's shape only. See § 4.2. - Which QoS level do devices actually publish at today? QoS 0 maps cleanly;
QoS 1 couples to persistence latency (R-6); QoS 2 has no equivalent. Blocks
MRDN-620.