Warpnet CRDT-Based Consensus Implementation

This implementation adds Conflict-free Replicated Data Type (CRDT) based consensus to Warpnet for managing distributed engagement statistics (reactions, retweets, replies, views and poll votes) across autonomous nodes without centralized coordination.

Overview

1. Core CRDT Statistics Store

  • Complete Merkle CRDT implementation using github.com/ipfs/go-ds-crdt

  • Five statistic types: Reactions (emoji; a heart is the default), Retweets, Replies, Views, Poll votes

  • PN-counter semantics: separate incr and decr sub-counters, aggregate clamped at zero

  • Delta broadcast via libp2p GossipSub on the /warpnet/stats/1.0.0 topic; Merkle-DAG blocks fetched over Bitswap with DHT provider lookup

  • Per-process generation nonce, so a node survives a full loss of local data without corrupting peer history

  • Graceful error handling and fallback mechanisms

2. Database Integration Layer

  • Maintains backward compatibility with existing code

  • Updates both local storage and CRDT counters inside the same call path

  • Prefers the network-wide aggregate; falls back to this node's own counter when the CRDT query fails, and takes the larger of the two for reactions and views

  • The isTransitive flag bumps the network counter only on the node where the action originated, so an action stored on both the actor's and the author's node is counted once

  • Only member nodes run the CRDT store; every repository accepts a nil stats store and degrades to purely local counting

3. Testing

  • core/crdt/stats_test.go — counter accumulation, zero-clamping, per-key isolation, prefix-sibling isolation, concurrent bumps, broadcast-on-every-write, generation uniqueness, codec round-trip, safe close

  • core/crdt/gossip-adapter_test.go — publish/subscribe wiring, context cancellation, full-buffer drop-oldest behaviour, no deadlock between Receive and Next

  • database/stats-repo_test.go and database/tweet-repo_test.go — datastore adapter, and the cross-stack rule that counters read the aggregated CRDT stat rather than only the local one

Key structure

Each statistic uses a Merkle CRDT with the following structure:

Key Pattern: /STATS/{incr|decr}/{dataKey}/{nodeID}/{generation} Value: 64-bit unsigned integer, big-endian — the running total of that sub-counter, not a delta

{dataKey} is the application key of the statistic, for example /REACTIONS/INCR/{tweetId}, /TWEETS/RETWEETSCOUNT/{tweetId}, /TWEETS/TWEETSCOUNT/{parentId}, /TWEETS/VIEWS/{tweetId} or /POLLS/VOTES/INCR/{tweetId}/{option}.

{generation} is a fresh 128-bit random nonce minted once per process start. /CRDT, which also appears on disk, is the local datastore namespace the CRDT lives in — not part of the CRDT key itself.

Key Design Decisions

1. Merkle CRDT Choice

  • Why: statistics need commutative, coordination-free merging across nodes that are frequently offline

  • Shape: a PN-counter — incr and decr are separate namespaces, and the aggregate is Σ incr − Σ decr, clamped at zero. A pure G-counter is not enough, because unreacting, un-retweeting and deleting a reply all have to take a count back down.

  • Benefits: simple aggregation, deterministic, commutative

2. Per-Node, Per-Generation Sub-Counters

  • Why: each (node, process lifetime) pair owns exactly one sub-counter and is its only writer. The value is kept in memory and bumped under a mutex, so the CRDT is never read back before a write — no read-modify-write hazard and no eventual-consistency window that could swallow an increment.

  • Benefits: true offline operation, no coordination needed; a node that loses its local database entirely simply starts a new sub-counter alongside the history peers replay back to it

  • Tradeoff: space grows with nodes × process restarts per key, not with nodes alone. Compaction is future work.

3. Hybrid Approach

  • CRDT: global aggregated counts

  • Local DB: user-specific actions (who reacted to what, who voted, who viewed)

  • Benefits: distributed consistency plus local tracking

4. Graceful Degradation

  • Fallback: fall back to local counts if the CRDT query fails, or if no CRDT store is configured at all

  • Benefits: resilience, backward compatibility

  • Tradeoff: temporary inconsistency during failures

Benefits Achieved

1. True Offline-First Operation

  • Nodes can react, retweet, reply and vote while disconnected

  • Changes sync automatically when reconnected

  • No data loss during network partitions

2. Automatic Conflict Resolution

  • No manual intervention needed

  • Deterministic outcomes

  • Commutative operations

3. Scalability

  • Linear scaling with nodes

  • No coordination bottleneck

  • Efficient delta-only sync

4. Fault Tolerance

  • No single point of failure

  • Nodes can join and leave freely

  • A node that loses its database re-converges from its peers' DAG

5. Consistency Without Coordination

  • Eventually consistent

  • Deterministic convergence

  • Strong mathematical guarantees

Transport

  • Delta broadcast: libp2p GossipSub, topic /warpnet/stats/1.0.0, through an adapter with a 100-message buffer and a drop-oldest, never-block policy

  • Block exchange: Bitswap over libp2p, with the DHT as the provider router; already-established connections are replayed into the Bitswap peer manager, otherwise small clusters never converge

  • Settings: rebroadcast interval 1 minute, DAG syncer timeout 1 minute, multi-head processing enabled

  • Blockstore: write-through, with an identity store in front

Performance Characteristics

Space Complexity

  • Per stat: O(n × r), where n is the number of nodes that updated it and r is how many times each of those processes restarted

  • Typical: a small constant factor for a popular tweet, accumulating over a node's lifetime

  • Worst case: unbounded over time until compaction lands

Time Complexity

  • Update: O(1) — single key write + PubSub broadcast

  • Query: O(n) — iterate and sum sub-counters

  • Sync: O(log n) — Merkle-DAG sync

Network Usage

  • Bandwidth: minimal — only deltas transmitted

  • Latency: depends on PubSub propagation (typically < 1s)

  • Overhead: 8 bytes of value per sub-counter, plus a key carrying the data key, a libp2p peer ID and a 32-character generation nonce

Future Enhancements

Near-Term

  1. Compaction: merge old (node, generation) sub-counters to bound storage growth across restarts

  2. Caching: LRU cache for frequently accessed stats

  3. Metrics: Prometheus instrumentation

Long-Term

  1. OR-Set: track who reacted or retweeted, not just the count

  2. LWW-Register: last-write-wins for timestamps

  3. Selective Sync: only sync stats for viewed tweets

  4. Sharding: partition stats across CRDT instances

Security Considerations

  • DoS Protection: counters are per-node, limiting attack surface

  • Data Integrity: CRDT deltas travel as raw payloads on GossipSub and carry no application-level signature of their own; they inherit libp2p-pubsub's default message signing, which binds each message to the publishing peer's key

  • Sybil Resistance: node IDs are tied to libp2p peer IDs

  • Spam Prevention: deduplication happens locally before any CRDT write — one view per viewer, one vote per voter, one reaction per user per tweet — and the network counter is only bumped on the node where the action originated

  • Rate Limiting: should be added in production (future work)

Note on other consensus mechanisms

CRDT counters are the only consensus mechanism in the codebase for engagement data. Moderation uses a separate, quorum-based voting protocol between moderator nodes — see the Moderation page. An earlier node-integrity challenge/validation scheme exists only as unused type definitions and is not implemented.

© 2026. All rights reserved. Legal information.

Donation

BTC: bc1quwwnec87tukn9j93spr4de7mctvexpftpwu09d

USDT (Tron): THXiCmfr6D4mqAfd4La9EQ5THCx7WsR143