Warpnet Moderation System

This document describes the current implementation of content moderation in Warpnet.

Overview

Warpnet implements a decentralized content moderation system using dedicated moderator nodes that run an AI safety classifier locally. Moderation is report-driven: nothing is scanned proactively. A report opens a voting round shared by all moderators, and a strict majority is required to act. There is no central control and no human reviewer anywhere in the loop.

Architecture

1. Moderator Nodes

Moderator nodes are specialized nodes that run a moderation engine. They:

  • Subscribe to the global reports topic and open a vote round per report

  • Fetch the reported object directly from the target node

  • Classify it with a local AI model and cast a signed vote

  • Tally the votes, and — for exactly one node per round — deliver the verdict

Anyone can run a moderator node. There is no application, no approval, and no way to choose which reports you judge.

2. Moderation Engine

The engine is built on llama.cpp bindings and provides:

  • LLM-based content classification

  • Binary decisions (OK or FAIL)

  • Reason generation for rejected content, derived from the model's hazard categories

  • Llama Guard 3 (1B, quantized GGUF) — model file Llama-Guard-3-1B-Q4_K_M.gguf

Engine configuration:

  • Context size: 4096 tokens

  • Output tokens: 64

  • Temperature: 0.0 (deterministic)

  • Top P: 0.9

  • Fixed seed: 42

  • Memory mapping enabled

  • GPU layers: 0 — inference runs on CPU, using all available cores

3. Reports Topic

Reports travel on a global gossip topic, /warpnet/reports/1.0.0. A report is broadcast to every moderator at once, so nobody can position themselves to intercept a particular report. A report carries the object type, the target user and node, the object id, a free-text reason (max 256 characters) and the reporter's identity, which is how the answer finds its way back.

Only posts and profiles can be reported. Reply-level and image-level reports are rejected by validation. Direct messages are never moderated — there is no mechanism for it.

4. Voting Rounds

Every report opens a round, identified by a SHA-256 digest of the report's own contents — so identical re-reports collapse into a single round and never trigger a second review.

  • Each moderator computes its own position in an unpredictable order derived from sha256(reportID | moderatorID). Nobody assigns judges and nobody can volunteer for a specific report.

  • The first three by that order (quorum target = 3) start immediately; the rest wait in 8-second increments and stay silent once the quorum is already served, so no inference is spent needlessly.

  • Votes are collected for a 30-second window, then tallied.

  • If an even number of votes arrived, the highest-ranked one is set aside by a rule every node computes identically, so a tie can never occur.

  • FAIL requires a strict majority. A tie, an even split, or no majority always falls in favour of the content.

  • Votes travel on their own topic, /warpnet/moderation/votes/1.0.0. The voter's identity is taken from the signature-verified gossip envelope; whatever the payload claims is discarded.

5. Announcing the Result

Exactly one moderator — the chair, chosen by the same ordering — delivers the outcome, so the reporter gets one answer rather than three. Every other voter stands by at a 10-second interval and takes over if no final announcement arrives, so a report is not lost when one machine dies.

6. Isolation Protocol (shadow ban)

Moderation is a shadow ban, not a deletion:

  • Only FAIL verdicts are broadcast, and only to the target's followers/observers topic.

  • The offending node never receives the verdict. The author's local view is unchanged; they are never notified.

  • Followers' apps mark the tweet with a moderation sidecar and drop it from the timeline; for a profile, the moderation flag makes clients hide bio, display name, website and url. Nothing is deleted from anyone's disk.

  • The reporter is notified separately, and — unlike the broadcast — for both outcomes, because silence on an OK verdict would read as a lost report.

7. Signatures

Every verdict is signed with the chair's Ed25519 node key over a canonical, length-prefixed byte stream. Receivers recover the verifying public key from the moderator's peer id itself, so a verdict is only valid for the peer that claims it. A verdict with a missing, malformed or borrowed signature is dropped before it can touch anything. Each verdict also carries the list of moderators whose votes entered the tally.

8. Mutual Audit

Every five minutes a moderator picks a random peer and challenges it to classify a specific text — one the network has already ruled on, so the expected answer is known. Challenges are drawn from live decided rounds rather than a fixed list, so they cannot be tabulated in advance. Answers are signed and accumulated per peer into a standing: probation, trusted, suspect or banned.

Being honest about the limit: standings are recorded and logged, but nothing is refused because of them. A check performed by a single machine is still a single machine's opinion, and letting one node disqualify another would create exactly the abuse it is meant to prevent. Enforcement requires several independent checkers to agree first, and that is not built yet.

Tweet Moderation

For each report of type "tweet":

  1. Fetch the tweet from the target node (3 attempts, 3 seconds apart)

  2. If it already carries a verdict, reuse it instead of re-running the model

  3. Otherwise classify the text with the local engine

  4. Broadcast a signed ballot to the votes topic

  5. After the window closes, tally; on FAIL the chair runs the isolation broadcast

If the content cannot be fetched, the moderator abstains with an "unavailable" result rather than voting against it. Nothing is decided and nobody is penalised.

User Profile Moderation

Profile moderation is fully implemented and follows the same round, vote and isolation path as posts. The text submitted to the model is built from the profile's public fields: username, bio, website, and any additional public metadata keys in sorted order.

Moderation Policy

The policy is the Llama Guard 3 hazard taxonomy, sent to the model in the prompt. The following categories are treated as violations:

  • S1 Violent Crimes

  • S2 Non-Violent Crimes

  • S3 Sex Crimes

  • S4 Child Exploitation

  • S7 Privacy

  • S9 Indiscriminate Weapons

  • S10 Hate

  • S11 Self-Harm

  • S14 Code Interpreter Abuse

The remaining Llama Guard categories are deliberately not enforced and are treated as non-violations: S5 Defamation, S6 Specialized Advice, S8 Intellectual Property, S12 Sexual Content, S13 Elections. If the model flags content using only those codes, the content is treated as safe.

Note: the taxonomy and its allow-list live in the Warp-net/moderation repository (prompt.go), which Warpnet consumes as a git submodule at core/moderation.

Prompt Engineering

The system uses the Llama Guard chat format, not a free-form instruction:

<|begin_of_text|><|start_header_id|>user<|end_header_id|> Task: Check if there is unsafe content in 'User' messages in conversations according our safety policy with the below categories. <BEGIN UNSAFE CONTENT CATEGORIES> S1: Violent Crimes. ... S13: Elections. <END UNSAFE CONTENT CATEGORIES> <BEGIN CONVERSATION> User: [content] <END CONVERSATION> Provide your safety assessment for ONLY THE LAST User message in the above conversation: - First line must read 'safe' or 'unsafe'. - If unsafe, a second line must include a comma-separated list of violated categories.<|eot_id|><|start_header_id|>assistant<|end_header_id|>

The model answers in one of two forms:

  • safe — content is acceptable

  • unsafe followed by a second line of comma-separated hazard codes, e.g. S9,S2

The codes are mapped to human-readable reasons through the allow-list above. Any other output is treated as an engine error, and the content is not condemned on it.

Limitations and Future Work

Current limitations:

  • Image and video content is not analysed; only text is classified

  • Single model (Llama Guard 3 1B); the model file and thread count are not configurable — the path is fixed and inference uses all CPU cores

  • No appeals and no review process; moderation decisions are final

  • Audit standings are recorded but never enforced

  • Coverage is only what is reported — nothing scans the network proactively

  • If no moderator node is online, a report is accepted and then simply goes nowhere, with no notification

  • A determined attacker running many moderator nodes can still try to swing votes; identity costs nothing today

Potential improvements:

  • Quorum-based enforcement of audit standings

  • Model identity attestation, so honest model diversity is not mistaken for dishonesty

  • Multi-model support, configurable policies, reputation weighting

  • Image and video content analysis

  • Appeals and review mechanism

Security Considerations

The moderation system:

  • Runs on dedicated nodes, separate from user content

  • Uses deterministic model settings (temperature 0, fixed seed) for reproducibility

  • Signs every verdict and every gossip envelope with Ed25519, verified against the sender's peer id

  • Requires a strict majority of an odd number of votes — no single moderator can decide alone

  • Broadcasts only FAIL verdicts, and only to the target's followers; an OK verdict reaches the reporter alone

  • Cannot delete content: each node applies verdicts voluntarily by following the protocol

  • Runs fully offline on the node — there are no external API calls

Performance

  • One object fetched and classified per report, not per peer scan

  • Typical end-to-end latency: well under a minute (~30 seconds in testing on a three-moderator network)

  • Around three moderators do the work for any given report; the rest spend nothing

  • A chair failure adds roughly 10 seconds

  • Inference time varies by hardware and is logged for monitoring

Network Protocol

Moderation uses standard Warpnet protocols over libp2p with protocol multiplexing:

Purpose Identifier Submit a report (local UI → node) /public/post/report/0.0.0 Reports gossip topic /warpnet/reports/1.0.0 Moderator votes gossip topic /warpnet/moderation/votes/1.0.0 Verdict delivery /public/post/moderate/result/0.0.0 Audit spot-check /public/get/moderate/challenge/0.0.0

Media Metadata Embedding Implementation

WarpNet embeds encrypted metadata into uploaded media to establish accountability and traceability for uploaded content, complementing text moderation.

Important Privacy Note: All uploaded images and videos carry embedded encrypted metadata including node information, the uploader's user record, and the hardware addresses of the uploading machine's network interfaces. The encryption is intentionally weak, so this metadata can be recovered by an entity willing to spend substantial computation. Hardware addresses in particular are persistent identifiers that can link accounts and devices.

What Metadata is Embedded

A single JSON object with three keys:

  • node — the uploading node's full NodeInfo (node id, network details, owner)

  • user — the uploader's full user record

  • MAC — the hardware addresses of all network interfaces on the uploading machine, comma-separated (not a single address)

Where It Is Embedded

  • Images: the encrypted blob is base64-encoded and written into the EXIF ImageDescription tag of IFD0. Uploads are re-encoded to JPEG at quality 100; PNG, JPEG and GIF are accepted as input, up to 50 MiB per file.

  • Videos: the same blob is embedded into the container metadata of ISO base media (MP4-family) files.

  • Imports: media brought in through the import pipeline goes through the same path as a fresh upload.

Files uploaded together share one blob — they carry identical metadata, so a per-file password would not raise the cost of recovering it.

Encryption Mechanism

The scheme is a deliberate "security through computational difficulty" design:

  • Algorithm: AES-256-GCM

  • Key derivation: Argon2id — always, for every blob. Parameters: time = 1, memory = 64 MiB, parallelism = 4, 32-byte key. Memory dominates on purpose: 64 MiB per guess is what denies an attacker the GPU/ASIC speed-up.

  • Password: a single-use value drawn uniformly at random from a deliberately small space of 110,000,000,000 possibilities (~2^36.7), encoded as 8 big-endian bytes. It is used once, zeroed in memory immediately afterwards, and never stored or logged. It is not derived from a timestamp or from any other predictable source.

  • Salt: 16 random bytes from the system CSPRNG, generated fresh per blob. It is public and travels with the file.

  • Nonce: 12 random bytes from the system CSPRNG, generated fresh per blob. It is public and travels with the file. It is not zero-filled.

  • Container layout: salt || nonce || ciphertext || tag

Cost of Recovery

At the parameters above, one guess costs roughly 19–22 ms. The expected exhaustive search is therefore on the order of 12 days on a 1000-core cluster and about 5 years on an eight-core laptop. Both figures scale with per-core speed, so treat them as an order of magnitude, not a promise.

Security Model Philosophy

  • Not designed for user decryption. Ordinary users cannot recover the metadata.

  • Designed for resourced decryption. Only an entity willing to spend cluster-scale computation can brute-force it.

  • Proof of ownership. The blob acts as evidence of origin without exposing anything during normal use.

  • Computational difficulty only. Security rests entirely on the cost of the search, not on any secret — the salt, the nonce and the algorithm are all public, and the password no longer exists anywhere.

How Metadata Embedding Prevents Harmful Content

  1. Attribution and accountability — every upload carries encrypted evidence of its origin, creating a chain of responsibility.

  2. Deterrence — users aware of the embedding are less likely to upload harmful content.

  3. Investigation support — when harmful content is reported, the metadata provides leads for authorized entities.

  4. Distributed accountability — in a decentralized network, metadata helps identify responsible parties without a central registry.

  5. Forensic evidence — the blob can serve as evidence of upload time, source node and user identity.

  6. Origin verification — helps distinguish original uploads from redistributed copies.

Limitations and Considerations

Privacy implications

  • All uploaded images and videos carry embedded encrypted user, node and hardware-address information

  • The metadata is recoverable with sufficient resources

  • Users should be aware that their uploads carry traceable information

Security limitations

  • Weak by design. The password space is deliberately small.

  • Batch correlation. Files uploaded together share one blob, so recovering one recovers the whole batch.

  • Metadata removal. A technically sophisticated user can strip EXIF or container metadata after the fact.

  • Not foolproof. A determined actor can circumvent the scheme.

Future enhancements

  • Image and video content analysis, extending moderation beyond text

  • Visible or invisible watermarking alongside the metadata blob

  • Richer forensic fields

WarpNet's media metadata embedding system balances privacy with accountability. Combined with LLM-based content moderation, it helps limit the spread of prohibited content on a decentralized social network — while remaining honest that the protection is a cost barrier, not a secret.

© 2026. All rights reserved. Legal information.

Donation

BTC: bc1quwwnec87tukn9j93spr4de7mctvexpftpwu09d

USDT (Tron): THXiCmfr6D4mqAfd4La9EQ5THCx7WsR143