Warpnet Moderation System
This document describes the current implementation of content moderation in Warpnet.
Overview
Warpnet implements a decentralized content moderation system using dedicated moderator nodes that run an AI safety classifier locally. Moderation is report-driven: nothing is scanned proactively. A report opens a voting round shared by all moderators, and a strict majority is required to act. There is no central control and no human reviewer anywhere in the loop.
Architecture
1. Moderator Nodes
Moderator nodes are specialized nodes that run a moderation engine. They:
Subscribe to the global reports topic and open a vote round per report
Fetch the reported object directly from the target node
Classify it with a local AI model and cast a signed vote
Tally the votes, and — for exactly one node per round — deliver the verdict
Anyone can run a moderator node. There is no application, no approval, and no way to choose which reports you judge.
2. Moderation Engine
The engine is built on llama.cpp bindings and provides:
LLM-based content classification
Binary decisions (OK or FAIL)
Reason generation for rejected content, derived from the model's hazard categories
Llama Guard 3 (1B, quantized GGUF) — model file Llama-Guard-3-1B-Q4_K_M.gguf
Engine configuration:
Context size: 4096 tokens
Output tokens: 64
Temperature: 0.0 (deterministic)
Top P: 0.9
Fixed seed: 42
Memory mapping enabled
GPU layers: 0 — inference runs on CPU, using all available cores
3. Reports Topic
Reports travel on a global gossip topic, /warpnet/reports/1.0.0. A report is broadcast to every moderator at once, so nobody can position themselves to intercept a particular report. A report carries the object type, the target user and node, the object id, a free-text reason (max 256 characters) and the reporter's identity, which is how the answer finds its way back.
Only posts and profiles can be reported. Reply-level and image-level reports are rejected by validation. Direct messages are never moderated — there is no mechanism for it.
4. Voting Rounds
Every report opens a round, identified by a SHA-256 digest of the report's own contents — so identical re-reports collapse into a single round and never trigger a second review.
Each moderator computes its own position in an unpredictable order derived from sha256(reportID | moderatorID). Nobody assigns judges and nobody can volunteer for a specific report.
The first three by that order (quorum target = 3) start immediately; the rest wait in 8-second increments and stay silent once the quorum is already served, so no inference is spent needlessly.
Votes are collected for a 30-second window, then tallied.
If an even number of votes arrived, the highest-ranked one is set aside by a rule every node computes identically, so a tie can never occur.
FAIL requires a strict majority. A tie, an even split, or no majority always falls in favour of the content.
Votes travel on their own topic, /warpnet/moderation/votes/1.0.0. The voter's identity is taken from the signature-verified gossip envelope; whatever the payload claims is discarded.
5. Announcing the Result
Exactly one moderator — the chair, chosen by the same ordering — delivers the outcome, so the reporter gets one answer rather than three. Every other voter stands by at a 10-second interval and takes over if no final announcement arrives, so a report is not lost when one machine dies.
6. Isolation Protocol (shadow ban)
Moderation is a shadow ban, not a deletion:
Only FAIL verdicts are broadcast, and only to the target's followers/observers topic.
The offending node never receives the verdict. The author's local view is unchanged; they are never notified.
Followers' apps mark the tweet with a moderation sidecar and drop it from the timeline; for a profile, the moderation flag makes clients hide bio, display name, website and url. Nothing is deleted from anyone's disk.
The reporter is notified separately, and — unlike the broadcast — for both outcomes, because silence on an OK verdict would read as a lost report.
7. Signatures
Every verdict is signed with the chair's Ed25519 node key over a canonical, length-prefixed byte stream. Receivers recover the verifying public key from the moderator's peer id itself, so a verdict is only valid for the peer that claims it. A verdict with a missing, malformed or borrowed signature is dropped before it can touch anything. Each verdict also carries the list of moderators whose votes entered the tally.
8. Mutual Audit
Every five minutes a moderator picks a random peer and challenges it to classify a specific text — one the network has already ruled on, so the expected answer is known. Challenges are drawn from live decided rounds rather than a fixed list, so they cannot be tabulated in advance. Answers are signed and accumulated per peer into a standing: probation, trusted, suspect or banned.
Being honest about the limit: standings are recorded and logged, but nothing is refused because of them. A check performed by a single machine is still a single machine's opinion, and letting one node disqualify another would create exactly the abuse it is meant to prevent. Enforcement requires several independent checkers to agree first, and that is not built yet.
Tweet Moderation
For each report of type "tweet":
Fetch the tweet from the target node (3 attempts, 3 seconds apart)
If it already carries a verdict, reuse it instead of re-running the model
Otherwise classify the text with the local engine
Broadcast a signed ballot to the votes topic
After the window closes, tally; on FAIL the chair runs the isolation broadcast
If the content cannot be fetched, the moderator abstains with an "unavailable" result rather than voting against it. Nothing is decided and nobody is penalised.
User Profile Moderation
Profile moderation is fully implemented and follows the same round, vote and isolation path as posts. The text submitted to the model is built from the profile's public fields: username, bio, website, and any additional public metadata keys in sorted order.
Moderation Policy
The policy is the Llama Guard 3 hazard taxonomy, sent to the model in the prompt. The following categories are treated as violations:
S1 Violent Crimes
S2 Non-Violent Crimes
S3 Sex Crimes
S4 Child Exploitation
S7 Privacy
S9 Indiscriminate Weapons
S10 Hate
S11 Self-Harm
S14 Code Interpreter Abuse
The remaining Llama Guard categories are deliberately not enforced and are treated as non-violations: S5 Defamation, S6 Specialized Advice, S8 Intellectual Property, S12 Sexual Content, S13 Elections. If the model flags content using only those codes, the content is treated as safe.
Note: the taxonomy and its allow-list live in the Warp-net/moderation repository (prompt.go), which Warpnet consumes as a git submodule at core/moderation.
Prompt Engineering
The system uses the Llama Guard chat format, not a free-form instruction:
<|begin_of_text|><|start_header_id|>user<|end_header_id|> Task: Check if there is unsafe content in 'User' messages in conversations according our safety policy with the below categories. <BEGIN UNSAFE CONTENT CATEGORIES> S1: Violent Crimes. ... S13: Elections. <END UNSAFE CONTENT CATEGORIES> <BEGIN CONVERSATION> User: [content] <END CONVERSATION> Provide your safety assessment for ONLY THE LAST User message in the above conversation: - First line must read 'safe' or 'unsafe'. - If unsafe, a second line must include a comma-separated list of violated categories.<|eot_id|><|start_header_id|>assistant<|end_header_id|>
The model answers in one of two forms:
safe — content is acceptable
unsafe followed by a second line of comma-separated hazard codes, e.g. S9,S2
The codes are mapped to human-readable reasons through the allow-list above. Any other output is treated as an engine error, and the content is not condemned on it.
Limitations and Future Work
Current limitations:
Image and video content is not analysed; only text is classified
Single model (Llama Guard 3 1B); the model file and thread count are not configurable — the path is fixed and inference uses all CPU cores
No appeals and no review process; moderation decisions are final
Audit standings are recorded but never enforced
Coverage is only what is reported — nothing scans the network proactively
If no moderator node is online, a report is accepted and then simply goes nowhere, with no notification
A determined attacker running many moderator nodes can still try to swing votes; identity costs nothing today
Potential improvements:
Quorum-based enforcement of audit standings
Model identity attestation, so honest model diversity is not mistaken for dishonesty
Multi-model support, configurable policies, reputation weighting
Image and video content analysis
Appeals and review mechanism
Security Considerations
The moderation system:
Runs on dedicated nodes, separate from user content
Uses deterministic model settings (temperature 0, fixed seed) for reproducibility
Signs every verdict and every gossip envelope with Ed25519, verified against the sender's peer id
Requires a strict majority of an odd number of votes — no single moderator can decide alone
Broadcasts only FAIL verdicts, and only to the target's followers; an OK verdict reaches the reporter alone
Cannot delete content: each node applies verdicts voluntarily by following the protocol
Runs fully offline on the node — there are no external API calls
Performance
One object fetched and classified per report, not per peer scan
Typical end-to-end latency: well under a minute (~30 seconds in testing on a three-moderator network)
Around three moderators do the work for any given report; the rest spend nothing
A chair failure adds roughly 10 seconds
Inference time varies by hardware and is logged for monitoring
Network Protocol
Moderation uses standard Warpnet protocols over libp2p with protocol multiplexing:
Purpose Identifier Submit a report (local UI → node) /public/post/report/0.0.0 Reports gossip topic /warpnet/reports/1.0.0 Moderator votes gossip topic /warpnet/moderation/votes/1.0.0 Verdict delivery /public/post/moderate/result/0.0.0 Audit spot-check /public/get/moderate/challenge/0.0.0
Media Metadata Embedding Implementation
WarpNet embeds encrypted metadata into uploaded media to establish accountability and traceability for uploaded content, complementing text moderation.
Important Privacy Note: All uploaded images and videos carry embedded encrypted metadata including node information, the uploader's user record, and the hardware addresses of the uploading machine's network interfaces. The encryption is intentionally weak, so this metadata can be recovered by an entity willing to spend substantial computation. Hardware addresses in particular are persistent identifiers that can link accounts and devices.
What Metadata is Embedded
A single JSON object with three keys:
node — the uploading node's full NodeInfo (node id, network details, owner)
user — the uploader's full user record
MAC — the hardware addresses of all network interfaces on the uploading machine, comma-separated (not a single address)
Where It Is Embedded
Images: the encrypted blob is base64-encoded and written into the EXIF ImageDescription tag of IFD0. Uploads are re-encoded to JPEG at quality 100; PNG, JPEG and GIF are accepted as input, up to 50 MiB per file.
Videos: the same blob is embedded into the container metadata of ISO base media (MP4-family) files.
Imports: media brought in through the import pipeline goes through the same path as a fresh upload.
Files uploaded together share one blob — they carry identical metadata, so a per-file password would not raise the cost of recovering it.
Encryption Mechanism
The scheme is a deliberate "security through computational difficulty" design:
Algorithm: AES-256-GCM
Key derivation: Argon2id — always, for every blob. Parameters: time = 1, memory = 64 MiB, parallelism = 4, 32-byte key. Memory dominates on purpose: 64 MiB per guess is what denies an attacker the GPU/ASIC speed-up.
Password: a single-use value drawn uniformly at random from a deliberately small space of 110,000,000,000 possibilities (~2^36.7), encoded as 8 big-endian bytes. It is used once, zeroed in memory immediately afterwards, and never stored or logged. It is not derived from a timestamp or from any other predictable source.
Salt: 16 random bytes from the system CSPRNG, generated fresh per blob. It is public and travels with the file.
Nonce: 12 random bytes from the system CSPRNG, generated fresh per blob. It is public and travels with the file. It is not zero-filled.
Container layout: salt || nonce || ciphertext || tag
Cost of Recovery
At the parameters above, one guess costs roughly 19–22 ms. The expected exhaustive search is therefore on the order of 12 days on a 1000-core cluster and about 5 years on an eight-core laptop. Both figures scale with per-core speed, so treat them as an order of magnitude, not a promise.
Security Model Philosophy
Not designed for user decryption. Ordinary users cannot recover the metadata.
Designed for resourced decryption. Only an entity willing to spend cluster-scale computation can brute-force it.
Proof of ownership. The blob acts as evidence of origin without exposing anything during normal use.
Computational difficulty only. Security rests entirely on the cost of the search, not on any secret — the salt, the nonce and the algorithm are all public, and the password no longer exists anywhere.
How Metadata Embedding Prevents Harmful Content
Attribution and accountability — every upload carries encrypted evidence of its origin, creating a chain of responsibility.
Deterrence — users aware of the embedding are less likely to upload harmful content.
Investigation support — when harmful content is reported, the metadata provides leads for authorized entities.
Distributed accountability — in a decentralized network, metadata helps identify responsible parties without a central registry.
Forensic evidence — the blob can serve as evidence of upload time, source node and user identity.
Origin verification — helps distinguish original uploads from redistributed copies.
Limitations and Considerations
Privacy implications
All uploaded images and videos carry embedded encrypted user, node and hardware-address information
The metadata is recoverable with sufficient resources
Users should be aware that their uploads carry traceable information
Security limitations
Weak by design. The password space is deliberately small.
Batch correlation. Files uploaded together share one blob, so recovering one recovers the whole batch.
Metadata removal. A technically sophisticated user can strip EXIF or container metadata after the fact.
Not foolproof. A determined actor can circumvent the scheme.
Future enhancements
Image and video content analysis, extending moderation beyond text
Visible or invisible watermarking alongside the metadata blob
Richer forensic fields
WarpNet's media metadata embedding system balances privacy with accountability. Combined with LLM-based content moderation, it helps limit the spread of prohibited content on a decentralized social network — while remaining honest that the protection is a cost barrier, not a secret.
Contacts
© 2026. All rights reserved. Legal information.
Donation
BTC: bc1quwwnec87tukn9j93spr4de7mctvexpftpwu09d
USDT (Tron): THXiCmfr6D4mqAfd4La9EQ5THCx7WsR143
