Pre-publish checks for MediaWiki and Git
Pre-publish checks for MediaWiki and Git is a project proposal for layered review before a wiki save or a Git push becomes durable. The problem is not only spam on a mutable website. Decentralized and permanent publishing, including Arweave, the permaweb, and content-addressed wiki exports, can make a bad paste hard to erase. A quiet secret, a malware link, or a doxxed address can outlive the editor who noticed too late. This page separates tools that already exist from ideas that are still proposals. It does not publish private implementation details, does not endorse a vendor, and does not claim a perfect filter.
Why checks before publish and push
MediaWiki keeps a revision history. Git keeps a commit history. Both make mistakes recoverable on a single host if operators act quickly. Permanent and content-addressed copies change the stakes. A snapshot written to a long-retention network, or a CID other people already mirrored, may remain fetchable after the live wiki deletes or redacts a revision. Pre-save and pre-push checks exist to catch the worst classes of mistake while the bytes are still local.
The useful goal is harm reduction, not zero risk. Filters miss things. Humans override filters. Attackers adapt. A proposal that promises absolute safety is not credible. A proposal that lists clear check classes, where they run, and what they cannot do is usable.
Threat classes to cover
Malware, scam, and spam links: outbound URLs that deliver malware, phishing, or bulk spam. Illegal threats and abuse: text that targets people with credible threats, or other material a site policy already forbids. Accidental secrets: .env fragments, API keys, private keys, tokens, and similar credentials pasted into a page or committed into a repo. Personally identifiable information (PII): names tied to private contact data, government ids, and other identifiers that should not go public without a clear reason and process.
These classes overlap. A spam link can also be malware. A secret can contain PII. The engineering point is that different detectors catch different shapes: URL lists, regex and entropy rules, token classifiers, and human review queues.
Existing pieces at a high level
MediaWiki already ships mature extensions for edit-time rules. AbuseFilter lets privileged users write conditions on edits and choose actions such as warn, disallow, throttle, or tag. SpamBlacklist blocks saves that add external links matching configured host patterns, often including shared community lists. Newer blocked-external-domain features related to AbuseFilter focus on whole domains with a clearer audit trail than large regex lists alone. These tools are rule systems. They are not full language understanding.
For Git, Gitleaks is a widely used open source scanner that looks for secrets in repositories, directories, and stdin, and that can run as a pre-commit or CI step. It is pattern and entropy oriented. It reduces accidental key leaks. It does not judge prose quality.
URLhaus, operated in the abuse.ch family of community threat feeds, publishes data about URLs observed in malware distribution. Downstream blocklists and filters consume those feeds. A wiki or CI job can compare newly added links against such a feed. Feeds lag, and false positives happen, so operators need an override path and a review habit.
Rampart is a local-first PII redaction system that pairs a small on-device token-classification model with deterministic validators, intended to redact identifiers in the browser before text leaves the device. Public materials describe open weights and a client-side design. It is harm reduction for typed PII, not a guarantee that every private fact is removed.
SetFit is an open few-shot classification framework from the Sentence Transformers ecosystem. Public hubs host spam-oriented SetFit or sentence-transformer checkpoints that teams can fine-tune on their own labels. A wiki can treat SetFit-style classifiers as a spam hint layer behind rules, not as a sole judge.
Jev, from TypeSafe AI, is a hosted decision model that returns structured answers such as yes or no probabilities for policy questions. Public write-ups describe content-moderation demos that ask whether text is safe to publish automatically. Available descriptions indicate a managed API rather than published self-host weights. This page therefore treats Jev as a candidate yes or no gate only where a team accepts a hosted dependency. If open weights appear later, that status should be re-checked. Until then, local open-weight classifiers remain the path for fully self-hosted gates.
None of the names above is a turnkey product endorsement. Versions, licenses, and threat models change. Operators should verify current docs before depending on any piece.
MediaWiki revision history versus Git
On MediaWiki, an edit creates a revision. Admins can delete, suppress, or revise under local policy, and readers of the live site see the current text. Mirrors, scrapes, exports, and permanent snapshots may still hold older bytes. Hashing and export workflows discussed elsewhere on this site make integrity checks possible. They also make silent disappearance harder once a copy left the building.
In Git, a commit that never left a laptop is cheap to amend. A commit that was pushed, cloned, and mirrored is expensive to erase. Force-push does not reach every fork. Pre-commit and pre-push hooks exist because the cost jumps at the network boundary. Secret scanning belongs before that boundary whenever possible.
The shared lesson is timing. Check while the author still has an undo that does not require global coordination.
Options for local and open-weight models
Teams that refuse hosted moderation can run open-weight classifiers for spam, toxicity-style gates, and PII token labels on their own hardware. Small models fit browser or edge runtimes for first-pass warnings. Larger models fit a server-side pre-save service. Local models improve privacy when raw text never leaves the organization. They still produce false positives and false negatives. They need labeled examples from the wiki's own domain, plus a human appeals path.
A practical pattern is layered: rules and URL feeds first, local classifiers second, human review for borderline scores, and hard blocks for high-confidence secrets or known-bad lists.
Pre-save browser warning and server check
For MediaWiki, a browser warning can run as the editor clicks save: scan the diff for secrets and PII, flag suspicious links, and ask the user to confirm. Client-side checks catch accidents early and can redact before upload when that is the policy. They are bypassable. A determined user can disable scripts.
A server check should still run on save: AbuseFilter and SpamBlacklist style rules, URL feed comparison, secret regexes, and optional model scores. The server is the enforcement point. The browser is the courtesy and privacy layer. Showing a clear message that names the matched class helps authors fix the edit without guessing.
Pre-commit and pre-push for Git
For Git, run secret scanning on staged changes at pre-commit, and again at pre-push against the range about to leave the machine. Block the action when high-confidence secrets appear. Warn on suspicious links in documentation commits if that matches team policy. Keep allowlists small and reviewed. Store scan logs with redaction so a detected key does not get re-logged in plain text.
CI can repeat the same scans on pull requests. CI is not a substitute for pre-push on unprotected branches.
Limits, false positives, and privacy
Every automated gate will wrongly block good edits and miss bad ones. Operators should measure both errors on their own corpus. Logs should redact secrets and PII so the safety system does not become a second leak. Access to full match details should be limited. Over-blocking educational discussion of malware or security can harm a technical wiki. Policies should distinguish quoting a bad URL for analysis from promoting it.
Checks also cannot replace law, trust and safety staff, or clear site rules. They reduce accident and casual abuse. They do not settle hard edge cases alone.
Separating verified tools from proposals
Verified today at a high level: MediaWiki AbuseFilter and SpamBlacklist patterns, Gitleaks for Git secrets, URLhaus-style malware URL intelligence, Rampart-style local PII redaction, SetFit-style trainable spam classifiers, and optional hosted yes or no gates such as Jev when weights are not self-hosted. Proposed as an integrated project: a documented pre-save and pre-push checklist that wires those classes together for wikis that also export to permanent or content-addressed stores, with browser warnings, server enforcement, redacted logs, and human review. Building that glue, evaluating false positive rates, and publishing open configs would be the project work. Shipping marketing claims without evaluation would not.
Conclusion
Permanent and content-addressed publishing raises the cost of a bad save or push. Pre-publish checks for MediaWiki and Git should cover malware and spam links, illegal threats and abuse under site policy, accidental secrets, and PII, using rule systems, threat feeds, secret scanners, and optional local or hosted models. Browser warnings and server checks belong on the wiki path. Pre-commit and pre-push belong on the Git path. Limits include false positives, privacy of logs, and the fact that permanence can outrun deletion. Treat named tools as building blocks. Treat the combined pre-publish pipeline as the proposal.