Jump to content

Content-Addressed Wikis: Permanence Without a Single Host

From IdeaWazaWiki
Revision as of 05:09, 3 October 2026 by Waza (talk | contribs) (Created page with "Content-addressed wikis store page bytes so that a cryptographic identifier, not a single website hostname, is the durable name of a version. A classic MediaWiki installation keeps articles in a database on one host. If that host goes offline, changes policy, or loses its disks, readers lose the pages even when many people valued them. Content addressing changes the failure mode. An IPFS content identifier (CID) or an Arweave transaction identifier names exact bytes. Any...")
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)

Content-addressed wikis store page bytes so that a cryptographic identifier, not a single website hostname, is the durable name of a version. A classic MediaWiki installation keeps articles in a database on one host. If that host goes offline, changes policy, or loses its disks, readers lose the pages even when many people valued them. Content addressing changes the failure mode. An IPFS content identifier (CID) or an Arweave transaction identifier names exact bytes. Anyone who still has a copy, or who can fetch one through a gateway or a pin, can check that those bytes match the identifier. This page describes how decentralized wikis use those identifiers, how gateways and pinning fit in, and how mutability and permanence divide the work. It is an architecture overview for publishers and archivists, not a vendor endorsement.

The single-host problem

A centralized wiki is easy to operate because one organization runs the software, the accounts, and the search index. That convenience concentrates risk. The organization pays for the server, applies the edits, and can remove pages, lock accounts, or shut the site down. Backups may exist, yet a backup that nobody can find, or that nobody has a license to republish, does not help a reader who only has a bookmark.

MediaWiki history can be exported, and exports matter. An export is still a file that someone must store. If the only copy sits on the same host that failed, the export fails with it. Content addressing starts after a snapshot exists. The snapshot is named by its bytes, replicated by parties who choose to keep it, and retrieved without asking the original host for permission to read a public object.

Content addressing in one pass

Content addressing means the identifier is a function of the content. Change one byte and the identifier changes. IPFS CIDs are hashes, with codec and multihash details, over the bytes or over a Merkle DAG of chunks. Arweave transaction identifiers name writes that the network accepts under its data protocol, and the bytes can be checked against the transaction data. In both cases, a citation can point at a version rather than at a URL path that a later editor may overwrite.

Caches are safe to share, because a gateway or a laptop can store the same CID and readers can verify the match. The label "latest" cannot live inside the hash. A signed pointer, a name record, or an application database must say which CID or transaction is current. Classic wikis hide that pointer inside the server. Content-addressed wikis make it explicit.

IPFS CIDs for wiki snapshots

IPFS fits wiki snapshots that many peers may host. A publisher exports a page, a namespace, or a full dump, adds the files to IPFS, and records the root CID. HTML, Markdown, wikitext, images, and MediaWiki XML dumps can all sit under that root. Readers who run a node, or who use a gateway, fetch the CID. If the publisher later revises an article, the new tree has a new CID. Old citations keep resolving to old bytes for as long as at least one pin or provider still offers them.

Pinning is the availability plan. IPFS does not promise that unpinned content stays online. A project that cares about a snapshot pays a pinning service, runs its own pinning nodes, or asks libraries and volunteers to pin the CID. Multiple independent pins reduce the chance that one budget cut erases the set. Pinning is an ongoing cost, unlike a one-time Arweave upload fee. Operators should treat pin renewal as a scheduled operating cost.

Gateways translate CIDs into ordinary HTTPS for browsers. A gateway can go down or refuse a CID, so projects list more than one gateway and keep a pin that does not depend on a single account. Where tools allow it, fetch the bytes, recompute the CID, and reject mismatches.

Arweave transaction identifiers and permaweb pages

Arweave is aimed at long-lived public data with payment up front. A wiki snapshot uploaded there receives a transaction identifier. After the write meets the confirmation bar the application requires, the identifier is a stable citation. New edits are new transactions. For most readers, the permaweb still looks like HTTPS through a gateway that maps identifiers and path manifests to bytes. See Permaweb_Basics:_How_Arweave_Makes_Content_Permanent for the storage model, and IPFS_and_Arweave_for_Agent_Memory_and_Knowledge_Bases for how the two networks differ when agents and knowledge bases are the workload.

Arweave fits releases that should stay citable for years, such as approved article versions and static HTML exports. It is a weak fit for every keystroke. Practical teams keep hot revisions in a mutable database and push signed releases on a schedule. Bundlers can group small objects. The team should know who holds signing keys, who pays the fee, and how a failed bundle is retried.

Mutability versus permanence

A wiki that never changes is a book. A wiki that can change under one administrator with no public history is a brochure. Content-addressed designs try to keep editability and a frozen record at the same time. The mutable layer holds accounts, recent drafts, search, and the pointer to the current release. The pinned or permanent layer holds snapshots that outsiders can mirror.

Edits inside the mutable layer can be reverted, renamed, or hidden under local policy. Once a snapshot is public under a CID or a transaction identifier, the publisher cannot guarantee that every copy on earth will disappear. Operators who must honor erasure duties should decide before they upload. Personal data and copyrighted media do not become lawful to distribute because the storage network is decentralized. Encryption before upload can protect private working notes. Encryption does not erase a plaintext copy that already circulated.

How this differs from a classic centralized wiki

Share-alike wiki text can be mirrored when attribution and revision history travel with the snapshot. A community split can publish two pointers. Each pointer can be valid. Readers still need a written rule for which pointer counts as current.

Classic MediaWiki offers live editing, talk pages, user permissions, categories, and search on one origin. Content-addressed publishing usually offers weaker live collaboration unless a separate application provides it. Search across public gateways is uneven. Account systems on a storage network are not wiki user accounts. Expect to run an editable front end that writes snapshots outward, or expect readers to tolerate a static mirror between releases.

Practical patterns

  • Export on a schedule: a static snapshot plus a manifest of titles and revision ids, published as a root CID or an Arweave transaction id.
  • Keep the live editor on ordinary hosting. Pin or permanently store releases, not every autosave.
  • Put the current identifier in a signed file and a changelog. Leave old identifiers in place so prior citations still resolve.
  • For IPFS, use at least two independent pinning setups. For Arweave, confirm the transaction on a gateway the project does not fully control, and keep a local replica.
  • Recompute CIDs on fetch. Check Arweave transaction data when the client allows it. Log mismatches.
  • Maintain a catalog. Permanent bytes are not a search ranking.
  • Keep non-public personal data off the public network. Encrypt private notes and store keys outside the snapshot. See Building_a_Second_Brain_on_the_Permaweb_Without_Losing_Ownership.
  • Write the steps to rebuild the mutable wiki from the latest snapshot.
  • Keep working notes in a mutable index and cited releases under content identifiers. See X402:_HTTP_402_Payments_for_AI_Agents and Open_Source_Freedom_Tech_Worth_Watching_in_Agent_Infrastructure.

What still fails

Content addressing does not pay editors, moderate abuse, or answer a reader who wants the page updated the same day. Gateways add availability risk even when the identifier is sound. IPFS content becomes unreachable when the last pin drops. Arweave retrieval still depends on gateways, caches, and network health. Courts and hosts can pressure gateways and naming layers even when copies exist somewhere else. A design that ignores those facts will fail during an ordinary outage and call the failure a surprise.

Conclusion

Content-addressed wikis keep public page versions from depending on one MediaWiki host. IPFS CIDs name snapshots that stay available while someone pins them. Arweave transaction identifiers name snapshots intended for long-term retrieval after an up-front fee, usually read through permaweb gateways. The live wiki remains a mutable application with an explicit pointer to the current identifier. Operators who publish on a schedule, replicate with more than one party, verify bytes, and keep non-public data off public networks get the durability the networks actually offer. They do not need to pretend that a hash replaces editors, policy, or search.

See also