How content hashing works
Content hashing is the practice of running data through a cryptographic hash function to get a short, fixed-length fingerprint, then using that fingerprint to check, name, and find the data. It quietly powers Git, IPFS, Arweave, software downloads, and the revision history of MediaWiki. This page explains the concept in plain language, shows how several real systems use it, and is honest about what a hash can and cannot prove.
For the bigger picture of why this matters to wikis, see Content-addressed wikis and History of decentralized wikis.
What a cryptographic hash is
A hash function takes any input, a single word or an entire encyclopedia, and produces an output of fixed length. A cryptographic hash function adds properties that make the output useful as a fingerprint:
- Deterministic. The same input always gives the same output, on any computer, today or decades from now.
- Fixed length. SHA-256 always produces 256 bits, usually written as 64 hexadecimal characters, whatever the input size.
- Avalanche effect. A tiny change in the input, even one letter, produces a completely different output.
- One-way. Given a hash, there is no practical way to work backward to the input.
- Collision resistant. It should be infeasible to find two different inputs with the same hash.
Here is a small example. These are the SHA-256 hashes of two words that differ only in the case of one letter:
IdeaWaza 32b043440df9cbdb89d9fa82bd58ed3a1cf1a6172a48fbf1efe1d475cd5da058 Ideawaza 1efa1c317f7cab94bdcbcf35d1b2c2f782dfafd54447492bdf63a592f48dbd47
The outputs share no visible pattern. Anyone who holds a copy of the data can recompute the hash and compare it to a published value. A match means the data is byte-for-byte identical. A mismatch means something changed.
Common hash functions
SHA-256 belongs to the SHA-2 family, designed by the United States National Security Agency and first published in 2001, then standardized by NIST. It is the workhorse of modern integrity checking.
SHA-1 is older, with an output of 160 bits, written as 40 hexadecimal characters. It was widely used for years, including by Git to name every object. Researchers found weaknesses over time, and on 23 February 2017 the SHAttered attack demonstrated a practical SHA-1 collision: two different files with the same SHA-1 hash. Git responded by switching to a hardened SHA-1 implementation by default starting with version 2.13.0, which detects that attack. In 2018 the Git project picked SHA-256 as the successor hash, and Git now supports repositories that use SHA-256 for object names.
The lesson is that hash functions age. Systems built for the long term should record which function they used, so a future reader knows how to check the data and when to migrate.
Content addressing
The web mostly uses location addressing. A URL says where to find something: this server, this path. If the server moves, or the page is quietly edited, the URL still looks the same.
Content addressing names data by its hash instead. The name says what the data is, regardless of where it lives. This has three useful effects:
- Any copy from any source can be verified against the name, so it no longer matters who serves it.
- Identical data gets the same name everywhere, which makes deduplication natural.
- An edit produces new data with a new name, so a name always refers to one exact version.
How Git uses hashes
At its core, Git is a content-addressable filesystem. It stores a few kinds of objects:
- A blob holds the contents of a file.
- A tree lists file names and points to the hashes of blobs and other trees, like a directory.
- A commit points to one tree, the snapshot of the whole project, plus the hash of its parent commit or commits, the author, a timestamp, and a message.
Each object is named by the hash of its type, its length, a null byte, and its content. You can reproduce this yourself. The Git blob name for a file containing the word hello and a newline is the SHA-1 of the bytes "blob 6", a null byte, and "hello" plus newline:
ce013625030ba8dba906f756967f9e9ca394464a
Running git hash-object on the file gives the same answer. Because each commit includes the hash of its parent, which in turn includes the hash of its own parent, a single commit hash covers the entire history behind it. Change one character in an old file and every hash from that point forward changes. This chain is a Merkle structure, in which each node commits to everything beneath it.
Diffs, patches, and hashes
A diff describes the difference between two versions. A patch is a diff packaged so someone else can apply it. Applying a patch to some content yields new content, and the new content has a new hash.
Git stores whole snapshots, not chains of diffs. When you ask for a diff, Git compares two snapshots and computes it on demand.
Patches can have identities too. The git patch-id command computes a "patch ID", which the documentation describes as a sum of SHA-1 hashes of the file diffs with line numbers ignored. Two commits that make the same change on different branches will usually share a patch ID even though their commit hashes differ, which makes the command handy for spotting likely duplicate commits.
IPFS CIDs and Merkle DAGs
IPFS names content with a content identifier, or CID. IPFS uses SHA-256 by default and supports other hash functions through a format called multihash. A CID also records how to interpret the data, using codec and encoding information.
Large files are split into blocks. The blocks are arranged in a directed acyclic graph (a DAG), and the root block contains links to the others. The CID is derived from the root block, so it commits to every block below it. This is why a CID usually does not match a plain SHA-256 checksum of the same file. Chunk size, layout, codec, and CID version all affect the result.
Arweave transaction ids
Arweave also leans on hashing. According to the Arweave HTTP API documentation, a transaction id is a SHA-256 hash of the transaction signature, a wallet address is a SHA-256 hash of the public key, and the data of a transaction is committed through a Merkle root of its chunks. Together these let anyone fetch data from any gateway and confirm it matches what was signed and stored. See Permaweb basics.
MediaWiki revisions
MediaWiki keeps a row for every edit in its revision table. Since MediaWiki 1.19, the rev_sha1 field stores a SHA-1 hash of the revision content, encoded in base 36. Since version 1.32, when a revision can hold several content slots, the value is a nested hash across all slots, and for a single slot it equals the hash of that content.
This gives wiki mirrors a cheap integrity check. Someone who exports pages, keeps a local Git mirror, or republishes revisions on permanent storage can recompute hashes and compare them with the source wiki. If every hash matches, the mirror holds the same text. Tools proposed in Pre-publish checks for MediaWiki and Git can use the same idea to confirm that what was reviewed is exactly what gets published.
What hashes cannot prove
Hashes are powerful but narrow.
- Integrity, not truth. A matching hash proves the data is unchanged. It says nothing about whether the content is accurate, fair, or lawful.
- Integrity, not authorship. Anyone can hash anything. To show who published a version, you need a digital signature, made with a private key and checked with the matching public key. Git supports signed commits and tags, and Arweave transactions are signed by the uploading wallet.
- Trust in the published hash. A checksum only helps if you got it from a source you trust. An attacker who can change both the file and the posted hash defeats the check.
- Aging algorithms. As SHA-1 showed, functions weaken over time, so record the algorithm and plan for migration.
Used with signatures, careful review, and good sources, content hashing lets a community share knowledge across many hosts while keeping a checkable record of exactly what was said, a solid foundation for Self-sustaining wikis and other long-lived knowledge projects.
See also
- Content-addressed wikis
- History of decentralized wikis
- Self-sustaining wikis
- IPFS and Arweave
- Permaweb basics
- Local-first wiki and Git mirrors
- Pre-publish checks for MediaWiki and Git