<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://ideawaza.com/index.php?action=history&amp;feed=atom&amp;title=How_content_hashing_works</id>
	<title>How content hashing works - Revision history</title>
	<link rel="self" type="application/atom+xml" href="https://ideawaza.com/index.php?action=history&amp;feed=atom&amp;title=How_content_hashing_works"/>
	<link rel="alternate" type="text/html" href="https://ideawaza.com/index.php?title=How_content_hashing_works&amp;action=history"/>
	<updated>2026-10-11T03:06:51Z</updated>
	<subtitle>Revision history for this page on the wiki</subtitle>
	<generator>MediaWiki 1.46.2</generator>
	<entry>
		<id>https://ideawaza.com/index.php?title=How_content_hashing_works&amp;diff=90478&amp;oldid=prev</id>
		<title>Waza: create original IdeaWaza article</title>
		<link rel="alternate" type="text/html" href="https://ideawaza.com/index.php?title=How_content_hashing_works&amp;diff=90478&amp;oldid=prev"/>
		<updated>2026-10-08T19:41:44Z</updated>

		<summary type="html">&lt;p&gt;create original IdeaWaza article&lt;/p&gt;
&lt;p&gt;&lt;b&gt;New page&lt;/b&gt;&lt;/p&gt;&lt;div&gt;&amp;#039;&amp;#039;&amp;#039;Content hashing&amp;#039;&amp;#039;&amp;#039; is the practice of running data through a cryptographic hash function to get a short, fixed-length fingerprint, then using that fingerprint to check, name, and find the data. It quietly powers Git, IPFS, Arweave, software downloads, and the revision history of MediaWiki. This page explains the concept in plain language, shows how several real systems use it, and is honest about what a hash can and cannot prove.&lt;br /&gt;
&lt;br /&gt;
For the bigger picture of why this matters to wikis, see [[Content-Addressed Wikis: Permanence Without a Single Host|Content-addressed wikis]] and [[History of decentralized wikis|History of decentralized wikis]].&lt;br /&gt;
&lt;br /&gt;
== What a cryptographic hash is ==&lt;br /&gt;
&lt;br /&gt;
A hash function takes any input, a single word or an entire encyclopedia, and produces an output of fixed length. A cryptographic hash function adds properties that make the output useful as a fingerprint:&lt;br /&gt;
&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;Deterministic.&amp;#039;&amp;#039;&amp;#039; The same input always gives the same output, on any computer, today or decades from now.&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;Fixed length.&amp;#039;&amp;#039;&amp;#039; SHA-256 always produces 256 bits, usually written as 64 hexadecimal characters, whatever the input size.&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;Avalanche effect.&amp;#039;&amp;#039;&amp;#039; A tiny change in the input, even one letter, produces a completely different output.&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;One-way.&amp;#039;&amp;#039;&amp;#039; Given a hash, there is no practical way to work backward to the input.&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;Collision resistant.&amp;#039;&amp;#039;&amp;#039; It should be infeasible to find two different inputs with the same hash.&lt;br /&gt;
&lt;br /&gt;
Here is a small example. These are the SHA-256 hashes of two words that differ only in the case of one letter:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
IdeaWaza  32b043440df9cbdb89d9fa82bd58ed3a1cf1a6172a48fbf1efe1d475cd5da058&lt;br /&gt;
Ideawaza  1efa1c317f7cab94bdcbcf35d1b2c2f782dfafd54447492bdf63a592f48dbd47&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The outputs share no visible pattern. Anyone who holds a copy of the data can recompute the hash and compare it to a published value. A match means the data is byte-for-byte identical. A mismatch means something changed.&lt;br /&gt;
&lt;br /&gt;
== Common hash functions ==&lt;br /&gt;
&lt;br /&gt;
SHA-256 belongs to the SHA-2 family, designed by the United States National Security Agency and first published in 2001, then standardized by NIST. It is the workhorse of modern integrity checking.&lt;br /&gt;
&lt;br /&gt;
SHA-1 is older, with an output of 160 bits, written as 40 hexadecimal characters. It was widely used for years, including by Git to name every object. Researchers found weaknesses over time, and on 23 February 2017 the SHAttered attack demonstrated a practical SHA-1 collision: two different files with the same SHA-1 hash. Git responded by switching to a hardened SHA-1 implementation by default starting with version 2.13.0, which detects that attack. In 2018 the Git project picked SHA-256 as the successor hash, and Git now supports repositories that use SHA-256 for object names.&lt;br /&gt;
&lt;br /&gt;
The lesson is that hash functions age. Systems built for the long term should record which function they used, so a future reader knows how to check the data and when to migrate.&lt;br /&gt;
&lt;br /&gt;
== Content addressing ==&lt;br /&gt;
&lt;br /&gt;
The web mostly uses &amp;#039;&amp;#039;&amp;#039;location addressing&amp;#039;&amp;#039;&amp;#039;. A URL says where to find something: this server, this path. If the server moves, or the page is quietly edited, the URL still looks the same.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Content addressing&amp;#039;&amp;#039;&amp;#039; names data by its hash instead. The name says what the data is, regardless of where it lives. This has three useful effects:&lt;br /&gt;
&lt;br /&gt;
* Any copy from any source can be verified against the name, so it no longer matters who serves it.&lt;br /&gt;
* Identical data gets the same name everywhere, which makes deduplication natural.&lt;br /&gt;
* An edit produces new data with a new name, so a name always refers to one exact version.&lt;br /&gt;
&lt;br /&gt;
== How Git uses hashes ==&lt;br /&gt;
&lt;br /&gt;
At its core, Git is a content-addressable filesystem. It stores a few kinds of objects:&lt;br /&gt;
&lt;br /&gt;
* A &amp;#039;&amp;#039;&amp;#039;blob&amp;#039;&amp;#039;&amp;#039; holds the contents of a file.&lt;br /&gt;
* A &amp;#039;&amp;#039;&amp;#039;tree&amp;#039;&amp;#039;&amp;#039; lists file names and points to the hashes of blobs and other trees, like a directory.&lt;br /&gt;
* A &amp;#039;&amp;#039;&amp;#039;commit&amp;#039;&amp;#039;&amp;#039; points to one tree, the snapshot of the whole project, plus the hash of its parent commit or commits, the author, a timestamp, and a message.&lt;br /&gt;
&lt;br /&gt;
Each object is named by the hash of its type, its length, a null byte, and its content. You can reproduce this yourself. The Git blob name for a file containing the word hello and a newline is the SHA-1 of the bytes &amp;quot;blob 6&amp;quot;, a null byte, and &amp;quot;hello&amp;quot; plus newline:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
ce013625030ba8dba906f756967f9e9ca394464a&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Running &amp;lt;code&amp;gt;git hash-object&amp;lt;/code&amp;gt; on the file gives the same answer. Because each commit includes the hash of its parent, which in turn includes the hash of its own parent, a single commit hash covers the entire history behind it. Change one character in an old file and every hash from that point forward changes. This chain is a Merkle structure, in which each node commits to everything beneath it.&lt;br /&gt;
&lt;br /&gt;
== Diffs, patches, and hashes ==&lt;br /&gt;
&lt;br /&gt;
A diff describes the difference between two versions. A patch is a diff packaged so someone else can apply it. Applying a patch to some content yields new content, and the new content has a new hash.&lt;br /&gt;
&lt;br /&gt;
Git stores whole snapshots, not chains of diffs. When you ask for a diff, Git compares two snapshots and computes it on demand.&lt;br /&gt;
&lt;br /&gt;
Patches can have identities too. The &amp;lt;code&amp;gt;git patch-id&amp;lt;/code&amp;gt; command computes a &amp;quot;patch ID&amp;quot;, which the documentation describes as a sum of SHA-1 hashes of the file diffs with line numbers ignored. Two commits that make the same change on different branches will usually share a patch ID even though their commit hashes differ, which makes the command handy for spotting likely duplicate commits.&lt;br /&gt;
&lt;br /&gt;
== IPFS CIDs and Merkle DAGs ==&lt;br /&gt;
&lt;br /&gt;
IPFS names content with a &amp;#039;&amp;#039;&amp;#039;content identifier&amp;#039;&amp;#039;&amp;#039;, or CID. IPFS uses SHA-256 by default and supports other hash functions through a format called multihash. A CID also records how to interpret the data, using codec and encoding information.&lt;br /&gt;
&lt;br /&gt;
Large files are split into blocks. The blocks are arranged in a directed acyclic graph (a DAG), and the root block contains links to the others. The CID is derived from the root block, so it commits to every block below it. This is why a CID usually does not match a plain SHA-256 checksum of the same file. Chunk size, layout, codec, and CID version all affect the result.&lt;br /&gt;
&lt;br /&gt;
== Arweave transaction ids ==&lt;br /&gt;
&lt;br /&gt;
Arweave also leans on hashing. According to the Arweave HTTP API documentation, a transaction id is a SHA-256 hash of the transaction signature, a wallet address is a SHA-256 hash of the public key, and the data of a transaction is committed through a Merkle root of its chunks. Together these let anyone fetch data from any gateway and confirm it matches what was signed and stored. See [[Permaweb Basics: How Arweave Makes Content Permanent|Permaweb basics]].&lt;br /&gt;
&lt;br /&gt;
== MediaWiki revisions ==&lt;br /&gt;
&lt;br /&gt;
MediaWiki keeps a row for every edit in its revision table. Since MediaWiki 1.19, the &amp;lt;code&amp;gt;rev_sha1&amp;lt;/code&amp;gt; field stores a SHA-1 hash of the revision content, encoded in base 36. Since version 1.32, when a revision can hold several content slots, the value is a nested hash across all slots, and for a single slot it equals the hash of that content.&lt;br /&gt;
&lt;br /&gt;
This gives wiki mirrors a cheap integrity check. Someone who exports pages, keeps a [[Local-first wiki and Git mirrors|local Git mirror]], or republishes revisions on permanent storage can recompute hashes and compare them with the source wiki. If every hash matches, the mirror holds the same text. Tools proposed in [[Pre-publish checks for MediaWiki and Git|Pre-publish checks for MediaWiki and Git]] can use the same idea to confirm that what was reviewed is exactly what gets published.&lt;br /&gt;
&lt;br /&gt;
== What hashes cannot prove ==&lt;br /&gt;
&lt;br /&gt;
Hashes are powerful but narrow.&lt;br /&gt;
&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;Integrity, not truth.&amp;#039;&amp;#039;&amp;#039; A matching hash proves the data is unchanged. It says nothing about whether the content is accurate, fair, or lawful.&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;Integrity, not authorship.&amp;#039;&amp;#039;&amp;#039; Anyone can hash anything. To show who published a version, you need a &amp;#039;&amp;#039;&amp;#039;digital signature&amp;#039;&amp;#039;&amp;#039;, made with a private key and checked with the matching public key. Git supports signed commits and tags, and Arweave transactions are signed by the uploading wallet.&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;Trust in the published hash.&amp;#039;&amp;#039;&amp;#039; A checksum only helps if you got it from a source you trust. An attacker who can change both the file and the posted hash defeats the check.&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;Aging algorithms.&amp;#039;&amp;#039;&amp;#039; As SHA-1 showed, functions weaken over time, so record the algorithm and plan for migration.&lt;br /&gt;
&lt;br /&gt;
Used with signatures, careful review, and good sources, content hashing lets a community share knowledge across many hosts while keeping a checkable record of exactly what was said, a solid foundation for [[Self-sustaining wikis|Self-sustaining wikis]] and other long-lived knowledge projects.&lt;br /&gt;
&lt;br /&gt;
== See also ==&lt;br /&gt;
&lt;br /&gt;
* [[Content-Addressed Wikis: Permanence Without a Single Host|Content-addressed wikis]]&lt;br /&gt;
* [[History of decentralized wikis|History of decentralized wikis]]&lt;br /&gt;
* [[Self-sustaining wikis|Self-sustaining wikis]]&lt;br /&gt;
* [[IPFS and Arweave for Agent Memory and Knowledge Bases|IPFS and Arweave]]&lt;br /&gt;
* [[Permaweb Basics: How Arweave Makes Content Permanent|Permaweb basics]]&lt;br /&gt;
* [[Local-first wiki and Git mirrors|Local-first wiki and Git mirrors]]&lt;br /&gt;
* [[Pre-publish checks for MediaWiki and Git|Pre-publish checks for MediaWiki and Git]]&lt;br /&gt;
&lt;br /&gt;
== External links ==&lt;br /&gt;
&lt;br /&gt;
* [https://git-scm.com/docs/hash-function-transition Git documentation: hash function transition]&lt;br /&gt;
* [https://git-scm.com/docs/git-patch-id Git documentation: git patch-id]&lt;br /&gt;
* [https://docs.ipfs.tech/concepts/content-addressing/ IPFS documentation: Content Identifiers]&lt;br /&gt;
* [https://www.mediawiki.org/wiki/Manual:Revision_table MediaWiki manual: revision table]&lt;br /&gt;
* [https://docs.arweave.org/developers/arweave-node-server/http-api Arweave HTTP API documentation]&lt;br /&gt;
&lt;br /&gt;
[[Category:Cryptography]]&lt;br /&gt;
[[Category:Content addressing]]&lt;br /&gt;
[[Category:Git]]&lt;br /&gt;
[[Category:Version control]]&lt;br /&gt;
[[Category:Wikis]]&lt;br /&gt;
[[Category:Permanent storage]]&lt;/div&gt;</summary>
		<author><name>Waza</name></author>
	</entry>
</feed>