Toolbit
Guides

MD5, SHA-256, and bcrypt: Which Hash Function to Use for What

Integrity, message authentication, and passwords need different tools: how MD5 and SHA-1 fell, why HMAC prevents length extension, and why passwords call for Argon2id or bcrypt.

By Javier VallejoPublished 6 min read

Three different jobs, three different tools

"Hashing" gets used to describe three jobs that have nothing to do with each other, and a good share of hash-related security bugs come from using the tool for one job to do another:

  1. Verifying integrity: checking that a file or message hasn't changed. Tool: a fast, collision-resistant hash such as SHA-256.
  2. Authenticating a message: checking that a message comes from someone holding a secret key and wasn't altered. Tool: HMAC.
  3. Storing passwords: being able to verify a password without storing it. Tool: a deliberately slow, salted function such as Argon2id or bcrypt.

This guide covers what each one guarantees, how MD5 and SHA-1 fell, and what to use where. To compute hashes or HMACs for text or a file, the hash generator runs right in your browser.

What a cryptographic hash guarantees

A cryptographic hash function has to withstand three kinds of attack, ordered from hardest to easiest for the attacker:

  • Preimage: given a hash, find any input that produces it.
  • Second preimage: given a specific input, find another one with the same hash.
  • Collision: find any two inputs with the same hash.

Collisions are the easiest to find for a statistical reason, the birthday paradox: with an n-bit hash, trying about 2^(n/2) inputs is enough to give a good chance of two matching. That's why collision resistance is half the hash size:

AlgorithmSizeTheoretical collision resistanceReal-world status
MD5128 bits2⁶⁴Collisions in seconds on an ordinary computer
SHA-1160 bits2⁸⁰Collisions demonstrated (2017) and chosen-prefix (2020)
SHA-256256 bits2¹²⁸No practical attacks
SHA-512512 bits2²⁵⁶No practical attacks
SHA3-256256 bits2¹²⁸No practical attacks; a different design (Keccak)

How MD5 and SHA-1 fell

MD5 was designed by Ron Rivest in 1991. The first theoretical weaknesses surfaced in 1996, and in 2004 Xiaoyun Wang's team published real collisions. From there the attacks improved fast: in 2008, a group of researchers used an MD5 collision to create a rogue certificate authority certificate that browsers accepted as valid. In 2012, the Flame espionage malware used a similar technique to forge a Microsoft signature and spread through Windows Update.

SHA-1 held out longer. A theoretical attack was published in 2005, and in 2017 Google and the CWI institute in Amsterdam unveiled SHAttered: two different PDF files with the same SHA-1 hash, produced after roughly 9 quintillion (9 × 10¹⁸) SHA-1 computations. That same year browsers stopped accepting SHA-1-signed certificates. In 2020 came the chosen-prefix collision, which lets you craft collisions from any two documents — the most dangerous kind of attack in practice. Git, which used SHA-1 to identify all its content, added collision detection and gained support for SHA-256 repositories.

The lesson isn't that MD5 and SHA-1 give "wrong" results: they still catch an accidentally damaged file perfectly well. What they lost is the guarantee against someone who chooses the content on purpose.

Integrity: download checksums

To verify that a downloaded file matches the published one, you compare its hash to the one the site publishes. With SHA-256, if they match, the file is identical bit for bit. With MD5, if they match, the file is identical unless someone deliberately prepared a different version with the same hash — which is feasible today.

Two practical notes. First, a CRC32 (the kind used by ZIP files or Ethernet) catches transmission errors but isn't cryptographic: forging a collision is trivial. Second, a checksum is only as trustworthy as the channel you got it from: if the attacker controls the server, they swap the file and the checksum together. That's why Linux distributions sign their SHA256SUMS file with GPG.

Authenticating messages: why HMAC and not hash(key + message)

Picture an API that signs each request by computing SHA-256(secret_key + message) and sending it along with the message. Seems safe: without the key, nobody can compute the signature. But MD5, SHA-1, SHA-256, and SHA-512 share an internal construction (Merkle–Damgård) that enables a length extension attack: anyone who knows the hash of key + message and the key's length can compute the hash of key + message + padding + extra_data without knowing the key. In 2009, Flickr's API was vulnerable to exactly this: an attacker could append parameters to a signed request and produce a valid signature.

The standard fix is HMAC (RFC 2104), which hashes twice with the key mixed in a way that neutralizes the attack. HMAC-SHA256 is what most platforms' webhooks use, what signs JWTs with HS256, and what authenticates many APIs. A test case from RFC 4231: with the key Jefe and the message what do ya want for nothing?, the HMAC-SHA256 is 5bdcc146bf60754e…64ec3843. You can check it in the generator by typing the key into the HMAC field.

When verifying a received HMAC, compare in constant time (crypto.timingSafeEqual in Node.js, hmac.compare_digest in Python). A regular comparison stops at the first differing character, and measuring that timing difference lets an attacker guess the signature one character at a time.

Passwords: why a fast hash is a mistake

To store passwords, you don't store them — you store something that lets you verify them: at login, you compute the same thing from the entered password and compare. The temptation is SHA-256(password). The problem is speed:

FunctionOrder of magnitude on a modern GPU
MD5over 100 billion guesses per second
SHA-256tens of billions per second
bcrypt (cost 10–12)thousands per second
Argon2id (19 MiB of memory)comparable to bcrypt, and it also demands 19 MiB of memory per guess

If a database of SHA-256 hashes leaks, an attacker runs dictionaries of billions of known passwords against it in seconds. Functions built for passwords add three defenses:

  • Salt: a random value unique to each user, stored alongside the hash. Two users with the same password get different hashes, and precomputed tables (rainbow tables) become useless.
  • Tunable cost: you can configure each computation to take, say, 100 ms on the server. Imperceptible to the user; for the attacker, it means millions of times fewer guesses.
  • Memory usage (Argon2id, scrypt): each guess requires tens of megabytes, which limits how many a GPU can run in parallel.

OWASP's current recommendations, in order of preference:

FunctionMinimum recommended parameters
Argon2id19 MiB of memory, 2 iterations, parallelism 1
scryptN = 2¹⁷, r = 8, p = 1
bcryptcost 10 or higher; has a 72-byte password limit
PBKDF2-HMAC-SHA256600,000 iterations (when FIPS-140 compliance is required)

How long a password holds out also depends on its entropy, covered in How to Create Strong Passwords: Entropy Explained. A slow function turns a 40-bit password from "minutes" into "years," but it won't save 123456, which tops every dictionary.

What to use for each job

JobUseAvoid
Verifying a download or fileSHA-256MD5 and SHA-1 if tampering is possible
Spotting duplicates or accidental changesSHA-256; MD5 is fine with no adversaryCRC32 if security matters
Signing messages with a shared keyHMAC-SHA256hash(key + message)
Storing passwordsArgon2id, scrypt, or bcryptAny fast hash, salted or not
Digital signatures and certificatesSHA-256 or strongerMD5 and SHA-1
Content addressing (Git, caches)SHA-256SHA-1 in new systems

Summary

A cryptographic hash condenses any input into a fixed-size fingerprint, and its value rests on nobody being able to forge two inputs with the same fingerprint: MD5 and SHA-1 lost that property, SHA-256 keeps it. To authenticate messages, a bare hash isn't enough — use HMAC. For passwords, the speed that makes SHA-256 useful becomes its worst flaw, and the answer is slow, salted functions: Argon2id, scrypt, or bcrypt.

Tools used in this guide

Related guides