Toolbit
Guides

Base64 Explained: How It Works, Padding, and Variants

How Base64 turns every 3 bytes into 4 characters, why = and == appear, what base64url changes, and the most common mistakes when using it from code.

By Javier VallejoPublished 6 min read

The problem Base64 solves

Early email systems were built to carry 7-bit ASCII text. Any byte above 127, or a control character in the wrong place, could get altered or dropped by some server along the way. Once people needed to send attachments — images, documents, programs, all of which are arbitrary byte sequences — there had to be a way to disguise bytes as harmless text. The MIME standard (RFC 2045, from 1996) adopted Base64 for exactly that, and the same trick became the standard answer whenever binary data has to squeeze through a text-only channel: HTTP headers, JSON, XML, URLs, config files.

This guide walks through how Base64 works bit by bit with worked examples, why = signs show up at the end, which variants exist, and the most common mistakes when using it from code. To encode or decode a specific value, the Base64 encoder does it right in your browser.

Encoding by hand: "Man" → "TWFu"

The textbook example is the word Man. That's three ASCII characters — three bytes:

CharacterDecimalBinary
M7701001101
a9701100001
n11001101110

Base64 lines up all 24 bits and slices them again, this time into four 6-bit groups:

01001101 01100001 01101110      ← 3 bytes (8 bits each)
010011 010110 000101 101110     ← 4 groups of 6 bits
  19     22      5     46       ← value of each group
  T      W       F     u        ← alphabet symbol

Each value from 0 to 63 maps through the alphabet: A–Z are 0–25, a–z are 26–51, 0–9 are 52–61, and + and / are 62 and 63. 19 is T; 22 is W; 5 is F; and 46 lands in the lowercase range: 46 − 26 = 20, and the letter at position 20 counting a as 0 is u. Result: TWFu.

If you want to double-check the binary-to-decimal step for each group, the base converter shows any value in all four bases at once.

Padding: why = and == appear

The method above needs complete 3-byte blocks. When the input length isn't a multiple of 3, the final block comes up short, and it's handled like this: zero bits are appended to complete the last 6-bit group, and any characters still missing to reach 4 are replaced with =.

Two bytes, "Ma": that's 16 bits. Two zeros bring it to 18 (three 6-bit groups), and a single = fills the block:

01001101 01100001 00            ← 16 bits + 2 padding zeros
010011 010110 000100            ← 3 groups
  T      W      E      =        ← TWE=

One byte, "M": that's 8 bits. Four zeros bring it to 12 (two groups), and == fills the block:

01001101 0000                   ← 8 bits + 4 padding zeros
010011 010000                   ← 2 groups
  T      Q      =      =        ← TQ==

The rule is easy to remember: == means 1 byte was left over, = means 2, and no padding means the input was a multiple of 3. Padding carries no information — it only rounds the length up to a multiple of 4 — which is why many implementations accept it missing and some variants drop it entirely.

Decoding by hand

The reverse is symmetrical: translate each character into its 6-bit value, concatenate, and slice into 8-bit bytes. With TWFu:

T=19   W=22   F=5    u=46
010011 010110 000101 101110
01001101 01100001 01101110  →  77 97 110  →  "Man"

If there's padding, the leftover filler bits are thrown away. Here's a handy sanity check: a Base64 string whose length (ignoring whitespace) leaves a remainder of 1 when divided by 4 is always invalid, because a single character only carries 6 bits — not enough to make a byte.

How much space it takes

Every 3 bytes become 4 characters, so size grows by 33% (4/3). In formats that wrap lines — like MIME, with 76 characters per line plus a \r\n — the real overhead is closer to 37%. That's why Base64 is a poor choice for shipping large files inside JSON: an API that takes multi-megabyte documents usually does better with multipart/form-data or a direct binary upload.

The variants

They all use the same 6-bit mechanism; what changes is the alphabet or the layout:

VariantDefined inDifference
Standard Base64RFC 4648, section 4Alphabet with + and /, padded with =
base64urlRFC 4648, section 5- and _ instead of + and /; padding usually omitted
MIMERFC 2045Standard alphabet, lines of at most 76 characters
PEMRFC 7468Standard alphabet, 64-character lines between -----BEGIN …----- and -----END …-----

base64url exists because +, /, and = all mean something in URLs. A standard Base64 token dropped straight into a query string can arrive altered: many servers read + as a space.

Where you'll run into it

HTTP Basic authentication. RFC 7617 uses the user Aladdin with the password open sesame as its example. The header the browser sends is:

Authorization: Basic QWxhZGRpbjpvcGVuIHNlc2FtZQ==

That value is just Aladdin:open sesame in Base64. Anyone who intercepts the request reads the password instantly, which is why Basic auth is only acceptable over HTTPS.

JSON Web Tokens. A JWT has three dot-separated parts. In a typical token, the first is eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9, which base64url-decodes to {"alg":"HS256","typ":"JWT"}. A JWT's header and payload are readable by anyone holding the token; all the signature guarantees is that nobody tampered with them.

Data URIs. data:image/svg+xml;base64,PHN2Zy... embeds a file inside HTML or CSS. It works well for small icons; for large images, the extra 33% and the loss of separate caching usually cost more than they save.

Base64 in code

EnvironmentEncode UTF-8 textDecode
Node.jsBuffer.from(text, "utf8").toString("base64")Buffer.from(b64, "base64").toString("utf8")
Node.js (base64url)Buffer.from(text).toString("base64url")Buffer.from(b64, "base64url")
Pythonbase64.b64encode(text.encode()).decode()base64.b64decode(b64).decode()
Python (base64url)base64.urlsafe_b64encode(data)base64.urlsafe_b64decode(b64)
Bashprintf '%s' "$text" | base64base64 -d

In the browser, btoa() and atob() have been around for decades, but they operate on "binary strings," where each character stands for one byte. With non-ASCII text they either fail or give a different result than you'd expect, so convert the text to bytes with TextEncoder first.

Common mistakes

  • Treating it as encryption. Storing passwords, API keys, or personal data "in Base64" protects nothing. It's as reversible as reading a number written in hex.
  • The echo newline. echo 'Hola' | base64 encodes Hola\n and produces SG9sYQo=. With printf or echo -n you get SG9sYQ==. This is the number one cause of "my Base64 doesn't match."
  • Latin-1 instead of UTF-8. btoa("ñ") returns 8Q==, which is ñ in Latin-1. In UTF-8 — what almost every modern system expects — ñ is w7E=. And btoa("€") simply throws an InvalidCharacterError.
  • Stripping padding and never restoring it. Many strict decoders, like Python's base64.b64decode, fail with Incorrect padding when the = is missing. When you receive unpadded base64url, pad it back to a multiple of 4 before decoding.
  • Mixing variants. A strict standard decoder rejects - and _, and a base64url decoder rejects + and /. When the source isn't clear, normalize first.
  • Encoding twice. SGVsbG8= encoded again becomes U0dWc2JHOD0=. If a decoded value still looks like Base64, it was probably encoded twice somewhere along the way.

Summary

Base64 turns every 3 bytes into 4 characters from a 64-symbol alphabet, so any binary data can travel over a text channel. = padding completes the final block, base64url swaps two symbols to make the output URL-safe, and the cost is 33% extra size. It protects nothing: it's transport, not security.

Tools used in this guide

Related guides