The problem Base64 solves
Early email systems were built to carry 7-bit ASCII text. Any byte above 127, or a control character in the wrong place, could get altered or dropped by some server along the way. Once people needed to send attachments — images, documents, programs, all of which are arbitrary byte sequences — there had to be a way to disguise bytes as harmless text. The MIME standard (RFC 2045, from 1996) adopted Base64 for exactly that, and the same trick became the standard answer whenever binary data has to squeeze through a text-only channel: HTTP headers, JSON, XML, URLs, config files.
This guide walks through how Base64 works bit by bit with worked examples, why = signs show up at the end, which variants exist, and the most common mistakes when using it from code. To encode or decode a specific value, the Base64 encoder does it right in your browser.
Encoding by hand: "Man" → "TWFu"
The textbook example is the word Man. That's three ASCII characters — three bytes:
| Character | Decimal | Binary |
|---|---|---|
| M | 77 | 01001101 |
| a | 97 | 01100001 |
| n | 110 | 01101110 |
Base64 lines up all 24 bits and slices them again, this time into four 6-bit groups:
01001101 01100001 01101110 ← 3 bytes (8 bits each)
010011 010110 000101 101110 ← 4 groups of 6 bits
19 22 5 46 ← value of each group
T W F u ← alphabet symbol
Each value from 0 to 63 maps through the alphabet: A–Z are 0–25, a–z are 26–51, 0–9 are 52–61, and + and / are 62 and 63. 19 is T; 22 is W; 5 is F; and 46 lands in the lowercase range: 46 − 26 = 20, and the letter at position 20 counting a as 0 is u. Result: TWFu.
If you want to double-check the binary-to-decimal step for each group, the base converter shows any value in all four bases at once.
Padding: why = and == appear
The method above needs complete 3-byte blocks. When the input length isn't a multiple of 3, the final block comes up short, and it's handled like this: zero bits are appended to complete the last 6-bit group, and any characters still missing to reach 4 are replaced with =.
Two bytes, "Ma": that's 16 bits. Two zeros bring it to 18 (three 6-bit groups), and a single = fills the block:
01001101 01100001 00 ← 16 bits + 2 padding zeros
010011 010110 000100 ← 3 groups
T W E = ← TWE=
One byte, "M": that's 8 bits. Four zeros bring it to 12 (two groups), and == fills the block:
01001101 0000 ← 8 bits + 4 padding zeros
010011 010000 ← 2 groups
T Q = = ← TQ==
The rule is easy to remember: == means 1 byte was left over, = means 2, and no padding means the input was a multiple of 3. Padding carries no information — it only rounds the length up to a multiple of 4 — which is why many implementations accept it missing and some variants drop it entirely.
Decoding by hand
The reverse is symmetrical: translate each character into its 6-bit value, concatenate, and slice into 8-bit bytes. With TWFu:
T=19 W=22 F=5 u=46
010011 010110 000101 101110
01001101 01100001 01101110 → 77 97 110 → "Man"
If there's padding, the leftover filler bits are thrown away. Here's a handy sanity check: a Base64 string whose length (ignoring whitespace) leaves a remainder of 1 when divided by 4 is always invalid, because a single character only carries 6 bits — not enough to make a byte.
How much space it takes
Every 3 bytes become 4 characters, so size grows by 33% (4/3). In formats that wrap lines — like MIME, with 76 characters per line plus a \r\n — the real overhead is closer to 37%. That's why Base64 is a poor choice for shipping large files inside JSON: an API that takes multi-megabyte documents usually does better with multipart/form-data or a direct binary upload.
The variants
They all use the same 6-bit mechanism; what changes is the alphabet or the layout:
| Variant | Defined in | Difference |
|---|---|---|
| Standard Base64 | RFC 4648, section 4 | Alphabet with + and /, padded with = |
| base64url | RFC 4648, section 5 | - and _ instead of + and /; padding usually omitted |
| MIME | RFC 2045 | Standard alphabet, lines of at most 76 characters |
| PEM | RFC 7468 | Standard alphabet, 64-character lines between -----BEGIN …----- and -----END …----- |
base64url exists because +, /, and = all mean something in URLs. A standard Base64 token dropped straight into a query string can arrive altered: many servers read + as a space.
Where you'll run into it
HTTP Basic authentication. RFC 7617 uses the user Aladdin with the password open sesame as its example. The header the browser sends is:
Authorization: Basic QWxhZGRpbjpvcGVuIHNlc2FtZQ==
That value is just Aladdin:open sesame in Base64. Anyone who intercepts the request reads the password instantly, which is why Basic auth is only acceptable over HTTPS.
JSON Web Tokens. A JWT has three dot-separated parts. In a typical token, the first is eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9, which base64url-decodes to {"alg":"HS256","typ":"JWT"}. A JWT's header and payload are readable by anyone holding the token; all the signature guarantees is that nobody tampered with them.
Data URIs. data:image/svg+xml;base64,PHN2Zy... embeds a file inside HTML or CSS. It works well for small icons; for large images, the extra 33% and the loss of separate caching usually cost more than they save.
Base64 in code
| Environment | Encode UTF-8 text | Decode |
|---|---|---|
| Node.js | Buffer.from(text, "utf8").toString("base64") | Buffer.from(b64, "base64").toString("utf8") |
| Node.js (base64url) | Buffer.from(text).toString("base64url") | Buffer.from(b64, "base64url") |
| Python | base64.b64encode(text.encode()).decode() | base64.b64decode(b64).decode() |
| Python (base64url) | base64.urlsafe_b64encode(data) | base64.urlsafe_b64decode(b64) |
| Bash | printf '%s' "$text" | base64 | base64 -d |
In the browser, btoa() and atob() have been around for decades, but they operate on "binary strings," where each character stands for one byte. With non-ASCII text they either fail or give a different result than you'd expect, so convert the text to bytes with TextEncoder first.
Common mistakes
- Treating it as encryption. Storing passwords, API keys, or personal data "in Base64" protects nothing. It's as reversible as reading a number written in hex.
- The
echonewline.echo 'Hola' | base64encodesHola\nand producesSG9sYQo=. Withprintforecho -nyou getSG9sYQ==. This is the number one cause of "my Base64 doesn't match." - Latin-1 instead of UTF-8.
btoa("ñ")returns8Q==, which is ñ in Latin-1. In UTF-8 — what almost every modern system expects — ñ isw7E=. Andbtoa("€")simply throws anInvalidCharacterError. - Stripping padding and never restoring it. Many strict decoders, like Python's
base64.b64decode, fail withIncorrect paddingwhen the=is missing. When you receive unpadded base64url, pad it back to a multiple of 4 before decoding. - Mixing variants. A strict standard decoder rejects
-and_, and a base64url decoder rejects+and/. When the source isn't clear, normalize first. - Encoding twice.
SGVsbG8=encoded again becomesU0dWc2JHOD0=. If a decoded value still looks like Base64, it was probably encoded twice somewhere along the way.
Summary
Base64 turns every 3 bytes into 4 characters from a 64-symbol alphabet, so any binary data can travel over a text channel. = padding completes the final block, base64url swaps two symbols to make the output URL-safe, and the cost is 33% extra size. It protects nothing: it's transport, not security.