Why URLs need encoding
A URL is text with structure: some characters are data, and others are delimiters that tell the browser and the server where one part ends and the next begins. Trouble starts when a piece of data contains one of those delimiters — an & in a product name, a / in a search term — or a character that simply isn't allowed in a URL, like a space or an accented letter. Percent-encoding handles both cases with a single mechanism. This guide covers what the standard says, what changes from one part of a URL to another, and why most bugs come from encoding too little, encoding too much, or reaching for the wrong function. To encode or decode a specific value, the URL encoder does it right in your browser.
Anatomy of a URL
RFC 3986 defines the generic syntax for URIs. Taking https://ana@shop.example.com:8443/products/café?color=red&size=M#reviews as an example:
| Part | Value in the example | Delimiter |
|---|---|---|
| Scheme | https | ends with : |
| User info | ana | ends with @ |
| Host | shop.example.com | starts after // |
| Port | 8443 | starts at : |
| Path | /products/café | segments separated by / |
| Query | color=red&size=M | starts at ? |
| Fragment | reviews | starts at # |
As written, that URL isn't valid per the RFC: café contains a non-ASCII character. What actually gets sent is /products/caf%C3%A9. Browsers show the readable version in the address bar, but the encoded one is what travels over the network.
Reserved and unreserved characters
RFC 3986 splits the allowed characters into groups:
| Group | Characters | Rule |
|---|---|---|
| Unreserved | A–Z a–z 0–9 - . _ ~ | Never need encoding |
| General delimiters | : / ? # [ ] @ | Separate the parts of the URL |
| Sub-delimiters | ! $ & ' ( ) * + , ; = | Separate data within a part (e.g., & and = in the query) |
The practical rule: a reserved character used as data, rather than as a delimiter, must be encoded. Anything outside these groups — spaces, accented letters, quotes, <, >, % as data — is always encoded.
How a character gets encoded
First, the character is converted to bytes with UTF-8, which is what RFC 3986 recommends for any new text. Then each byte is written as % followed by two hex digits, uppercase per the standard's recommendation:
| Character | UTF-8 bytes | Encoded |
|---|---|---|
| space | 20 | %20 |
& | 26 | %26 |
é | C3 A9 | %C3%A9 |
ñ | C3 B1 | %C3%B1 |
€ | E2 82 AC | %E2%82%AC |
🙂 | F0 9F 99 82 | %F0%9F%99%82 |
So café becomes caf%C3%A9. The host is the exception: domain names with non-ASCII characters use Punycode instead of percent-encoding, so münchen.de resolves as xn--mnchen-3ya.de.
Each part has its own rules
Path. / separates segments, so if a segment contains a slash as data (a file named report 1/2.pdf), it has to be encoded as %2F. A + in the path is a literal plus sign, not a space. Spaces go in as %20.
Query string. By convention, parameters are written as name=value separated by &, so inside a name or value you must encode &, =, #, and +. The RFC allows unencoded / and ? in the query, but encoding them never breaks anything. This is where the great space confusion lives: HTML forms use the application/x-www-form-urlencoded format, where a space is written as +, while RFC 3986 writes it as %20. In the query, almost every server accepts both — which is exactly why a + that's meant as data has to go in as %2B, or it'll be read as a space.
Fragment. Everything after # never reaches the server: the browser uses it locally, to jump to a section of the page or for single-page app routing. It has its own encoding rules, but sensitive data in the fragment doesn't land in server logs — and the server can't read it either.
Classic mistakes
Concatenating instead of encoding
const url = "https://api.example.com/search?q=" + "salt & pepper";
// https://api.example.com/search?q=salt & pepper
// The server receives q = "salt " plus an empty parameter named " pepper".
The robust fix is to stop building the query by hand:
const url = new URL("https://api.example.com/search");
url.searchParams.set("q", "salt & pepper");
// https://api.example.com/search?q=salt+%26+pepper
Using the wrong function
encodeURI("salt & pepper") returns salt%20&%20pepper: the spaces get encoded but the & doesn't, because encodeURI assumes it's handed a complete URL and preserves its delimiters. For a value, the right tool is always encodeURIComponent (or URLSearchParams).
Encoding twice
When an already-encoded value goes through another encoder, every % becomes %25: a b → a%20b → a%2520b. It usually happens when code encodes a parameter and then hands it to an HTTP library that encodes it again. The tell is a %25 followed by two hex digits in a URL.
Decoding twice
The reverse mistake has security consequences. If a server validates a path and then decodes it a second time, input like %252e%252e%252f passes validation (it contains no ../) and, after the second decode, becomes ../. A number of path traversal vulnerabilities have relied on exactly this. The rule is to decode once, at a well-defined point, and validate afterward.
The same operation across languages
Every language has one function for the form format (space as +) and another for RFC 3986 style (space as %20), and mixing them up is a common source of bugs:
| Language | Space as %20 (RFC 3986) | Space as + (forms) |
|---|---|---|
| JavaScript | encodeURIComponent(s) | new URLSearchParams({ q: s }) |
| Python | urllib.parse.quote(s, safe="") | urllib.parse.quote_plus(s), urlencode(dict) |
| PHP | rawurlencode($s) | urlencode($s) |
| Go | url.PathEscape(s) | url.QueryEscape(s) |
| Java | URI via its multi-argument constructors | URLEncoder.encode(s, UTF_8) |
| curl | — | --data-urlencode "q=salt & pepper" |
Two specific traps: Python's quote() doesn't encode / by default (its safe parameter defaults to "/"), so for a value you need safe="". And Java's URLEncoder.encode() produces form encoding, with +, despite what the name suggests — using it for path segments is a common mistake.
URL encoding and Base64
Sometimes binary data has to go into a URL (a token, a signature). Base64-encoding it and then percent-encoding the result works, but +, /, and = turn into %2B, %2F, and %3D, and the URL balloons. That's why the base64url variant exists: it uses - and _ and drops the padding, so the output is already URL-safe. It's covered in Base64 Explained: How It Works, Padding, and Variants.
Summary
Percent-encoding turns each character into its UTF-8 bytes written as %XX. What needs encoding depends on the part of the URL: inside a value, every delimiter is data and gets encoded. For values, use encodeURIComponent or URLSearchParams — never concatenation, never encodeURI. Encode once, decode once, and remember that + means a space only in a form-style query string.