Toolbit
Guides

URL Encoding Explained: Percent-Encoding, %20, and +

Which characters need encoding in each part of a URL per RFC 3986, why a space is sometimes %20 and sometimes +, and the classic bugs: concatenation, double encoding, and double decoding.

By Javier VallejoPublished 6 min read

Why URLs need encoding

A URL is text with structure: some characters are data, and others are delimiters that tell the browser and the server where one part ends and the next begins. Trouble starts when a piece of data contains one of those delimiters — an & in a product name, a / in a search term — or a character that simply isn't allowed in a URL, like a space or an accented letter. Percent-encoding handles both cases with a single mechanism. This guide covers what the standard says, what changes from one part of a URL to another, and why most bugs come from encoding too little, encoding too much, or reaching for the wrong function. To encode or decode a specific value, the URL encoder does it right in your browser.

Anatomy of a URL

RFC 3986 defines the generic syntax for URIs. Taking https://ana@shop.example.com:8443/products/café?color=red&size=M#reviews as an example:

PartValue in the exampleDelimiter
Schemehttpsends with :
User infoanaends with @
Hostshop.example.comstarts after //
Port8443starts at :
Path/products/cafésegments separated by /
Querycolor=red&size=Mstarts at ?
Fragmentreviewsstarts at #

As written, that URL isn't valid per the RFC: café contains a non-ASCII character. What actually gets sent is /products/caf%C3%A9. Browsers show the readable version in the address bar, but the encoded one is what travels over the network.

Reserved and unreserved characters

RFC 3986 splits the allowed characters into groups:

GroupCharactersRule
UnreservedA–Z a–z 0–9 - . _ ~Never need encoding
General delimiters: / ? # [ ] @Separate the parts of the URL
Sub-delimiters! $ & ' ( ) * + , ; =Separate data within a part (e.g., & and = in the query)

The practical rule: a reserved character used as data, rather than as a delimiter, must be encoded. Anything outside these groups — spaces, accented letters, quotes, <, >, % as data — is always encoded.

How a character gets encoded

First, the character is converted to bytes with UTF-8, which is what RFC 3986 recommends for any new text. Then each byte is written as % followed by two hex digits, uppercase per the standard's recommendation:

CharacterUTF-8 bytesEncoded
space20%20
&26%26
éC3 A9%C3%A9
ñC3 B1%C3%B1
€E2 82 AC%E2%82%AC
🙂F0 9F 99 82%F0%9F%99%82

So café becomes caf%C3%A9. The host is the exception: domain names with non-ASCII characters use Punycode instead of percent-encoding, so münchen.de resolves as xn--mnchen-3ya.de.

Each part has its own rules

Path. / separates segments, so if a segment contains a slash as data (a file named report 1/2.pdf), it has to be encoded as %2F. A + in the path is a literal plus sign, not a space. Spaces go in as %20.

Query string. By convention, parameters are written as name=value separated by &, so inside a name or value you must encode &, =, #, and +. The RFC allows unencoded / and ? in the query, but encoding them never breaks anything. This is where the great space confusion lives: HTML forms use the application/x-www-form-urlencoded format, where a space is written as +, while RFC 3986 writes it as %20. In the query, almost every server accepts both — which is exactly why a + that's meant as data has to go in as %2B, or it'll be read as a space.

Fragment. Everything after # never reaches the server: the browser uses it locally, to jump to a section of the page or for single-page app routing. It has its own encoding rules, but sensitive data in the fragment doesn't land in server logs — and the server can't read it either.

Classic mistakes

Concatenating instead of encoding

const url = "https://api.example.com/search?q=" + "salt & pepper";
// https://api.example.com/search?q=salt & pepper
// The server receives q = "salt " plus an empty parameter named " pepper".

The robust fix is to stop building the query by hand:

const url = new URL("https://api.example.com/search");
url.searchParams.set("q", "salt & pepper");
// https://api.example.com/search?q=salt+%26+pepper

Using the wrong function

encodeURI("salt & pepper") returns salt%20&%20pepper: the spaces get encoded but the & doesn't, because encodeURI assumes it's handed a complete URL and preserves its delimiters. For a value, the right tool is always encodeURIComponent (or URLSearchParams).

Encoding twice

When an already-encoded value goes through another encoder, every % becomes %25: a b → a%20b → a%2520b. It usually happens when code encodes a parameter and then hands it to an HTTP library that encodes it again. The tell is a %25 followed by two hex digits in a URL.

Decoding twice

The reverse mistake has security consequences. If a server validates a path and then decodes it a second time, input like %252e%252e%252f passes validation (it contains no ../) and, after the second decode, becomes ../. A number of path traversal vulnerabilities have relied on exactly this. The rule is to decode once, at a well-defined point, and validate afterward.

The same operation across languages

Every language has one function for the form format (space as +) and another for RFC 3986 style (space as %20), and mixing them up is a common source of bugs:

LanguageSpace as %20 (RFC 3986)Space as + (forms)
JavaScriptencodeURIComponent(s)new URLSearchParams({ q: s })
Pythonurllib.parse.quote(s, safe="")urllib.parse.quote_plus(s), urlencode(dict)
PHPrawurlencode($s)urlencode($s)
Gourl.PathEscape(s)url.QueryEscape(s)
JavaURI via its multi-argument constructorsURLEncoder.encode(s, UTF_8)
curl—--data-urlencode "q=salt & pepper"

Two specific traps: Python's quote() doesn't encode / by default (its safe parameter defaults to "/"), so for a value you need safe="". And Java's URLEncoder.encode() produces form encoding, with +, despite what the name suggests — using it for path segments is a common mistake.

URL encoding and Base64

Sometimes binary data has to go into a URL (a token, a signature). Base64-encoding it and then percent-encoding the result works, but +, /, and = turn into %2B, %2F, and %3D, and the URL balloons. That's why the base64url variant exists: it uses - and _ and drops the padding, so the output is already URL-safe. It's covered in Base64 Explained: How It Works, Padding, and Variants.

Summary

Percent-encoding turns each character into its UTF-8 bytes written as %XX. What needs encoding depends on the part of the URL: inside a value, every delimiter is data and gets encoded. For values, use encodeURIComponent or URLSearchParams — never concatenation, never encodeURI. Encode once, decode once, and remember that + means a space only in a form-style query string.

Tools used in this guide

Related guides