URL encoding explained: percent-encoding, encodeURIComponent, and + vs %20
How percent-encoding works, reserved vs unreserved characters, encodeURI vs encodeURIComponent, and why double-encoding breaks URLs.
Published 2026-09-25
Why URLs need encoding at all
A URL's syntax uses specific characters — /, ?, #, &, =, : — to separate its parts: scheme, host, path, query string, fragment. RFC 3986, the URI specification, calls these reserved characters, because a parser relies on them to know where one part of the URL ends and the next begins. The problem: real data — a search query, a filename, a piece of user input — often contains those same characters as data, not as structure. If you drop a literal & into a query parameter's value, the URL parser reads it as the start of a new parameter, not as part of your value.
Percent-encoding solves this by giving every byte a way to appear as data without being misread as syntax: % followed by two hex digits representing that byte's value. Per RFC 3986 §2.1, %20 represents a space (0x20 in hex), %26 represents &, %2F represents /, and so on.
Sarah & Sons → Sarah%20%26%20Sons
Reserved vs unreserved characters
RFC 3986 §2.3 defines the unreserved set — the only characters that never need encoding because they carry no structural meaning anywhere in a URL:
A–Z a–z 0–9 - . _ ~
Everything else that isn't already part of the URL's intended structure should be percent-encoded when it appears as data. The reserved set (§2.2) — : / ? # [ ] @ as general delimiters, plus ! $ & ' ( ) * + , ; = as sub-delimiters — is reserved specifically because those characters have a defined structural meaning in at least some part of a URL. Whether a given reserved character needs encoding depends on where it sits: / is required as-is between path segments but must be encoded if it's literally part of a filename value; = and & are structural inside a query string but need encoding if they're literally part of a parameter's value.
encodeURI vs encodeURIComponent
JavaScript ships two built-in encoders that get confused for each other constantly, and the difference matters:
encodeURI()is for encoding a whole URL you're about to use as one. It leaves reserved delimiter characters (: / ? # @ ! $ & ' ( ) * + , ; =) alone (it does encode[and]), because encoding them would break the URL's own structure — it assumes what you're encoding already has its slashes and question marks in the right structural places.encodeURIComponent()is for encoding a single piece of a URL — a query parameter's value, a path segment — that will be inserted into a larger URL afterward. It encodes those same delimiters too (all except! ' ( ) *, which it leaves as is), because inside a single component, a literal&or=or/isn't structure, it's just data that happens to look like structure.
const query = "C++ & Rust?";
encodeURI(query);
// "C++%20&%20Rust?" — & and ? survive: wrong if this is a query VALUE
encodeURIComponent(query);
// "C%2B%2B%20%26%20Rust%3F" — correct for inserting as one parameter's value
The practical rule: if you're building a query string parameter-by-parameter, use encodeURIComponent() on each value. Using encodeURI() on a value that's going into a query parameter is a common source of bugs — an & in user-entered text silently becomes a second, unintended parameter instead of part of the first one's value.
Spaces: %20 vs +
This is a real inconsistency, not a bug in any one tool. RFC 3986, the generic URI specification, only recognizes %20 for a space — it says nothing about +. The +-for-space convention comes from a different, older specification: the application/x-www-form-urlencoded content type used by HTML forms, where + represents a space and a literal + in the data must itself be percent-encoded as %2B to avoid ambiguity.
That means the correct encoding of a space depends on context:
- In a URL path or most of a query string per RFC 3986:
%20. - In an HTML form submission body (
application/x-www-form-urlencoded) or a query string generated by that same convention:+.
Since many web frameworks build query strings using the form-encoded convention, + for space in a query string is common in practice and usually decodes correctly — but mixing the two conventions without knowing which one a given decoder expects is a frequent source of a literal + showing up in decoded text instead of a space, or vice versa. When in doubt, %20 is the RFC 3986-correct choice for anything outside an actual form body.
Double-encoding: the bug that looks like the tool is broken
Double-encoding happens when a value that's already percent-encoded gets run through an encoder a second time. % itself is not in the unreserved set, so encoding an already-encoded string encodes the % characters too:
original: Sarah & Sons
encoded once: Sarah%20%26%20Sons
encoded again (bug): Sarah%2520%2526%2520Sons
Decoding that double-encoded string once gives you back Sarah%20%26%20Sons — still encoded, not the original text — which looks like a decoding failure but is actually an encoding-side bug: something in the pipeline encoded a value that was already encoded. This commonly happens when a URL is built by concatenating a query string that a library already encoded, then the whole concatenated URL gets passed through encodeURI() or an HTTP client that encodes automatically, encoding it a second time.
The fix is architectural, not a smarter decoder: encode each raw value exactly once, at the point where you insert it into the URL, and never re-encode a string you're not sure is still raw. If you're debugging garbled query parameters, decoding the value twice with this hub's URL decode tool and getting readable text back is a strong signal that double-encoding is the culprit.
A practical checklist
- Encoding a full URL you already assembled correctly?
encodeURI(). Encoding one value that's going into a URL?encodeURIComponent(), via the URL encode tool or your language's equivalent. - Building a query string by hand? Encode each parameter value individually before concatenating — never encode the whole assembled string afterward.
- Seeing
%2520or similar in a decoded value? That's double-encoding — trace back to where a value got encoded twice, don't just decode twice and move on. - Space showing up as a literal
+? Check whether the source used form-urlencoded conventions (+) rather than RFC 3986's%20, and decode accordingly.