URL encoding explained: percent-encoding, encodeURIComponent, and + vs %20

How percent-encoding works, reserved vs unreserved characters, encodeURI vs encodeURIComponent, and why double-encoding breaks URLs.

Published 2026-09-25

Why URLs need encoding at all

A URL's syntax uses specific characters — /, ?, #, &, =, : — to separate its parts: scheme, host, path, query string, fragment. RFC 3986, the URI specification, calls these reserved characters, because a parser relies on them to know where one part of the URL ends and the next begins. The problem: real data — a search query, a filename, a piece of user input — often contains those same characters as data, not as structure. If you drop a literal & into a query parameter's value, the URL parser reads it as the start of a new parameter, not as part of your value.

Percent-encoding solves this by giving every byte a way to appear as data without being misread as syntax: % followed by two hex digits representing that byte's value. Per RFC 3986 §2.1, %20 represents a space (0x20 in hex), %26 represents &, %2F represents /, and so on.

Sarah & Sons  →  Sarah%20%26%20Sons

Reserved vs unreserved characters

RFC 3986 §2.3 defines the unreserved set — the only characters that never need encoding because they carry no structural meaning anywhere in a URL:

A–Z  a–z  0–9  -  .  _  ~

Everything else that isn't already part of the URL's intended structure should be percent-encoded when it appears as data. The reserved set (§2.2) — : / ? # [ ] @ as general delimiters, plus ! $ & ' ( ) * + , ; = as sub-delimiters — is reserved specifically because those characters have a defined structural meaning in at least some part of a URL. Whether a given reserved character needs encoding depends on where it sits: / is required as-is between path segments but must be encoded if it's literally part of a filename value; = and & are structural inside a query string but need encoding if they're literally part of a parameter's value.

encodeURI vs encodeURIComponent

JavaScript ships two built-in encoders that get confused for each other constantly, and the difference matters:

  • encodeURI() is for encoding a whole URL you're about to use as one. It leaves reserved delimiter characters (: / ? # @ ! $ & ' ( ) * + , ; =) alone (it does encode [ and ]), because encoding them would break the URL's own structure — it assumes what you're encoding already has its slashes and question marks in the right structural places.
  • encodeURIComponent() is for encoding a single piece of a URL — a query parameter's value, a path segment — that will be inserted into a larger URL afterward. It encodes those same delimiters too (all except ! ' ( ) *, which it leaves as is), because inside a single component, a literal & or = or / isn't structure, it's just data that happens to look like structure.
const query = "C++ & Rust?";

encodeURI(query);
// "C++%20&%20Rust?"   — & and ? survive: wrong if this is a query VALUE

encodeURIComponent(query);
// "C%2B%2B%20%26%20Rust%3F"   — correct for inserting as one parameter's value

The practical rule: if you're building a query string parameter-by-parameter, use encodeURIComponent() on each value. Using encodeURI() on a value that's going into a query parameter is a common source of bugs — an & in user-entered text silently becomes a second, unintended parameter instead of part of the first one's value.

Spaces: %20 vs +

This is a real inconsistency, not a bug in any one tool. RFC 3986, the generic URI specification, only recognizes %20 for a space — it says nothing about +. The +-for-space convention comes from a different, older specification: the application/x-www-form-urlencoded content type used by HTML forms, where + represents a space and a literal + in the data must itself be percent-encoded as %2B to avoid ambiguity.

That means the correct encoding of a space depends on context:

  • In a URL path or most of a query string per RFC 3986: %20.
  • In an HTML form submission body (application/x-www-form-urlencoded) or a query string generated by that same convention: +.

Since many web frameworks build query strings using the form-encoded convention, + for space in a query string is common in practice and usually decodes correctly — but mixing the two conventions without knowing which one a given decoder expects is a frequent source of a literal + showing up in decoded text instead of a space, or vice versa. When in doubt, %20 is the RFC 3986-correct choice for anything outside an actual form body.

Double-encoding: the bug that looks like the tool is broken

Double-encoding happens when a value that's already percent-encoded gets run through an encoder a second time. % itself is not in the unreserved set, so encoding an already-encoded string encodes the % characters too:

original:            Sarah & Sons
encoded once:         Sarah%20%26%20Sons
encoded again (bug):  Sarah%2520%2526%2520Sons

Decoding that double-encoded string once gives you back Sarah%20%26%20Sons — still encoded, not the original text — which looks like a decoding failure but is actually an encoding-side bug: something in the pipeline encoded a value that was already encoded. This commonly happens when a URL is built by concatenating a query string that a library already encoded, then the whole concatenated URL gets passed through encodeURI() or an HTTP client that encodes automatically, encoding it a second time.

The fix is architectural, not a smarter decoder: encode each raw value exactly once, at the point where you insert it into the URL, and never re-encode a string you're not sure is still raw. If you're debugging garbled query parameters, decoding the value twice with this hub's URL decode tool and getting readable text back is a strong signal that double-encoding is the culprit.

A practical checklist

  1. Encoding a full URL you already assembled correctly? encodeURI(). Encoding one value that's going into a URL? encodeURIComponent(), via the URL encode tool or your language's equivalent.
  2. Building a query string by hand? Encode each parameter value individually before concatenating — never encode the whole assembled string afterward.
  3. Seeing %2520 or similar in a decoded value? That's double-encoding — trace back to where a value got encoded twice, don't just decode twice and move on.
  4. Space showing up as a literal +? Check whether the source used form-urlencoded conventions (+) rather than RFC 3986's %20, and decode accordingly.

Tools mentioned in this guide