Base64 explained: what it does, what it doesn't, and where it breaks

What Base64 actually is, standard vs URL-safe alphabets, padding rules, and why it is not encryption.

Published 2026-09-25

What Base64 actually does

Base64 takes arbitrary bytes — text, an image, a binary key, anything — and re-represents them using only 64 printable ASCII characters: A–Z, a–z, 0–9, plus two more symbols. It exists because a lot of systems (email, older protocols, certain text fields, JSON strings) only reliably carry printable text, not raw binary. Base64 is defined by RFC 4648, and the mechanism is simple: it groups input bytes three at a time (24 bits), splits that into four 6-bit chunks, and maps each 6-bit chunk to one of 64 characters.

That 3-bytes-in, 4-characters-out ratio is why Base64 output is always about 33% larger than the original — three bytes of input never take fewer than four characters of output.

It is encoding, not encryption

This is the single most common misunderstanding. Encoding is a reversible, keyless transformation — anyone who has the Base64 string can decode it back to the original bytes with nothing else required. There is no secret, no password, no key. If you Base64-encode a value and someone intercepts it, they can read the original content in one step with any decoder, including this hub's Base64 decode tool.

Encryption is different: it requires a key, and without that key the ciphertext is (ideally) computationally infeasible to reverse. Seeing something that looks like Base64 — letters, digits, maybe a trailing = — tells you nothing about whether the underlying data is sensitive plaintext or encrypted ciphertext; Base64 only tells you how the bytes are represented, not how they're protected. A password, API key, or credit card number that's "obfuscated" by Base64 alone is still fully readable to anyone who copies the string.

Standard vs URL-safe alphabets

RFC 4648 defines two alphabets that agree on the first 62 characters (A–Z, a–z, 0–9) and differ on the last two:

| Position | Standard (§4) | URL/filename-safe (§5) | |---|---|---| | 62nd character | + | - | | 63rd character | / | _ | | Padding character | = | = (often omitted) |

The reason the URL-safe variant exists: + and / both have special meaning inside a URL (+ can mean a space in query strings, / is a path separator), so putting standard Base64 straight into a URL or filename can break parsing or require percent-encoding. Swapping in - and _, which have no special meaning in URLs, avoids that entirely — this is stated directly in RFC 4648 §5.

If you don't know which alphabet a given string uses, look for + or / (standard) versus - or _ (URL-safe) — a string can't validly contain both alphabets' distinguishing characters at once.

Padding, and why it sometimes goes missing

Because groups are 3 bytes → 4 characters, an input length that isn't a multiple of 3 bytes leaves a partial group at the end. Base64 pads that partial group with = characters so every encoded block is a multiple of 4 characters long — one = if 2 bytes were left over, two = if 1 byte was left over, none if the input divided evenly by 3.

Padding is optional in a lot of real-world usage, and its presence or absence is a common source of "invalid Base64" errors:

  • The URL-safe variant is very often used without padding, since the decoder can usually infer how many characters are missing from the string's length, and dropping = avoids needing to percent-encode it in a URL.
  • Standard Base64 in strict contexts (MIME email, many library defaults) expects padding to be present.
  • Concatenating multiple Base64 strings without accounting for padding, or truncating a string that ends in =, produces "invalid length" errors on decode.

A decoder that's lenient about this — accepting both alphabets and adding missing padding automatically — saves a lot of manual fiddling, which is what this hub's Base64 decode tool does.

Where Base64 shows up: JWTs

JSON Web Tokens are the most common place a developer runs into Base64 directly. A JWT has three dot-separated parts — header, payload, signature — and per RFC 7519, each part is base64url-encoded (the URL-safe alphabet), and per the JWS spec it references, without padding. That's why pasting a JWT payload segment into a strict, padding-required Base64 decoder often fails with a length error, while a URL-safe-aware decoder handles it directly.

To read a JWT's payload by hand: split the token on ., take the middle segment, and Base64-URL-decode it — you'll get the claims as plain JSON. The signature segment, by contrast, is not meant to be human-readable; it's the raw bytes of an HMAC or signature, and decoding it just gives you binary data, not text.

Common bugs and what they mean

  • "Invalid character" errors usually mean the string mixes the two alphabets, or has been through something that added stray characters — line breaks inserted by an email client or a fixed-width terminal are a frequent culprit.
  • Wrong output length / garbled text after decoding almost always means a character was dropped or altered somewhere upstream — because 4 characters map to exactly 3 bytes, losing even one character shifts every subsequent group and corrupts the rest of the decode.
  • A UTF-8 decode error after Base64-decoding doesn't mean the Base64 was invalid — it means the original bytes weren't text at all. Base64 is used for images, PDFs, encryption keys and other binary formats just as often as for text, and decoding those correctly gives you binary bytes that aren't valid UTF-8. That's expected, not a bug in the tool.

Base64 is a small, well-specified piece of plumbing — the trouble it causes almost always comes from assuming it does more (or less) than represent bytes as text.

Tools mentioned in this guide