What is Base64?
Base64 rewrites bytes using sixty-four readable characters so that data survives a channel built for text: an email body, a JSON string, an XML attribute, a URL. Every three bytes become four characters, which is where the name comes from — the alphabet holds sixty-four symbols and each one carries exactly six bits.
It is a transport encoding, not a security measure and not a compression step. The output is roughly a third larger than the input, and anybody who has the string can turn it straight back. If you want to know whether data was altered, that is a hash (Hash Generator); if you want it to stay secret, that is encryption. Base64 does neither.
The one number to remember: three input bytes become four output characters. Everything else on this page follows from it. Four characters carry twenty-four bits, and twenty-four bits is exactly three bytes, so no data is ever lost or invented — except at the very end, where the input may not fill the last group of three.
The five alphabets this page can use
The first two rows are standards and the last three are conventions that grew up around Unix. All five carry the same six bits per character and differ only in which sixty-four characters do the carrying, so the same bytes become five different strings. Nothing inside a string says which alphabet produced it, which is why the alphabet selector defaults to standard and why the URL-safe checkbox is a shortcut for the second row.
How the page picks a candidate: the selector is the default for both directions. When decoding, the panel on the right also reads the string itself — - and _ can only come from the URL-safe alphabet, + or / can only come from the standard one, and a / or . near the start points at the two Unix alphabets. Every alphabet that can read the current input is listed there, with the reason each one is, or is not, a candidate.
Line wrapping: 76 and 64 characters
MIME asked for Base64 that fits in an email without any line growing unreasonably long, and settled on 76 characters per line separated by CRLF. PEM files, which wrap certificates and keys, use 64. The line breaks are not part of the data: a decoder removes them and joins the lines back together before doing anything else, so a wrapped string and an unwrapped one decode to exactly the same bytes. Every line except the last holds a whole number of four-character groups, which is what makes joining them back safe.
This page leaves the output on one line unless you pick a width, and the decoder accepts input with or without line breaks. The breaks this page writes are plain line feeds, which is what a text box can hold; MIME mail bodies traditionally use a carriage return before each line feed, and every decoder in use accepts both.
Padding, and why it can be dropped
When the input is not a multiple of three bytes the last group of four characters is incomplete, and one or two = characters fill it out. They carry no data at all: they only say how many of the four characters are real. A decoder can work that out from the length instead, which is why URL-safe Base64 traditionally leaves the padding off — a = is awkward in a URL or a file name, and some systems percent-encode it.
The checkbox in the options row keeps the padding when the URL-safe alphabet is selected. Nothing else changes; the characters that carry data are identical either way.
Common misunderstandings
- "Base64 hides the data." It does not. It is a substitution with no key, and the whole point is that any reader can undo it. Wrapping a password or a token in Base64 changes nothing about who can read it.
- "Base64 compresses." The opposite: every three input bytes become four output characters, so the text is a third larger than the data it carries.
- "The line breaks are part of the value." They are presentation. Adding or removing them never changes the bytes that come back out.
- "URL-safe means a different encoding." It is the same encoding with two characters swapped:
+ becomes - and / becomes _, so the string can sit in a URL path or a query value without escaping.
- "A string that decodes must be valid Base64." Any string of alphabet characters is valid Base64, because every combination of six-bit values is a valid sequence of bytes. A string only turns out to be wrong at the next step, when those bytes are checked against the format they were supposed to be.
- "Padding tells you which alphabet was used." It does not, and in most cases the characters do not either: the standard and URL-safe alphabets agree on sixty-two of their sixty-four symbols, so a short string often reads the same way in both.
Questions that come up
Why does a Base64 string sometimes end with one = and sometimes two? One means the last group carries three real characters, which is two input bytes; two means it carries two, which is one input byte. No = means the input length was a multiple of three.
Can Base64 represent arbitrary binary files? Yes, and that is its main use. This page treats the input as UTF-8 text; the dedicated File to Base64 tool reads raw bytes when you want to encode something that is not text.
Which alphabet should I use in an API? Whatever the receiving end documents. If it says URL-safe, use that; if it says nothing and the value travels inside a URL, URL-safe with the padding dropped is the safe choice.
Is the output case-sensitive? Completely. a and A are different values, which is why Base64 survives a mistyped character far less gracefully than Base32 or Crockford Base32 do.