Base91 Encode / Decode

Convert text or binary bytes with basE91: 13 bits of input per 2 printable characters, which makes it the densest of the everyday printable encodings — about 23% overhead against Base64's 33%. Text is handled as UTF-8 bytes; a hexadecimal byte mode is there for binary data. Also: Base64 | Base32 / Base58 | Ascii85.
All processing happens locally in your browser — your data is never uploaded to any server.

Never heard of Base91? What is Base91 — explained in plain English (with a worked example, the alphabet and the traps).
How it works: The input is turned into bytes, then packed into 13 bit groups that become two characters each, with one 14 bit group wherever the short form would be ambiguous.
Input is:
How it works: Whitespace is ignored, and every character is looked up in the alphabet printed further down this page. An unknown character is reported with its position.
Output as:
Ctrl+Enter in the input runs the active mode.
Pick Encode or Decode and click its action button. The self-test runs 130 encode/decode round trips and reports the count here.
The alphabet, printed in full: Base91 has no standards body behind it, so the alphabet is not guaranteed across implementations. This page uses the alphabet of the reference implementation, and every value below was produced with exactly the 91 characters listed in the table.
Green cells are characters that can sit safely in HTML text or a URL query; the orange ones must be escaped or percent-encoded (see the notes below the table).
Measured sizes on the same 92 byte input:
Character counts measured here on the string . Base91 wins because 912 = 8281 covers the 8192 values of 13 bits, while Base64 spends 4 characters on only 24 bits.
What is Base91?

Base91 is a way of writing arbitrary bytes using only printable ASCII characters, the same job Base64 does, but it fits more bits into each character. Instead of chopping the input into 6 bit groups and printing one character per group, it collects bits until it has at least 13 of them, then prints two characters that together name the value of those 13 (or 14) bits.

Two characters out of an alphabet of 91 can name 91 × 91 = 8281 different values, and 213 = 8192, so two characters always have room for 13 bits — with 89 values to spare. That is the whole trick: 13 bits per 2 characters is 6.5 bits per character, against Base64's 6, and it buys about 23% overhead instead of 33%.

The 30 second version: test becomes fPNKd, and Hello, World! becomes >OwJh>}AQ;r@@Y?F. Decoding is the exact reverse, and it needs the same 91 character alphabet — a different alphabet gives different output from the same bytes.
Where the 13 and the 14 bits come from

91 × 91 = 8281 values can be named by a pair of characters, which is more than 213 = 8192 but less than 214 = 16384. So a pair of characters can always carry 13 bits, and sometimes 14. The encoder tries 13 bits first: if the 13 bit value is greater than 88 it prints that value as two characters; otherwise it takes a 14th bit and prints the resulting 14 bit value instead.

The boundary at 88 is not arbitrary. The largest value two characters can name is 90 + 91 × 90 = 8280, and 8192 + 89 = 8281 is already past it. So no 14 bit group can ever have a low 13 bit window of 89 or more, which means the decoder can tell the two cases apart by looking at nothing but the value itself:

What the decoder seesBits it consumedWhy
value & 8191 is 89 to 819113 bitsValues of 89 and up can only come from the 13 bit branch, because a 14 bit group with a low 13 bit window that high would have to be at least 8281 — which no pair of characters can name.
value & 8191 is 0 to 8814 bitsThe encoder only falls back to 14 bits when the 13 bit reading was 88 or less, so this branch is the fallback and nothing else.

Whatever is left over at the end — anything from 1 to 13 bits — is printed as one character, plus a second character when the leftover really needed more than 7 bits. That is why the encoded length grows in steps of two characters and sometimes ends on an odd count.

The alphabet, and the three characters it drops

Printable ASCII holds 94 characters, from ! (0x21) to ~ (0x7E). Base91 uses 91 of them, and the three it leaves out are the hyphen -, the apostrophe ' and the backslash \. Letters, digits and almost all punctuation are in the alphabet exactly once, in the order shown in the table above; the 92nd through 94th printable characters are simply not addressable.

Because the alphabet is a fixed order, the character for a value is a lookup and not a formula. Two implementations that order the alphabet differently will both be "Base91" and will not be able to read each other's output, which is the single most common source of confusing results. This page prints its alphabet so the output can always be traced back to it.

Base91 next to the other printable encodings
EncodingBits per characterSize against the inputNotes
Base16 (hex)4200%Two characters per byte, trivially readable, no alphabet to look up.
Base325160%Case-insensitive and safe for human dictation, but the most wasteful of the group.
Base58~5.86~137%Chosen for addresses, not for density: it avoids look-alike characters.
Base646133%The default everywhere: 4 characters per 3 bytes, with the 6 unused bits of the last character wasted.
Ascii85 / Base856.4125%5 characters per 4 bytes, no lookup table needed, but it divides by 85 — see the Ascii85 converter.
Base91~6.5~123%The densest of these: 13 bits per 2 characters, at the cost of a 14 bit group here and there.

The percentages are the asymptotic cost of each encoding, rounded; the measured character counts for one real 92 byte input are in the table further up this page. Base91 is close to the best any encoding can do with printable ASCII: 7 bits per character would mean using 128 distinct characters, which the printable range cannot provide.

What Base91 is not
  • Not a standard. There is no RFC. The reference implementation is basE91 0.6.0 by Joachim Henke, released under the BSD 3 clause licence; "Base91" is also used for alphabets that are not this one.
  • Not encryption. Every character is a straight re-shaping of the input bits, and anyone with the alphabet can reverse it in one pass.
  • Not URL safe. The alphabet contains #, %, &, /, ? and ", all of which change meaning in a URL, so Base91 output has to be percent-encoded before it goes in a query string.
  • Not HTML safe. It also contains <, >, & and ", so a Base91 string placed into markup has to be escaped with the HTML entity converter.
  • Not fixed length. The output length depends on the bit pattern, not only on the byte count, so two inputs of the same size can print a character apart.
Traps that actually bite
  • The alphabet in this page is the reference one. If another tool produces different output for the same input, compare its alphabet before assuming either side is broken. The 91 characters must appear in the same order.
  • A single wrong character can change the whole tail. Group boundaries are decided by the values that came before, so an edit early in the string shifts everything after it, and the tail is decoded with the leftover bits of the last group only.
  • Odd output length is normal. It does not mean the string was cut short; the final group prints one character when its leftover value fits in 7 bits or fewer.
  • A published example can carry a newline you did not type. A shell command that adds a line ending is the usual reason a well known Base91 value is one character longer than the same text produces here: echo "Hello, World!" hands the encoder 14 bytes and prints >OwJh>}AQ;r@@Y?FF, while the 13 bytes of the text alone print >OwJh>}AQ;r@@Y?F. Both are correct, for different byte counts, and the count under the output box says which one you are holding.
  • Decoding binary data as text will not work. If the decoded bytes are not valid UTF-8 there is no sensible text to show, which is what the hexadecimal output mode is for.
Quick answers

“Why does this page show a value that is one character different from another Base91 tool?”

Almost always a different alphabet. The second most common cause is a tool that only uses the 13 bit branch and never the 14 bit fallback; that variant is smaller in code but produces different output from the reference algorithm used here.

There is also one exact case that can be settled by counting bytes, and it explains the most widely quoted example of all: a trailing newline. Most published vectors were produced by a shell command, and the shell adds the line ending that finishes the line. The 13 bytes of Hello, World! print 16 characters here, while the 14 bytes of Hello, World! followed by a newline print the 17 character >OwJh>}AQ;r@@Y?FF. Paste either value into Decode: this page reports the byte count, so the two readings can be told apart instead of argued about. The Sample button loads whichever of the two belongs to the reader you are using, and it follows the Encode / Decode switch and the Input is / Output as selector: the text reader gets the 16 character value, the hexadecimal reader gets the 17 character one with its trailing 0A byte in plain sight, and the encoder gets the text to encode rather than a Base91 string.

“Is Base91 smaller than Base64?”

Yes, by about 8% of the encoded size: 13 bits per 2 characters against 6 bits per character. On 92 bytes of text, Base64 prints 124 characters and Base91 prints 114.

“Can I decode this with a Base64 decoder, or the other way round?”

No. The alphabets overlap in the letters and digits, so a Base64 decoder will accept a Base91 string as "valid" characters and produce silent garbage. The two agree on nothing except the shape of the scheme.

“Does the output end in padding like Base64?”

No. Base91 has no padding character at all: the leftover bits at the end are printed as the final group, so nothing has to be added or stripped.

“How do I know the result on this page is right?”

Press Self-test 0 to 64 bytes. It encodes and decodes 130 byte patterns covering every length where the 13 and 14 bit split changes, and compares the bytes that come back with the bytes that went in. The vectors on the Sample button are additionally cross-checked against a second implementation written for this site's test tooling.

Standards and credit

This page implements the algorithm of the basE91 reference implementation (version 0.6.0) by Joachim Henke, distributed under the BSD 3 clause licence, with the alphabet printed above. Decoding accepts and ignores whitespace, as the reference tool does.

The page text was written for this site; the algorithm was not copied from anywhere. The Sample button values are the one example that ships with the reference package (marked "reference example") plus values produced by a second, independent implementation kept in this site's test tooling (marked "cross-check").