Percent Encoding Inspector

Percent encoding is not one rule but several, and they disagree on about a dozen characters. This page runs the same text through encodeURIComponent, encodeURI and the application/x-www-form-urlencoded serialisation side by side, marks the characters where they disagree, prints a row per character with its UTF-8 bytes and its RFC 3986 class, and says how many times a string has been encoded. The decode side reads a string with a strict or a lenient rule and tells you which escape is broken and where, instead of handing back a quietly wrong result.
All processing happens locally in your browser — your data is never uploaded to any server.

Not sure what percent encoding is? Percent encoding explained in plain English — what it is, why a space becomes %20 or +, and what it means when a URL shows %25.
What this does:
How to read the escapes:
encodeURIComponent
encodeURI
form (application/x-www-form-urlencoded)
Ctrl+Enter in the input runs the active mode.
Pick a mode and click its action button. Encode & inspect shows the three encodings of your text plus a row per character; Decode reads the escapes back and counts the layers. The Self-test button re-runs every documented sample, every printable ASCII character and the layer chains recorded in this page's data file, and reports the count here.
Where the encoders disagree: only the characters below are treated differently by the three methods. Everything else goes through all three of them unchanged, or is escaped identically by all three.
Every character, one row each: the UTF-8 bytes a character turns into, the class RFC 3986 puts it in, and the escape each context writes. The four context columns are the WHATWG URL percent-encode sets, which are what a browser uses when it builds a path, a query or a fragment; a shaded row is a character the three encoders treat differently.
The exact set this page escapes for each part of a URL: printed from the encoder itself, so this table and the code that fills the columns above cannot drift apart. Every set also escapes the C0 control characters (U+0000 to U+001F) and every code point above U+007E, which is why non-ASCII text is always escaped in every context.
These sets are derived from the WHATWG percent-encode sets and differ from them in two deliberate places, both of which make hand assembly safe: a hash is escaped inside a query, because a raw one would start the fragment; and the component set is pinned to exactly the characters encodeURIComponent escapes, apostrophe and all, while the form set is pinned to what URLSearchParams escapes. The two columns in the table above are therefore the platform's own behaviour rather than a reimplementation of it. That list also explains a result people find surprising: a percent sign is escaped by encodeURIComponent and by encodeURI, but it is in none of the WHATWG context sets, so the path, query and fragment columns of the per-character table leave it alone. A raw percent sign in a URL you assembled by hand is still a risk, because the standard expects the input to have been validated first: %41 inside a query is a single A to plenty of server frameworks, and this page's own strict reading treats a lone % as a broken escape. Escaping it yourself as %25 is the safer habit, and the per-character table shows both spellings side by side.
How many times this string has been encoded: each round of decoding is one layer. A value that has been encoded twice starts with %25 where a single encoding would leave %20, and a value that has been encoded three times starts with %2525.
The characters a URL may carry as they are: RFC 3986 splits the ASCII range into three groups, and the escape rules only make sense against that split.
The reserved characters are not forbidden, they are taken: they mean something to the URL parser. The rule that follows from the list is short — if a character is reserved and you mean it as data, escape it; if it is unreserved, you never need to.
The three methods, and what each one leaves alone:
The lists are the exact sets each method leaves unescaped in the ASCII range; the Self-test checks every printable ASCII character against them, so the table cannot drift away from the encoders it describes.
How large an input this page accepts, and why: percent encoding is a single pass over the text, so unlike the radix converters on this site the cost grows in a straight line rather than as a square: there is no length at which the arithmetic itself becomes the problem. The caps below exist because the per-character table is drawn in the document, and a table with a million rows in it is what actually stops a browser.
What is percent encoding?

Percent encoding — also called URL encoding, and in older documents escaping — is the rule that lets a URL carry characters it is not allowed to contain. A percent escape is three ASCII characters: a percent sign and two hexadecimal digits, standing for one byte. %20 is the byte 0x20, which is a space. %2F is the byte 0x2F, which is a slash. Nothing else is involved: the escape names a byte, not a character.

Because it names bytes, the text has to be turned into bytes first, and the byte encoding is UTF-8. That is where the long escapes come from: one Chinese character is three UTF-8 bytes and therefore three escapes, and one emoji is four bytes and twelve ASCII characters.

The 30 second version: a space is %20 as a URL value but + in form data; # is %23; a literal percent sign is %25. The three encoders on this page agree on the letters, the digits and -_.~, and disagree on about a dozen punctuation characters — which is exactly why encodeURIComponent and encodeURI produce different strings for the same text, and why the wrong one of the two can split a value into two query parameters.
Why a URL needs it at all

A URI is defined as a sequence of ASCII characters (RFC 3986 section 2.1), and inside that ASCII range some characters are reserved: they are the punctuation that gives the URL its structure. A colon and two slashes separate the scheme, an at sign separates the user information from the host, a question mark starts the query, a hash starts the fragment, an ampersand separates query parameters, and a percent sign introduces an escape.

That creates one problem and one loophole. The problem: a value that itself contains one of those characters — a search box containing "a&b", a file name containing a slash, a password containing a colon — would otherwise change the structure of the URL it travels in. The loophole: the percent sign itself, so the encoder can simply extend its own alphabet. % is written as %25, and the two hexadecimal digits after it are ordinary unreserved characters.

Worked examples

These are the exact outputs of the three encoders this page runs, on inputs chosen because they behave differently from each other. The page recomputes them in front of you, and its Self-test re-checks every row against the platform's own implementations.

The three methods, and when each is the right one

The short version that saves most of the bugs: anything going into a single parameter value goes through encodeURIComponent (or through a form serialiser, which does the same job with a different safe set), and a complete URL that is already assembled goes through encodeURI if it needs escaping at all. Reaching for the wrong one of those two is the single most common percent encoding mistake in web code.

The characters a URL may carry, in full
Two plus signs, two meanings

A space is the one character whose encoding depends on where the text is going, and the plus sign is the reason.

The trap sits in the middle of that table. In form data a plus sign means a space, so a value that genuinely contains a plus sign — a phone number written as +44 20 7946 0958, an arithmetic expression, an email address with a plus tag — has to travel as %2B. If it travels as a literal plus sign, the receiver reads it back as a space, and +44 arrives as " 44".

Encoded twice, and how to recognise it

A double encoding happens when a string is already escaped and is then escaped again. The percent sign of the first escape is a character like any other, so it becomes %25, and the layer underneath survives as visible text.

Recognising it is a matter of counting: decode the string once and look at the result. If it still contains escapes that look deliberate — a valid %XX or a form plus sign where a space is expected — there is another layer. If it contains %25 followed by two hexadecimal digits, the next layer is certainly an escape.

Two practical notes. First, decoding twice is not harmless: it is a well known way to get a security filter and the code behind it to see different strings, so a filter that runs before a second decode is filtering the wrong text. Second, the layer count here is a property of the string, not a promise about intent: a legitimate value can be a single layer deep and a broken one can be two.

The mistakes this page is built to catch
  • A value encoded with encodeURI. The separators survive, so a value containing & or = splits into several parameters on the other side. Compare the two columns at the top of this page on the input a&b=c: the component column escapes both characters, the encodeURI column keeps them.
  • A literal plus sign sent as form data. It arrives as a space. Use %2B.
  • A value escaped twice "to be safe". It is not safer. The receiver decodes once and gets a string that still contains percent escapes, which is the same failure as a wrong character set, one layer down.
  • Decoding what was never encoded. A literal percent sign in text — 100% sure — is not an escape. In a strict reading it is a broken one, at the position this page reports; in a lenient reading it survives, which is why lenient is the default here and not the other way round.
  • Assuming the escape has to be upper case. Decoding treats %2f and %2F as the same byte, and both are valid. Upper case is only the canonical spelling that RFC 3986 recommends, and the one other tools will produce.
  • Assuming percent encoding makes text safe. It makes text unambiguous inside a URL. It does not stop injection, it is not encryption, and it is not a sanitiser: the escaping has to be the last step before assembly, for the context the text is going into.
Quick answers

“Is a plus sign a space in a URL?”

Only in form data, and only in the query part of a form submission. In a path, in a fragment and in the query of a plain link, a plus sign is a literal plus sign. That is why this page has a switch for it instead of guessing.

“Why does my link show %25?”

Because something encoded a percent sign. Either a literal percent sign went through an encoder — 100% sure becomes 100%25%20sure — or a string that already contained an escape was encoded a second time. Decode the string once and look at what comes back: if the result still contains %XX, there is another layer.

“Does it matter whether a slash inside a query value is %2F or /?”

To the URL parser, no: both reach the server as a slash, and the WHATWG query set leaves a slash alone. To interoperability, sometimes yes: plenty of server frameworks run their own decoding over a raw query string, and a few of them disagree about characters that were never escaped. When a value is assembled by hand rather than by a serialiser, escaping the reserved characters explicitly is the safer habit.

“Why does the form column escape a tilde that the other two leave alone?”

Because the form encoder's safe set is narrower: only letters, digits and -_.*. The WHATWG form encoder escapes ~ ! ' ( ) even though RFC 3986 lists ~ as unreserved. Both spellings decode to the same value; the difference only matters when comparing strings byte for byte.

“Can the same characters appear unescaped in a browser address bar?”

Yes, and that is the address bar, not the URL. Browsers display a readable form of a URL and encode it when they send it, so 中文 in a link becomes %E4%B8%AD%E6%96%87 on the wire. Copy the URL out of the address bar and you may get either form depending on the browser.

“Every escape here looks valid, but the character that comes out is not the one I expected. Why?”

Two different failures hide behind "it decoded wrong", and the page reports them differently. The first is a broken escape: a percent sign that is not followed by two hexadecimal digits, like a trailing 100% or a %zz. That is a syntax problem, and the strict reading refuses it and says where it sits while the lenient reading keeps it as literal text and lists it. The second is an escape whose byte is perfectly legal but whose bytes do not spell a valid UTF-8 character, like %FF on its own or a three byte character cut short as %E4%B8. That is not a syntax problem, so both readings return the text, replace the bytes that cannot be decoded with the replacement character U+FFFD ("�"), and list a warning that says so. That warning is the useful part: text that arrives as replacement characters usually means the sender used a different character set, or that the string was escaped once too often. Percent encoding names bytes, and this page, like the URL standard, always reads those bytes as UTF-8.

“What does this page not do?”

It does not split a URL into scheme, host, path and query — that is what URL Parser is for. It does not check whether a link is safe to open, it does not contact anything, and it does not guess the character set of the bytes behind an escape: it always reads them as UTF-8, which is what the URL standard requires.

“Which fallback does decoding need, and how do I know it is right?”

Press Self-test. It compares the page's own encoders against the platform's encodeURIComponent, encodeURI and URLSearchParams for every printable ASCII character, re-runs every example in the tables above, and re-checks the recorded layer chains, then reports the number of checks here.

Standards and scope

Percent encoding is specified in RFC 3986 section 2.1, with the definition of reserved and unreserved characters in sections 2.2 and 2.3 and the case rule in section 6.2.2.1. The sets a browser actually uses when it builds a URL are in the WHATWG URL Standard, which splits them by context: one set for a fragment, one for a query, one for a special query, one for a path, one for user information, and the narrower one used for a single component. Form submission, including the plus sign, is in the HTML Standard under application/x-www-form-urlencoded. The two functions encodeURI and encodeURIComponent are part of ECMAScript itself.

This page fills the path, query and fragment columns from its own per context sets, and prints the exact list for every context next to the result. Those sets are derived from the WHATWG ones and differ from them in two deliberate places: a hash is escaped inside a query, because a raw hash would start the fragment, and the component set is pinned to exactly the characters encodeURIComponent escapes. Where a platform function is the definition of the answer — encodeURIComponent, encodeURI and the form serialisation — that function is called directly, so those columns are the browser's own behaviour rather than a reimplementation of it.

The page text was written for this site. Every value in the tables above is produced by the page's own code at load time and was verified against the platform's implementations in this site's test tooling; the Self-test repeats that comparison in your browser.