Escape text as JavaScript \uXXXX / \u{HEX} sequences, decode mixed escapes back to text, and inspect every character as a
code point: U+ hex, decimal, UTF-8 bytes, UTF-16 units and JS escapes. Astral code points (emoji) are never split into surrogate halves.
All processing happens locally in your browser — your data is never uploaded to any server.
Escape style:
Encode: Text is split into Unicode code points (so emoji stay intact) and every one is written as a JavaScript escape sequence.
Decode: Every \uXXXX, \u{hex} or U+hex sequence (upper or lower case) is converted to its character. Adjacent surrogate-pair escapes are merged into one code point. Unescaped text passes through unchanged.
Inspector: Each code point becomes one tab-separated row: character, U+HEX, decimal, UTF-8 bytes, UTF-16 units and JS escape. The whole table can be copied from the output or downloaded.
Input
Output
Ctrl+Enter in the input runs the active mode.
Pick Encode, Decode or Inspector and run it. All processing stays in your browser.
Reading the inspector columns: U+HEX is the code point in hexadecimal; decimal is the same number in base 10. UTF-8 bytes shows the one-to-four bytes
the character occupies when written as UTF-8 (e.g. 中 = E4 B8 AD). UTF-16 shows the one or two 16-bit units the character uses inside a
JavaScript string (emoji beyond U+FFFF need two units, a “surrogate pair”: e.g. U+1F600 = D83D DE00). JS escape is the recommended
source-code form of the character.
What the Sample button loads:
The example that belongs to the mode you are in, never one fixed text, because one text cannot serve all three.
Encode and Inspector both take plain characters, so they get Hello 中文! í ¼í¾‰ — ASCII, CJK, a fullwidth
mark and one emoji. Decode takes escape sequences, and a plain sentence holds none: it would be echoed back word for word
and the mode would have nothing to show, so Decode gets Say \u4e2d\u6587 \u{1F600} U+0041 \uD83C\uDF89 OK, which carries one of
every form the decoder accepts (a \uXXXX pair in lower case, a \u{HEX} brace, a U+hex form,
an adjacent surrogate pair, and words that are not escapes). The line under the boxes states the result to look for; in Encode it
also follows the escape style and the ASCII checkbox, so the three other styles can be seen by pressing Sample again.
What is Unicode?
Unicode is a single catalogue that assigns every character of every writing system a number, its code point, from U+0000 to U+10FFFF — 1,114,112 possible slots arranged in 17 planes of 65,536. The first plane, the Basic Multilingual Plane (BMP), holds the characters of most living languages; most emoji and many historic scripts live on the higher (“astral”) planes.
UTF-8, UTF-16 and UTF-32 are not rival character sets but three ways of serialising those numbers into bytes — the Inspector mode shows all three for the same character, next to the JavaScript escape forms.
Unicode names and numbers the characters; UTF-8/UTF-16/UTF-32 and \uXXXX escapes are just different ways of writing the same numbers down.