Unicode Escape / Unescape Converter

Turn any text into \uXXXX escapes, and read them back again.

Unicode Escape / Unescape Converter

Unicode Escape / Unescape converts characters to safe escape sequences such as \uXXXX, \u{XXXXX} and \UXXXXXXXX,
and turns those sequences back into real characters. Emoji and rare scripts are handled as proper surrogate pairs or code points.
Everything runs in your browser, so your text never leaves your device.


Load an example

Settings

The notation used when escaping. Surrogate pairs are used when the format needs them.

Non-ASCII leaves plain English and punctuation untouched; All characters gives pure-ASCII output.

Short escapes apply only in All characters mode; otherwise control characters use their numeric form.

Off keeps the good output and flags the first malformed sequence. On, any bad sequence clears the result.

Input: 0 characters · Output: 0 characters
Ready to escape or unescape Unicode.


Overview

The Unicode Escape / Unescape Converter is a free, browser-based tool for moving text safely between plain characters and backslash escape sequences. Escape mode rewrites a character as \uXXXX, \u{XXXXX} or \UXXXXXXXX using the number of its Unicode code point; Unescape mode reads those sequences and rebuilds the original characters. It is the opposite job to reading raw bytes: it changes how text is spelled in source code, not how it is encoded on disk.

This matters whenever your text has to survive a place that only accepts plain ASCII: a JavaScript or JSON string, a Java .properties file, a C string literal, a configuration value, a log line or a shell command. Non-English names, accented letters, emoji and rare scripts can all be lost, mangled or rejected in those channels. Writing them as escapes keeps them intact no matter how the file is copied or transported.

The page is written in plain HTML, CSS and JavaScript using this site's shared Toolkit helper, and the escape and unescape engines are implemented directly in the browser. There is no server, no upload and no third-party library, so you can paste sensitive strings with confidence. It handles the full Unicode range, including characters above U+FFFF, which are emitted as UTF-16 surrogate pairs in the four-digit format and as a single code point in the braced and Python formats.


What This Tool Does

Escape: walks the input one code point at a time and replaces the characters you asked for with an escape sequence. By default only characters above ASCII (code point 127 and up) are escaped, so ordinary English stays readable; switch the scope to All characters to turn every character into an escape.

Format: the same code point can be written in more than one notation, and the Escape format setting chooses which one. \uXXXX is the four-digit form used by JavaScript, JSON and Java. \u{XXXXX} is the modern JavaScript form that can hold a whole code point directly. \UXXXXXXXX is Python's eight-digit form.

Unescape: scans the input for escape sequences and turns them back into characters. It understands \uXXXX, \u{XXXXX}, \UXXXXXXXX and \xXX, joins surrogate pairs back into a single character, and treats \n, \t, \r, \b and \f as their control characters. A doubled backslash \\ is read as a literal backslash, so already-escaped backslashes are not double processed.

Error handling: a broken sequence such as a truncated \u00 does not ruin the whole result. In the default lenient mode the rest of the text is converted and the first problem is reported with its position; turn on Strict mode to make any malformed sequence clear the output instead.


Features

Three escape formats: choose the notation that matches your target language. The four-digit form is the safest for JSON and older JavaScript; the braced form is shorter for emoji; the Python form is fixed at eight digits.

Example: the rocket U+1F680 is \ud83d\ude80 in the four-digit form, \u{1f680} in the braced form and \U0001f680 in the Python form.

Escape scope: Non-ASCII only is the everyday choice and keeps the output readable by escaping just the characters that cause trouble. All characters produces a string that contains no raw Unicode at all, which is what you need for strict ASCII-only pipelines.

Example: the word Hi! stays Hi! in Non-ASCII mode, but becomes \u0048\u0069\u0021 in All characters mode.

Hex case and short control escapes: the hex digits can be lowercase or uppercase to match your project's style. When you escape All characters you can also write the common control characters in their short form instead of the numeric one.

Example: a line break becomes \u000a by default, or \n when short escapes are on.

Surrogate pairs handled correctly: characters above U+FFFF, such as most emoji, cannot fit in four hex digits. In the four-digit format they are split into a valid high and low surrogate pair, and unescaping joins the pair back into one character.

Example: the grinning face U+1F600 escapes to \ud83d\ude00, and that pair unescapes back to the single emoji.

Lenient or strict unescape: lenient mode converts everything it can and reports the first malformed sequence with its character position, which is far more useful on a long file than losing the lot. Strict mode is there when you would rather fail fast and fix the source.

Live statistics, file upload, copy and download: the character counts for input and output update as you work, so you can see how much the escapes add; you can also load a .txt file, copy the result and download it as a file.


Examples

Click Try it on any example to load it into the tool. Each example also selects the format, scope and options it demonstrates, and the text on the right is the exact output the tool produces.

Example 1: Accented letters and a symbol (non-ASCII)
Input
Café ☕
Output (Escape, \uXXXX, lowercase)
Caf\u00e9 \u2615

Only the accented é and the coffee symbol are rewritten; the plain Latin letters and the space pass through unchanged. This is the default and the form you want inside a JSON string.

Example 2: An emoji becomes a surrogate pair
Input
🚀 Launch
Output (Escape, \uXXXX, lowercase)
\ud83d\ude80 Launch

The rocket U+1F680 is above U+FFFF, so the four-digit format splits it into the high and low halves \ud83d and \ude80. The English word is left alone.

Example 3: The same emoji in ES6 braced form
Input
Hello 😀
Output (Escape, \u{XXXXX})
Hello \u{1f600}

The braced format keeps the whole code point in one escape, which is shorter and easier to read than a surrogate pair. Use it only where the target understands ES6 braces, not in strict JSON.

Example 4: Pure-ASCII output by escaping everything
Input
Hi!
Output (Escape, \uXXXX, All characters)
\u0048\u0069\u0021

With the scope set to All characters even plain letters and punctuation are escaped, giving a string that is guaranteed to contain only ASCII. This mode also makes an escape then unescape round trip safe.

Example 5: Unescaping mixed sequences back to text
Input
\u4f60\u597d, \u4e16\u754c! \u{1f680}
Output (Unescape)
你好, 世界! 🚀

Unescape reads four-digit escapes, the braced form and control shortcuts together and rebuilds the original characters, including the multi-byte Chinese characters and the emoji.


FAQs

FAQ 1: What is a Unicode escape sequence?
Answer: It is a way of writing a character using the number of its code point instead of the character itself. The sequence starts with a backslash followed by a letter that marks the notation and then hexadecimal digits, for example \u00e9 for é. The Unicode standard gives every character a code point, so any of them can be written this way.

FAQ 2: What is the difference between the three escape formats?
Answer: \uXXXX is four hexadecimal digits and is understood by JavaScript, JSON and Java; characters above U+FFFF must be split into a surrogate pair. \u{XXXXX} is the ES6 JavaScript form and holds a whole code point, so it does not need a pair. \UXXXXXXXX is Python's eight-digit form. They all describe the same character, only the notation differs.

FAQ 3: Why does an emoji turn into two escapes?
Answer: The four-digit format can only hold values up to U+FFFF, but many emoji and rarer scripts sit above that. Such a character is stored as a UTF-16 surrogate pair, two 16-bit halves, so it becomes two escapes such as \ud83d\ude00. This tool does that automatically and rejoins the pair when unescaping. If you prefer a single escape, choose the braced or Python format.

FAQ 4: Should I escape everything or only non-ASCII characters?
Answer: Non-ASCII only keeps the result readable and is what most people need, because plain ASCII text rarely breaks. Escaping all characters produces a string that contains no raw Unicode at all, which is required when a toolchain refuses anything non-ASCII, and it is the only mode that round trips safely when the input already contains backslashes.

FAQ 5: What happens if the input contains a broken escape sequence?
Answer: In the default lenient mode the tool converts everything it can and reports the first malformed sequence together with its character position, leaving that fragment as literal text. Turn on Strict mode if you would rather the conversion stop and clear the output, so a bad source file cannot pass unnoticed.

FAQ 6: Is my text sent to a server?
Answer: No. Escaping and unescaping both run in your browser using plain JavaScript, so the text never leaves your device and nothing is stored anywhere.


Tools Worth Trying

TextToolz

The Ultimate Text Tools

TextToolz works seamlessly to let you convert and design your text. It is fast, reliable and secure. Trusted by thousands of users.