ASCII table and UTF-8 converter
Type text to see its UTF-8 bytes, code points and UTF-16 units, or paste bytes to decode them and learn exactly why a sequence is invalid. A full ASCII table with the names and meanings of the control characters is below. Everything runs in your browser.
How to use
- Choose what you are starting from. Text encodes the characters you type into UTF-8. Code points takes Unicode scalar values such as
U+1F600(also0xor\u{...}forms), separated by spaces or commas. UTF-8 bytes in hex or binary decodes bytes into text. - For text and code points the tool lists, per character, its code point, UTF-8 bytes in hex and binary, byte count and UTF-16 code units, then draws the bit layout for the first few distinct characters: the template (
110xxxxx 10xxxxxx) with your payload bits in place of the x's. - For bytes the tool decodes strictly, following the well-formed byte ranges in Unicode Table 3-7 (reproduced below). Every problem is listed with its offset, the bytes involved and the reason, and each bad byte or maximal bad run is replaced by U+FFFD the way the Unicode standard recommends.
- Use the table at the bottom of the page for control characters (names and RFC 20 meanings), printable characters, and decimal, hex and binary codes.
Worked examples
Each result is recomputed by an automated test from the tool's engine. The bit-level derivations below were done by hand from RFC 3629 and compared with the browser's TextEncoder in the tests.
One, three and four bytes: A, euro sign, grinning face
A is U+0041, below U+0080, so it is the single byte 41. The euro sign U+20AC is 0010 0000 1010 1100 in 16 bits. The three-byte template is 1110xxxx 10xxxxxx 10xxxxxx, so the bits split as 0010 | 000010 | 101100, giving 11100010 10000010 10101100 = E2 82 AC. The grinning face U+1F600 is 21 bits, 0 0001 1111 0110 0000 0000; the four-byte template 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx gives 11110000 10011111 10011000 10000000 = F0 9F 98 80.
Start from Text
Input: A€😀
Output: 41 E2 82 AC F0 9F 98 80
Two spellings of café
The precomposed é is U+00E9 = 000 1110 1001 in 11 bits; the two-byte template 110xxxxx 10xxxxxx gives 11000011 10101001 = C3 A9.
Start from Text
Input: café
Output: 63 61 66 C3 A9
Typed as the letter e plus the combining acute accent U+0301 it looks identical but is different data: 65 for the e, and U+0301 = 011 0000 0001 gives 11001100 10000001 = CC 81.
Start from Text
Input: café
Output: 63 61 66 65 CC 81
From code points
Start from Code points
Input: U+1F600 U+0301
Output: F0 9F 98 80 | CC 81
From bytes to text
Start from UTF-8 bytes (hex)
Input: C3 A9
Result: valid UTF-8: U+00E9
Reference: the well-formed UTF-8 byte sequences (Unicode Table 3-7)
A byte sequence is valid UTF-8 only if it matches one row. The tool's decoder is tested against this table and against the replacement examples in Tables 3-8 to 3-11.
| Code points | First byte | Second byte | Third byte | Fourth byte |
|---|---|---|---|---|
U+0000..U+007F | 00..7F | | | |
U+0080..U+07FF | C2..DF | 80..BF | | |
U+0800..U+0FFF | E0 | A0..BF | 80..BF | |
U+1000..U+CFFF | E1..EC | 80..BF | 80..BF | |
U+D000..U+D7FF | ED | 80..9F | 80..BF | |
U+E000..U+FFFF | EE..EF | 80..BF | 80..BF | |
U+10000..U+3FFFF | F0 | 90..BF | 80..BF | 80..BF |
U+40000..U+FFFFF | F1..F3 | 80..BF | 80..BF | 80..BF |
U+100000..U+10FFFF | F4 | 80..8F | 80..BF | 80..BF |
Reading the rows: after E0 the second byte must be A0 to BF (anything lower would be an overlong form), after ED it must be 80 to 9F (anything higher would be a surrogate), after F0 it must be 90 to BF, and after F4 it must be 80 to 8F (anything higher would be beyond U+10FFFF).
Control characters (0 to 31) and DEL (127)
Names and meanings are from RFC 20 sections 4 and 5.2 (shortened and reworded). The caret notation and C escape columns are common conventions, not part of RFC 20.
| Dec | Hex | Binary | Abbr. | Name | Caret | C escape | Meaning (RFC 20) |
|---|---|---|---|---|---|---|---|
| 0 | 00 | 0000000 | NUL | Null | ^@ | \0 | The all-zeros character, used for time fill and media fill. Ends a string in C and most Unix conventions. |
| 1 | 01 | 0000001 | SOH | Start of Heading | ^A | Begins a heading: routing information at the start of a transmission. | |
| 2 | 02 | 0000010 | STX | Start of Text | ^B | Starts the text that follows the heading and ends the heading. | |
| 3 | 03 | 0000011 | ETX | End of Text | ^C | Ends a text that began with STX. | |
| 4 | 04 | 0000100 | EOT | End of Transmission | ^D | Marks the end of a transmission that may have held several texts. | |
| 5 | 05 | 0000101 | ENQ | Enquiry | ^E | A request for a reply from a remote station ("Who are you?"). | |
| 6 | 06 | 0000110 | ACK | Acknowledge | ^F | A receiver’s positive reply to a sender. | |
| 7 | 07 | 0000111 | BEL | Bell | ^G | \a | Calls for human attention; may sound an alarm. |
| 8 | 08 | 0001000 | BS | Backspace | ^H | \b | Moves the print position one place back on the same line. |
| 9 | 09 | 0001001 | HT | Horizontal Tabulation | ^I | \t | Moves to the next of a series of preset positions on the line. |
| 10 | 0A | 0001010 | LF | Line Feed | ^J | \n | Moves to the next line (RFC 20 also allows the meaning "New Line" by agreement). The line ending on Unix, Linux and macOS. |
| 11 | 0B | 0001011 | VT | Vertical Tabulation | ^K | \v | Moves to the next of a series of preset lines. |
| 12 | 0C | 0001100 | FF | Form Feed | ^L | \f | Moves to the first line of the next page or form. |
| 13 | 0D | 0001101 | CR | Carriage Return | ^M | \r | Moves to the first print position of the same line. CR LF is the line ending on Windows and in many Internet protocols. |
| 14 | 0E | 0001110 | SO | Shift Out | ^N | Says the codes that follow lie outside the standard code table until Shift In. | |
| 15 | 0F | 0001111 | SI | Shift In | ^O | Says the codes that follow follow the standard table again. | |
| 16 | 10 | 0010000 | DLE | Data Link Escape | ^P | Changes the meaning of a few following characters, only for data communication controls. | |
| 17 | 11 | 0010001 | DC1 | Device Control 1 | ^Q | Turns an auxiliary device on or off. Often called XON (resume output). | |
| 18 | 12 | 0010010 | DC2 | Device Control 2 | ^R | Turns an auxiliary device on or off. | |
| 19 | 13 | 0010011 | DC3 | Device Control 3 | ^S | Turns an auxiliary device on or off. Often called XOFF (pause output). | |
| 20 | 14 | 0010100 | DC4 | Device Control 4 | ^T | Turns an auxiliary device on or off; RFC 20 prefers it as the single "stop". | |
| 21 | 15 | 0010101 | NAK | Negative Acknowledge | ^U | A receiver’s negative reply to a sender. | |
| 22 | 16 | 0010110 | SYN | Synchronous Idle | ^V | Sent by a synchronous link when there is nothing else, to keep or regain synchronism. | |
| 23 | 17 | 0010111 | ETB | End of Transmission Block | ^W | Marks the end of a block of data for communication purposes. | |
| 24 | 18 | 0011000 | CAN | Cancel | ^X | Says the data sent with it is wrong or to be ignored. | |
| 25 | 19 | 0011001 | EM | End of Medium | ^Y | Marks the physical end of the medium or the end of the wanted part of it. | |
| 26 | 1A | 0011010 | SUB | Substitute | ^Z | Stands in for a character found to be invalid or in error. Ctrl+Z (^Z) was conventionally used as an end-of-file marker; the name "EOF" stuck to it. | |
| 27 | 1B | 0011011 | ESC | Escape | ^[ | \e | Starts a code extension: it changes the meaning of the characters that follow it. Starts ANSI escape sequences in terminals (ESC [ ...). |
| 28 | 1C | 0011100 | FS | File Separator | ^\ | Information separator: the most inclusive level (file). | |
| 29 | 1D | 0011101 | GS | Group Separator | ^] | Information separator: group level, inside a file. | |
| 30 | 1E | 0011110 | RS | Record Separator | ^^ | Information separator: record level, inside a group. | |
| 31 | 1F | 0011111 | US | Unit Separator | ^_ | Information separator: the least inclusive level (unit). | |
| 127 | 7F | 1111111 | DEL | Delete | ^? | Used to erase or obliterate a wrong character on perforated tape. RFC 20 notes it is not strictly a control character. |
Printable characters (32 to 126)
The same 95 characters, in table order. Names follow RFC 20 section 4.2 (a few reworded, for example "Slant" is the slash).
| Dec | Hex | Binary | Char | Name |
|---|---|---|---|---|
| 32 | 20 | 0100000 | SP | Space |
| 33 | 21 | 0100001 | ! | Exclamation Point |
| 34 | 22 | 0100010 | " | Quotation Marks |
| 35 | 23 | 0100011 | # | Number Sign |
| 36 | 24 | 0100100 | $ | Dollar Sign |
| 37 | 25 | 0100101 | % | Percent |
| 38 | 26 | 0100110 | & | Ampersand |
| 39 | 27 | 0100111 | ' | Apostrophe |
| 40 | 28 | 0101000 | ( | Opening Parenthesis |
| 41 | 29 | 0101001 | ) | Closing Parenthesis |
| 42 | 2A | 0101010 | * | Asterisk |
| 43 | 2B | 0101011 | + | Plus |
| 44 | 2C | 0101100 | , | Comma |
| 45 | 2D | 0101101 | - | Hyphen (Minus) |
| 46 | 2E | 0101110 | . | Period (Decimal Point) |
| 47 | 2F | 0101111 | / | Slant |
| 48 | 30 | 0110000 | 0 | Digit 0 |
| 49 | 31 | 0110001 | 1 | Digit 1 |
| 50 | 32 | 0110010 | 2 | Digit 2 |
| 51 | 33 | 0110011 | 3 | Digit 3 |
| 52 | 34 | 0110100 | 4 | Digit 4 |
| 53 | 35 | 0110101 | 5 | Digit 5 |
| 54 | 36 | 0110110 | 6 | Digit 6 |
| 55 | 37 | 0110111 | 7 | Digit 7 |
| 56 | 38 | 0111000 | 8 | Digit 8 |
| 57 | 39 | 0111001 | 9 | Digit 9 |
| 58 | 3A | 0111010 | : | Colon |
| 59 | 3B | 0111011 | ; | Semicolon |
| 60 | 3C | 0111100 | < | Less Than |
| 61 | 3D | 0111101 | = | Equals |
| 62 | 3E | 0111110 | > | Greater Than |
| 63 | 3F | 0111111 | ? | Question Mark |
| 64 | 40 | 1000000 | @ | Commercial At |
| 65 | 41 | 1000001 | A | Capital letter A |
| 66 | 42 | 1000010 | B | Capital letter B |
| 67 | 43 | 1000011 | C | Capital letter C |
| 68 | 44 | 1000100 | D | Capital letter D |
| 69 | 45 | 1000101 | E | Capital letter E |
| 70 | 46 | 1000110 | F | Capital letter F |
| 71 | 47 | 1000111 | G | Capital letter G |
| 72 | 48 | 1001000 | H | Capital letter H |
| 73 | 49 | 1001001 | I | Capital letter I |
| 74 | 4A | 1001010 | J | Capital letter J |
| 75 | 4B | 1001011 | K | Capital letter K |
| 76 | 4C | 1001100 | L | Capital letter L |
| 77 | 4D | 1001101 | M | Capital letter M |
| 78 | 4E | 1001110 | N | Capital letter N |
| 79 | 4F | 1001111 | O | Capital letter O |
| 80 | 50 | 1010000 | P | Capital letter P |
| 81 | 51 | 1010001 | Q | Capital letter Q |
| 82 | 52 | 1010010 | R | Capital letter R |
| 83 | 53 | 1010011 | S | Capital letter S |
| 84 | 54 | 1010100 | T | Capital letter T |
| 85 | 55 | 1010101 | U | Capital letter U |
| 86 | 56 | 1010110 | V | Capital letter V |
| 87 | 57 | 1010111 | W | Capital letter W |
| 88 | 58 | 1011000 | X | Capital letter X |
| 89 | 59 | 1011001 | Y | Capital letter Y |
| 90 | 5A | 1011010 | Z | Capital letter Z |
| 91 | 5B | 1011011 | [ | Opening Bracket |
| 92 | 5C | 1011100 | \ | Reverse Slant |
| 93 | 5D | 1011101 | ] | Closing Bracket |
| 94 | 5E | 1011110 | ^ | Circumflex |
| 95 | 5F | 1011111 | _ | Underline |
| 96 | 60 | 1100000 | ` | Grave Accent |
| 97 | 61 | 1100001 | a | Small letter a |
| 98 | 62 | 1100010 | b | Small letter b |
| 99 | 63 | 1100011 | c | Small letter c |
| 100 | 64 | 1100100 | d | Small letter d |
| 101 | 65 | 1100101 | e | Small letter e |
| 102 | 66 | 1100110 | f | Small letter f |
| 103 | 67 | 1100111 | g | Small letter g |
| 104 | 68 | 1101000 | h | Small letter h |
| 105 | 69 | 1101001 | i | Small letter i |
| 106 | 6A | 1101010 | j | Small letter j |
| 107 | 6B | 1101011 | k | Small letter k |
| 108 | 6C | 1101100 | l | Small letter l |
| 109 | 6D | 1101101 | m | Small letter m |
| 110 | 6E | 1101110 | n | Small letter n |
| 111 | 6F | 1101111 | o | Small letter o |
| 112 | 70 | 1110000 | p | Small letter p |
| 113 | 71 | 1110001 | q | Small letter q |
| 114 | 72 | 1110010 | r | Small letter r |
| 115 | 73 | 1110011 | s | Small letter s |
| 116 | 74 | 1110100 | t | Small letter t |
| 117 | 75 | 1110101 | u | Small letter u |
| 118 | 76 | 1110110 | v | Small letter v |
| 119 | 77 | 1110111 | w | Small letter w |
| 120 | 78 | 1111000 | x | Small letter x |
| 121 | 79 | 1111001 | y | Small letter y |
| 122 | 7A | 1111010 | z | Small letter z |
| 123 | 7B | 1111011 | { | Opening Brace |
| 124 | 7C | 1111100 | | | Vertical Line |
| 125 | 7D | 1111101 | } | Closing Brace |
| 126 | 7E | 1111110 | ~ | Overline (Tilde) |
What goes wrong
Real byte sequences that break decoders. The results are the tool's exact output, and the replacement counts for the first four match the examples printed in the Unicode Standard (Tables 3-8 to 3-11), which the tests also run.
C0 AF, an overlong slash
C0 can only begin an overlong two-byte form; AF is then a continuation byte with no lead. Unicode Table 3-8 lists exactly two U+FFFD for C0 AF.
Start from UTF-8 bytes (hex)
Input: C0 AF
Result: 2 errors: U+FFFD U+FFFD
Latin-1 bytes read as UTF-8
The Latin-1 text "é café" is E9 20 63 61 66 E9. E9 would start a three-byte sequence, but the next byte (20) is not a continuation byte, so E9 is replaced alone; the same happens at the end, where E9 is truncated.
Start from UTF-8 bytes (hex)
Input: E9 20 63 61 66 E9
Result: 2 errors: U+FFFD U+0020 U+0063 U+0061 U+0066 U+FFFD
An emoji cut off after three bytes
F0 9F 98 is the first three bytes of F0 9F 98 80. It is the start of a valid sequence, so the standard replaces the whole truncated prefix with a single U+FFFD, not three (Table 3-11).
Start from UTF-8 bytes (hex)
Input: F0 9F 98
Result: 1 error: U+FFFD
A surrogate encoded as three bytes
ED A0 80 is what a naive encoder writes for the surrogate code point U+D800. UTF-8 forbids surrogates (RFC 3629; after ED the second byte must be 80 to 9F), so each of the three bytes becomes its own U+FFFD (Table 3-9).
Start from UTF-8 bytes (hex)
Input: ED A0 80
Result: 3 errors: U+FFFD U+FFFD U+FFFD
A lone surrogate in text
JavaScript strings can hold a lone surrogate code unit, but it is not a Unicode scalar value, so it cannot be written in UTF-8. TextEncoder (which takes a USVString) substitutes U+FFFD, bytes EF BF BD, and so does this tool; it also counts it so you notice.
Start from Text
Input: a\uD800b
Result: 61 EF BF BD 62 (1 lone surrogate replaced by EF BF BD)
A code point beyond Unicode
Start from Code points
Input: U+110000
Error: Code point 1 (U+110000) is above U+10FFFF, the last Unicode code point.
Mojibake in the other direction
The UTF-8 bytes of é are C3 A9. Reading them as windows-1252 or Latin-1 shows é; re-encoding that text as UTF-8 gives C3 83 C2 A9, four bytes, which is the classic "double-encoded" form. In the browser, new TextDecoder("windows-1252").decode(new Uint8Array([0xC3, 0xA9])) returns é.
Limits & gotchas
- UTF-8 only. The tool encodes and decodes UTF-8 (and shows UTF-16 code units for reference). Latin-1, windows-1252, Shift_JIS and other legacy encodings are not converted; the page only explains how reading UTF-8 bytes with them goes wrong.
- Code points, not glyphs. The tool lists code points. What a reader sees as one "character" may be several (a letter plus a combining accent, or an emoji sequence joined with U+200D). Each code point is shown separately on purpose.
- No normalisation. The tool does not convert between composed and decomposed forms; it shows the bytes of exactly what you typed. Names and categories are not looked up for every code point; only the ASCII table has names (from RFC 20).
- Input size. At most 20,000 characters or 20,000 bytes. The per-character table shows the first 200 characters and the byte-by-byte table the first 300 entries.
- Pasted text may change. Browsers and editors can convert line endings or replace characters when you paste. If a byte count looks off, paste into the "UTF-8 bytes" mode instead, or check for a hidden trailing newline (LF, 0A).
- Replacement follows Unicode's recommended practice. Other decoders may produce a different number of U+FFFD characters for the same bad bytes; the Unicode Standard says it does not require this practice for conformance. This tool's decoder agrees with the browser's
TextDecoderon 20,000 random byte strings in the tests. - Browser. Tested in a current Chrome only.
FAQ
Why does one emoji take four bytes, and why does my string length say 2?
UTF-8 uses 1 byte for U+0000 to U+007F, 2 bytes up to U+07FF, 3 bytes up to U+FFFF and 4 bytes above that (RFC 3629). The grinning face U+1F600 is above U+FFFF, so it is the four bytes F0 9F 98 80. JavaScript strings are sequences of 16-bit code units (ECMAScript), and a code point above U+FFFF takes two of them (a surrogate pair, D83D DE00 here), so "😀".length is 2 while [..."😀"].length is 1.
What is the difference between ASCII and UTF-8?
ASCII is a 7-bit code with 128 characters (RFC 20): codes 0 to 31 and 127 are control characters, 32 to 126 are space and printable characters. UTF-8 encodes all of Unicode and is designed so that code points 0 to 127 use exactly the ASCII byte values, so plain ASCII text is already valid UTF-8. Every other character takes two to four bytes, each with the top bit set. The names "ASCII" and "ANSI" are often used loosely for 8-bit code pages such as Latin-1 or windows-1252, which are neither.
Why do I see é as é (or a question-mark box)?
Because bytes were read with the wrong encoding. UTF-8 stores é (U+00E9) as the two bytes C3 A9. Read as Latin-1 or windows-1252, C3 is "Ã" and A9 is "©", giving "é". The reverse mistake is the bytes E9 20 63 61 66 E9 (the Latin-1 text "é café"): they are not valid UTF-8, so a decoder shows the replacement character U+FFFD in place of each bad byte. The tool's "What goes wrong" section shows the exact byte-by-byte result.
Why are there two different byte sequences for "café"?
Unicode can write é as the single code point U+00E9 (bytes C3 A9) or as the letter e followed by the combining acute accent U+0301 (bytes 65 CC 81). They look the same and the Unicode database records the first as canonically equivalent to the second (UnicodeData.txt lists U+00E9 as decomposing to 0065 0301). But as bytes they differ, so a hash, a file name or a string comparison sees different data until the text is normalised. The tool shows both for "café".
What does C0 AF mean, and why is it rejected?
C0 starts a two-byte sequence that could only encode a code point below U+0080, which UTF-8 must write as one byte. C0 AF would be an "overlong" encoding of "/" (U+002F). RFC 3629 says decoders must not accept such sequences: they are a security risk, because a naive decoder could let a "/" or NUL slip past a filter that only looks for the one-byte form. Unicode 16 Table 3-8 shows the replacement for C0 AF as two U+FFFD characters, and the tool produces exactly that.
Sources
- IETF: RFC 3629: UTF-8, a transformation format of ISO 10646 Used for: The byte layouts for one to four bytes, the code point range U+0000 to U+10FFFF, surrogates not allowed, and that decoders must protect against decoding invalid sequences such as the overlong C0 80.
- Unicode Consortium: The Unicode Standard 16.0, chapter 3 (Conformance): Table 3-7 and Tables 3-8 to 3-11 Used for: The well-formed UTF-8 byte ranges (Table 3-7), the definition of ill-formed sequences, and the U+FFFD replacement examples (Tables 3-8 to 3-11) the decoder is tested against.
- Unicode Consortium: UnicodeData.txt, Unicode 16.0.0 Used for: Official character names and the canonical decompositions of U+00E9 (0065 0301) and U+00DC (0055 0308), used for the e + combining acute example.
- IETF: RFC 20: ASCII format for network interchange Used for: The 7-bit code table, the control character names and meanings, and the definition of each control code.
- Wikipedia: ASCII Used for: C and Unix use the null character to terminate strings; Control-Z (SUB) was used as an end-of-file marker; caret notation for control characters.
- WHATWG: Encoding Standard: TextEncoder and TextDecoder Used for: TextEncoder.encode takes a USVString, and TextDecoder returns a USVString, so lone surrogates cannot be encoded. The labels "latin1", "iso-8859-1" and "us-ascii" all name the windows-1252 encoding on the web platform (the table of names and labels).
- WHATWG: Web IDL: the USVString type Used for: USVString corresponds to scalar value strings, which exclude surrogate code points.
- Ecma International / TC39: ECMAScript Language Specification: Number type, bitwise shift operators and BigInt operations Used for: Number is IEEE 754 binary64; Number shifts convert with ToInt32 or ToUint32 and take the count modulo 32; BigInt left and right shift act on an infinite two's complement string; BigInt::unsignedRightShift throws a TypeError; strings are sequences of 16-bit code units.
- MDN Web Docs: TextEncoder Used for: TextEncoder always encodes with UTF-8. Used in the unit tests as an independent oracle for the UTF-8 encoder.
Every document above was opened and read on 2026-10-02. Documentation changes; if a page here disagrees with the current docs, trust the docs and tell us.
The ASCII table's names and control-character meanings are shortened and reworded from RFC 20 sections 4 and 5.2. Caret notation and C escapes (such as \n) are common conventions, not part of RFC 20, and are labelled so.