ASCII table and UTF-8 converter

Type text to see its UTF-8 bytes, code points and UTF-16 units, or paste bytes to decode them and learn exactly why a sequence is invalid. A full ASCII table with the names and meanings of the control characters is below. Everything runs in your browser.

Start from
Try:

How to use

  1. Choose what you are starting from. Text encodes the characters you type into UTF-8. Code points takes Unicode scalar values such as U+1F600 (also 0x or \u{...} forms), separated by spaces or commas. UTF-8 bytes in hex or binary decodes bytes into text.
  2. For text and code points the tool lists, per character, its code point, UTF-8 bytes in hex and binary, byte count and UTF-16 code units, then draws the bit layout for the first few distinct characters: the template (110xxxxx 10xxxxxx) with your payload bits in place of the x's.
  3. For bytes the tool decodes strictly, following the well-formed byte ranges in Unicode Table 3-7 (reproduced below). Every problem is listed with its offset, the bytes involved and the reason, and each bad byte or maximal bad run is replaced by U+FFFD the way the Unicode standard recommends.
  4. Use the table at the bottom of the page for control characters (names and RFC 20 meanings), printable characters, and decimal, hex and binary codes.

Worked examples

Each result is recomputed by an automated test from the tool's engine. The bit-level derivations below were done by hand from RFC 3629 and compared with the browser's TextEncoder in the tests.

One, three and four bytes: A, euro sign, grinning face

A is U+0041, below U+0080, so it is the single byte 41. The euro sign U+20AC is 0010 0000 1010 1100 in 16 bits. The three-byte template is 1110xxxx 10xxxxxx 10xxxxxx, so the bits split as 0010 | 000010 | 101100, giving 11100010 10000010 10101100 = E2 82 AC. The grinning face U+1F600 is 21 bits, 0 0001 1111 0110 0000 0000; the four-byte template 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx gives 11110000 10011111 10011000 10000000 = F0 9F 98 80.

Start from Text

Input: A€😀

Output: 41 E2 82 AC F0 9F 98 80

Two spellings of café

The precomposed é is U+00E9 = 000 1110 1001 in 11 bits; the two-byte template 110xxxxx 10xxxxxx gives 11000011 10101001 = C3 A9.

Start from Text

Input: café

Output: 63 61 66 C3 A9

Typed as the letter e plus the combining acute accent U+0301 it looks identical but is different data: 65 for the e, and U+0301 = 011 0000 0001 gives 11001100 10000001 = CC 81.

Start from Text

Input: café

Output: 63 61 66 65 CC 81

From code points

Start from Code points

Input: U+1F600 U+0301

Output: F0 9F 98 80 | CC 81

From bytes to text

Start from UTF-8 bytes (hex)

Input: C3 A9

Result: valid UTF-8: U+00E9

Reference: the well-formed UTF-8 byte sequences (Unicode Table 3-7)

A byte sequence is valid UTF-8 only if it matches one row. The tool's decoder is tested against this table and against the replacement examples in Tables 3-8 to 3-11.

Code pointsFirst byteSecond byteThird byteFourth byte
U+0000..U+007F00..7F
U+0080..U+07FFC2..DF80..BF
U+0800..U+0FFFE0A0..BF80..BF
U+1000..U+CFFFE1..EC80..BF80..BF
U+D000..U+D7FFED80..9F80..BF
U+E000..U+FFFFEE..EF80..BF80..BF
U+10000..U+3FFFFF090..BF80..BF80..BF
U+40000..U+FFFFFF1..F380..BF80..BF80..BF
U+100000..U+10FFFFF480..8F80..BF80..BF

Reading the rows: after E0 the second byte must be A0 to BF (anything lower would be an overlong form), after ED it must be 80 to 9F (anything higher would be a surrogate), after F0 it must be 90 to BF, and after F4 it must be 80 to 8F (anything higher would be beyond U+10FFFF).

Control characters (0 to 31) and DEL (127)

Names and meanings are from RFC 20 sections 4 and 5.2 (shortened and reworded). The caret notation and C escape columns are common conventions, not part of RFC 20.

DecHexBinaryAbbr.NameCaretC escapeMeaning (RFC 20)
0000000000NULNull ^@\0The all-zeros character, used for time fill and media fill. Ends a string in C and most Unix conventions.
1010000001SOHStart of Heading ^ABegins a heading: routing information at the start of a transmission.
2020000010STXStart of Text ^BStarts the text that follows the heading and ends the heading.
3030000011ETXEnd of Text ^CEnds a text that began with STX.
4040000100EOTEnd of Transmission ^DMarks the end of a transmission that may have held several texts.
5050000101ENQEnquiry ^EA request for a reply from a remote station ("Who are you?").
6060000110ACKAcknowledge ^FA receiver’s positive reply to a sender.
7070000111BELBell ^G\aCalls for human attention; may sound an alarm.
8080001000BSBackspace ^H\bMoves the print position one place back on the same line.
9090001001HTHorizontal Tabulation ^I\tMoves to the next of a series of preset positions on the line.
100A0001010LFLine Feed ^J\nMoves to the next line (RFC 20 also allows the meaning "New Line" by agreement). The line ending on Unix, Linux and macOS.
110B0001011VTVertical Tabulation ^K\vMoves to the next of a series of preset lines.
120C0001100FFForm Feed ^L\fMoves to the first line of the next page or form.
130D0001101CRCarriage Return ^M\rMoves to the first print position of the same line. CR LF is the line ending on Windows and in many Internet protocols.
140E0001110SOShift Out ^NSays the codes that follow lie outside the standard code table until Shift In.
150F0001111SIShift In ^OSays the codes that follow follow the standard table again.
16100010000DLEData Link Escape ^PChanges the meaning of a few following characters, only for data communication controls.
17110010001DC1Device Control 1 ^QTurns an auxiliary device on or off. Often called XON (resume output).
18120010010DC2Device Control 2 ^RTurns an auxiliary device on or off.
19130010011DC3Device Control 3 ^STurns an auxiliary device on or off. Often called XOFF (pause output).
20140010100DC4Device Control 4 ^TTurns an auxiliary device on or off; RFC 20 prefers it as the single "stop".
21150010101NAKNegative Acknowledge ^UA receiver’s negative reply to a sender.
22160010110SYNSynchronous Idle ^VSent by a synchronous link when there is nothing else, to keep or regain synchronism.
23170010111ETBEnd of Transmission Block ^WMarks the end of a block of data for communication purposes.
24180011000CANCancel ^XSays the data sent with it is wrong or to be ignored.
25190011001EMEnd of Medium ^YMarks the physical end of the medium or the end of the wanted part of it.
261A0011010SUBSubstitute ^ZStands in for a character found to be invalid or in error. Ctrl+Z (^Z) was conventionally used as an end-of-file marker; the name "EOF" stuck to it.
271B0011011ESCEscape ^[\eStarts a code extension: it changes the meaning of the characters that follow it. Starts ANSI escape sequences in terminals (ESC [ ...).
281C0011100FSFile Separator ^\Information separator: the most inclusive level (file).
291D0011101GSGroup Separator ^]Information separator: group level, inside a file.
301E0011110RSRecord Separator ^^Information separator: record level, inside a group.
311F0011111USUnit Separator ^_Information separator: the least inclusive level (unit).
1277F1111111DELDelete ^?Used to erase or obliterate a wrong character on perforated tape. RFC 20 notes it is not strictly a control character.

Printable characters (32 to 126)

The same 95 characters, in table order. Names follow RFC 20 section 4.2 (a few reworded, for example "Slant" is the slash).

DecHexBinaryCharName
32200100000SPSpace
33210100001!Exclamation Point
34220100010"Quotation Marks
35230100011#Number Sign
36240100100$Dollar Sign
37250100101%Percent
38260100110&Ampersand
39270100111'Apostrophe
40280101000(Opening Parenthesis
41290101001)Closing Parenthesis
422A0101010*Asterisk
432B0101011+Plus
442C0101100,Comma
452D0101101-Hyphen (Minus)
462E0101110.Period (Decimal Point)
472F0101111/Slant
483001100000Digit 0
493101100011Digit 1
503201100102Digit 2
513301100113Digit 3
523401101004Digit 4
533501101015Digit 5
543601101106Digit 6
553701101117Digit 7
563801110008Digit 8
573901110019Digit 9
583A0111010:Colon
593B0111011;Semicolon
603C0111100<Less Than
613D0111101=Equals
623E0111110>Greater Than
633F0111111?Question Mark
64401000000@Commercial At
65411000001ACapital letter A
66421000010BCapital letter B
67431000011CCapital letter C
68441000100DCapital letter D
69451000101ECapital letter E
70461000110FCapital letter F
71471000111GCapital letter G
72481001000HCapital letter H
73491001001ICapital letter I
744A1001010JCapital letter J
754B1001011KCapital letter K
764C1001100LCapital letter L
774D1001101MCapital letter M
784E1001110NCapital letter N
794F1001111OCapital letter O
80501010000PCapital letter P
81511010001QCapital letter Q
82521010010RCapital letter R
83531010011SCapital letter S
84541010100TCapital letter T
85551010101UCapital letter U
86561010110VCapital letter V
87571010111WCapital letter W
88581011000XCapital letter X
89591011001YCapital letter Y
905A1011010ZCapital letter Z
915B1011011[Opening Bracket
925C1011100\Reverse Slant
935D1011101]Closing Bracket
945E1011110^Circumflex
955F1011111_Underline
96601100000`Grave Accent
97611100001aSmall letter a
98621100010bSmall letter b
99631100011cSmall letter c
100641100100dSmall letter d
101651100101eSmall letter e
102661100110fSmall letter f
103671100111gSmall letter g
104681101000hSmall letter h
105691101001iSmall letter i
1066A1101010jSmall letter j
1076B1101011kSmall letter k
1086C1101100lSmall letter l
1096D1101101mSmall letter m
1106E1101110nSmall letter n
1116F1101111oSmall letter o
112701110000pSmall letter p
113711110001qSmall letter q
114721110010rSmall letter r
115731110011sSmall letter s
116741110100tSmall letter t
117751110101uSmall letter u
118761110110vSmall letter v
119771110111wSmall letter w
120781111000xSmall letter x
121791111001ySmall letter y
1227A1111010zSmall letter z
1237B1111011{Opening Brace
1247C1111100|Vertical Line
1257D1111101}Closing Brace
1267E1111110~Overline (Tilde)

What goes wrong

Real byte sequences that break decoders. The results are the tool's exact output, and the replacement counts for the first four match the examples printed in the Unicode Standard (Tables 3-8 to 3-11), which the tests also run.

C0 AF, an overlong slash

C0 can only begin an overlong two-byte form; AF is then a continuation byte with no lead. Unicode Table 3-8 lists exactly two U+FFFD for C0 AF.

Start from UTF-8 bytes (hex)

Input: C0 AF

Result: 2 errors: U+FFFD U+FFFD

Latin-1 bytes read as UTF-8

The Latin-1 text "é café" is E9 20 63 61 66 E9. E9 would start a three-byte sequence, but the next byte (20) is not a continuation byte, so E9 is replaced alone; the same happens at the end, where E9 is truncated.

Start from UTF-8 bytes (hex)

Input: E9 20 63 61 66 E9

Result: 2 errors: U+FFFD U+0020 U+0063 U+0061 U+0066 U+FFFD

An emoji cut off after three bytes

F0 9F 98 is the first three bytes of F0 9F 98 80. It is the start of a valid sequence, so the standard replaces the whole truncated prefix with a single U+FFFD, not three (Table 3-11).

Start from UTF-8 bytes (hex)

Input: F0 9F 98

Result: 1 error: U+FFFD

A surrogate encoded as three bytes

ED A0 80 is what a naive encoder writes for the surrogate code point U+D800. UTF-8 forbids surrogates (RFC 3629; after ED the second byte must be 80 to 9F), so each of the three bytes becomes its own U+FFFD (Table 3-9).

Start from UTF-8 bytes (hex)

Input: ED A0 80

Result: 3 errors: U+FFFD U+FFFD U+FFFD

A lone surrogate in text

JavaScript strings can hold a lone surrogate code unit, but it is not a Unicode scalar value, so it cannot be written in UTF-8. TextEncoder (which takes a USVString) substitutes U+FFFD, bytes EF BF BD, and so does this tool; it also counts it so you notice.

Start from Text

Input: a\uD800b

Result: 61 EF BF BD 62 (1 lone surrogate replaced by EF BF BD)

A code point beyond Unicode

Start from Code points

Input: U+110000

Error: Code point 1 (U+110000) is above U+10FFFF, the last Unicode code point.

Mojibake in the other direction

The UTF-8 bytes of é are C3 A9. Reading them as windows-1252 or Latin-1 shows é; re-encoding that text as UTF-8 gives C3 83 C2 A9, four bytes, which is the classic "double-encoded" form. In the browser, new TextDecoder("windows-1252").decode(new Uint8Array([0xC3, 0xA9])) returns é.

Limits & gotchas

  • UTF-8 only. The tool encodes and decodes UTF-8 (and shows UTF-16 code units for reference). Latin-1, windows-1252, Shift_JIS and other legacy encodings are not converted; the page only explains how reading UTF-8 bytes with them goes wrong.
  • Code points, not glyphs. The tool lists code points. What a reader sees as one "character" may be several (a letter plus a combining accent, or an emoji sequence joined with U+200D). Each code point is shown separately on purpose.
  • No normalisation. The tool does not convert between composed and decomposed forms; it shows the bytes of exactly what you typed. Names and categories are not looked up for every code point; only the ASCII table has names (from RFC 20).
  • Input size. At most 20,000 characters or 20,000 bytes. The per-character table shows the first 200 characters and the byte-by-byte table the first 300 entries.
  • Pasted text may change. Browsers and editors can convert line endings or replace characters when you paste. If a byte count looks off, paste into the "UTF-8 bytes" mode instead, or check for a hidden trailing newline (LF, 0A).
  • Replacement follows Unicode's recommended practice. Other decoders may produce a different number of U+FFFD characters for the same bad bytes; the Unicode Standard says it does not require this practice for conformance. This tool's decoder agrees with the browser's TextDecoder on 20,000 random byte strings in the tests.
  • Browser. Tested in a current Chrome only.

FAQ

Why does one emoji take four bytes, and why does my string length say 2?

UTF-8 uses 1 byte for U+0000 to U+007F, 2 bytes up to U+07FF, 3 bytes up to U+FFFF and 4 bytes above that (RFC 3629). The grinning face U+1F600 is above U+FFFF, so it is the four bytes F0 9F 98 80. JavaScript strings are sequences of 16-bit code units (ECMAScript), and a code point above U+FFFF takes two of them (a surrogate pair, D83D DE00 here), so "😀".length is 2 while [..."😀"].length is 1.

What is the difference between ASCII and UTF-8?

ASCII is a 7-bit code with 128 characters (RFC 20): codes 0 to 31 and 127 are control characters, 32 to 126 are space and printable characters. UTF-8 encodes all of Unicode and is designed so that code points 0 to 127 use exactly the ASCII byte values, so plain ASCII text is already valid UTF-8. Every other character takes two to four bytes, each with the top bit set. The names "ASCII" and "ANSI" are often used loosely for 8-bit code pages such as Latin-1 or windows-1252, which are neither.

Why do I see é as é (or a question-mark box)?

Because bytes were read with the wrong encoding. UTF-8 stores é (U+00E9) as the two bytes C3 A9. Read as Latin-1 or windows-1252, C3 is "Ã" and A9 is "©", giving "é". The reverse mistake is the bytes E9 20 63 61 66 E9 (the Latin-1 text "é café"): they are not valid UTF-8, so a decoder shows the replacement character U+FFFD in place of each bad byte. The tool's "What goes wrong" section shows the exact byte-by-byte result.

Why are there two different byte sequences for "café"?

Unicode can write é as the single code point U+00E9 (bytes C3 A9) or as the letter e followed by the combining acute accent U+0301 (bytes 65 CC 81). They look the same and the Unicode database records the first as canonically equivalent to the second (UnicodeData.txt lists U+00E9 as decomposing to 0065 0301). But as bytes they differ, so a hash, a file name or a string comparison sees different data until the text is normalised. The tool shows both for "café".

What does C0 AF mean, and why is it rejected?

C0 starts a two-byte sequence that could only encode a code point below U+0080, which UTF-8 must write as one byte. C0 AF would be an "overlong" encoding of "/" (U+002F). RFC 3629 says decoders must not accept such sequences: they are a security risk, because a naive decoder could let a "/" or NUL slip past a filter that only looks for the one-byte form. Unicode 16 Table 3-8 shows the replacement for C0 AF as two U+FFFD characters, and the tool produces exactly that.

Sources

Every document above was opened and read on 2026-10-02. Documentation changes; if a page here disagrees with the current docs, trust the docs and tell us.

The ASCII table's names and control-character meanings are shortened and reworded from RFC 20 sections 4 and 5.2. Caret notation and C escapes (such as \n) are common conventions, not part of RFC 20, and are labelled so.