Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

ASCII is a 7-bit character code, Unicode defines the characters and code points computers can represent, and UTF-8 is a way to encode Unicode code points as bytes. Nic Barker’s “UTF-8, Explained Simply,” covered by Hackaday on January 22, 2026, walks through how those pieces fit together, from ASCII compatibility to multi-byte characters and grapheme clusters.

ASCII, Unicode, and UTF-8 are different things

Text requires more than storing a character’s appearance: a computer needs a numeric value for it and a rule for representing that value in memory or in a file. ASCII, Unicode, and UTF-8 address different parts of that problem.

  • ASCII is a character code with 128 possible values, numbered 0 through 127. It uses seven bits and covers basic Latin letters, digits, punctuation, and control codes.
  • Unicode is the standard repertoire of characters and their numeric identifiers, called code points. A code point is written in the form U+ followed by hexadecimal digits, such as U+0041 for “A.”
  • UTF-8 is an encoding form that converts Unicode code points into sequences of bytes. It is not a separate character repertoire.

This distinction matters when diagnosing garbled text: the characters a system intends to represent, the encoding used to store them, and the software interpreting those bytes must agree. Barker’s presentation is described in Hackaday’s coverage; the formal rules are in the Unicode 16.0.0 Core Specification.

How UTF-8 turns code points into bytes

UTF-8 uses one, two, three, or four bytes for a code point. The bit pattern of the first byte indicates how long the sequence is; any following bytes in that sequence are continuation bytes, identified by a distinct leading-bit pattern. This lets software distinguish a sequence’s start from its continuation bytes rather than treating every byte as an independent character.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
UTF-8 sequence length Code point range Byte pattern
1 byte U+0000–U+007F 0xxxxxxx
2 bytes U+0080–U+07FF 110xxxxx 10xxxxxx
3 bytes U+0800–U+FFFF 1110xxxx 10xxxxxx 10xxxxxx
4 bytes U+10000–U+10FFFF 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx

The x marks bits carrying the code point’s value. The table summarizes the UTF-8 byte patterns described by the Unicode 16.0.0 Core Specification. The Unicode Standard’s term “code unit” means the basic unit an encoding form uses; UTF-8’s code units are 8-bit bytes.

Leading and continuation bytes

A leading byte begins a UTF-8 sequence and signals how many bytes it contains. A continuation byte belongs to a sequence already begun. Because their roles have distinguishable patterns, a decoder can find a character boundary even if it starts at an arbitrary byte position: UTF-8 is self-synchronizing, and the Unicode specification says it can search back at most four bytes to find a character start. That property helps with parsing and recovery; it does not make every byte position a valid place to begin decoding.

Why some emoji use four bytes

A code point above U+FFFF is supplementary and takes four bytes in UTF-8. Many emoji are in supplementary Unicode ranges, so an individual emoji code point can require four bytes. But “emoji” does not mean “one code point”: some displayed emoji are sequences of multiple code points joined or modified together, so their UTF-8 representation can use more than four bytes.

Why UTF-8 is compatible with ASCII

UTF-8 encodes every code point from U+0000 through U+007F as the same single byte, 0x00 through 0x7F, used by ASCII. As a result, ASCII text is byte-for-byte unchanged when represented in UTF-8. This transparency supports compatibility with software and formats that expect ASCII characters in that range, while allowing UTF-8 to represent the broader Unicode repertoire. The Unicode specification documents this rule in its UTF-8 encoding definition.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A character is not always a code point

A Unicode code point is a numeric identifier, not necessarily a complete user-perceived character. A visible character can be made from multiple code points—for example, a base letter followed by a combining mark—or an emoji sequence. A grapheme cluster is a sequence of one or more code points treated as a single unit for text interaction, such as cursor movement or selection.

Therefore, counting bytes does not count code points, and counting code points does not always count the characters a reader sees. Software that moves a cursor, measures displayed characters, or truncates text should use Unicode-aware text handling rather than assume one byte or one code point equals one visible character.

UTF-8, UTF-16, and UTF-32 compared

Encoding form Storage unit and length Supplementary code points Practical considerations
UTF-8 8-bit bytes; one to four bytes per code point Four bytes ASCII-transparent, self-synchronizing, and has no byte-order issue. Often preferred for HTML and similar internet protocols.
UTF-16 16-bit code units; one or two per code point Two code units, called a surrogate pair Variable-width handling requires care. Its storage can be smaller than UTF-8 for some Asian writing systems.
UTF-32 32-bit code units; one per code point One code unit Fixed-width code points can simplify direct indexing, but use more storage than narrower forms.

These are encoding forms for Unicode, not different character sets. The comparison follows the Unicode 16.0.0 Core Specification, which notes that UTF-8 can be smaller for ASCII-heavy and Western-language text, while UTF-16 can be smaller for some Asian writing systems.

Should you use UTF-8 or UTF-16?

For web pages and general text interchange, UTF-8 is usually the practical default: ASCII bytes remain unchanged, byte boundaries are recoverable, and the Unicode Consortium describes UTF-8 as typically preferred for HTML and similar protocols, particularly on the internet. Use UTF-16 when a system or protocol specifically requires it, or when its characteristics suit a known workload and its software handles surrogate pairs correctly. Neither form makes indexing by visible character automatic; text operations still need to account for code points and grapheme clusters.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What a UTF-8 BOM does—and when to use one

A byte order mark (BOM) is a signature that may appear at the beginning of a text file. UTF-8 is a byte sequence, so unlike UTF-16 or UTF-32 it has no endianness to signal. In UTF-8, the BOM is therefore optional metadata, not a byte-order instruction, as explained in the Unicode FAQ.

Whether to include one depends on the file format and the programs that consume it. A BOM can interfere when a format expects specific ASCII characters at the very beginning of the file; the Unicode FAQ gives a Unix shell script beginning with the “#!” marker as an example. Follow the requirements of the target protocol or application rather than adding a BOM to every UTF-8 file by default.

Learn more about Barker’s explanation

Hackaday’s January 22, 2026 report points to Nic Barker’s presentation “UTF-8, Explained Simply” and to the follow-up reading “Understanding And Using Unicode.”

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.