ASCII is a 7-bit character code, Unicode defines the characters and code points computers can represent, and UTF-8 is a way to encode Unicode code points as bytes. Nic Barker’s “UTF-8, Explained Simply,” covered by Hackaday on January 22, 2026, walks through how those pieces fit together, from ASCII compatibility to multi-byte characters and grapheme clusters.
ASCII, Unicode, and UTF-8 are different things
Text requires more than storing a character’s appearance: a computer needs a numeric value for it and a rule for representing that value in memory or in a file. ASCII, Unicode, and UTF-8 address different parts of that problem.
- ASCII is a character code with 128 possible values, numbered 0 through 127. It uses seven bits and covers basic Latin letters, digits, punctuation, and control codes.
- Unicode is the standard repertoire of characters and their numeric identifiers, called code points. A code point is written in the form U+ followed by hexadecimal digits, such as U+0041 for “A.”
- UTF-8 is an encoding form that converts Unicode code points into sequences of bytes. It is not a separate character repertoire.
This distinction matters when diagnosing garbled text: the characters a system intends to represent, the encoding used to store them, and the software interpreting those bytes must agree. Barker’s presentation is described in Hackaday’s coverage; the formal rules are in the Unicode 16.0.0 Core Specification.
How UTF-8 turns code points into bytes
UTF-8 uses one, two, three, or four bytes for a code point. The bit pattern of the first byte indicates how long the sequence is; any following bytes in that sequence are continuation bytes, identified by a distinct leading-bit pattern. This lets software distinguish a sequence’s start from its continuation bytes rather than treating every byte as an independent character.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware match#1 Best Overall
| UTF-8 sequence length | Code point range | Byte pattern |
|---|---|---|
| 1 byte | U+0000–U+007F | 0xxxxxxx |
| 2 bytes | U+0080–U+07FF | 110xxxxx 10xxxxxx |
| 3 bytes | U+0800–U+FFFF | 1110xxxx 10xxxxxx 10xxxxxx |
| 4 bytes | U+10000–U+10FFFF | 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx |
The x marks bits carrying the code point’s value. The table summarizes the UTF-8 byte patterns described by the Unicode 16.0.0 Core Specification. The Unicode Standard’s term “code unit” means the basic unit an encoding form uses; UTF-8’s code units are 8-bit bytes.
Leading and continuation bytes
A leading byte begins a UTF-8 sequence and signals how many bytes it contains. A continuation byte belongs to a sequence already begun. Because their roles have distinguishable patterns, a decoder can find a character boundary even if it starts at an arbitrary byte position: UTF-8 is self-synchronizing, and the Unicode specification says it can search back at most four bytes to find a character start. That property helps with parsing and recovery; it does not make every byte position a valid place to begin decoding.
Rank #2
Why some emoji use four bytes
A code point above U+FFFF is supplementary and takes four bytes in UTF-8. Many emoji are in supplementary Unicode ranges, so an individual emoji code point can require four bytes. But “emoji” does not mean “one code point”: some displayed emoji are sequences of multiple code points joined or modified together, so their UTF-8 representation can use more than four bytes.
Why UTF-8 is compatible with ASCII
UTF-8 encodes every code point from U+0000 through U+007F as the same single byte, 0x00 through 0x7F, used by ASCII. As a result, ASCII text is byte-for-byte unchanged when represented in UTF-8. This transparency supports compatibility with software and formats that expect ASCII characters in that range, while allowing UTF-8 to represent the broader Unicode repertoire. The Unicode specification documents this rule in its UTF-8 encoding definition.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #3
A character is not always a code point
A Unicode code point is a numeric identifier, not necessarily a complete user-perceived character. A visible character can be made from multiple code points—for example, a base letter followed by a combining mark—or an emoji sequence. A grapheme cluster is a sequence of one or more code points treated as a single unit for text interaction, such as cursor movement or selection.
Therefore, counting bytes does not count code points, and counting code points does not always count the characters a reader sees. Software that moves a cursor, measures displayed characters, or truncates text should use Unicode-aware text handling rather than assume one byte or one code point equals one visible character.
UTF-8, UTF-16, and UTF-32 compared
| Encoding form | Storage unit and length | Supplementary code points | Practical considerations |
|---|---|---|---|
| UTF-8 | 8-bit bytes; one to four bytes per code point | Four bytes | ASCII-transparent, self-synchronizing, and has no byte-order issue. Often preferred for HTML and similar internet protocols. |
| UTF-16 | 16-bit code units; one or two per code point | Two code units, called a surrogate pair | Variable-width handling requires care. Its storage can be smaller than UTF-8 for some Asian writing systems. |
| UTF-32 | 32-bit code units; one per code point | One code unit | Fixed-width code points can simplify direct indexing, but use more storage than narrower forms. |
These are encoding forms for Unicode, not different character sets. The comparison follows the Unicode 16.0.0 Core Specification, which notes that UTF-8 can be smaller for ASCII-heavy and Western-language text, while UTF-16 can be smaller for some Asian writing systems.
Should you use UTF-8 or UTF-16?
For web pages and general text interchange, UTF-8 is usually the practical default: ASCII bytes remain unchanged, byte boundaries are recoverable, and the Unicode Consortium describes UTF-8 as typically preferred for HTML and similar protocols, particularly on the internet. Use UTF-16 when a system or protocol specifically requires it, or when its characteristics suit a known workload and its software handles surrogate pairs correctly. Neither form makes indexing by visible character automatic; text operations still need to account for code points and grapheme clusters.
Best Value
- Used Book in Good Condition
What a UTF-8 BOM does—and when to use one
A byte order mark (BOM) is a signature that may appear at the beginning of a text file. UTF-8 is a byte sequence, so unlike UTF-16 or UTF-32 it has no endianness to signal. In UTF-8, the BOM is therefore optional metadata, not a byte-order instruction, as explained in the Unicode FAQ.
Whether to include one depends on the file format and the programs that consume it. A BOM can interfere when a format expects specific ASCII characters at the very beginning of the file; the Unicode FAQ gives a Unix shell script beginning with the “#!” marker as an example. Follow the requirements of the target protocol or application rather than adding a BOM to every UTF-8 file by default.
Learn more about Barker’s explanation
Hackaday’s January 22, 2026 report points to Nic Barker’s presentation “UTF-8, Explained Simply” and to the follow-up reading “Understanding And Using Unicode.”
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →

