Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Encoding is the rule that turns text values into bytes for storage or transmission—and lets software turn those bytes back into text. Unicode provides the shared set of text values; UTF-8, UTF-16, and UTF-32 are different ways to represent them. For new web content and interchange formats, UTF-8 is generally the right choice.

What does encoding mean in computing?

An encoding defines a mapping between a sequence of text values and a sequence of bytes. An encoder applies that mapping when writing or sending data; a decoder applies the corresponding mapping when reading or receiving it. As the W3C Encoding specification puts it, “An encoding defines a mapping from a scalar value sequence to a byte sequence (and vice versa).”

Text is not stored as the appearance of letters on screen. A Unicode code point is a numeric value for a character in the standard’s repertoire, and an encoding specifies how such values are represented in code units and bytes. The application then interprets those values and displays them using available fonts and rendering rules.

Encoding is not encryption or compression: it does not conceal text or necessarily make it smaller. It establishes how values correspond to bytes so different software can interpret the same data consistently.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Unicode and UTF-8 are not the same thing

Unicode is the universal character encoding standard for written characters and text. It supplies the shared repertoire and code points. UTF-8, UTF-16, and UTF-32 are encoding forms that represent Unicode values in different-sized code units. They can all represent the full Unicode range; they differ in representation, not in the character set they support.

A code point is not always identical to one visible symbol: some displayed text can involve multiple code points. The encoding determines the bytes for the code point sequence, not how a font draws the result.

How UTF-8, UTF-16, and UTF-32 differ

Encoding form Code-unit size and length ASCII compatibility Interchange considerations
UTF-8 Variable length: one to four 8-bit code units per encoded value. ASCII characters retain their familiar byte values. W3C and WHATWG identify UTF-8 as the appropriate choice for Unicode interchange; W3C requires new protocols and formats that expose an encoding label to use UTF-8 exclusively.
UTF-16 One or two 16-bit code units per encoded value. Not ASCII-compatible at the byte level in the way UTF-8 is. Represents the full Unicode range, but byte-level interchange must agree on the encoding and byte interpretation.
UTF-32 One 32-bit code unit per encoded value. Not ASCII-compatible at the byte level in the way UTF-8 is. Represents the full Unicode range; storage and runtime trade-offs depend on the data and implementation.

These widths describe code units, not a universal file size or performance ranking. A file’s size depends on its text and encoding; memory use and speed also depend on the software implementation. UTF-8 often uses one byte for ASCII text and more for other values, while UTF-16 uses one or two 16-bit units and UTF-32 uses a 32-bit unit for each encoded value.

Should you use UTF-8 or UTF-16?

For new web pages, APIs, and files intended to move between systems, choose UTF-8 unless a specific protocol, format, or runtime requirement calls for something else. Its ASCII-compatible byte values ease interoperability with software built around ASCII, while its variable-length representation covers the full Unicode range. W3C calls UTF-8 the most appropriate encoding for interchange of Unicode, and WHATWG likewise identifies it as the appropriate interchange encoding.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

UTF-16 can be appropriate where an existing interface or system specifically expects it. UTF-32 may fit a particular internal representation, but its fixed 32-bit code unit should not be mistaken for a universally faster or simpler interchange format. The standards define the encodings; they do not establish a one-size-fits-all memory or speed outcome.

Why does text become garbled after decoding?

Garbled text often means the bytes were decoded using a different encoding from the one used to create them. The bytes themselves may be intact, but the decoder maps them to the wrong values. Re-decoding the same bytes with the correct encoding can restore the intended text; changing the displayed font will not fix a byte-decoding mismatch.

  1. Identify the producer’s encoding. Check the protocol header, file metadata, or explicit format declaration from the system that created or sent the bytes.
  2. Configure the consumer to use that encoding. The decoder must match the producer’s actual encoding, not merely a guessed setting.
  3. Check for invalid sequences. A byte sequence may not be valid under the declared encoding, or the data may have been altered or truncated.
  4. Choose an error policy deliberately. A replacement mode substitutes a replacement value for malformed input, which can keep processing but obscure the original problem. Fatal handling reports the decoding error instead, making it easier to detect invalid data.

When the bytes have already been replaced or lost, selecting a different decoder cannot reliably reconstruct the original text. Preserve the original byte sequence when diagnosing the problem.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What encoding should new formats declare?

For new web and interchange formats, declare and use UTF-8 consistently. W3C specifies UTF-8 for new protocols and formats that expose an encoding label, and the WHATWG Encoding Standard defines browser-facing encoding behavior and JavaScript APIs. The key is not just picking an encoding: producers and consumers must agree on it, and malformed input needs a deliberate handling policy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.