Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A binary string is just a sequence of bits or bytes; it does not identify readable text by itself. To display it as text, software needs an encoding rule that maps values to characters. Choose a different rule and the same bytes can show different characters—or fail to decode as text.

What a binary string contains—and what it does not

Bits record values, and text data is often grouped into bytes, or octets. Neither bits nor bytes inherently say “this is text.” A byte sequence might represent text, an image, compressed data, or something else; its meaning comes from the format and context that define how to interpret it.

CBOR makes this distinction explicit: it defines byte strings for unstructured bytes separately from text strings, which contain UTF-8 text. Trying to display arbitrary bytes does not turn them into text. RFC 8949: Concise Binary Object Representation

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Unicode and an encoding are different layers

Unicode assigns code points to characters. An encoding form specifies how Unicode text is represented as code units. UTF-8, UTF-16, and UTF-32 use different representations, so the bytes or code units for a text depend on the encoding form. When that representation is serialized as bytes, details such as byte order or a byte-order mark can also matter. Unicode FAQ: UTF-8, UTF-16, UTF-32 & BOM

In short, Unicode identifies characters; UTF-8, UTF-16, and UTF-32 describe ways to encode Unicode text. Unicode is not itself one universal byte layout.

Why the same bytes can show different characters

Consider the character “é” (U+00E9). Its UTF-8 representation is the two bytes C3 A9. If software instead decodes those byte values as Latin-1, they display as “é”. The bytes did not change; the decoding rule did. A mismatch can produce garbled text, and some byte sequences are invalid under a particular encoding. RFC 3629: UTF-8

How UTF-8, UTF-16, and UTF-32 represent text

Encoding form Units used What to know when reading serialized data
UTF-8 One to four 8-bit bytes per Unicode scalar value. Values U+0000 through U+007F use the same single-byte values as ASCII. UTF-8 does not need a byte-order choice because its units are bytes.
UTF-16 One or two 16-bit code units per Unicode scalar value. When serialized as bytes, byte order can matter; a byte-order mark may be relevant.
UTF-32 One 32-bit code unit per Unicode scalar value. When serialized as bytes, byte order can matter; a byte-order mark may be relevant.

These are representation differences, not different sets of characters. UTF-8 is variable-width, and its ASCII compatibility means ASCII text has the same byte values in ASCII and UTF-8. The Unicode Consortium describes UTF-8 as “the byte-oriented encoding form of Unicode.” Unicode FAQ: UTF-8, UTF-16, UTF-32 & BOM · Unicode 16.0.0, Chapter 2

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to tell whether bytes are UTF-8

Use the file format, protocol, or application that produced the bytes as your first clue: a defined format may specify its text encoding. If that context is missing, check whether the bytes form valid UTF-8, but treat validity as evidence rather than proof. A sequence that is valid UTF-8 could also be meaningful under another interpretation; bytes alone do not carry a label naming their intended encoding.

If decoding as UTF-8 fails, do not silently assume that another encoding is correct. Check the source or format specification, then choose a compatible decoder. A wrong guess may turn valid text into mojibake or make an invalid sequence appear to be something else.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why text looks garbled after opening a file

Garbled characters commonly indicate that the application decoded the file using a different encoding than the one used to write it. The file may also not contain text at all. Confirm what kind of data the file is and what encoding its format or producer specifies before changing or re-saving it; saving with the wrong interpretation can preserve the wrong characters rather than repair the original bytes.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.