Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

ASCII and EBCDIC assign different byte values to characters; ISO/IEC 646 carried the 7-bit character-set tradition into an international standard, while Unicode created a much larger shared repertoire. For new systems and open interchange, Unicode encoded as UTF-8 is generally the practical choice. EBCDIC and UTF-EBCDIC still matter when data must work with EBCDIC-based IBM host environments.

How the character codes developed

Early character codes were designed around the machines and communication systems they served. ASCII standardized a compact 7-bit set. EBCDIC, developed by IBM for its systems, used 8-bit bytes with a different arrangement. ISO/IEC 646 established an international standard in the same 7-bit tradition. Unicode later addressed a broader problem: representing text from scripts and symbol systems around the world in a coordinated repertoire.

  • 1963: The U.S. chronology records approval of ASA X3.4-1963 in June. A 1965 revision assigned characters to all 128 positions and incorporated compatibility changes connected to ISO and CCITT work.
  • 1960s onward: IBM’s EBCDIC family served mainframe systems, using 8-bit bytes and assignments that differ from ASCII.
  • 1991: ISO published ISO/IEC 646:1991, a 128-character, 7-bit standard for Latin-script information interchange.
  • Late 1980s to 1991: Unicode work grew from discussions involving Xerox and Apple engineers. Unicode, Inc. was incorporated in California in January 1991.
  • 1991 to 1993: Unicode and ISO worked toward a shared universal repertoire. Unicode’s account of the standards’ history says ISO/IEC 10646-1:1993 and Unicode 1.1 had precisely the same encoded characters and names.

What ASCII, EBCDIC, ISO 646, and Unicode mean

System What it defines Where it fits
ASCII A 7-bit code with 128 numeric values. IBM describes 33 values as reserved for special functions. A foundational character set whose design influenced many later systems. The values do not define every character needed for writing worldwide.
EBCDIC An IBM-designed family of 8-bit character sets. Its byte assignments and character ordering differ from ASCII; particular assignments can depend on the EBCDIC code page. Especially associated with IBM mainframes and related environments. Its organization reflects the constraints of punch-card and mainframe design.
ISO/IEC 646 The international 7-bit standard lineage. The 1991 edition specifies 128 control and graphic characters for information interchange. Standardized and localized the ASCII-era repertoire for international use. Some national variants changed selected character positions, so a reference to “ISO 646” may need a specific variant.
Unicode A coordinated character repertoire with code points, alongside encoding forms including UTF-8, UTF-16, and UTF-32. Designed to cover scripts and symbols well beyond ASCII’s Latin-focused set and to support broad text interchange.

Unicode is not simply another 8-bit character set. A code point identifies a character in the repertoire; UTF-8, UTF-16, and UTF-32 are different ways to encode Unicode code points as units for storage or transmission. The Unicode Standard describes its design as drawing on ASCII’s simplicity and consistency while extending far beyond ASCII’s limited ability to encode the Latin alphabet.

Is ISO 646 the same as ASCII?

Not exactly. They belong to the same 7-bit standards lineage, and the international reference version of ISO 646 is closely associated with ASCII. But ISO/IEC 646 also provided for national variants that reassigned certain positions to suit local requirements. A system that assumes every ISO 646 variant has identical character assignments can therefore misread data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Also distinguish ASCII itself from later 8-bit character sets. “ASCII” is often used informally to mean a compatible legacy encoding, but the original ASCII code has 128 values. ISO-8859 families and vendor code pages add or assign further characters; they are not just “8-bit ASCII.” The IANA character-set registry lists aliases for the ASCII lineage, including ANSI_X3.4-1968, ANSI X3.4-1986, and ISO_646.irv:1991.

Why Unicode was needed

A collection of small, machine-oriented character sets cannot provide one dependable representation for text spanning many scripts and symbols. The same byte may mean different things under different code pages, and a character needed in one language may be absent from another system’s set. Unicode was organized to provide a shared repertoire rather than require every exchange to negotiate among incompatible local assignments.

Unicode and ISO/IEC 10646 converged on a synchronized repertoire and code-point assignments. This does not mean every document has the same byte sequence: the chosen Unicode encoding form still determines how code points are represented in bytes or code units. It means the standards share the character identities and assignments that encodings represent.

UTF-8, UTF-16, UTF-32, and UTF-EBCDIC

UTF-8

UTF-8 encodes Unicode using variable-length byte sequences. Its key practical advantage for legacy text is that the Unicode values U+0000 through U+007F use the same bytes as ASCII. ASCII text can therefore remain byte-compatible when encoded as UTF-8, while UTF-8 can also represent the wider Unicode repertoire.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

UTF-16 and UTF-32

UTF-16 and UTF-32 are also Unicode encoding forms, but they organize code units differently from UTF-8. They are alternatives for representing Unicode, not different character repertoires. Software must know which form it is reading rather than infer character meaning from an arbitrary sequence of bytes.

UTF-EBCDIC

UTF-EBCDIC is a specialized transformation intended for EBCDIC-oriented environments. It creates an intermediate variable-length sequence from Unicode scalar values, then applies a reversible byte mapping that follows EBCDIC conventions for controls and invariant characters. IBM documentation notes that base EBCDIC and control characters can remain single-byte values while other characters use multiple bytes, which can help some legacy applications preserve data they do not recognize.

Unicode Technical Report #16 explicitly says UTF-EBCDIC and its intermediate form, UTF-8-Mod, are not intended for open interchange; it describes them as useful in homogeneous EBCDIC systems and networks. That makes UTF-EBCDIC a specialist bridge, not the normal choice for new Internet-facing or cross-platform text exchange.

Which encoding should you use?

Situation Practical choice Reason
New software, files, or Internet interchange Unicode with UTF-8 It supports the Unicode repertoire and preserves ASCII byte compatibility for the ASCII range.
Existing IBM host data or an EBCDIC-based application boundary The required EBCDIC code page; use UTF-EBCDIC only where the EBCDIC environment specifically calls for it Byte mappings and host expectations differ from ASCII-compatible encodings, so compatibility at the boundary matters.
Older system labeled “ASCII” or “ISO 646” Identify the actual variant or code page before converting Those labels can refer to different mappings, especially when national variants or 8-bit extensions are involved.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to avoid corrupting text during conversion

Treat the encoding and code page as part of the data’s meaning, not as a cosmetic setting. A byte stream alone does not reliably reveal which character mapping produced it. Converting bytes under the wrong mapping can change punctuation, control characters, or national characters even when ordinary letters appear intact.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Find the source mapping. Check the producing application, file or protocol metadata, system configuration, and—on IBM host data—the specific EBCDIC code page. Do not assume that “ASCII” names the complete encoding.
  2. Decode before re-encoding. Interpret the original bytes using the identified source mapping, then encode the resulting characters as the target format, commonly UTF-8 for open interchange.
  3. Check conversion behavior. Confirm how the converter handles characters absent from the source or target repertoire, invalid byte sequences, and control values. Avoid silent replacement when preserving the original data matters.
  4. Validate representative records. Inspect punctuation, language-specific letters, controls, and any bytes that are meaningful to the receiving application. Test both directions if data will travel back to the host.

These checks are particularly important at system boundaries: two formats may both use one byte for a character without assigning that character the same byte value.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.