Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

ANSEL is the Extended Latin Alphabet Coded Character Set for Bibliographic Use, a character set used in bibliographic data. In MARC-8, ASCII graphics are the default G0 set and ANSEL graphics are the default G1 set. ANSEL supplies extended Latin letters, symbols, and combining marks that complement ASCII; it is not another name for Unicode or for the whole MARC-8 encoding environment.

What ANSEL means

The name ANSEL refers to the “Extended Latin Alphabet Coded Character Set for Bibliographic Use,” identified as ANSI Z39.47 in the Library of Congress’s MARC 21 character-set overview. It is intended for bibliographic character data, including extended Latin letters, symbols, and combining marks.

ANSEL is one component of MARC-8, not a complete encoding by itself. MARC-8 combines graphic character sets and defines how they are used in a record. The Library of Congress describes ASCII graphics as the default G0 set and ANSEL graphics as the default G1 set in its MARC-8 Encoding Environment.

How ANSEL works in MARC-8

G0 and G1 are graphic-set designations in the MARC-8 environment. ASCII is the default G0 set; ANSEL is the default G1 set. The MARC-8 rules invoke ANSEL G1 for code values A1 through FE hexadecimal. Thus, a byte value cannot always be interpreted as a character on its own: its meaning depends on the encoding environment and the active set.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a specific character, consult the Library of Congress Extended Latin (ANSEL) table. Its columns distinguish MARC-8 code values from UCS/Unicode code points and UTF-8 representations, alongside character forms and names. Use the MARC-8 value in the MARC-8 context; do not treat it as though it were automatically the character’s Unicode value or its UTF-8 bytes. The Library of Congress also says to use only MARC-8 code points included in its official tables; its MARC-8 Code Tables overview describes the published mappings.

How to tell whether a MARC 21 record uses ANSEL or Unicode

  1. Check Leader position 9. It identifies the character coding scheme for that record as MARC-8 or Unicode, as explained in the Library of Congress’s character-set introduction.
  2. If it is MARC-8, interpret extended characters in the MARC-8 context. Use the official ANSEL mapping table and account for the applicable graphic-set context rather than reading a code value as a Unicode code point.
  3. For character-set details in a non-Unicode record, consult field 066 where applicable. The field communicates character-set information; the Library of Congress documentation says default ANSEL need not be identified when it is the primary extended set. See the MARC 21 Authority Data field 066 documentation and General Character Set Issues.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

ANSEL and Unicode are different encoding environments

MARC 21 records use one character coding scheme at a time: MARC-8 or Unicode, as indicated in Leader position 9. ANSEL is the default extended G1 graphic set within MARC-8; Unicode is the alternative character coding environment for a record. The Library of Congress notes that conversion to Unicode has taken place in many large library systems, but does not give a count or establish a present-day prevalence figure.

When converting a MARC-8 record, preserve the character’s identity by using the official MARC-8-to-UCS/Unicode mappings rather than assuming that a MARC-8 code, a Unicode code point, and UTF-8 bytes are interchangeable. The mapping table records changes to particular entries, including additions for Eszett and Euro in June 2004 and mapping changes for ligature, double tilde, and Alif during 2004–2005. Those dates describe the change history noted in the table, not its latest update date.

What ANSEL does not tell you

  • It does not mean every MARC-8 character is ANSEL. MARC-8 uses multiple graphic character sets.
  • It does not identify Unicode or UTF-8. The official table lists distinct MARC-8, UCS/Unicode, and UTF-8 representations.
  • It does not establish current usage levels. The Library of Congress standards and tables define encoding and mappings, but do not report current counts of ANSEL records, institutions, or software products.
  • It does not identify a particular converter. The cited documentation supplies mappings and rules, not a current vendor-tool comparison or tested converter recommendation.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.