Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

To delete every character in a Python string that ASCII cannot represent, encode it as ASCII with the ignore error handler, then decode the result back into a string:

text = "café — 東京"
clean = text.encode("ascii", "ignore").decode("ascii")
print(clean)  # caf 

This removes é, the em dash, and the Japanese characters. It does not convert them to similar-looking or equivalent text.

What the encode-and-decode method does

Python strings are Unicode text, while ASCII can represent only a limited set of characters. str.encode("ascii", "ignore") encodes the characters ASCII can represent and silently omits those it cannot. The result of encode() is bytes, so .decode("ascii") converts those bytes back into a Python str.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Python Software Foundation explains in its Unicode HOWTO that str.encode() returns a bytes representation of a Unicode string in the requested encoding. The full expression is useful when the cleaned result needs to remain a string:

text = "café — 東京"
encoded = text.encode("ascii", "ignore")
clean = encoded.decode("ascii")

print(encoded)  # b'caf  '
print(clean)    # caf  

In the output, the spaces around the removed characters remain because they were ASCII characters in the input.

Deletion is not transliteration

The ignore handler deletes characters that cannot be encoded; it does not replace them with approximate ASCII spellings. For example, it will not turn é into e, 東京 into Tokyo, or ß into ss. If you need those kinds of substitutions, choose a transliteration library or define explicit mappings appropriate to the text. Do not use this encoding approach when preserving the original meaning or spelling matters.

Remove non-ASCII characters without encoding

A character-map translation or a simple character filter can delete non-ASCII characters directly from a string. Python’s character-map translation API permits mapping a character’s code point to None to delete it; characters omitted from the map pass through unchanged. See the Python Unicode C API documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a translation table for selected characters

This table is built from the non-ASCII characters in the example input:

text = "café — 東京"
remove_non_ascii = {ord(ch): None for ch in text if not ch.isascii()}
clean = text.translate(remove_non_ascii)

print(clean)  # caf  

Because the table is derived from text, build or define an appropriate map for each input or for the specific characters you want to handle. Characters not included in a translation table are left alone.

Use a character filter for a reusable helper

def remove_non_ascii(text: str) -> str:
    return "".join(ch for ch in text if ch.isascii())

This keeps each character for which isascii() is true and joins those characters into a new string.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choose how encoding failures should appear

ignore is only one of Python’s encoding error handlers. Choose based on whether your output should discard, mark, or expose characters ASCII cannot encode:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Handler Effect when encoding as ASCII Use it when
ignore Deletes unencodable characters. You intentionally want only the encodable characters and accept information loss.
replace Inserts ? for encoding errors. You want a visible marker in place of omitted characters.
backslashreplace Writes escaped code-point forms for encoding errors. You want to see which characters could not be encoded.
xmlcharrefreplace Writes numeric character references for unencodable characters. You need those characters represented as numeric references.

For example, these alternatives still produce bytes because they are passed to encode(); decode the result if you need a string:

text = "café"
marked = text.encode("ascii", "replace").decode("ascii")
escaped = text.encode("ascii", "backslashreplace").decode("ascii")
reference = text.encode("ascii", "xmlcharrefreplace").decode("ascii")

Python documents these handlers in its codecs reference. Select the behavior deliberately: dropping or substituting characters can change data, so retain the original string when that loss would be unacceptable.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.