Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Apache Spark’s normalize function converts strings among Unicode normalization forms. Use it to make canonically equivalent text consistent for comparisons or keys—not as a general cleanup function. Spark 4.4.0 documents the API, which defaults to NFC and also accepts NFD, NFKC, and NFKD.

What Unicode normalization does

A character that looks the same on screen can have different Unicode encodings. For example, an accented letter may be stored as one precomposed code point or as a base letter followed by a combining mark. Unicode defines these representations as canonically equivalent. Normalization puts equivalent strings into a consistent representation, including ordering combining marks canonically.

The Unicode Consortium advises: “Programs should always compare canonical-equivalent Unicode strings as equal”. If equality matching or key generation should treat those equivalent encodings alike, normalize inputs consistently before comparing or generating keys. See the Unicode Consortium’s normalization FAQ.

Normalization does not decide every kind of text equivalence. Case folding, transliteration, whitespace and punctuation policies, and language-specific rewriting are separate choices. Spark’s function performs Unicode normalization, not those other forms of cleanup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which normalization form should you choose?

The key choice is whether to preserve compatibility distinctions, and whether downstream systems expect composed or decomposed text. NFC and NFD handle canonical normalization. NFKC and NFKD also apply compatibility normalization, which can collapse distinctions. Choose according to the data contract for your identifiers, search keys, stored values, and downstream consumers.

Form What it does When it may fit
NFC Composes canonically equivalent sequences where a composed form exists. Use when a composed canonical representation is wanted. This is Spark’s default.
NFD Uses canonical decomposition. Use when a decomposed canonical representation is wanted.
NFKC Applies compatibility normalization as well as canonical normalization. Use only when compatibility distinctions should be folded. Spark’s example converts the ligature “fi” to “fi”.
NFKD Applies compatibility decomposition. Use when compatibility decomposition is required by the data contract.

NFKC and NFKD are not universally safer: compatibility normalization may erase distinctions that matter to an application. Check the expectations of every system that consumes the normalized text before changing stored values.

Use normalize in Spark

Spark documents normalize in SQL, Scala DataFrame functions, and PySpark, including classic PySpark and Spark Connect. The API is marked since Spark 4.4.0; check the documentation for the exact Spark release deployed in your environment because versioned documentation can differ.

SQL

SELECT normalize(name);          -- NFC default
SELECT normalize(name, 'NFD');

The form name is case-insensitive. The documented choices are NFC, NFD, NFKC, and NFKD.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scala DataFrame API

import org.apache.spark.sql.functions.normalize

normalize(col("name"))
normalize(col("name"), "NFD")

The one-argument Scala function defaults to NFC; the two-argument form lets you select another accepted form.

PySpark

from pyspark.sql.functions import normalize

normalize("name")
normalize("name", form="NFD")

PySpark documents pyspark.sql.functions.normalize(str, form=None). Spark’s change record lists SQL, Scala DataFrame functions, and PySpark support for both classic PySpark and Spark Connect.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Reproducibility and performance considerations

Spark documents that normalize uses bundled ICU4J rather than the JVM’s own Unicode data, which Spark says provides stable results across JVM vendors and versions. This does not establish that outputs are identical across every Spark release: the bundled library can change. For pipelines where normalized strings are persisted or used in joins, record the Spark release alongside the transformation contract.

The cited API documentation does not establish a workload-specific speed advantage over a user-defined function. Avoid assuming a numerical performance gain; measure the actual workload if performance is a deciding factor.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.