Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

PySpark is Apache Spark’s Python API for distributed data processing. Install it with Python 3.10 or newer and Java 17 or newer, create a SparkSession, and use DataFrames as the default structured API. This cheat sheet covers installation, DataFrame syntax, lazy execution, joins, aggregations, windows, SQL, UDFs, RDDs and ways to run PySpark locally or on a cluster.

Install PySpark

The current Apache Spark installation documentation lists Python 3.10 or newer and Java 17 or newer. Java must be available through a correctly configured JAVA_HOME.

python -m venv .venv
source .venv/bin/activate
pip install pyspark

On Windows PowerShell, activate the environment with ..venvScriptsActivate.ps1. Install only the optional extra that matches your workload:

pip install "pyspark[sql]"
pip install "pyspark[pandas_on_spark]"
pip install "pyspark[connect]"
pip install "pyspark[ml]"

Use the standard package for core DataFrame work; the extras add dependencies for SQL support, pandas API on Spark, Spark Connect or MLlib.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
LAPGEAR Home Office Pro Lap Desk - Black Carbon, Fits 15.6” Laptops
  • Spacious Design: Measuring 21.1" wide and 14.1" deep, our lap desk comfortably fits most laptops up to 15.6". Extra room for accessories ensures convenience.
  • Enhanced Functionality: Packed with handy features, including a 5x9" precision tracking mouse pad and a built-in phone slot for seamless work or video calls. Plus, enjoy ergonomic support with the integrated cushioned wrist rest.
  • Cool Comfort: Enjoy a stable surface with our lap desk's dual bolster cushion, designed for comfort and airflow, keeping your lap cool during extended use.
  • Durable Surface: Work with confidence on our lap desk's solid surface, featuring a sleek black carbon color, ensuring optimal air circulation to prevent your laptop from overheating.
  • On-the-Go Convenience: With an integrated handle and lightweight design (2.8 lbs), our lap desk is portable for travel or moving around the house, offering flexibility in any space.

Create a Spark application

from pyspark.sql import SparkSession

spark = (
    SparkSession.builder
    .appName("example")
    .getOrCreate()
)

getOrCreate() reuses an existing session when one is available, which is convenient in notebooks and tests. Stop a standalone application when its work is complete:

spark.stop()

Create and inspect DataFrames

Create from Python rows

from pyspark.sql import Row

rows = [
    Row(id=1, category="a", value=10),
    Row(id=2, category="b", value=20),
]
df = spark.createDataFrame(rows)

createDataFrame also accepts common row-like Python structures, pandas DataFrames and RDDs. Supply a schema explicitly when stable names and data types matter, especially for empty data or production pipelines.

from pyspark.sql.types import StructType, StructField, LongType, StringType

schema = StructType([
    StructField("id", LongType(), nullable=False),
    StructField("category", StringType(), nullable=True),
    StructField("value", LongType(), nullable=True),
])
df = spark.createDataFrame([(1, "a", 10), (2, "b", 20)], schema)

Inspect rows, columns and types

df.printSchema()
df.show()
df.show(5, truncate=False)
df.columns
df.dtypes
df.select("id", "value").show()

show, like count and collect, is an action: it causes Spark to evaluate the required plan.

Transformations and actions

Transformations describe a new DataFrame and are evaluated lazily. Actions request a result, write output or otherwise trigger execution. A chain of select, filter, withColumn, join and groupBy calls therefore builds a plan without immediately processing every row.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
Transformations (lazy) Actions (execute the plan)
select, filter, where, withColumn, drop, join, groupBy, orderBy show, count, collect, first, take, write

Keep transformations readable and trigger an action only when you need to inspect, return or persist the result. Calling collect() brings all returned rows to the driver, so a large result can exhaust driver memory.

Common DataFrame transformations

from pyspark.sql import functions as F

clean = (
    df
    .filter(F.col("value") > 0)
    .withColumn("value_doubled", F.col("value") * 2)
    .select("id", "category", "value_doubled")
)

clean.show()

Select and rename

df.select("id", F.col("value").alias("amount"))
df.withColumnRenamed("category", "group")

Filter rows

df.filter(F.col("value") >= 10)
df.where((F.col("category") == "a") | F.col("category").isNull())

Use column expressions rather than Python and, or and not; combine Spark conditions with &, | and ~.

Handle nulls and duplicates

df.fillna({"category": "unknown", "value": 0})
df.dropna(subset=["id"])
df.dropDuplicates(["id"])

Sort and limit

df.orderBy(F.col("value").desc()).limit(10)

Group, aggregate and pivot

summary = (
    clean
    .groupBy("category")
    .agg(
        F.count("*").alias("rows"),
        F.avg("value_doubled").alias("avg_value"),
        F.sum("value_doubled").alias("total_value"),
    )
)
summary.show()

Other frequently used aggregates include min, max, countDistinct, first and collect_list. A pivot turns category values into columns:

df.groupBy("category").pivot("status").agg(F.sum("value"))

Joins

joined = left.join(right, on="id", how="left")

The key named in on must exist on both sides. Choose the join type according to which unmatched rows you need:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Yilador Webcam Cover 3 Pack, 0.03 inch Ultra Thin Laptop Camera Cover Slide
  • Note: Not suitable for MacBooks released after 2023 or devices with a protruding front camera; Not applicable to full-screen or notch-style tempered glass screen protectors; Do not use on the rear camera of the phone.
  • 💻 Why Do You Need a Webcam Cover Slide? — Safeguard your privacy by covering your webcam with our reliable webcam cover when not in use. Don't let anyone secretly watch you. Stay protected!
  • ✅ Thin & Stylish — Enhance your laptop's functionality and aesthetics with our 0.027" ultra-thin webcam covers. Seamlessly close your laptop while adding a touch of sophistication.
  • ✅ Fits Most Devices — Compatible with laptops, phones, tablets, desktops! Keep your privacy intact on Ap/ple, Mac/Book, iPh/one, iP/ad, H/P, L/novo, De/ll, Ac/er, As/us, Sa/msung devices.
  • ✅ 365 Days Protection — Our upgraded 3.0 adhesive ensures a strong hold that won't damage your equipment. Experience reliable, long-term privacy protection day in and day out.
Join type Result
inner Only keys present in both DataFrames
left Every row from the left, with matching right-side values when available
right Every row from the right, with matching left-side values when available
full All keys from both sides
left_semi Left rows that have a match; no right columns
left_anti Left rows with no match

For different key names or compound conditions, use an expression:

joined = left.alias("l").join(
    right.alias("r"),
    (F.col("l.customer_id") == F.col("r.id")) & (F.col("l.region") == F.col("r.region")),
    "inner",
)

Window functions

Windows calculate values across related rows without collapsing them into one row per group.

from pyspark.sql.window import Window

w = Window.partitionBy("category").orderBy(F.col("value").desc())
ranked = df.withColumn("rank", F.row_number().over(w))
ranked.show()
w = Window.partitionBy("category").orderBy("event_time")
with_running_total = df.withColumn(
    "running_total",
    F.sum("value").over(w.rowsBetween(Window.unboundedPreceding, Window.currentRow)),
)

Common window functions include row_number, rank, dense_rank, lag, lead, sum and avg. Define both the partition and ordering deliberately; ordering ties can produce nondeterministic row-number assignments.

Use Spark SQL with DataFrames

DataFrame expressions and Spark SQL use the same execution engine and can be mixed. Register a temporary view, run SQL, and continue with the DataFrame API:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
AboveTEK Portable Laptop Lap Desk w/Retractable Left/Right Mouse Pad Tray, Non-Slip Heat Shield Tablet Notebook Computer Stand Table w/Sturdy Stable Work Surface for Bed Sofa Couch or Travel
  • Anti-Slip Surface - Transform your laptop into a mobile workstation with the AboveTEK portable laptop lap desk. The anti-slip surface provides a strong grip for laptops up to 15.6 inches(Diagonal), while the double rubber strip on the bottom ensures a stable display or typing experience on your lap, couch, or bed.
  • Retractable Mouse Pad - Retractable laptop mouse pad extends on both directions for the left/right handed with elevation along the edges for stopping mouse from falling off. The size of laptop tray is 14" X 9.7" and the size of mouse pad is 7.4" X 6.1".
  • Effective Heat Shield - The effective heat shield made of sturdy and thick material protects your laptop from overheating. Prioritizes your comfort and safety, an ideal lap pad or board for working anywhere.
  • EASY to Carry and Store - With an ergonomic and simplistic design, the lap desk is portable to store in a backpack. Only 15" in size, 2.2 lb of weight and with slim 0.6 inch thickness, it is ready to be easily carried around.
  • Widely Applicable - The smooth platform accommodates laptops and tablets up to 15.6 inches(Diagonal), making it a versatile accessory and one of the best gifts for mom, dad, students and professionals. Perfect for use as a laptop bed tray or tablet holder anywhere at home, library, or park.
df.createOrReplaceTempView("items")

result = spark.sql("""
    SELECT category,
           COUNT(*) AS rows,
           AVG(value) AS avg_value
    FROM items
    GROUP BY category
""")

result.show()

Temporary views last for the lifetime of the session. SQL is often convenient for analyst-authored queries, while the DataFrame API makes Python control flow and reusable expressions easier to compose.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Built-in functions, Python UDFs and pandas UDFs

Prefer built-in functions from pyspark.sql.functions whenever they express the operation. Spark can analyze and optimize these expressions directly.

df.withColumn("normalized", F.lower(F.trim(F.col("category"))))

Use a Python UDF only when the required logic cannot be expressed with built-ins. UDFs add Python serialization and dependency considerations, so keep them narrow and test their behavior with nulls and unexpected types.

from pyspark.sql.functions import udf
from pyspark.sql.types import StringType

def label(value):
    return "high" if value and value > 100 else "normal"

label_udf = udf(label, StringType())
df.withColumn("label", label_udf("value"))

Pandas UDFs and mapInPandas process batches through pandas and are useful for vectorized custom logic when supported by your environment. They still require compatible Python dependencies on the workers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
LAPGEAR Home Office Lap Desk – Pink, Fits 15.6” Laptops
  • Spacious Design: Measuring 21.1" wide and 12" deep, our lap desk comfortably fits most laptops up to 15.6". Extra room for accessories ensures convenience.
  • Enhanced Functionality: Packed with handy features, including a 5x9" precision tracking mouse pad and a built-in phone slot for seamless work or video calls. Plus, enjoy laptop support with the integrated device ledge.
  • Cool Comfort: Enjoy a stable surface with our lap desk's dual bolster cushion, designed for comfort and airflow, keeping your lap cool during extended use.
  • Durable Surface: Work with confidence on our lap desk's solid surface, featuring a blush pink color, ensuring optimal air circulation to prevent your laptop from overheating.
  • On-the-Go Convenience: With an integrated handle and lightweight design (2.14 lbs), our lap desk is portable for travel or moving around the house, offering flexibility in any space.

DataFrame versus RDD

Choice Best fit Trade-off
DataFrame Structured rows, joins, aggregations and SQL-style processing Requires expressing work through columns and schemas
RDD Lower-level distributed collections or operations that need explicit object-level control Less schema information and fewer optimizer opportunities than DataFrames
Spark SQL SQL text over registered tables and views Requires SQL syntax and view/table management

DataFrames are implemented on top of RDDs, but the official quickstart presents DataFrames as the main starting point for structured data. Choose RDDs when the lower-level abstraction is genuinely required.

Read and write data

events = spark.read.parquet("data/events")
events.write.mode("overwrite").parquet("output/events")

csv = spark.read.option("header", True).option("inferSchema", True).csv("data/file.csv")
csv.write.mode("append").option("header", True).csv("output/csv")

For production pipelines, define schemas instead of relying on inference when input types must remain stable. Spark writes a directory containing part files and metadata rather than one ordinary local file.

Run locally, with Spark Connect or on a cluster

Local development

The PyPI package is suitable for local development and testing. Start a session with the default local configuration, inspect a small sample, and keep collected results bounded.

Spark Connect

Spark Connect separates the client from the Spark driver. Install the Connect extra when needed and configure the client for the server endpoint documented for your Spark deployment. This adds endpoint, authentication and dependency choices beyond a local session.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cluster deployment

On a cluster, the driver coordinates executors that process partitions. Make Python code and required packages available to workers, use paths visible to the cluster, and size driver memory carefully when collecting or broadcasting data. Exact submission commands and configuration depend on the cluster manager and deployment environment.

Advanced PySpark APIs

  • Structured Streaming: DataFrame-style processing for continuously arriving data.
  • Pandas API on Spark: pandas-like syntax backed by distributed Spark execution.
  • Spark Connect: Client-server access to a remote Spark session.
  • MLlib: Distributed machine-learning algorithms and utilities.

These APIs share Spark concepts but have separate dependency, state, deployment or model-management concerns; consult the API documentation for their feature-specific configuration.

Fast troubleshooting checklist

  • Java gateway or startup error: verify Java 17 or later and that JAVA_HOME points to that installation.
  • Python version error: use Python 3.10 or newer in the active virtual environment.
  • Unexpected type or null behavior: print the schema and provide an explicit schema when creating the DataFrame.
  • Slow custom logic: replace a Python UDF with a built-in expression where possible.
  • Driver out-of-memory failure: avoid unbounded collect(); aggregate or write distributed results instead.
  • Missing worker dependency: install or distribute the same Python packages and code to executors.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.