Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →PySpark is Apache Spark’s Python API for distributed data processing. Install it with Python 3.10 or newer and Java 17 or newer, create a SparkSession, and use DataFrames as the default structured API. This cheat sheet covers installation, DataFrame syntax, lazy execution, joins, aggregations, windows, SQL, UDFs, RDDs and ways to run PySpark locally or on a cluster.
Install PySpark
The current Apache Spark installation documentation lists Python 3.10 or newer and Java 17 or newer. Java must be available through a correctly configured JAVA_HOME.
python -m venv .venv
source .venv/bin/activate
pip install pyspark
On Windows PowerShell, activate the environment with ..venvScriptsActivate.ps1. Install only the optional extra that matches your workload:
pip install "pyspark[sql]"
pip install "pyspark[pandas_on_spark]"
pip install "pyspark[connect]"
pip install "pyspark[ml]"
Use the standard package for core DataFrame work; the extras add dependencies for SQL support, pandas API on Spark, Spark Connect or MLlib.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
- Spacious Design: Measuring 21.1" wide and 14.1" deep, our lap desk comfortably fits most laptops up to 15.6". Extra room for accessories ensures convenience.
- Enhanced Functionality: Packed with handy features, including a 5x9" precision tracking mouse pad and a built-in phone slot for seamless work or video calls. Plus, enjoy ergonomic support with the integrated cushioned wrist rest.
- Cool Comfort: Enjoy a stable surface with our lap desk's dual bolster cushion, designed for comfort and airflow, keeping your lap cool during extended use.
- Durable Surface: Work with confidence on our lap desk's solid surface, featuring a sleek black carbon color, ensuring optimal air circulation to prevent your laptop from overheating.
- On-the-Go Convenience: With an integrated handle and lightweight design (2.8 lbs), our lap desk is portable for travel or moving around the house, offering flexibility in any space.
Create a Spark application
from pyspark.sql import SparkSession
spark = (
SparkSession.builder
.appName("example")
.getOrCreate()
)
getOrCreate() reuses an existing session when one is available, which is convenient in notebooks and tests. Stop a standalone application when its work is complete:
spark.stop()
Create and inspect DataFrames
Create from Python rows
from pyspark.sql import Row
rows = [
Row(id=1, category="a", value=10),
Row(id=2, category="b", value=20),
]
df = spark.createDataFrame(rows)
createDataFrame also accepts common row-like Python structures, pandas DataFrames and RDDs. Supply a schema explicitly when stable names and data types matter, especially for empty data or production pipelines.
from pyspark.sql.types import StructType, StructField, LongType, StringType
schema = StructType([
StructField("id", LongType(), nullable=False),
StructField("category", StringType(), nullable=True),
StructField("value", LongType(), nullable=True),
])
df = spark.createDataFrame([(1, "a", 10), (2, "b", 20)], schema)
Inspect rows, columns and types
df.printSchema()
df.show()
df.show(5, truncate=False)
df.columns
df.dtypes
df.select("id", "value").show()
show, like count and collect, is an action: it causes Spark to evaluate the required plan.
Transformations and actions
Transformations describe a new DataFrame and are evaluated lazily. Actions request a result, write output or otherwise trigger execution. A chain of select, filter, withColumn, join and groupBy calls therefore builds a plan without immediately processing every row.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #2
- 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
- Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
- Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
- HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
- What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
| Transformations (lazy) | Actions (execute the plan) |
|---|---|
select, filter, where, withColumn, drop, join, groupBy, orderBy |
show, count, collect, first, take, write |
Keep transformations readable and trigger an action only when you need to inspect, return or persist the result. Calling collect() brings all returned rows to the driver, so a large result can exhaust driver memory.
Common DataFrame transformations
from pyspark.sql import functions as F
clean = (
df
.filter(F.col("value") > 0)
.withColumn("value_doubled", F.col("value") * 2)
.select("id", "category", "value_doubled")
)
clean.show()
Select and rename
df.select("id", F.col("value").alias("amount"))
df.withColumnRenamed("category", "group")
Filter rows
df.filter(F.col("value") >= 10)
df.where((F.col("category") == "a") | F.col("category").isNull())
Use column expressions rather than Python and, or and not; combine Spark conditions with &, | and ~.
Handle nulls and duplicates
df.fillna({"category": "unknown", "value": 0})
df.dropna(subset=["id"])
df.dropDuplicates(["id"])
Sort and limit
df.orderBy(F.col("value").desc()).limit(10)
Group, aggregate and pivot
summary = (
clean
.groupBy("category")
.agg(
F.count("*").alias("rows"),
F.avg("value_doubled").alias("avg_value"),
F.sum("value_doubled").alias("total_value"),
)
)
summary.show()
Other frequently used aggregates include min, max, countDistinct, first and collect_list. A pivot turns category values into columns:
df.groupBy("category").pivot("status").agg(F.sum("value"))
Joins
joined = left.join(right, on="id", how="left")
The key named in on must exist on both sides. Choose the join type according to which unmatched rows you need:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #3
- Note: Not suitable for MacBooks released after 2023 or devices with a protruding front camera; Not applicable to full-screen or notch-style tempered glass screen protectors; Do not use on the rear camera of the phone.
- 💻 Why Do You Need a Webcam Cover Slide? — Safeguard your privacy by covering your webcam with our reliable webcam cover when not in use. Don't let anyone secretly watch you. Stay protected!
- ✅ Thin & Stylish — Enhance your laptop's functionality and aesthetics with our 0.027" ultra-thin webcam covers. Seamlessly close your laptop while adding a touch of sophistication.
- ✅ Fits Most Devices — Compatible with laptops, phones, tablets, desktops! Keep your privacy intact on Ap/ple, Mac/Book, iPh/one, iP/ad, H/P, L/novo, De/ll, Ac/er, As/us, Sa/msung devices.
- ✅ 365 Days Protection — Our upgraded 3.0 adhesive ensures a strong hold that won't damage your equipment. Experience reliable, long-term privacy protection day in and day out.
| Join type | Result |
|---|---|
inner |
Only keys present in both DataFrames |
left |
Every row from the left, with matching right-side values when available |
right |
Every row from the right, with matching left-side values when available |
full |
All keys from both sides |
left_semi |
Left rows that have a match; no right columns |
left_anti |
Left rows with no match |
For different key names or compound conditions, use an expression:
joined = left.alias("l").join(
right.alias("r"),
(F.col("l.customer_id") == F.col("r.id")) & (F.col("l.region") == F.col("r.region")),
"inner",
)
Window functions
Windows calculate values across related rows without collapsing them into one row per group.
from pyspark.sql.window import Window
w = Window.partitionBy("category").orderBy(F.col("value").desc())
ranked = df.withColumn("rank", F.row_number().over(w))
ranked.show()
w = Window.partitionBy("category").orderBy("event_time")
with_running_total = df.withColumn(
"running_total",
F.sum("value").over(w.rowsBetween(Window.unboundedPreceding, Window.currentRow)),
)
Common window functions include row_number, rank, dense_rank, lag, lead, sum and avg. Define both the partition and ordering deliberately; ordering ties can produce nondeterministic row-number assignments.
Use Spark SQL with DataFrames
DataFrame expressions and Spark SQL use the same execution engine and can be mixed. Register a temporary view, run SQL, and continue with the DataFrame API:
Recommended Free Tools
Rank #4
- Anti-Slip Surface - Transform your laptop into a mobile workstation with the AboveTEK portable laptop lap desk. The anti-slip surface provides a strong grip for laptops up to 15.6 inches(Diagonal), while the double rubber strip on the bottom ensures a stable display or typing experience on your lap, couch, or bed.
- Retractable Mouse Pad - Retractable laptop mouse pad extends on both directions for the left/right handed with elevation along the edges for stopping mouse from falling off. The size of laptop tray is 14" X 9.7" and the size of mouse pad is 7.4" X 6.1".
- Effective Heat Shield - The effective heat shield made of sturdy and thick material protects your laptop from overheating. Prioritizes your comfort and safety, an ideal lap pad or board for working anywhere.
- EASY to Carry and Store - With an ergonomic and simplistic design, the lap desk is portable to store in a backpack. Only 15" in size, 2.2 lb of weight and with slim 0.6 inch thickness, it is ready to be easily carried around.
- Widely Applicable - The smooth platform accommodates laptops and tablets up to 15.6 inches(Diagonal), making it a versatile accessory and one of the best gifts for mom, dad, students and professionals. Perfect for use as a laptop bed tray or tablet holder anywhere at home, library, or park.
df.createOrReplaceTempView("items")
result = spark.sql("""
SELECT category,
COUNT(*) AS rows,
AVG(value) AS avg_value
FROM items
GROUP BY category
""")
result.show()
Temporary views last for the lifetime of the session. SQL is often convenient for analyst-authored queries, while the DataFrame API makes Python control flow and reusable expressions easier to compose.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Built-in functions, Python UDFs and pandas UDFs
Prefer built-in functions from pyspark.sql.functions whenever they express the operation. Spark can analyze and optimize these expressions directly.
df.withColumn("normalized", F.lower(F.trim(F.col("category"))))
Use a Python UDF only when the required logic cannot be expressed with built-ins. UDFs add Python serialization and dependency considerations, so keep them narrow and test their behavior with nulls and unexpected types.
from pyspark.sql.functions import udf
from pyspark.sql.types import StringType
def label(value):
return "high" if value and value > 100 else "normal"
label_udf = udf(label, StringType())
df.withColumn("label", label_udf("value"))
Pandas UDFs and mapInPandas process batches through pandas and are useful for vectorized custom logic when supported by your environment. They still require compatible Python dependencies on the workers.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallBest Value
- Spacious Design: Measuring 21.1" wide and 12" deep, our lap desk comfortably fits most laptops up to 15.6". Extra room for accessories ensures convenience.
- Enhanced Functionality: Packed with handy features, including a 5x9" precision tracking mouse pad and a built-in phone slot for seamless work or video calls. Plus, enjoy laptop support with the integrated device ledge.
- Cool Comfort: Enjoy a stable surface with our lap desk's dual bolster cushion, designed for comfort and airflow, keeping your lap cool during extended use.
- Durable Surface: Work with confidence on our lap desk's solid surface, featuring a blush pink color, ensuring optimal air circulation to prevent your laptop from overheating.
- On-the-Go Convenience: With an integrated handle and lightweight design (2.14 lbs), our lap desk is portable for travel or moving around the house, offering flexibility in any space.
DataFrame versus RDD
| Choice | Best fit | Trade-off |
|---|---|---|
| DataFrame | Structured rows, joins, aggregations and SQL-style processing | Requires expressing work through columns and schemas |
| RDD | Lower-level distributed collections or operations that need explicit object-level control | Less schema information and fewer optimizer opportunities than DataFrames |
| Spark SQL | SQL text over registered tables and views | Requires SQL syntax and view/table management |
DataFrames are implemented on top of RDDs, but the official quickstart presents DataFrames as the main starting point for structured data. Choose RDDs when the lower-level abstraction is genuinely required.
Read and write data
events = spark.read.parquet("data/events")
events.write.mode("overwrite").parquet("output/events")
csv = spark.read.option("header", True).option("inferSchema", True).csv("data/file.csv")
csv.write.mode("append").option("header", True).csv("output/csv")
For production pipelines, define schemas instead of relying on inference when input types must remain stable. Spark writes a directory containing part files and metadata rather than one ordinary local file.
Run locally, with Spark Connect or on a cluster
Local development
The PyPI package is suitable for local development and testing. Start a session with the default local configuration, inspect a small sample, and keep collected results bounded.
Spark Connect
Spark Connect separates the client from the Spark driver. Install the Connect extra when needed and configure the client for the server endpoint documented for your Spark deployment. This adds endpoint, authentication and dependency choices beyond a local session.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteCluster deployment
On a cluster, the driver coordinates executors that process partitions. Make Python code and required packages available to workers, use paths visible to the cluster, and size driver memory carefully when collecting or broadcasting data. Exact submission commands and configuration depend on the cluster manager and deployment environment.
Advanced PySpark APIs
- Structured Streaming: DataFrame-style processing for continuously arriving data.
- Pandas API on Spark: pandas-like syntax backed by distributed Spark execution.
- Spark Connect: Client-server access to a remote Spark session.
- MLlib: Distributed machine-learning algorithms and utilities.
These APIs share Spark concepts but have separate dependency, state, deployment or model-management concerns; consult the API documentation for their feature-specific configuration.
Quick Recap
Fast troubleshooting checklist
- Java gateway or startup error: verify Java 17 or later and that
JAVA_HOMEpoints to that installation. - Python version error: use Python 3.10 or newer in the active virtual environment.
- Unexpected type or null behavior: print the schema and provide an explicit schema when creating the DataFrame.
- Slow custom logic: replace a Python UDF with a built-in expression where possible.
- Driver out-of-memory failure: avoid unbounded
collect(); aggregate or write distributed results instead. - Missing worker dependency: install or distribute the same Python packages and code to executors.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

