Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Use SparkSession.createDataFrame() to build a PySpark DataFrame from Python data, or use spark.read to load one from a file. The examples below cover both approaches, show how to control the schema, and explain how to inspect and troubleshoot the result.
What is a PySpark DataFrame?
A PySpark DataFrame is Spark’s table-like abstraction: rows are organized into named columns, and a schema describes each column’s data type and whether it can be null. Spark DataFrames are equivalent to relational tables in Spark SQL (Apache Spark DataFrame API).
A PySpark DataFrame is not a pandas DataFrame. Spark represents data and computation for distributed execution; pandas primarily operates on data in local memory. Spark transformations such as filter and select are generally evaluated lazily, while actions such as show and count request results and can trigger computation.
For structured data, DataFrames are usually a better starting point than manually manipulating RDDs: they expose named columns and schemas and integrate with Spark SQL. RDDs remain supported and can be useful when data already exists as an RDD or a specific use case calls for one.
#1 Best Overall
- Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
- Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
- Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
- Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
- 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
Start or reuse a SparkSession
SparkSession is the entry point for Spark functionality (Spark SQL getting started). In a standalone local script, create or reuse a session like this:
from pyspark.sql import SparkSession
spark = (
SparkSession.builder
.master("local[*]")
.appName("Create DataFrame")
.getOrCreate()
)
local[*] is for local development and uses the available local cores; cluster deployments use different settings. In a notebook or application, create or reuse one session rather than creating a new one repeatedly inside each function. The PySpark shell normally provides a spark session already.
If Spark fails to start, the problem may be the local Java, Python, or PySpark setup rather than DataFrame code. These commands help identify the versions in use:
python --version
python -c "import pyspark; print(pyspark.__version__)"
java -version
Create a DataFrame from Python data
Tuples with column names
For fixed tabular records, a list of tuples is a straightforward starting point. Pass the column names in the same order as the values in every tuple:
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsdata = [
("Alice", 29),
("Bob", 35),
("Charlie", 41),
]
df = spark.createDataFrame(data, ["name", "age"])
df.show()
df.printSchema()
The first value in each tuple becomes name and the second becomes age. A row with the wrong number of values, or values incompatible with the schema, can cause an exception.
show() displays a tabular preview. printSchema() reveals the types Spark assigned; always check it rather than assuming values were interpreted as intended.
Lists of lists
Lists of lists also work when each inner list represents a row and you supply the column names:
data = [
["Alice", 29],
["Bob", 35],
["Charlie", 41],
]
df = spark.createDataFrame(data, ["name", "age"])
Tuples are often more idiomatic for fixed records; dictionaries and Row objects make field names more visible in the input.
Rank #2
- 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
- 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
- Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
- 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
- What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.
Dictionaries
A list of dictionaries is convenient when each record is easier to read as named fields:
data = [
{"name": "Alice", "age": 29},
{"name": "Bob", "age": 35},
{"name": "Charlie", "age": 41},
]
df = spark.createDataFrame(data)
df.show()
Keep the records’ fields and value types compatible. Do not treat dictionary key order as the schema contract; specify a schema when column definitions and ordering need to be controlled. Missing keys may result in null values or schema-related problems depending on the input and Spark version.
Named Row objects
Row attaches field names directly to each record:
from pyspark.sql import Row
data = [
Row(name="Alice", age=29),
Row(name="Bob", age=35),
Row(name="Charlie", age=41),
]
df = spark.createDataFrame(data)
df.show()
Row is readable in small examples; use an explicit schema when types and nullability must be controlled precisely. The PySpark DataFrame user guide covers dictionaries, pandas DataFrames, and other supported inputs.
Define and inspect the schema
When Spark infers types, the values determine the result. For example, if ages are supplied as quoted text, the column can be a string even though the values look numeric:
data = [("Alice", "29"), ("Bob", "35")]
df = spark.createDataFrame(data, ["name", "age"])
df.printSchema()
For repeatable pipelines, declare types and nullability explicitly with a StructType:
from pyspark.sql.types import StructType, StructField, StringType, IntegerType
schema = StructType([
StructField("name", StringType(), nullable=False),
StructField("age", IntegerType(), nullable=True),
])
data = [("Alice", 29), ("Bob", 35), ("Charlie", 41)]
df = spark.createDataFrame(data, schema=schema)
df.printSchema()
df.show()
The schema output shows name as a non-nullable string and age as a nullable integer. Input values must be compatible with the schema. For a short example, a schema string is more compact:
df = spark.createDataFrame(data, schema="name string, age int")
Use StructType when a schema will be reused, documented, nested, or built programmatically. The createDataFrame API accepts column-name lists, Spark data types or schemas, and schema strings. Its documented signature also includes samplingRatio and verifySchema; for RDD input, samplingRatio relates to schema inference.
When inference is suitable
Inference is convenient for tiny examples and exploratory work, but it is not a data-quality check. Mixed values, null-only columns, empty input, or inconsistent records can produce surprising types or prevent Spark from determining a schema. For production ETL and stable nested data, an explicit schema makes the expected structure clearer.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #3
- Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
- Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
- Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
- Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
- What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
Create an empty DataFrame
Spark cannot infer column types from an empty collection. Supply a schema:
from pyspark.sql.types import StructType, StructField, StringType, IntegerType
schema = StructType([
StructField("name", StringType(), True),
StructField("age", IntegerType(), True),
])
empty_df = spark.createDataFrame([], schema)
empty_df.show()
empty_df.printSchema()
Create DataFrames from pandas or an RDD
From a pandas DataFrame
For data already in pandas, pass the pandas DataFrame to createDataFrame:
import pandas as pd
pdf = pd.DataFrame({
"name": ["Alice", "Bob", "Charlie"],
"age": [29, 35, 41],
})
df = spark.createDataFrame(pdf)
df.show()
This conversion workflow requires the pandas data to fit in driver-side memory, so it is not a way to ingest arbitrarily large data. Pandas and Spark types do not map perfectly in every case. Arrow can improve conversion performance in supported configurations, but adds dependency and compatibility considerations; the API documents Arrow-related differences in schema verification. For large sources, read them directly with Spark instead of loading them into pandas first.
From an RDD
If the records already exist as an RDD, Spark can construct a DataFrame from it:
rdd = spark.sparkContext.parallelize([
("Alice", 29),
("Bob", 35),
("Charlie", 41),
])
df = spark.createDataFrame(rdd, ["name", "age"])
You can also pass an explicit schema: spark.createDataFrame(rdd, schema=schema). For new structured data, prefer creating the DataFrame directly from the collection unless the RDD already exists or the use case specifically requires one. Spark’s SQL getting-started guide also demonstrates applying a schema to RDD records.
Read DataFrames from files
createDataFrame constructs a DataFrame from in-memory Python data or an RDD. spark.read loads one from an external source. Use the latter when the data is in files rather than already held in Python.
CSV
For a CSV file with a header row, enable the header and, for exploration, ask Spark to infer column types:
df = spark.read.csv(
"people.csv",
header=True,
inferSchema=True
)
df.show()
df.printSchema()
Equivalent reader options can be chained. For example, nullValue tells the reader which text marker represents a null value:
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #4
- Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
- Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
- Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
- Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
- Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
df = (
spark.read
.option("header", True)
.option("inferSchema", True)
.option("sep", ",")
.option("nullValue", "NA")
.csv("people.csv")
)
For a repeatable pipeline, provide the schema instead of relying on inference:
df = (
spark.read
.schema(schema)
.option("header", True)
.csv("people.csv")
)
CSV data is text, so columns can remain strings when inference is disabled or unsuccessful. inferSchema=True concerns type inference; it does not repair malformed records or validate data quality.
JSON
Read JSON with spark.read.json. A common input is newline-delimited JSON, with one object per line:
{"name": "Alice", "age": 29}
{"name": "Bob", "age": 35}
df = spark.read.json("people.json")
df.show()
df.printSchema()
Nested objects can remain structured as fields. For example, given a record such as {"name": "Alice", "address": {"city": "Boston"}}, select the nested city with df.select("name", "address.city").show() rather than flattening every nested field immediately.
Parquet
Read Parquet with the DataFrame reader:
df = spark.read.parquet("people.parquet")
Unlike plain CSV text, Parquet carries schema information with the data, which is useful for repeated analytical workloads.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Inspect, query, and validate a DataFrame
These methods answer different questions:
df.show()displays a preview.df.printSchema()prints column names, types, and nullability.df.columnsreturns the column names.df.dtypesreturns column names and type names.df.count()computes the row count and can trigger work.
Use DataFrame operations to select or filter rows:
df.select("name").show()
df.filter(df.age > 30).show()
df.describe().show() provides a basic summary, including numeric summaries where applicable. Actions such as count() can trigger computation because Spark evaluates transformations lazily.
Use collect() only when the full result is small enough for the driver: it transfers every row to driver memory and can exhaust it for a large DataFrame. For inspection, prefer df.show(20, truncate=False), df.limit(20).collect(), or df.take(20). The PySpark DataFrame quickstart demonstrates displaying and inspecting DataFrames and explains driver-side collection.
Query with SQL
Register a temporary view to query a DataFrame with Spark SQL:
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallBest Value
- 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
- Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
- Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
- HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
- What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
df.createOrReplaceTempView("people")
result = spark.sql("""
SELECT name, age
FROM people
WHERE age >= 30
""")
result.show()
A temporary view makes the DataFrame available to SQL in the session; registering it does not create a permanent table or write data to storage. See the Spark SQL getting-started guide for temporary-view usage.
Troubleshoot common creation errors
“Can not infer schema from empty dataset”
An empty collection has no values from which Spark can infer types. Pass an explicit schema, as in spark.createDataFrame([], schema).
“Some of types cannot be determined”
Null-only or ambiguous columns may leave Spark without enough information to infer a type. Provide a StructType, ensure representative non-null values are available, or normalize the Python values before creating the DataFrame.
Row length or value types do not match
Check that each record has the same number of fields as the schema or column-name list, and that values match the expected types. For example, a three-field tuple cannot fit a two-column schema. Clean inconsistent values such as an integer in one row and non-numeric text in another, or define and apply an appropriate conversion before creation.
Recommended Free Tools
Numeric-looking values are strings
Quoted values such as "29" are text, so Spark may infer a string column. Convert values before creation where possible, or cast an existing column:
from pyspark.sql.functions import col
df = df.withColumn("age", col("age").cast("int"))
CSV columns have the wrong types
CSV is text. Enable inference for exploration with inferSchema=True, or set a schema explicitly for a repeatable load. Neither approach repairs malformed input records.
pandas conversion fails or is slow
Check that pandas is installed and inspect pdf.dtypes for ambiguous object columns, nullable integers, or dates. If Arrow optimization is enabled, a missing or incompatible PyArrow dependency may be involved; temporarily disabling Arrow can help isolate the conversion path. For a large pandas object, use Spark’s reader on the source instead of sending the entire object through the driver.
Spark startup fails before DataFrame creation
Check that the installed Java, Python, and PySpark versions work together and that the local Java configuration is valid. Compare the version reported by your environment with the documentation for the release you actually installed; the API page describes the currently documented method, not necessarily an older local installation.
Free tools Windows power users keep installed
One-click scans. No signup required.
Practical habits for reliable DataFrames
- Use
SparkSessionas the standard entry point. - Use inference for small examples or exploration; define schemas for repeatable pipelines.
- Inspect
printSchema()after creation or loading. - Keep input fields and types consistent, and define a schema for empty data.
- Read large sources directly with Spark instead of first collecting them into pandas.
- Limit previews and avoid collecting large results to the driver.
- In a standalone script, stop the session when the application is finished; do not stop it after every notebook cell.
The standard entry point is SparkSession.createDataFrame(data, schema) for in-memory records and spark.read for files. The currently documented API lists support for RDDs, iterables, pandas DataFrames, and NumPy arrays; Apache Arrow table input was added in Spark 4.0.0. The method dates to Spark 2.0.0, and Spark Connect support was added in 3.4.0, so check the API documentation for your installed version before relying on newer input types.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

