Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Yes, but not as a native one-step Spark import. Spark 2.0.1 can read and write HDFS through its Hadoop client libraries, while Excel workbook parsing requires a separate reader or a CSV export. The dependable workflow is to parse the .xls or .xlsx file, create and validate a Spark DataFrame, and write that DataFrame to a new HDFS path—preferably as Parquet for later Spark jobs.

What “directly” means in Spark 2.0.1

The reviewed Spark 2.0.1 SQL documentation demonstrates built-in sources such as JSON and Parquet; it does not list Excel as a native input format. Therefore, a call such as spark.read.format("excel") must not be treated as part of Spark 2.0.1 itself.

Excel ingestion and HDFS persistence are two separate operations:

  1. Read workbook bytes with an Excel-capable library, connector, or a controlled CSV export.
  2. Convert the selected sheet and range into rows with an explicit or checked schema.
  3. Create a Spark DataFrame and validate its contents.
  4. Write the DataFrame to HDFS in a format supported by Spark 2.0.1.

This design also makes workbook-specific decisions—sheet selection, headers, formulas, dates, and error cells—visible instead of hiding them behind an assumed data-source implementation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check the legacy runtime before moving data

Spark 2.0.1 is an old release, so dependencies from current Spark tutorials are not automatically compatible. Its overview specifies Java 7 or newer and Scala 2.11.x for Scala applications, and Spark uses Hadoop client libraries to access HDFS and YARN.

  • Identify the exact Spark 2.0.1 distribution running on the cluster.
  • Record the cluster’s Hadoop distribution and client-library versions.
  • Use a Scala 2.11 build when writing a Scala ingestion application.
  • Ensure the Excel reader or connector is compiled for the same Spark and Scala binary versions.
  • Test the dependency set on a small workbook before submitting a production job.

Do not copy a dependency version from a modern Spark guide without checking these constraints.

Choose an Excel ingestion route

Route Workbook support Advantages Risks and checks
Excel-to-Spark connector Depends on the specific release; connector examples commonly expose sheet, range, schema, and cell-handling options. Can create a DataFrame with less custom parsing code. Compatibility with Spark 2.0.1, Scala 2.11, and the cluster’s Hadoop libraries is not established here and must be verified before adoption.
Apache POI in custom Java or Scala code HSSF reads older binary .xls workbooks; XSSF reads Excel 2007 OOXML .xlsx workbooks. POI also provides an event model for read-only processing. Fine-grained control over sheets, cells, types, formulas, and validation. The simpler user model uses more memory. POI notes that XSSF’s XML handling uses more memory than HSSF’s older binary format; large workbooks need an appropriate streaming or event-based design.
CSV intermediate Only the exported sheet and values, not the workbook structure. Uses Spark’s standard file and DataFrame APIs and is practical for a simple, single-sheet extract. Does not preserve formatting, formulas, or multi-sheet semantics. Delimiters, quoting, nulls, encoding, and types must be specified deliberately.

Apache POI describes XSSF as its pure-Java implementation of the Excel 2007 OOXML (.xlsx) format. Select HSSF or XSSF according to the actual extension, and do not assume that renaming a file changes its format.

Inventory the workbook before parsing

Write down the workbook assumptions that will become part of the ingestion contract:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • File extension and approximate size.
  • Intended sheet name or index.
  • Header row and the first and last data rows.
  • Merged cells, blank rows, hidden rows, and repeated headers.
  • Date columns, identifiers with leading zeroes, mixed-type columns, and expected nulls.
  • Whether formulas should be imported as cached results or formula text.
  • How Excel error cells such as #N/A are represented by the chosen reader.

These choices affect the resulting dataset more than the final HDFS write command does.

Create and validate the DataFrame

Prefer an explicit schema for important columns

Inference can turn an identifier into a number, remove leading zeroes, or choose an unexpected type for a column containing both numbers and text. Supply a schema, or compare an inferred schema with an expected one before writing. Pay particular attention to dates, identifiers, decimal values, and columns containing mixed cells.

Make sheet and range selection explicit

Read only the intended sheet and range when the reader supports those options. A workbook can contain multiple tabs, title blocks, notes, and trailing formatting that should not become records. If the parser cannot reliably select a range, filter and validate the rows in the ingestion layer before constructing the DataFrame.

Check values, not just job completion

After creating the DataFrame, inspect its column names and schema, count rows, and sample representative records. Verify dates, nulls, formula results, error cells, and values with leading zeroes. Compare these checks with the original workbook or a trusted export.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Write the validated dataset to HDFS

Spark 2.0.1’s DataFrame writer can persist a DataFrame to an HDFS destination URI. For structured data that will be processed again by Spark, Parquet is a sensible default because it preserves a typed table layout. CSV is appropriate when another system requires text interchange, but document its delimiter, quoting, encoding, null representation, and type limitations.

  1. Choose a new destination such as hdfs://<nameservice>/data/incoming/workbook-name/ rather than writing over an important existing directory.
  2. Write the DataFrame in the selected format, for example Parquet through the Spark SQL/DataFrame writer available in 2.0.1.
  3. Read the written path back with Spark and verify the schema, row count, columns, nulls, and sample values.
  4. Publish or promote the path only after those checks succeed.

The exact URI depends on the cluster’s configured nameservice and filesystem settings; use the HDFS URI already configured for that deployment.

Handle save modes safely

Spark 2.0.1 documents error, append, overwrite, and ignore save modes. Save modes do not provide locking or atomicity, and overwrite deletes existing data before writing. Consequently:

  • Use a fresh, versioned path for an initial load.
  • Use append only when duplicate or repeated batches are acceptable and separately controlled.
  • Use overwrite only with an explicit replacement and recovery procedure.
  • Do not assume that a failed overwrite leaves the previous dataset intact.

A practical migration sequence

  1. Inspect: record workbook format, sheets, range, headers, formulas, dates, errors, and size.
  2. Align runtimes: confirm Spark 2.0.1, Java, Scala 2.11, Hadoop libraries, and reader dependencies.
  3. Select a reader: use a verified connector, Apache POI, or CSV for a simple sheet.
  4. Parse: select the intended sheet and range and define how each cell type is handled.
  5. Build the DataFrame: apply or validate the schema and reject malformed records according to an explicit policy.
  6. Validate: compare row counts, columns, nulls, dates, formula results, error cells, and representative values.
  7. Persist: write to a new HDFS path in Parquet or a documented interchange format.
  8. Re-read: load the HDFS output with Spark and repeat the essential checks.

Common failure modes

“Excel” format is not recognized

This is expected when only Spark 2.0.1 is installed. Add and verify an external reader, implement parsing with POI, or export the required sheet to CSV.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Dependency or class-version errors

Check Spark and Scala binary versions, the Hadoop client libraries supplied by the distribution, and the reader’s published compatibility. Remove conflicting copies from the driver and executor classpaths.

Rows or columns are unexpectedly missing

Check the selected sheet, range, header-row setting, merged cells, blank-row policy, and whether trailing formatted cells were interpreted as data.

Identifiers or dates changed type

Replace unconstrained inference with an explicit schema and inspect the parser’s conversion rules. Preserve identifiers as strings when formatting is significant.

The output path is damaged after a retry

Because overwrite is destructive and save modes are not atomic, write to a new path, validate it, and switch consumers only after success. Restore the prior path from your operational backup or published version if necessary.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Frequently Asked Questions

Can Spark 2.0.1 read an .xlsx file without another library?

No. Excel parsing is not documented as a built-in Spark 2.0.1 SQL source; use a compatible connector, Apache POI, or a CSV export.

Should the HDFS result be CSV or Parquet?

Use Parquet for typed data consumed by Spark. Use CSV only when text interchange is required and its delimiter, quoting, encoding, null, and type rules are documented.

The Bottom Line

The safe migration is parse, validate, and then write: Spark 2.0.1 supplies the HDFS and DataFrame machinery, but an Excel-capable reader supplies workbook interpretation. Verify every dependency against the legacy Spark/Scala/Hadoop stack, write to a fresh path, and validate the HDFS dataset before replacing existing data.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.