Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Moving an AWS Glue job to OCI Data Flow is two migrations at once: a Spark application port and a cloud-service migration. The PySpark transformation logic can often move with little change. What does not move by itself is everything Glue provides around that code: managed job state such as bookmarks, the Glue Data Catalog and connections, IAM and network access, dependency packaging, scheduling, and monitoring. Each of those needs an explicit replacement or a redesign. A job that launches successfully on OCI has not yet been shown to be a correct port. That proof comes from comparing its outputs with the Glue job’s outputs on representative workloads.

What changes besides the PySpark code

Glue Spark jobs run in an AWS-managed Spark environment. Data Flow runs ordinary Spark applications in Oracle’s environment. The table below lists the areas that usually need a decision, based on what each platform’s documentation describes.

Area AWS Glue behavior OCI Data Flow behavior What you must decide
Runtime and versions Job uses a Glue version with its own Spark and Python versions; AWS’s migration guidance covers version differences (AWS, “Migrate Apache Spark programs to AWS Glue”) Runtime is selected per application; Oracle’s tutorial notes that the Spark session is created before your application starts (Oracle, “Migrating Spark Applications to Oracle Cloud Infrastructure Data Flow”) Which Data Flow Spark and Python versions match the Glue job closely enough to test against
Incremental state Job bookmarks store state that depends on job initialization, commit, and a consistent transformation context (AWS, “Using job bookmarks”) No documented path carries Glue bookmark state into Data Flow How the new job selects new input and records progress
Metadata Data Catalog tables and DynamicFrame access Can use a Hive-compatible metastore (Oracle, “Set Up Data Flow”) Whether catalog metadata is recreated, or the job reads paths directly
Connections and network Glue connections place private access in a selected VPC subnet with security groups (AWS, “Setting up network access to data stores”) Private access to on-premises or private systems is designed around OCI private endpoints and existing network connectivity (Oracle, “Importing an Apache Spark Application to the Oracle Cloud”) Routing, DNS, firewall rules, and reachability for every store
Identity IAM role and connection credentials, depending on the source Runs use permissions of the user who starts them for IAM-compatible services (Oracle, “Security”) Which principal runs jobs, and how non-IAM credentials are supplied
Configuration Job parameters are passed as arguments (AWS, “Using job parameters in AWS Glue jobs”) Run arguments and parameters, plus a supported set of Spark properties; environment variables cannot be set Where each current setting will live
Monitoring Your existing Glue job monitoring and alerting Run output, statistics, driver and executor logs, and Spark UI access (Oracle, “Run Applications”) Where logs must be centralized, and what replaces each current alert

Step 1: Inventory the Glue job before touching the code

Start by writing down what the job actually does, including the parts that are not in the script. A migration plan built only from the PySpark file will miss triggers, arguments, and state.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Record the job definition

  • Job type: batch or streaming. AWS notes that some Spark job features do not apply to streaming ETL jobs (AWS, “AWS Glue Spark and PySpark jobs”). Confirm that your job type is supported on Data Flow before planning further.
  • Glue version, Spark version, and Python version.
  • Script entry point, and every library or extra file the job loads.
  • Job arguments, with their defaults.
  • Retry settings, output behavior (append, overwrite, or partitioned writes), and failure handling.
  • Triggers, workflows, schedules, and event sources that start the job.
  • Monitoring and alert destinations.

Find the Glue-specific code

Search the code and job definitions for these items. Each one is either a Glue service call or Glue-managed behavior that has no direct equivalent in plain Spark:

  • GlueContext and DynamicFrame usage
  • Data Catalog references
  • Glue connection references
  • job.init and job.commit
  • transformation_ctx values
  • Bookmark-related arguments
  • Glue-specific transforms
  • AWS SDK calls made from inside the job

Plain Spark transformations are the part most likely to port directly. Everything in the list above is the part that needs a replacement decision.

Document each source and sink

For every input and output, record the format, schema or catalog dependency, authentication method, network path, read or write mode, partitioning, failure behavior, and whether the job is incremental. For each bookmark, write down the key or source-selection logic and how the job currently prevents duplicate or missed output. AWS states that user-defined JDBC bookmark keys must be strictly monotonic, and that changing a source or its transformation context can invalidate earlier bookmark behavior (AWS, “Using job bookmarks”).

Step 2: Establish the OCI application baseline

Build the OCI side as a working Data Flow application before you move the migrated code into it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Choose a Data Flow Spark runtime. Compare its Spark and Python versions with the Glue job’s versions. Where they differ, plan to test the differences rather than assume compatibility.
  2. Check every custom Spark setting. Test each property against Data Flow’s supported-property list. Oracle’s migration tutorial identifies Spark properties that cannot be set or overridden, because Data Flow creates the Spark session before the application starts (Oracle, “Migrating Spark Applications to Oracle Cloud Infrastructure Data Flow”).
  3. Remove environment-variable dependencies. Oracle’s tutorial states: “You can’t set environment variables in Data Flow jobs.” Move each value your code reads from the environment into a run argument or into application-level configuration that the code loads at startup.
  4. Run a minimal application first. Confirm that Data Flow can start an application, read a small input, write a small output, and produce logs. Only then add the migrated transformation code. This separates platform problems from logic problems.
  5. Upload artifacts to OCI Object Storage. Data Flow hosts applications in Object Storage. Confirm that the run principal can read the application file and every asset it references.
  6. Package dependencies for the language. For Java or Scala, bundle dependencies into an uber or assembly JAR, and check for runtime library conflicts. Oracle’s guidance on shading applies where it is relevant. For Python, follow Data Flow’s Spark-submit and package mechanism for third-party packages. Do not treat a zip of your project as a runnable application; the entry file and dependency layout must follow Oracle’s documented structure.

Step 3: Replace the Glue integrations

DynamicFrames and the Glue Data Catalog

Locate every DynamicFrame and catalog reference. For each one, decide whether to convert it to a Spark DataFrame, and whether the table metadata should be recreated in an OCI metastore or replaced by direct path reads. Data Flow can use a Hive-compatible metastore, and Oracle’s setup documentation identifies the storage buckets used for managed and external tables (Oracle, “Set Up Data Flow”). Catalog metadata and Glue connection objects do not become OCI resources automatically. Each one must be recreated or replaced on purpose.

Connections and secrets

Translate each endpoint and authentication path explicitly. Glue connections supply both the data access settings and the network configuration for a source, so each connection needs an equivalent on the OCI side. Data Flow runs obtain permissions from the user who starts them for IAM-compatible services. For services that do not accept IAM, Oracle’s security documentation points to credential and key management (Oracle, “Security”). Do not embed secrets in code or pass them as application arguments.

Bookmarks and retries

No documented path carries Glue bookmark state into Data Flow. Treat the job’s progress state as something you must migrate or rebuild. The usual approach is to design explicit checkpointing or incremental selection on the OCI side, and then test the cases that expose gaps:

  • A restart after a mid-run failure
  • A full rerun of the same input
  • Late-arriving data that falls on an earlier boundary
  • Duplicate output produced by a retry

Record who owns the checkpoint, and what the first run does. A first run that reprocesses all history is a backfill decision, not a default.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Triggers and orchestration

List every Glue trigger, workflow, schedule, event source, retry rule, and alert. Map each one to the orchestration system you use outside Spark, and test ordering and retry behavior. Do not assume a one-to-one trigger conversion. The correct mapping depends on your orchestration platform and on how dependent jobs are chained, and that is outside what Oracle or AWS documentation can answer for you.

Step 4: Move data and establish network access

Data Flow is designed around OCI Object Storage. Oracle states that access is highly performant when the application and the data are in the same OCI region. Data Flow can also read other Spark-supported sources, such as relational databases. For on-premises systems, Oracle’s import guide describes private endpoint access using an existing FastConnect configuration (Oracle, “Importing an Apache Spark Application to the Oracle Cloud”).

On the AWS side, a Glue job that reaches a private JDBC source uses elastic network interfaces in the selected subnet, and every JDBC store it accesses must be reachable from that subnet (AWS, “Setting up network access to data stores”). The OCI equivalent is a different design. Do not treat VPC security groups and OCI network policy as interchangeable settings.

For each source and sink, confirm the following before cutover:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • The connector and driver that Data Flow will use
  • The credentials and how they are supplied
  • Routing and DNS from the Data Flow run to the store
  • Firewall rules on both ends
  • The OCI region of the store and of the application
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Step 5: Map job configuration and resources

Translate Glue job arguments into Data Flow run arguments and parameters. Then choose the driver and executor shapes and counts again. The sizing units on the two services are different, and the sources reviewed for this article do not establish a general worker-count conversion, so do not carry Glue worker counts over as if they were equivalent. Start from the workload, and benchmark with representative input sizes, data skew, shuffle volume, and output patterns. Confirm every custom Spark property against the supported list before the first full run.

Step 6: Validate output correctness and operations

A successful launch does not establish parity. Compare the Glue job and the Data Flow job on the same inputs, and check the following:

  • Small, normal, and peak input sizes
  • Schema and null handling
  • Partition counts and file layout
  • Ordering assumptions in downstream reads
  • Incremental boundaries and checkpoint positions
  • Restart behavior and failure and retry behavior
  • Runtime and resource use compared with the Glue baseline

Check the operational side as well. Data Flow reports run output, run statistics, and driver and executor logs, and gives access to the Spark UI (Oracle, “Running an Application”). If logs must be centralized, configure OCI Logging policies and destinations as described in Oracle’s application logging documentation (Oracle, “Data Flow Application Logging”).

Review Data Flow’s maximum run duration for batch jobs against the length of your longest run. Oracle documents automatic stopping for long-running batch jobs, with different maximum periods depending on whether the run uses delegation tokens or resource principals. These limits are volatile, so confirm the current rule for your chosen configuration and region before cutover.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Step 7: Cut over safely

  1. Run the Glue job and the Data Flow job in parallel on bounded, controlled inputs where practical.
  2. Reconcile outputs and state between the two, including checkpoint positions.
  3. Define rollback conditions in advance, such as a row-count mismatch, a schema change, or a failed restart test.
  4. Assign ownership of checkpoints and of the first-run and backfill policy.
  5. Switch schedules only after reconciliation passes, and keep the Glue job available until the rollback window closes.

The right cutover method depends on data volume, how mutable the source is, which downstream consumers depend on the output, and how much duplicate or missing data the business can accept. Those inputs belong to your team, not to a generic runbook.

Deciding whether to migrate a given job

Before committing, compare the job against each of these dimensions. Current documentation supports the technical comparisons, but it does not establish a universal cost or performance winner, so the decision must rest on your own measurements.

  • Supported workload type: batch or streaming
  • Spark and Python version fit
  • Glue-specific feature use, especially bookmarks and Catalog dependence
  • Data source and connector coverage
  • Private network reachability
  • Identity and secrets handling
  • Catalog and file formats
  • Incremental and checkpoint semantics
  • Dependency packaging
  • Available Spark properties
  • Runtime and throughput under representative loads
  • Maximum run duration
  • Logging and alerting
  • Orchestration and retries
  • Regional data placement
  • Total operating cost

A job that scores well on the first dimensions but depends heavily on bookmark state or private connections usually needs the most migration effort, and that is the place to start planning.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.