Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A production AWS Glue pipeline can run faster when you optimize the work that actually limits it—not simply by adding Spark capacity. In a case study published September 26, 2026, engineer Kiran Gunturu reports reducing an Oracle-to-S3 pipeline’s end-to-end runtime by about 56% across four workstreams: ingestion orchestration, output publishing, ZIP compression, and runtime choice for SFTP transfer. Two investigations also exposed a JDBC partitioning problem and a missing socket timeout. These are results from one environment, not AWS benchmarks or guaranteed outcomes.
What changed, and what did the case study measure?
Gunturu describes an AWS Glue and PySpark pipeline that ingests data from Oracle to S3, then packages, encrypts, and transfers files to a downstream analytics platform. His figures come from his own runs; he says environment-specific details were genericized. They have not been independently validated. The reported runtime improvements apply to the specific workloads and configurations below, rather than indicating that any single setting will deliver the same gain elsewhere.
| Workstream | Reported result | What changed |
|---|---|---|
| Ingestion orchestration | About 13 minutes 55 seconds to about 8 minutes | Step Functions Map concurrency increased from 30 to 50 for 63 tables, reducing the number of execution waves from three to two. |
| Publishing output | About 25 minutes to 7 minutes | Final bytes were written once from executors, avoiding separate full-data rename, newline, and standardization passes; reconciliation was parallelized. |
| ZIP compression | About 9 minutes to 4 minutes | Compression was threaded and DEFLATE lowered to level 1, trading CPU time for somewhat larger archives. |
| SFTP transfer runtime | About $63 to $11 estimated annual cost at the author’s stated rate | A single-stream transfer moved from Glue Spark at six DPU on G.2X to Glue Python Shell at one DPU, with observed throughput retained in that setup. |
Across these four workstreams, Gunturu reports about a 56% end-to-end runtime reduction. The AWS tuning principle that generalizes is to set a performance goal, measure, find the bottleneck, reduce its impact, and measure again—not to assume that more parallelism or more compute is automatically faster. See AWS Glue performance tuning.
Why did changing concurrency speed up ingestion?
The workload had 63 tables launched through a Step Functions Map. At concurrency 30, the work required three waves; the author reports a total wall-clock time of about 13 minutes 55 seconds. Raising concurrency to 50 let the same set complete in two waves, in about eight minutes.
#1 Best Overall
This is a wave-boundary effect in that workload: overall completion waits for the last task, so reducing a whole wave can matter more than a small improvement to an individual task. It is not evidence that raising concurrency always reduces runtime. Task duration, service limits, source capacity, and uneven table sizes can change the result. Gunturu also suggests starting larger tables early, so heavy work is less likely to be left in the final wave.
How did one final write shorten the publishing stage?
The original publishing flow wrote output, then reread and rewrote it in separate passes for renaming, newline handling, and standardization. Gunturu reports cutting this stage from about 25 minutes to seven by producing the final bytes once from executors and parallelizing reconciliation.
Rank #2
The useful diagnostic is whether a stage repeatedly reads and writes the same data to perform transformations that could happen during production. If it does, consolidating compatible transformations can reduce full-data passes. Confirm that the output still meets downstream naming, formatting, and reconciliation requirements before removing intermediate steps.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
When does parallel compression help—and when can it run out of memory?
Thread ZIP compression when CPU is the bottleneck
The outbound job zipped files, GPG-encrypted each ZIP, and uploaded the results with SSE-KMS. Although it ran on Glue Spark with six DPUs on G.2X, the compression and encryption work was serial on the driver, leaving Spark executors mostly idle. Threading ZIP compression and lowering DEFLATE to level 1 reportedly reduced compression from about nine minutes to four. Archives became slightly larger, so the tradeoff depends on whether reduced CPU time is worth additional storage and transfer bandwidth.
Rank #3
Avoid recompressing ZIP files with GPG
Gunturu found that GPG’s default compression recompressed already-compressed ZIP input. In one reported batch, the output grew from 5,246 MB to 5,312 MB. Setting --compress-algo none avoided that redundant compression in his configuration.
Do not parallelize encryption by copying multi-gigabyte files into driver memory
In this case, parallelizing encryption caused driver out-of-memory failures: each multi-gigabyte ZIP and an armored copy were held in memory. The author describes file streaming, isolated GPG home directories, and binary rather than armored output as a safer design, but does not report a measured production result for that proposed design. Streaming and isolation are ways to avoid a memory-heavy driver bottleneck; they still need validation with the actual file sizes, key setup, and downstream requirements.
Rank #4
Why move an SFTP transfer from Glue Spark to Python Shell?
The SFTP task sent large files to an external endpoint and was characterized by Gunturu as bandwidth-bound. He reports about 11 MB/s for files around 5 GB, with no throughput improvement from greater connector concurrency. In his setup, moving the single-stream transfer from Glue Spark at six DPU on G.2X to a one-DPU Glue Python Shell job retained observed throughput.
Recommended Free Tools
Using the stated rate of $0.44 per DPU-hour, the author estimated annual cost falling from about $63 to $11. Those are his environment-specific estimates, not generally applicable AWS prices or a cost forecast for another schedule. The broader decision is whether the job needs distributed compute: a serial, I/O-bound transfer may gain little from Spark executors when the external connection is the limiting factor.
Best Value
What did the JDBC partitioning investigation uncover?
Gunturu investigated a discrepancy in Oracle ingestion: about 8.9 million rows took eight minutes in a smaller test, while 25 million rows in production took two hours 45 minutes and continued climbing. He initially suspected skew; a bucket-distribution query reportedly showed near-uniform distribution instead.
His reported diagnosis was that a computed-expression partition column could not use an Oracle index. As a result, each JDBC connection scanned the full time window to find its own slice. He proposed partitioning on a column the source can prune, but explicitly says that fix had not yet been implemented and measured. Treat this as a case-specific diagnosis and unverified remedy, not a demonstrated speedup.
For the general principles, AWS explains that pushdown can apply filters closer to the data source to reduce scanning and processing. Its guidance also covers partition pruning, parallelism, and related tuning tradeoffs. Glue’s JDBC pushdown documentation describes custom SQL and parallel sample queries; it is context for configuring Glue, not proof of what happened in this Oracle system:
Free tools Windows power users keep installed
One-click scans. No signup required.
- AWS Glue pushdown
- AWS Prescriptive Guidance: parallelize tasks
- AWS Big Data Blog: optimize memory management in AWS Glue
- AWS Glue JDBC connections
How can a missing socket timeout hold up a Spark stage?
A separate run reportedly hung for two hours 42 minutes, then completed in 14 minutes on an identical rerun. Gunturu attributes the delay to a JDBC socket read without a timeout: one connection silently died, but its task did not return, preventing Spark from completing the stage. This is his diagnosis of that incident, not a general guarantee about JDBC failure behavior.
His recommendation is to configure query and Oracle JDBC read/connect timeouts and enable Spark speculation so a slow or stuck read can be retried. He calls out a configuration trap: oracle.net.CONNECT_TIMEOUT is in milliseconds as a connection property but seconds when supplied bare in the URL. Verify the units and syntax for the Oracle JDBC driver version and connection method in your own Glue job before deploying settings; a timeout that is too short can also interrupt legitimate slow connections.
Quick Recap
How to apply the lessons to another Glue pipeline
- Measure each stage separately. Establish an end-to-end goal and capture where wall-clock time is spent before changing worker counts or parallelism.
- Check for wave boundaries and skew. For orchestration fan-out, compare task counts, concurrency, and task durations. Determine whether a slow final wave or a few large tasks control completion.
- Count full-data passes. Identify repeated reads and writes in publishing or transformation stages, then combine only operations that preserve required output behavior.
- Match parallelism to the resource. Distinguish executor work from driver-side CPU or memory work, and distinguish compute-bound tasks from transfers limited by network bandwidth.
- Inspect source-side query behavior. Confirm partition predicates can be pruned or indexed by the source database; do not assume evenly distributed keys mean efficient scans.
- Set and validate failure bounds. Check query, connection, and socket timeouts, including units, and test retry or speculation behavior under the JDBC driver and Glue/Spark versions you use.
- Measure again under representative conditions. Compare runtime, resource use, output size, data correctness, and cost assumptions; a proposed fix is not a win until the relevant workload verifies it.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

